A Beginner’s Guide to Protein Crystallization and X-ray Data Processing

The ability to determine and visualize the atomic structure of proteins is invaluable in modern drug discovery. Understanding how proteins are arranged in three dimensions not only helps researchers investigate biological function but also supports the rational design of new therapeutic molecules that can interact with specific targets.

One of the most widely used techniques for determining protein structures is X-ray crystallography. However, obtaining a protein structure is far from a single-step process. Success depends on two critical stages: generating high-quality protein crystals and extracting meaningful information from the resulting X-ray diffraction data. Both stages require specialist expertise and, often, a great deal of patience.

Side-by-side circular microscope images under polarised light. The left image shows two elongated, diamond-shaped crystals within a pink and grey field. The right image shows a single rod-shaped crystal against a colourful iridescent background with shades of orange, pink, purple and blue.
Figure 1. Examples of protein crystals obtained during crystallization screening. Growing well-ordered crystals is often the rate-limiting step in X-ray crystallography and typically requires screening a wide range of experimental conditions.

As its name suggests, X-ray crystallography relies on the availability of well-ordered three-dimensional crystals of the target protein. For many projects, producing these crystals is the rate-limiting step in the entire structure determination process.

Unlike other experimental techniques, there is currently no reliable way to predict the optimal crystallization conditions for a protein based solely on its amino acid sequence. Instead, crystallization remains an empirical process that relies on screening large numbers of conditions and systematically optimizing promising hits. This trial-and-error approach has historically given protein crystallization a reputation as something of a “black art” among non-specialists.

Researchers may screen thousands of different crystallization conditions while varying factors such as pH, temperature, salts, precipitants, and protein concentration in an effort to identify conditions that produce suitable crystals. When successful, these crystals provide the highly ordered molecular arrangement required for X-ray diffraction experiments.

Although extensive literature exists on protein crystallization methodologies there remains a need for accessible introductions aimed at scientists who may be considering structural studies for the first time.

Obtaining protein crystals is only the beginning of the structure determination process. Once suitable crystals have been generated, they are exposed to an X-ray beam to produce diffraction images.

When X-rays interact with the highly ordered lattice of protein molecules within the crystal, they generate a distinctive diffraction pattern that contains information about the arrangement of atoms within the protein.

The quality of these diffraction images is heavily dependent on crystal quality. Consequently, the significant time and resources invested in identifying optimal crystallization conditions are justified by the potential to generate higher-resolution structural information.

However, diffraction images alone do not provide an immediately interpretable protein structure. The raw experimental data must first undergo a series of computational processing steps before it can be used for structure determination and model building.

Given the effort required to obtain suitable crystals, it is important to extract as much information as possible from every diffraction experiment.

Although most synchrotron facilities now perform initial data processing automatically, manual intervention is often still required. This is particularly true for newly obtained crystal forms where there may be uncertainty regarding the correct space group or when more complex crystallographic phenomena, such as crystal twinning, are present.

For this reason, even non-experts can benefit from developing a basic understanding of how X-ray diffraction data are processed.

Greyscale X-ray diffraction image featuring numerous small dark spots arranged in concentric circular rings around a bright central beam. A vertical beam stop extends from the top of the image to the centre, with faint grid lines visible across the detector.
Figure 2. Raw X-ray diffraction image collected from a protein crystal. The diffraction spots contain information about the arrangement of atoms within the protein.
Software interface for crystallographic data analysis showing a diffraction pattern with blue reflection spots distributed across a detector image. The left panel contains a processing workflow and spot-finding settings, while the main panel displays the indexed diffraction data.
Figure 3. Example diffraction data being analysed using the DIALS software package. Data processing identifies and measures diffraction spots before generating the datasets used for structure determination.

Today’s crystallographers have access to several powerful software packages capable of processing diffraction data efficiently and accurately.

Among the most widely used are:

  • DIALS
  • XDS
  • MOSFLM

These programs perform a range of computational tasks, including:

  • Identifying diffraction spots
  • Determining crystal lattice parameters
  • Indexing reflections
  • Integrating reflection intensities
  • Scaling and merging datasets
  • Assessing overall data quality

Each stage contributes to the creation of a refined dataset that accurately represents the experimental observations and can be used for subsequent structure solution and model refinement.

Protein crystallization and X-ray data processing are often viewed as separate disciplines, but they are fundamentally linked.

High-quality crystals are essential for producing high-quality diffraction data, while sophisticated data processing is required to maximize the value of that experimental data. Weaknesses at either stage can limit the quality of the final structural model.

Together, these techniques form the foundation of X-ray crystallography and enable researchers to answer key questions about protein function, ligand binding, and molecular mechanisms of action.

Protein crystallization and X-ray data processing are often viewed as separate technical disciplines, but in practice they are interconnected steps within a broader structural biology workflow. Success depends on expertise at every stage, from designing the right protein construct and identifying crystallization conditions to collecting high-quality diffraction data and generating robust structural models.

Sygnature Discovery’s Protein Sciences and Structural Biology teams provide integrated support across this entire process. Our scientists have extensive experience tackling challenging targets, developing crystallization strategies, processing complex datasets, and delivering high-resolution structures that support decision-making across drug discovery programmes.

By combining specialist structural biology expertise with Sygnature Discovery’s wider drug discovery capabilities, we help clients move beyond simply generating structures to understanding how those structures can be used to accelerate medicinal chemistry, reduce project risk, and advance promising programs towards the clinic.