Cryoelectron microscope density map refinement method based on mixed attention mechanism
By combining a Transformer network based on a hybrid attention mechanism and a pre-trained protein large language model with a cross-attention mechanism of density maps and structural features, the problems of long-range dependence and lack of structural constraints in cryo-electron microscopy density map refinement are solved, and high-precision density map refinement is achieved.
Patent Information
- Application Number
- CN202511383670.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing cryo-electron microscopy density map refinement methods suffer from insufficient long-range dependency modeling, limitations in low resolution, and lack of structural constraints. Furthermore, existing deep learning models struggle to capture the long-range spatial dependencies of density maps and neglect the geometric and physical constraints of atomic structures, resulting in poor density map refinement performance.
A Transformer network based on a hybrid attention mechanism is adopted, which is combined with a pre-trained protein large language model to extract atomic structure feature embeddings. The density map and structural features are fused through a cross-attention mechanism, and the network parameters are optimized using mean squared error and structural similarity loss functions to generate a high-precision density map.
It effectively improves the accuracy and robustness of density maps, can handle medium and low resolution density maps, captures long-range spatial dependencies and enhances structural constraints, and improves the interpretability and refinement of density maps.
Smart Images

Figure CN120876304A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of bioinformatics, structural biology detection, and computer applications, and in particular relates to a cryo-electron microscopy density map refinement method based on a hybrid attention mechanism. Background Technology
[0002] In recent years, with the development of cryo-electron microscopy (cryo-EM) hardware and image processing technologies, cryo-EM has become a core technology for resolving the structures of large biomolecular complexes. The basic principle of cryo-EM is to acquire two-dimensional electron microscopic images of samples after rapid freezing, extract particle images from these images, and reconstruct a three-dimensional electron density map by aligning particle images from different orientations, thereby constructing a high-precision atomic structure model. Compared to traditional structural biology methods, such as X-ray crystallography and nuclear magnetic resonance, cryo-EM has unique advantages in resolving the structures of large complexes. However, due to factors such as molecular flexible motion, conformational heterogeneity, and noise and errors during the imaging process, the original reconstructed density map often suffers from imaging artifacts, low contrast, and high background noise, thus affecting the accurate resolution and construction of the structural model.
[0003] In traditional methods, density map sharpening techniques are mainly divided into two categories: global sharpening and local sharpening. Global sharpening typically enhances the overall image contrast by applying a uniform B-factor. While simple and efficient, it struggles to adjust for resolution differences in local regions, potentially leading to over- or under-sharpening in some areas. To overcome this limitation, local sharpening methods have emerged. These methods, based on local resolution estimation, adaptive filtering, or reference structural information, apply differentiated processing to different regions, thereby more precisely improving density map quality and structural interpretability. Although local sharpening is highly effective in improving local details, it often relies on high-quality prior structural models or accurate mask information, which are often difficult to obtain in practical applications. Therefore, the key challenge in current density map refinement is how to achieve efficient and reliable refinement of complex cryo-electron microscopy density maps while minimizing reliance on prior information. Existing single deep learning models such as UNet and Transformer are limited by local convolution operations, making it difficult to capture the long-range spatial dependencies of density maps, resulting in incomplete detail recovery. While Transformer can capture global relationships, its computational complexity increases quadratically with the number of voxels, and it does not optimize the importance of channel features for 3D density map data. In addition, existing deep learning methods only refine medium-to-high resolution cryo-EM, and cannot refine medium-to-low resolution density maps or even lower resolution cryo-electron tomography (cryo-ET) density maps. Furthermore, existing methods only use density map voxel values as supervision signals, ignoring the implicit geometric and physical constraints in the corresponding atomic structures (such as bond lengths, bond angles, dihedral angles, side chain orientations, secondary structures, etc.), making it difficult for the network to learn the inherent physicochemical properties of the 3D structure of biological macromolecules. Summary of the Invention
[0004] To overcome the shortcomings and insufficient accuracy of existing cryo-electron microscopy density map refinement methods, this invention proposes a cryo-electron microscopy density map refinement method based on a hybrid attention mechanism with higher accuracy and stronger robustness, thereby solving the technical problems of insufficient long-range dependency modeling, low-resolution limitations, and lack of structural constraints in existing methods.
[0005] The technical solution adopted by this invention to solve its technical problem is: A cryo-electron microscopy density map refinement method based on a hybrid attention mechanism, the method comprising the following steps: 1) Construct a training dataset containing experimental density maps and their corresponding atomic structures, use a pre-trained protein large language model to extract feature embedding representations of atomic structures, and generate simulated density maps based on atomic structures; 2) Perform feature processing and adjustment on the feature embedding representations of the atomic structures of different protein complexes, and output fixed-dimensional feature embedding representations to adapt to the input requirements of downstream networks; 3) Preprocess and slice each experimental density map and its corresponding simulated density map to obtain multiple experimental density map blocks and simulated density map blocks respectively; 4) Input the experimental density map patches and the corresponding fixed-dimensional feature embedding representations into the hybrid attention Transformer network for training; during training, linear projection is performed on the fixed-dimensional feature embedding representations to obtain structural features with the same feature dimensions as the density map. Cross-modal fusion of density map features and structural features is achieved through the cross-attention mechanism to generate the corresponding predicted density map patches. 5) Calculate the loss value between each group of predicted density patches and the corresponding simulated density patches using a composite loss function composed of mean squared error loss and structural similarity loss, and update the network parameters by backpropagating the loss value; 6) The cryo-electron microscopy density map to be processed is divided into multiple density patches, which are then input into the trained network for prediction. The output predicted density patches are then assembled to obtain the final refined density map.
[0006] Furthermore, in step 2), the structural feature embedding representation process includes the following sub-steps: 2.1) Extract the atomic coordinates of the protein backbone from the atomic structure, retain only the standard residues with complete backbone information, and remove the data of missing or non-standard residues; 2.2) The extracted skeleton coordinates are input into the pre-trained protein big language model to generate an initial embedding representation for each protein chain. For a protein structure, the protein big language model will generate a structural embedding representation. , ,in C Let be the number of chains. R The number of residues in a single strand. D For the embedded dimension.
[0007] Furthermore, in step 2), the structural feature embedding representation process also includes the following sub-steps: 2.3) For each chain residue embedding Perform average pooling to obtain the chain-level embedding representation. , ; 2.4) For the chain Chain-level embedding representation and Chain Chain-level embedding representation Perform dot product operations to construct a chain-level similarity matrix. , ; 2.5) Chain-level similarity matrix Normalize the matrix column-wise to generate a chain-level attention weight matrix. And the chain-level attention weight matrix The global weight of each chain is obtained by averaging the values of each row. ; 2.6) Based on the global weight of each chain For residue-level embedding By performing a weighted summation, a multi-chain-level embedding that incorporates information from all chains is obtained. , .
[0008] Furthermore, in step 2), the structural feature embedding representation process also includes the following sub-steps: 2.7) For residues residue embedding representation and residues residue embedding representation Perform dot product operations to construct a residue-level similarity matrix. , ; 2.8) For residue-level similarity matrix Normalization is performed column-wise to generate a residue-level attention weight matrix. And the residue-level attention weight matrix The global weight of each residue is obtained by averaging the values of each row. ; 2.9) Embedding multi-chain levels After uniformly adjusting to a fixed length L, the output will be adjusted. If the actual number of residues is greater than L, the output will be determined based on the global weights of the residues. The top L residues are selected; otherwise, they are sorted by weight and padded cyclically up to L to ensure consistent output dimensionality. The final output is a fixed-dimensional feature embedding representation. , As input to the network module.
[0009] The technical concept of this invention is as follows: First, a training dataset containing multiple experimental density maps and their corresponding atomic structures is constructed, and corresponding simulated density maps are generated based on the atomic structures. Then, each experimental density map and its corresponding simulated density map are segmented to obtain multiple experimental density map patches and simulated density map patches. Unlike traditional methods that rely solely on density map voxel values, this invention extracts feature embedding representations of atomic structures as cross-modal supervision signals using a pre-trained protein large language model. The feature embedding representations of atomic structures of different protein complexes are processed to output fixed-length feature embedding representations. Next, these density map patches and their corresponding fixed-length features are... The length of the feature embedding is input into the Transformer deep learning network based on a hybrid attention mechanism to generate corresponding predicted density patches. Subsequently, the loss difference between each group of predicted density patches and the corresponding simulated density patches is calculated using the mean squared error loss function and the structural similarity loss function. The network parameters are updated by backpropagating the difference, thereby completing the training of the deep learning model. Finally, in the inference stage, there is no need to input the feature embedding representation again. It is only necessary to divide the cryo-electron microscopy density map to be processed into multiple density patches, input them into the trained network for prediction, and assemble the output predicted density patches to obtain the final refined density map.
[0010] The beneficial effects of this invention are as follows: First, this invention adopts a Transformer network architecture based on a hybrid attention mechanism and uses simulated density maps generated from atomic structures as training targets, effectively avoiding noise interference introduced by using experimental density maps as targets. Second, it extracts feature embedding representations of atomic structures (such as bond lengths, bond angles, dihedral angles, side chain directions, secondary structures, etc.) through a pre-trained protein large language model as auxiliary modal enhancements to structural feature representations. In addition, the network structure integrates channel attention and self-attention mechanisms, utilizing the global information perception capability of channel attention to compensate for the receptive field limitations of local self-attention. At the same time, it dynamically weights key feature channels through a compression excitation mechanism to suppress noise-related channels and introduces a grid attention module to enhance the network's ability to learn cross-regional features, effectively overcoming the limitation of traditional sliding windows in capturing long-distance dependencies. By combining fixed window and sliding window attention, it balances computational efficiency and global perception capability. Finally, by jointly minimizing the mean squared error loss and maximizing the structural similarity loss, the network can fully capture the structural correlation between experimental and simulated density maps. Attached Figure Description
[0011] Figure 1 This is a basic flowchart of a cryo-electron microscopy density map refinement method based on a hybrid attention mechanism.
[0012] Figure 2 This is the cryo-electron microscopy structure of the protein complex 8iaj.
[0013] Figure 3 This is a cryo-electron microscopy simulation density map of the protein complex 8iaj.
[0014] Figure 4 This is a cryo-electron microscopy density map of the protein complex 8iaj.
[0015] Figure 5 This is the density map of the protein complex 8iaj after processing by a cryo-electron microscopy density map refinement method based on a hybrid attention mechanism.
[0016] Figure 6 This is a schematic diagram of a cryo-electron microscopy density map refinement method based on a hybrid attention mechanism. Detailed Implementation
[0017] The present invention will now be further described with reference to the accompanying drawings.
[0018] Reference Figures 1-5 A cryo-electron microscopy density map refinement method based on a hybrid attention mechanism includes the following steps: 1) Construct a non-redundant cryo-electron microscopy 3D density map dataset from the EMDB and PDB databases, and extract feature embeddings of atomic structures using a pre-trained protein large language model, including the following sub-steps: 1.1) Download all single-particle cryo-electron microscopy density maps with resolutions ranging from 3 Å to 8 Å from the EMDB and PDB databases (e.g., Figure 2 ) and the corresponding atomic structures as the initial dataset (e.g. Figure 4 The initial dataset was filtered based on density map resolution, orthogonal axes, and the completeness of the corresponding atomic structures. The filtered data was then clustered based on sequence similarity, and all cluster centers were selected to construct the final dataset. The process is as follows: 1.1.1) Filter and remove data that does not meet the requirements for resolution, orthogonality axes, and the completeness of the corresponding atomic structure; 1.1.2) Filter and remove initial data in the initial dataset where CC_mask is less than the first threshold (0.75) and CC_box is less than the second threshold (0.6) in the experimental density map. CC_mask and CC_box are structural correlation coefficients, which are calculated by inputting the cryo-electron microscopy density map and atomic structure using the phenix.map_model_cc tool. 1.1.3) The clustering method in MMseqs2 is used to cluster the remaining initial dataset after filtering in 1.1.2) to remove redundant data. Specifically, if the sequence similarity between any chain in the atomic structure of the initial data and any chain in the atomic structure of another initial data is greater than a set percentage (30%), then the two initial data are grouped into one class. The data of all class centers are selected as the training dataset, and finally 1128 pairs of experimental density maps with resolution in the range of 3 Å-8 Å and the corresponding atomic structures are obtained. 1.2) A set number (1012) of data points were randomly selected from the 1128 data points in the initial dataset as the training dataset for the deep learning model used in this embodiment; 1.3) Based on the atomic structures from the training data set specified in 1.2) (take 10¹²), a noise-free simulated density map is obtained, such as... Figure 3 As shown, the grid size of the experimental density maps and simulated density maps with different grid sizes is standardized to 1 Å, negative density values are turned into 0, and the density values of the experimental density maps are normalized using the 99.999 percentile density value normalization strategy. 1.4) Input the atomic structures from the number (1012) training data points set in 1.2) into the pre-trained Protein Large Language Model (ESM), and extract the feature embedding representation of the atomic structures through the pre-trained Protein Large Language Model; 2) Perform feature processing and adjustment on the feature embedding representations of different protein complex atomic structures, and output fixed-dimensional feature embedding representations to adapt to the input requirements of downstream networks. The structural feature embedding representation processing includes the following sub-steps: 2.1) Extracting the atomic coordinates (including N, C) of the protein backbone from its atomic structure. (C atoms), retaining only standard residues with complete backbone information and discarding data with missing or non-standard residues; 2.2) The extracted skeleton coordinates are input into the pre-trained protein big language model to generate an initial embedding representation for each protein chain. For a protein structure, the protein big language model generates a structural embedding representation. , ,in C Let be the number of chains. R The number of residues in a single strand. D For the embedded dimension; 2.3) For each chain residue embedding Perform average pooling to obtain the chain-level embedding representation. : ; 2.4) For the chain Chain-level embedding representation and Chain Chain-level embedding representation Perform a dot product operation to construct a chain-level similarity matrix: ; in, This represents the similarity between chain i and chain j; 2.5) Chain-level similarity matrix Normalize the matrix column-wise to generate a chain-level attention weight matrix. And the chain-level attention weight matrix The global weight of each chain is obtained by averaging the values of each row. ; 2.6) Based on the global weight of each chain For residue-level embedding By performing a weighted summation, a multi-chain-level embedding that incorporates information from all chains is obtained. , ; 2.7) For residues residue embedding representation and residues residue embedding representation Perform dot product operations to construct a residue-level similarity matrix. : ; in, Indicates the similarity between residue i and residue j; 2.8) Similarly, for the residue-level similarity matrix Normalization is performed column-wise to generate a residue-level attention weight matrix. And the residue-level attention weight matrix The global weight of each residue is obtained by averaging the values of each row. ; 2.9) Embedding multi-chain levels After uniformly adjusting to a fixed length L, the output will be adjusted. If the actual number of residues is greater than L, the output will be determined based on the global weights of the residues. The top L residues are selected; otherwise, they are sorted by weight and padded cyclically up to L to ensure consistent output dimensionality. The final output is a fixed-dimensional feature embedding representation. , As input to the network module; 3) Using a sliding window with a step size of 24 grid cells, the density map and the corresponding simulated density map are cut into density map patches of a preset grid size (64×64×64). To reduce the risk of overfitting in the network model, the following steps are used for data augmentation during data loading: including the following sub-steps: 3.1) Randomly crop the experimental density map and the corresponding simulated density map of the preset grid size (64×64×64) to three-quarters of the preset grid size (48×48×48) density map, and rotate them randomly by multiples of 90 degrees. 3.2) Add random Gaussian noise to the experimental density plot: ; in, This represents a density patch after adding random Gaussian noise. Represents the original density map patch. This indicates that the mean is 0 and the standard deviation is 0. ( Random Gaussian noise with a probability of 0.1 (=0.1) is added, with an addition probability of 20%. 3.3) Perform random Gaussian kernel blur convolution on the experimental density map patches: ; in, , These represent the spatial offsets of the point to be blurred relative to the center of the Gaussian kernel in the horizontal, vertical, and depth directions, respectively. Indicates in three-dimensional space ( , () represents the Gaussian kernel value of the coordinates. The standard deviation of the Gaussian kernel is used to control the degree of blurring. Its value range is set to a random value of (0, 0.5). The Gaussian kernel is used to perform a three-dimensional spatial convolution operation on the density map to achieve a random blurring effect. The addition probability is set to 20%. 4) Input the density map patches processed in 3) and the fixed-dimensional feature embedding representations processed in 2) into a Transformer network based on a hybrid attention mechanism for training; during training, linear projection is performed on the fixed-dimensional feature embedding representations to obtain structural features consistent with the feature dimensions of the density map. Cross-modal fusion of density map features and structural features is achieved through a cross-attention mechanism to generate corresponding predicted density map patches; the Transformer network architecture includes 4 layers of channel attention modules and a hybrid attention module. Each channel attention module contains 2 convolutional blocks and 1 compression activation module. The hybrid attention module includes a fixed window multi-head self-attention mechanism module, a sliding window attention mechanism module, a grid attention mechanism module, a cross-attention module, and 1 feature fusion module; The cross-attention module achieves multimodal feature fusion through the following steps: linearly projecting the fixed-dimensional feature embedding representation output and processed by the pre-trained protein large language model to obtain structural features consistent with the feature dimensions of the density map, and then querying the matrix... (From density map features) and bond matrix The similarity is calculated by multiplying the dot product of the final structural features output from the pre-trained protein large language model to characterize the strength of feature association between modalities, and then divided by... To prevent gradient vanishing, by The function generates attention weights, and a weighted summation matrix. (From the final structural features), the output is the fused cross-modal representation. The cross-attention implementation is as follows: ; in The density map feature matrix, , The structural feature projection matrix, This is the scaling factor; 5) Using the mean squared error loss function (MSE) and the structural similarity loss function (SSIM), calculate the difference in density value, contrast similarity, and structural similarity between each pair of predicted density patches and their corresponding simulated density patches; weighted sum the two loss function values to obtain the composite loss function, and update the network parameters through backpropagation of the loss function value, thereby completing the training of the deep learning model; the mean squared error loss function is shown below: ; in, Indicates the first The true value of each sample Indicates the first The predicted value for each sample, Indicates the total number of samples; The structural similarity loss function is calculated as follows: ; in , These represent the predicted sample (predicted density patch) and the target sample (simulated density patch), respectively. This represents the structural similarity function, with values ranging from 0 to 1. The closer the value is to 1, the more similar the images are. and They represent the predicted samples respectively. With target sample The variance reflects the contrast similarity of the samples; The covariance of two samples is used to measure their structural similarity. It is a very small positive number 1e-5, used to prevent the denominator from being zero and to ensure the stability of numerical calculations; The composite loss function is calculated as follows: ; This is a weighting coefficient, ranging from 0 to 1. In this invention, it is set to 0.5 to balance the impact of the mean squared error loss function and the structural similarity loss function on the training process. The larger the value, the more sensitive it is to errors between density map values; conversely, the smaller the value, the more sensitive it is to errors in structural similarity. During training, a set percentage (20%) of the density patches obtained in step 3) is randomly selected as the validation set, and the remaining 80% is used for training. The training rounds are set to 250 rounds. Each batch contains a set batch size (16) of density patches. The cosine annealing strategy in the Adam optimizer is used to automatically decay the learning rate. The initial learning rate is set to 5e-5, and the minimum learning rate is set to 1e-5. Training is stopped after the preset number of rounds, resulting in multiple deep learning network models. The target network model with the smallest loss function value on the validation set is taken as the final model after training is completed. 6) During the inference process, the cryo-electron microscopy density map to be processed is divided into multiple density patches of a preset three-quarters grid size (48×48×48) with a step size of 12 as the second preset step size. These patches are then input into the trained model for prediction. Since the step size is smaller than the window size, the resulting density patches overlap to a certain extent in space. These overlapping density patches are input into the trained network for prediction. Finally, the prediction results for the overlapping areas are averaged, and the output density patches are fused together to reassemble them into a complete density map, such as... Figure 5 As shown.
[0019] The experimental structure and corresponding density map of protein complex 8iaj are used as examples, referring to... Figure 6 A cryo-electron microscopy density map refinement method based on a hybrid attention mechanism is described below, and the deep learning framework used is summarized as follows: a. Construction of the training dataset and segmentation of density blocks: Specifically, such as Figure 6 As shown in Figure a, cryo-electron microscopy density maps and their corresponding PDB atomic structures were downloaded from the EMDB and PDB databases, respectively. Simulated density maps were generated based on these atomic structures. Simultaneously, a pre-trained protein large language model was used to extract feature embeddings of the atomic structures. These feature embeddings were processed for different protein complex atomic structures, and fixed-dimensional feature embeddings were output to adapt to the input requirements of the downstream network. Subsequently, after normalizing the experimental and simulated density maps, a sliding window with a stride of 24 was used to cut the density maps and corresponding simulated density maps into 64×64×64 grid-sized density patches. During data loading, the 64×64×64 density patches were randomly cropped to 48×48×48 density patches and rotated randomly by multiples of 90 degrees. Random Gaussian noise was added to the experimental density patches, and random Gaussian kernel blur convolution was performed.
[0020] b. Training process of deep learning network: In each round of training, 16 experimental density blocks in a single batch are input into the deep learning network based on hybrid attention Transformer. During training, structural features with the same feature dimension as the density map are fused with density map features through cross attention to achieve cross-modal fusion. The density blocks output by the model are compared with the corresponding simulated density blocks. By combining the mean squared error loss function and the structural similarity loss function, the parameters of the deep learning network model are updated using backpropagation. Specifically, this invention uses a Transformer-based network with hybrid attention to refine cryo-electron microscopy density maps. Figure 6 b describes a schematic diagram of the network structure, which consists of multiple convolutional blocks, multi-layer hybrid multi-attention modules, and a feature fusion module. Each layer of hybrid multi-attention modules includes a channel attention module and a hybrid attention Transformer module, and the feature fusion module is composed of multi-layer perceptrons. The channel attention module consists of a fusion convolutional module with a compression activation mechanism. First, the number of input channels is expanded to 6 times through a 1×1×1 convolution. Then, after processing by the compression activation mechanism, the number of channels is compressed again by a 1×1×1 convolution to restore the original dimension. The compression activation mechanism sequentially executes global average pooling, fully connected layer, ReLU activation, fully connected layer, and Sigmoid function. Finally, the channel weights are multiplied with the original input feature map channel by channel to obtain a weighted feature map. The channel weights are learned through the compression activation mechanism to enhance key feature channels and suppress invalid noise channels.
[0021] The hybrid attention Transformer module consists of a Swin-Transformer layer, multiple attention mechanisms, and a feature fusion module. It integrates the multi-head self-attention mechanism and sliding window mechanism from the Swin Transformer to jointly enable information interaction within and between windows. Before partitioning features along the channel dimension, a SwinTransformer layer is first introduced to extract local and non-local spatial features in parallel, serving as the initial representation for the downstream attention mechanism. To further expand the receptive field and enhance long-range dependency modeling capabilities, we introduce a grid attention mechanism module based on sparse attention while retaining the original window attention structure. This strategy divides the feature map into multiple non-overlapping grid regions and randomly perturbs them, enabling sparse attention computation to be performed within each grid, thereby achieving non-local feature interaction across spatial locations while maintaining computational efficiency. Compared with traditional local window mechanisms, this sparse attention design is more effective at capturing distant but structurally related repeating units and globally similar structural features in the density map. Linear projection is performed on the fixed-dimensional feature embeddings obtained from the pre-trained protein large language model to obtain structural features consistent with the feature dimensions of the density map. Multimodal feature fusion is achieved using a cross-attention module and the query matrix is then used. (From density map features) and bond matrix The similarity is calculated by multiplying the dot product of the final structural features output from the pre-trained protein large language model to characterize the strength of feature association between modalities, and then divided by... To prevent gradient vanishing, by The function generates attention weights, and a weighted summation matrix. (From the final structural features), output the fused cross-modal representation; The single-batch input of the entire Transformer deep learning network architecture based on hybrid attention is a density patch of 16 grids with a grid spacing of 1 Å and a size of 48×48×48. The output of the network is a predicted density patch of the same size. The mean squared error loss and structural similarity loss of the predicted density patch and the simulated density patch are calculated and backpropagated to update the parameters of the deep learning model.
[0022] c. The inference process of deep learning networks: Cryo-electron microscopy density maps of a given protein complex 8iaj, such as... Figure 2 As shown, it is first cut into multiple density blocks of 48×48×48 grid size. All density blocks are input into the trained deep learning model for refinement. Then, the output density blocks are reassembled to generate the refined density map of the protein complex, as shown in Figure 8iaj. Figure 5 As shown.
[0023] Taking the cryo-electron microscopy density map of protein complex 8iaj as an example, its FSC-0.5 cutoff resolution is 3.61 Å. The density map obtained after refinement using this method is as follows: Figure 5 As shown, its FSC-0.5 cutoff resolution has been improved to 2.27 Å.
[0024] The above description describes preferred embodiments of the present invention and is not intended to limit the present invention. Various changes can be made to it without departing from the basic spirit of the present invention and without exceeding the scope of the present invention.
Claims
1. A cryo-electron microscopy density map refinement method based on a hybrid attention mechanism, characterized in that, The method includes the following steps: 1) Construct a training dataset containing experimental density maps and their corresponding atomic structures, use a pre-trained protein large language model to extract feature embedding representations of atomic structures, and generate simulated density maps based on atomic structures; 2) Perform feature processing and adjustment on the feature embedding representations of the atomic structures of different protein complexes, and output fixed-dimensional feature embedding representations to adapt to the input requirements of downstream networks; 3) Preprocess and slice each experimental density map and its corresponding simulated density map to obtain multiple experimental density map blocks and simulated density map blocks respectively; 4) Input the experimental density map patches and the corresponding fixed-dimensional feature embedding representations into the hybrid attention Transformer network for training; during training, linear projection is performed on the fixed-dimensional feature embedding representations to obtain structural features with the same feature dimensions as the density map. Cross-modal fusion of density map features and structural features is achieved through the cross-attention mechanism to generate the corresponding predicted density map patches. 5) Calculate the loss value between each group of predicted density patches and the corresponding simulated density patches using a composite loss function composed of mean squared error loss and structural similarity loss, and update the network parameters by backpropagating the loss value; 6) The cryo-electron microscopy density map to be processed is divided into multiple density patches, which are then input into the trained network for prediction. The output predicted density patches are then assembled to obtain the final refined density map.
2. The cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in claim 1, characterized in that, In step 2), the structural feature embedding representation processing includes the following sub-steps: 2.1) Extract the atomic coordinates of the protein backbone from the atomic structure, retain only the standard residues with complete backbone information, and remove the data of missing or non-standard residues; 2.2) The extracted skeleton coordinates are input into the pre-trained protein big language model to generate an initial embedding representation for each protein chain. For a protein structure, the protein big language model will generate a structural embedding representation. , ,in C Let be the number of chains. R The number of residues in a single strand. D For the embedded dimension.
3. The cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in claim 2, characterized in that, In step 2), the structural feature embedding representation process further includes the following sub-steps: 2.3) For each chain residue embedding Perform average pooling to obtain the chain-level embedding representation. , ; 2.4) For the chain Chain-level embedding representation and Chain Chain-level embedding representation Perform dot product operations to construct a chain-level similarity matrix. , ; 2.5) Chain-level similarity matrix Normalize the matrix column-wise to generate a chain-level attention weight matrix. And the chain-level attention weight matrix The global weight of each chain is obtained by averaging the values of each row. ; 2.6) Based on the global weight of each chain For residue-level embedding By performing a weighted summation, a multi-chain-level embedding that incorporates information from all chains is obtained. , .
4. The cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in claim 3, characterized in that, In step 2), the structural feature embedding representation process further includes the following sub-steps: 2.7) For residues residue embedding representation and residues residue embedding representation Perform dot product operations to construct a residue-level similarity matrix. , ; 2.8) For residue-level similarity matrix Normalization is performed column-wise to generate a residue-level attention weight matrix. And the residue-level attention weight matrix The global weight of each residue is obtained by averaging the values of each row. ; 2.9) Embedding multi-chain levels After uniformly adjusting to a fixed length L, the output will be adjusted. If the actual number of residues is greater than L, the output will be determined based on the global weights of the residues. The top L residues are selected; otherwise, they are sorted by weight and padded cyclically up to L to ensure consistent output dimensionality. The final output is a fixed-dimensional feature embedding representation. , As input to the network module.
5. A cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in any one of claims 1 to 4, characterized in that, In step 1), the construction process of the training dataset is as follows: filtering and removing data whose resolution, orthogonal axis, and the completeness of the corresponding atomic structure do not meet the requirements; filtering and removing initial data in the experimental density map of the initial dataset where CC_mask is less than the set first threshold and CC_box value is lower than the set second threshold; using the clustering method in MMseqs2 to cluster the remaining initial dataset after filtering to remove redundant data, and finally obtaining non-redundant experimental density maps and corresponding atomic structures with a resolution in the range of 3Å-8Å.
6. A cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in any one of claims 1 to 4, characterized in that, In section 4), the Transformer network architecture includes the following modules: a multi-layer channel attention module and a hybrid attention module. Each channel attention module includes a convolutional block and a compression activation module. The hybrid attention module includes a fixed window multi-head self-attention mechanism module, a sliding window attention mechanism module, a grid attention mechanism module, a feature fusion module, and a cross attention module.
7. The cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in claim 6, characterized in that, In the channel attention module, the number of input channels is expanded by 1×1×1 convolution, then processed by a compression activation mechanism, and then the number of channels is compressed by 1×1×1 convolution to restore the original dimension; finally, the channel weights are multiplied with the original input feature map channel by channel to obtain a weighted feature map. The channel weights are learned through the compression activation mechanism to enhance key feature channels and suppress invalid noise channels.
8. The cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in claim 6, characterized in that, Before feature partitioning by channel dimension in the hybrid attention module, a Swin-Transformer layer is first introduced to extract local and non-local spatial features in parallel, serving as the initial representation for the downstream attention mechanism.
9. The cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in claim 6, characterized in that, In the grid attention mechanism module, the feature map is divided into multiple non-overlapping grid regions and randomly perturbed, so that sparse attention computation can be performed within each grid.
10. The cryo-electron microscopy density map refinement method based on a hybrid attention mechanism as described in claim 6, characterized in that, In the cross-attention module, the fixed-dimensional feature embedding representation output and processed by the pre-trained protein large language model is... Linear projection is performed to obtain the final structural features with the same dimension as the density map features. The density map features are used as the query matrix, with keys and values derived from the final structural features. Cross-attention is calculated to achieve cross-modal fusion of density map features and structural features.
Citation Information
Patent Citations
Freezing electron microscope three-dimensional density map post-processing method and device based on deep learning
CN114841898A
Infrared-visible light image fusion method based on mask image modeling pre-training
CN120410878A
Hybrid embedded attention time-stop target tracking method based on Transform
CN120598998A
Cited By
A protein cryo-em structure quality evaluation method based on multi-feature fusion
CN122392613A