A method and apparatus for optimizing molecular structure based on three-dimensional atomic density maps
By generating Cα atom and backbone atom density maps using 3D image recognition models and combining them with molecular dynamics simulation techniques for local and global optimization, the problem of residue mismatch caused by insufficient density map annotation was solved, thus improving the accuracy and efficiency of three-dimensional structure optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2026-04-03
AI Technical Summary
In the process of optimizing molecular three-dimensional structures based on cryo-electron microscopy, the lack of labeling in density maps in existing technologies makes it difficult to distinguish adjacent regions and easily leads to residue mismatch problems.
A 3D image recognition model is used to perform semantic recognition on the three-dimensional atomic density map, generating a Cα atomic density map and a backbone atomic density map. Local and global structure optimization is performed using molecular dynamics simulation technology. The initial structure is labeled and optimized using the Cα atomic density map and the backbone atomic density map.
It improves the accuracy and efficiency of three-dimensional structure optimization, avoids residue mismatch problems, and achieves more precise molecular structure optimization.
Smart Images

Figure CN115691658B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a processing method and apparatus for optimizing molecular structures based on three-dimensional atomic density maps. Background Technology
[0002] Single-particle cryo-electron microscopy (cryo-electron microscopy) can analyze and optimize the corresponding three-dimensional molecular structures based on the three-dimensional density maps obtained by cryo-electron microscopy. Molecular dynamics simulations are commonly used to optimize the three-dimensional structure. However, since the density maps themselves are not labeled, it may be difficult to distinguish the density maps of adjacent regions, which can easily lead to residue mismatch problems during optimization. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by providing a method, apparatus, electronic device, and computer-readable storage medium for optimizing molecular structures based on three-dimensional atomic density maps. The invention utilizes a 3D image recognition model to perform semantic recognition on the molecular three-dimensional density map, resulting in two feature maps: a Cα atomic density map identifying all Cα atom features and a backbone atomic density map identifying the features of the protein's core atoms (C, Cα, N). Multiple residue annotation fragments are extracted from the Cα atomic density map using the Cα atomic density map and a known protein sequence. Using these residue annotation fragments as targets, a priori three-dimensional initial structure is locally optimized using molecular dynamics simulation techniques. Furthermore, the core atomic density map is used as the target for global structural optimization of the same three-dimensional initial structure using molecular simulation techniques. This invention, by using the priori three-dimensional initial structure as the optimization object and the Cα atomic density map and backbone atomic density map derived from the three-dimensional density map as targets for global and local optimization, avoids residue mismatch problems caused by the lack of annotations in the density map itself, improves the accuracy of three-dimensional structure optimization, and enhances the efficiency of three-dimensional structure optimization.
[0004] To achieve the above objectives, a first aspect of the present invention provides a method for optimizing molecular structures based on three-dimensional atomic density maps, the method comprising:
[0005] Obtain the first 3D atomic density map and the corresponding first protein sequence and first 3D initial structure;
[0006] Based on a preset 3D image recognition model, the first 3D atomic density map is processed to generate the corresponding first Cα atomic density map and first backbone atomic density map.
[0007] Multiple first labeled fragments are generated by identifying residue labeled fragments based on the first Cα atom density map and the first protein sequence;
[0008] Based on all the first annotated fragments and the first backbone atomic density map, the first 3D initial structure is subjected to three-dimensional molecular structure optimization processing to generate the corresponding first optimized structure.
[0009] Preferably, the first 3D atomic density map is an electron microscope three-dimensional atomic density map; the shape of the first 3D atomic density map is H0×W0×D0, where H0, W0 and D0 are the height, width and number of channels of the first 3D atomic density map, respectively.
[0010] The first protein sequence is a residue type sequence of a protein molecule corresponding to the first 3D atomic density map;
[0011] The first 3D initial structure is a standard three-dimensional protein molecular structure corresponding to the first protein sequence;
[0012] The first Cα atom density map includes a plurality of first Cα atoms; the first Cα atom includes first Cα atom coordinates, one or more first peptide bond orientations, and a plurality of first residue type probabilities; the number of first residue type probabilities is the same for each first Cα atom;
[0013] The first backbone atom density map includes multiple first backbone atom density regions, and the backbone atoms include C atoms, Cα atoms and N atoms;
[0014] The first labeled segment includes first segment feature data; the first segment feature data includes first segment number, first segment type sequence and first segment starting coordinates.
[0015] Preferably, the 3D image recognition model is implemented based on the 3D Unet model.
[0016] Preferably, the step of performing target recognition processing on the first 3D atomic density map based on a preset 3D image recognition model to generate the corresponding first Cα atomic density map and first backbone atomic density map specifically includes:
[0017] Based on the 3D image recognition model, residue feature recognition is performed on the first 3D atomic density map to obtain a corresponding first feature map; Cα atom feature recognition is performed on the first 3D atomic density map to obtain a corresponding second feature map; N atom feature recognition is performed on the first 3D atomic density map to obtain a corresponding third feature map; C atom feature recognition is performed on the first 3D atomic density map to obtain a corresponding fourth feature map; Cα atom feature fusion is performed on the first 3D atomic density map, the first feature map, and the second feature map to generate a corresponding first Cα atomic density map; and backbone atomic density regions are fused on the first 3D atomic density map, the second feature map, the third feature map, and the fourth feature map to generate a corresponding first backbone atomic density map.
[0018] in,
[0019] The first feature map includes multiple first residue targets, each first residue target including one or more first peptide bond orientations and multiple first residue type probabilities; the second feature map includes multiple first Cα atom targets, each first Cα atom target including first Cα atom coordinates; the third feature map includes multiple first N atom targets, each first N atom target including first N atom coordinates; the fourth feature map includes multiple first C atom targets, each first C atom target including first C atom coordinates; the shapes of the first, second, third, and fourth feature maps are H1×W1×D1, H2×W2×D2, H3×W3×D3, and H4×W4×D4, respectively, where H1, H2, H3, and H4 are the heights of the corresponding feature maps, H1 = H2 = H3 = H4 = H0, W1, W2, W3, and W4 are the widths of the corresponding feature maps, W1 = W2 = W3 = W4 = W0, and D1, D2, D3, and D4 are the number of channels in the corresponding feature maps;
[0020] The first Cα atom density map includes a plurality of first Cα atoms; each first Cα atom corresponds one-to-one with the first Cα atom target in the second feature map; each first Cα atom corresponds to a set of first Cα atom feature data, the first Cα atom feature data including the first Cα atom coordinates of the first Cα atom target corresponding to the first Cα atom target on the second feature map, the first density map feature matching the corresponding first Cα atom coordinates on the first 3D atom density map, one or more first peptide bond orientations and multiple first residue type probabilities of the first residue target matching the corresponding first Cα atom coordinates on the first feature map;
[0021] The first backbone atom density map includes multiple first backbone atom density regions; each first backbone atom density region includes one or more second Cα atoms, or one or more first N atoms, or one or more first C atoms; each second Cα atom corresponds one-to-one with the first Cα atom target in the second feature map, each first N atom corresponds one-to-one with the first N atom target in the third feature map, and each first C atom corresponds one-to-one with the first C atom target in the fourth feature map; each second Cα atom, first N atom, and first C atom corresponds to a set of first backbone atom feature data; the first... The backbone atom feature data includes a first backbone atom type, first backbone atom coordinates, and a first backbone atom density map feature; the first backbone atom type includes Cα atom type, N atom type, and C atom type; the first backbone atom coordinates are the first Cα atom coordinates of the first Cα atom target corresponding to the second feature map, or the first N atom coordinates of the first N atom target corresponding to the third feature map, or the first C atom coordinates of the first C atom target corresponding to the fourth feature map; the first backbone atom density map feature is the density map feature on the first 3D atom density map that matches the first backbone atom coordinates.
[0022] Preferably, the step of generating multiple first labeled fragments by identifying residue-labeled fragments based on the first Cα atom density map and the first protein sequence specifically includes:
[0023] Based on one or more of the first peptide bond directions of each of the first Cα atoms in the first Cα atom density map, adjacent Cα atoms are linked to obtain multiple first Cα atom chains; and each first Cα atom chain is regarded as a corresponding first residue fragment; the first Cα atom chain is composed of multiple first Cα atoms linked together, and each pair of linked first Cα atoms in the first Cα atom chain has a first peptide bond direction that overlaps with each other, and the linking relationship between the two is formed by the pair of overlapping first peptide bond directions and the first Cα atom coordinates of the two first Cα atoms respectively;
[0024] The number of first residue type probabilities of the first Cα atom on the first Cα atom density map is statistically analyzed to generate the corresponding total number of first residue types M; the residue type sequence lengths of the first protein sequence are statistically analyzed to generate the corresponding first sequence length L; and the number of first Cα atoms in each first residue fragment is statistically analyzed to generate the corresponding first fragment length L. x ;
[0025] Based on the total number M of the first residue types, the length L of the first sequence, and the length L of each of the first residue fragments.x The residue score and residue start position of each first residue fragment in the first protein sequence are predicted to generate the corresponding first residue fragment score and first residue fragment start position.
[0026] The first residue fragment whose score exceeds a preset score threshold is recorded as the corresponding first pre-selected residue fragment; and based on the first protein sequence and the first fragment length L corresponding to each first pre-selected residue fragment... x The fragment residue type of each of the first pre-selected residue fragments is corrected according to the starting position of the first residue fragment to generate a corresponding first fragment residue type sequence; and the average fragment probability of each of the first pre-selected residue fragments is statistically analyzed to generate a corresponding first fragment average probability; the first fragment residue type sequence includes multiple second residue types;
[0027] The length L of the first segment x The first pre-selected residue fragments that exceed a preset length threshold and whose average probability exceeds a preset probability threshold are designated as the corresponding first labeled fragments; and the obtained multiple first labeled fragments are sorted according to the order of their corresponding first residue fragment start positions, with the sorting index serving as the corresponding first fragment number; the first fragment residue type sequence corresponding to each first labeled fragment is designated as the corresponding first fragment type sequence; and the first Cα atom coordinate of the first first Cα atom of each first labeled fragment is designated as the corresponding first fragment start coordinate; and the first fragment number, the first fragment type sequence, and the first fragment start coordinate of each first labeled fragment constitute the corresponding first fragment feature data.
[0028] Furthermore, the step involves determining the total number of the first residue types M, the length of the first sequence L, and the length of the first fragment L of each of the first residue fragments. x Predicting the residue score and residue start position of each first residue fragment in the first protein sequence to generate the corresponding first residue fragment score and first residue fragment start position, specifically including:
[0029] Based on the total number M of the first residue types and the first sequence length L, the first protein sequence is subjected to one-hot matrix encoding to obtain a corresponding first matrix vector F of shape M×L. * The first matrix vector F * It includes M*L vector data elements, and the first matrix vector F * Each row corresponds to a residue type, and the first matrix vector F *The columns are sorted according to the residue type order of the first protein sequence; the first matrix vector F * In each column of M vector data elements, the vector data element whose code matches the residue type corresponding to the current column is set to 1, and the code of the remaining M-1 vector data elements is set to 0.
[0030] Based on the total number M of the first residue type and the length L of the first fragment of the first residue fragment. x The first residue fragment is subjected to a two-dimensional matrix transformation to obtain a corresponding shape of M×L. x The second matrix vector G * The second matrix vector G * Including M*L x There are vector data elements, and the second matrix vector G * Each row corresponds to a residue type, and the second matrix vector G * The columns are sorted according to the order of the first Cα atom of the first residue fragment; the second matrix vector G * The M vector data elements in each column represent the M probabilities of the first residue type of the first Cα atom corresponding to the current column;
[0031] Let the sliding window's sliding step size be 1 column at a time, and let the sliding window's width be L. x The column; and based on the set sliding step size and sliding window width, the first matrix vector F is... * Starting from the first column, a sliding window process is performed to obtain (LL) x +1) Shapes of M×L x The first sliding window matrix vector; and perform one-dimensional vector transformation on each of the first sliding window matrix vectors to generate a corresponding length (M*L) x The first one-dimensional vector of ); and obtained from (LL x +1) of the first one-dimensional vectors form a shape corresponding to (LL) x +1)×(M*L x The third matrix vector F;
[0032] For the second matrix vector G * Perform one-dimensional vector transformation to generate the corresponding vector of length (M*L) x The second one-dimensional vector is then subjected to a two-dimensional matrix upscaling process to obtain the corresponding shape (M*L). x The fourth matrix vector G is 1×1;
[0033] Performing a cross product operation on the third matrix vector F and the fourth matrix vector G generates a corresponding shape of (LL). xThe fifth matrix vector S, S = F × G, is obtained by performing a one-dimensional vector reduction process on the fifth matrix vector S to generate a corresponding vector of length (LL). x +1) the third one-dimensional vector; and from the (LL) of the third one-dimensional vector x The maximum value among the +1) vector data elements is selected as the corresponding score of the first residue fragment; and the sorting index of the first residue fragment score in the third one-dimensional vector is used as the starting position of the corresponding first residue fragment.
[0034] Furthermore, the first fragment length L corresponding to the first protein sequence and each of the first preselected residue fragments... x The process involves performing residue type correction on each of the first pre-selected residue fragments based on the start position of the first residue fragment to generate a corresponding first fragment residue type sequence, specifically including:
[0035] Using the start position of the first residue fragment as the extraction start position, and the length L of the first fragment as the extraction start position... x To extract the length, a subsequence is extracted from the first protein sequence as the corresponding first subsequence; the first subsequence includes L x Each residue type;
[0036] L of the first preselected residue fragment x Each of the first Cα atoms is traversed sequentially; during traversal, the sorting index of the currently traversed first Cα atom in the first pre-selected residue fragment is used as the corresponding first index; the residue type corresponding to the highest probability among the M first residue type probabilities corresponding to the currently traversed first Cα atom is used as the corresponding second residue type; the residue type whose sorting index matches the first index in the first sub-sequence is used as the corresponding first target type; and the matching between the second residue type and the first target type is identified; if they match, the traversal continues to the next first Cα atom until the last first Cα atom is traversed; if they do not match, the second residue type is reset using the first target type, and the traversal continues to the next first Cα atom until the last first Cα atom is traversed; and at the end of the traversal, L is obtained. x The second residue type is sorted to form the corresponding first fragment residue type sequence.
[0037] Furthermore, the step of statistically calculating the average probability of each of the first preselected residue fragments to generate the corresponding first fragment average probability specifically includes:
[0038] L of the first preselected residue fragment xEach of the first Cα atoms is traversed sequentially; and during the traversal, the maximum probability among the M probabilities of the first residue type corresponding to the currently traversed first Cα atom is taken as the corresponding first maximum probability; and at the end of the traversal, the obtained L x The average probability of the first segment is generated by averaging the first maximum probability.
[0039] Preferably, the step of performing three-dimensional molecular structure optimization processing on the first 3D initial structure based on all the first labeled fragments and the first backbone atomic density map to generate the corresponding first optimized structure specifically includes:
[0040] The first 3D initial structure is locally optimized using all the first annotated segments as targets to generate a new first 3D initial structure;
[0041] Using the first backbone atomic density map as the target, the first 3D initial structure is globally optimized according to a preset global optimization mode to generate a new first 3D initial structure; the global optimization mode includes a first mode and a second mode;
[0042] The first 3D initial structure, after completing local and global structural optimization, is output as the corresponding first optimized structure.
[0043] Furthermore, the step of performing local optimization processing on the first 3D initial structure targeting all the first labeled fragments to generate a new first 3D initial structure specifically includes:
[0044] In the three-dimensional space of the first 3D initial structure, the first Cα atoms of all the first labeled fragments are marked as the corresponding first target points; and the Cα atoms on the first 3D initial structure corresponding to each of the first target points are recorded as the corresponding first initial points; and the first 3D initial structure is iteratively optimized based on molecular dynamics simulation technology with all the first target points as the optimization targets. During the iteration process, the point distance between each of the first initial points and the corresponding first target points on the first process optimized structure obtained in each iteration is calculated to generate the corresponding first point distance. When all the first point distances are lower than the preset point distance threshold, the iterative optimization is stopped and the latest first process optimized structure is output as the new first 3D initial structure.
[0045] Furthermore, the step of performing global optimization processing on the first 3D initial structure using the first backbone atomic density map as the target and according to a preset global optimization mode to generate a new first 3D initial structure specifically includes:
[0046] Identify the global optimization pattern;
[0047] When the global optimization mode is the first mode, based on the selected first force field and the first force field potential energy function, the potential energy of the first backbone atom density map is calculated to generate the corresponding first potential energy; and a corresponding first force field target potential energy function is constructed according to the first potential energy and the first force field potential energy function; the first iteration counter is initialized to 0; and based on molecular simulation technology, the first 3D initial structure is optimized for A iterations according to a preset iteration number threshold A, with the goal of minimizing the first force field target potential energy function. The first iteration counter is incremented by 1 during each iteration optimization, and the newly obtained second process is optimized every preset iteration interval X starting when the count value of the first iteration counter equals the preset initial iteration number threshold B. The structure is saved once to obtain a first number Y of the second process optimized structures at the end of A iterations of optimization, where Y = int[(AB) / X] + 1, and int[] is the floor function; and the second process optimized structures whose structural potential energy is lower than a preset potential energy threshold among the first number Y of the second process optimized structures are recorded as the corresponding third process optimized structures; and the three-dimensional electron microscopy density map conversion processing is performed on each of the third process optimized structures to generate the corresponding first electron microscopy density map; and the correlation between each of the first electron microscopy density maps and the first backbone atom density map is calculated to generate the corresponding first correlation, and the third process optimized structure corresponding to the largest first correlation is output as the new first 3D initial structure;
[0048] When the global optimization mode is the second mode, the second iteration counter is initialized to 0; and based on molecular simulation technology under the simulation conditions of a selected force field and a selected potential energy function, the first 3D initial structure is iteratively optimized with the goal of maximizing the correlation between the 3D electron microscopy density map corresponding to the process-optimized structure and the first electron microscopy density map. The second iteration counter is incremented by 1 during each iteration optimization. Starting from when the count value of the second iteration counter equals the preset initial iteration number threshold E, the latest obtained fourth process-optimized structure is converted into a 3D electron microscopy density map every preset iteration interval Z to generate the corresponding second electron microscopy density map. The correlation between the second electron microscopy density map obtained in this iteration and the first backbone atom density map is calculated to generate the corresponding second correlation. When the second correlation exceeds the preset correlation threshold, the iterative optimization is stopped and the latest fourth process-optimized structure is output as the new first 3D initial structure.
[0049] A second aspect of the present invention provides an apparatus for implementing the processing method for optimizing molecular structure based on three-dimensional atomic density maps as described in the first aspect above. The apparatus includes: an acquisition module, an image recognition module, a fragment annotation module, and a structure optimization module.
[0050] The acquisition module is used to acquire the first 3D atomic density map and the corresponding first protein sequence and first 3D initial structure;
[0051] The image recognition module is used to perform target recognition processing on the first 3D atomic density map based on a preset 3D image recognition model to generate the corresponding first Cα atomic density map and first backbone atomic density map;
[0052] The fragment labeling module is used to generate multiple first labeled fragments by performing residue labeling fragment identification processing based on the first Cα atom density map and the first protein sequence.
[0053] The structure optimization module is used to perform three-dimensional molecular structure optimization processing on the first 3D initial structure based on all the first labeled fragments and the first backbone atomic density map to generate the corresponding first optimized structure.
[0054] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0055] The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above;
[0056] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0057] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.
[0058] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for optimizing molecular structures based on three-dimensional atomic density maps. A 3D image recognition model is used to perform semantic recognition on the molecular three-dimensional density map, resulting in two feature maps: a Cα atomic density map identifying all Cα atom features and a backbone atomic density map identifying the features of the protein's core atoms (C, Cα, N). Multiple residue-labeled fragments are extracted from the Cα atomic density map using the Cα atomic density map and a known protein sequence. Local structural optimization of a priori three-dimensional initial structure is performed using molecular dynamics simulation techniques, targeting these residue-labeled fragments. Global structural optimization of the same three-dimensional initial structure is then performed using molecular simulation techniques, targeting the backbone atomic density map. This invention, by using the priori three-dimensional initial structure as the optimization object and the Cα atomic density map and backbone atomic density map from the three-dimensional density map as targets for global and local optimization, avoids residue mismatch problems caused by unlabeled density maps, improves the accuracy of three-dimensional structure optimization, and enhances the optimization efficiency. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of a method for optimizing molecular structure based on three-dimensional atomic density maps provided in Embodiment 1 of the present invention;
[0060] Figure 2 This is a module structure diagram of a processing device for optimizing molecular structure based on three-dimensional atomic density maps, provided in Embodiment 2 of the present invention.
[0061] Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0063] Embodiment 1 of the present invention provides a method for optimizing molecular structure based on three-dimensional atomic density maps, such as... Figure 1 The schematic diagram shows a method for optimizing molecular structure based on three-dimensional atomic density maps provided in Embodiment 1 of the present invention. This method mainly includes the following steps:
[0064] Step 1: Obtain the first 3D atomic density map and the corresponding first protein sequence and first 3D initial structure;
[0065] The first 3D atomic density map is a three-dimensional atomic density map obtained by electron microscopy; the shape of the first 3D atomic density map is H0×W0×D0, where H0, W0 and D0 are the height, width and number of channels of the first 3D atomic density map, respectively; the first protein sequence is the residue type sequence of the protein molecule corresponding to the first 3D atomic density map; and the first 3D initial structure is the standard three-dimensional protein molecule structure corresponding to the first protein sequence.
[0066] Here, based on the well-known imaging principle of cryo-electron microscopy, it is known that cryo-electron microscopy can capture multi-angle two-dimensional projection images of a single protein particle (hereinafter referred to as a protein macromolecule), i.e., multi-angle two-dimensional atomic density maps. Then, based on the multi-angle two-dimensional atomic density maps, a three-dimensional atomic density map can be obtained by constructing a three-dimensional atomic density map. In this embodiment of the invention, this three-dimensional atomic density map is regarded as an electron microscope three-dimensional atomic density map output by cryo-electron microscopy. As is known from the working principle of cryo-electron microscopy, the density map features of each point on the electron microscope three-dimensional atomic density map are actually an electron coulomb potential. The first 3D atomic density map of this embodiment of the invention is the electron microscope three-dimensional atomic density map, and the density map features of each point on the first 3D atomic density map are actually the density map features of each point on the electron microscope three-dimensional atomic density map. The first protein sequence is the protein sequence of the protein macromolecule corresponding to the first 3D atomic density map. The first 3D initial structure is the prior structure of the protein macromolecule corresponding to the first 3D atomic density map.
[0067] Step 2: Based on the preset 3D image recognition model, target recognition processing is performed on the first 3D atomic density map to generate the corresponding first Cα atomic density map and first backbone atomic density map;
[0068] The first Cα atom density map includes multiple first Cα atoms; each first Cα atom includes first Cα atom coordinates, a density map feature, one or more first peptide bond orientations, and multiple first residue type probabilities; the number of first residue type probabilities for each first Cα atom is the same; the first backbone atom density map includes multiple first backbone atom density regions, and the backbone atoms include carbon atoms (i.e., C atoms), alpha carbon atoms (i.e., Cα atoms), and nitrogen atoms (i.e., N atoms);
[0069] Specifically, it includes: Step 21, performing residue feature recognition on the first 3D atomic density map based on a 3D image recognition model to obtain the corresponding first feature map;
[0070] The 3D image recognition model can be implemented based on the 3D Unet model; the first feature map includes multiple first residue targets, each first residue target including one or more first peptide bond orientations and multiple first residue type probabilities; the shape of the first feature map is H1×W1×D1, where H1, W1 and D1 are the height, width and number of channels of the first feature map, respectively, and H1=H0, W1=W0;
[0071] Here, the 3D image recognition model of this invention can be implemented based on the 3D U-Net model. The implementation mechanism of the 3D image recognition model based on the 3D U-Net model can perform multi-level downsampling and hierarchical feature extraction on the input first 3D atomic density map. The model structure of the 3D U-Net model is described in detail in the paper "3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation", and will not be repeated here.
[0072] In the hierarchical feature extraction process of downsampling, the 3D image recognition model of this invention performs target feature identification on amino acid residues (referred to as residues) in the density map. When identifying target features on the amino acid residues (referred to as residues) in the density map, the 3D image recognition model identifies multiple residue objects, i.e., first residue targets, based on the basic density characteristics of the residues. It also predicts the residue type probability (the number of which is consistent with the preset total number of residue types) of each first residue target based on the individual density characteristics of each type of residue. Furthermore, as is generally known, peptide bonds in amino acid residues have orientation. Therefore, by identifying the density change trend of the density map region (also called the density region) corresponding to each first residue target, the peptide bond orientation of each residue can also be identified. Since there may be more than one peptide bond orientation for an amino acid residue, one or more peptide bond orientations, i.e., first peptide bond orientations, can also be identified when performing residue feature identification on the first 3D atomic density map. After obtaining the first feature map containing each first residue target, the 3D image recognition model utilizes 3D... The upsampling mechanism of the Unet model restores the height and width dimensions of the first feature map to be consistent with the height and width dimensions of the first 3D atomic density map, so H1 = H0 and W1 = W0 of the first feature map;
[0073] Step 22, and perform Cα atom feature recognition on the first 3D atomic density map to obtain the corresponding second feature map;
[0074] The second feature map includes multiple first Cα atom targets, each of which includes the coordinates of a first Cα atom. The shape of the second feature map is H2×W2×D2, where H2, W2, and D2 are the height, width, and number of channels of the second feature map, respectively, and H2 = H0 and W2 = W0.
[0075] Here, in the 3D image recognition model of this embodiment, particle target feature identification is performed on the Cα atoms of the residues in the density map during the downsampling hierarchical feature extraction process. When particle target feature identification is performed on the Cα atoms of the residues in the density map, although the atomic radii of Cα atoms and C atoms are basically the same, the bonding structure features of Cα atoms and C atoms in a residue are different. This will make the neighborhood density features of Cα atoms and C atoms different. Therefore, when Cα feature identification is performed on the first 3D atomic density map, each first Cα atom target can be identified based on the C atom radius feature and the Cα atom neighborhood density feature. The three-dimensional coordinates of each first Cα atom target in the first 3D atomic density map are the corresponding first Cα atom coordinates. After obtaining the second feature map with each first Cα atom target, the 3D image recognition model uses the upsampling processing mechanism of the 3D Unet model to restore the height and width dimensions of the second feature map to be consistent with the height and width dimensions of the first 3D atomic density map. Therefore, H2 = H0 and W2 = W0 of the second feature map.
[0076] Step 23, and perform N-atom feature recognition on the first 3D atomic density map to obtain the corresponding third feature map;
[0077] The third feature map includes multiple first N atom targets, each first N atom target including the coordinates of the first N atoms; the shape of the third feature map is H3×W3×D3, where H3, W3 and D3 are the height, width and number of channels of the third feature map, respectively, and H3=H0, W3=W0;
[0078] Here, in the hierarchical feature extraction process of downsampling, the 3D image recognition model of this embodiment performs particle target feature recognition on the N atoms of the residues in the density map. When performing particle target feature recognition on the N atoms of the residues in the density map, because the radius of the N atom is significantly different from that of the Cα atom and the C atom, each first N atom target can be identified directly based on the radius of the N atom. The three-dimensional coordinates of each first N atom target in the first 3D atomic density map are the coordinates of the corresponding first N atom. After obtaining the third feature map with each first N atom target, the 3D image recognition model uses the upsampling processing mechanism of the 3D Unet model to restore the height and width dimensions of the third feature map to be consistent with the height and width dimensions of the first 3D atomic density map. Therefore, H3 = H0 and W3 = W0 of the third feature map.
[0079] Step 24, and perform C atom feature recognition on the first 3D atom density map to obtain the corresponding fourth feature map;
[0080] The fourth feature map includes multiple first C atom targets, each first C atom target including the coordinates of a first C atom; the shape of the fourth feature map is H4×W4×D4, where H4, W4 and D4 are the height, width and number of channels of the fourth feature map, respectively, and H4 = H0 and W4 = W0;
[0081] Here, in the 3D image recognition model of this embodiment, particle target feature identification is performed on the C atoms of the residues in the density map during the downsampling hierarchical feature extraction process. When particle target feature identification is performed on the C atoms of the residues in the density map, although the atomic radii of C atoms and Cα atoms are basically the same, the bonding structure features of C atoms and Cα atoms in a residue are different. This will make the neighborhood density features of C atoms and Cα atoms different. Therefore, when C feature identification is performed on the first 3D atomic density map, each first C atom target can be identified based on the C atom radius feature and the C atom neighborhood density feature. The three-dimensional coordinates of each first C atom target in the first 3D atomic density map are the corresponding first C atom coordinates. After obtaining the fourth feature map with each first C atom target, the 3D image recognition model uses the upsampling processing mechanism of the 3D Unet model to restore the height and width dimensions of the second feature map to be consistent with the height and width dimensions of the first 3D atomic density map. Therefore, H4 = H0 and W4 = W0 of the fourth feature map.
[0082] Step 25, and perform Cα atom feature fusion on the first 3D atomic density map, the first feature map and the second feature map to generate the corresponding first Cα atom density map;
[0083] The first Cα atom density map includes multiple first Cα atoms; each first Cα atom corresponds one-to-one with the first Cα atom target in the second feature map; each first Cα atom corresponds to a set of first Cα atom feature data, which includes the first Cα atom coordinates of the corresponding first Cα atom target on the second feature map, the first density map features matching the corresponding first Cα atom coordinates on the first 3D atom density map, and one or more first peptide bond orientations and multiple first residue type probabilities of the first residue target matching the corresponding first Cα atom coordinates on the first feature map.
[0084] Here, when the 3D image recognition model of this embodiment of the invention performs Cα atom feature fusion on the first 3D atomic density map, the first feature map, and the second feature map, it first counts the number of first peptide bond directions of each first residue target on the first feature map and takes the maximum value as the maximum number of peptide bonds a, and counts the total number of residue types as the total number of residue types b; then, it adds (1+a+b) feature channels on the second feature map that only has Cα atom coordinate features and initializes them to the default value, and uses the second feature map with the newly added channels as the initialized first Cα atom density map. At this time, the first Cα atom on the first Cα atom density map actually corresponds one-to-one with the first Cα atom target on the second feature map; then, it extracts one density map feature corresponding to the matching position of each first Cα atom coordinate on the first 3D atomic density map, one or more (at most a) first peptide bond directions and b first residue type probabilities of the first residue target corresponding to the matching position of each first Cα atom coordinate on the first feature map, and fuses them into the newly added feature channels of each first Cα atom on the first Cα atom density map to output the final fused feature map, i.e., the first Cα atom density map;
[0085] Step 26, and perform backbone atomic density region fusion on the first 3D atomic density map, the second feature map, the third feature map and the fourth feature map to generate the corresponding first backbone atomic density map;
[0086] The first backbone atomic density map includes multiple first backbone atomic density regions; each first backbone atomic density region includes one or more second Cα atoms, or one or more first N atoms, or one or more first C atoms; each second Cα atom corresponds one-to-one with the first Cα atom target in the second feature map, each first N atom corresponds one-to-one with the first N atom target in the third feature map, and each first C atom corresponds one-to-one with the first C atom target in the fourth feature map; each second Cα atom, first N atom, and first C atom corresponds to a set of first backbone atomic feature data; the first backbone atomic feature data includes the first backbone atom type, the first backbone atom coordinates, and the first backbone atomic density map features; the first backbone atom types include Cα atom type, N atom type, and C atom type; the first backbone atom coordinates are the first Cα atom coordinates of the corresponding first Cα atom target in the second feature map, or the first N atom coordinates of the corresponding first N atom target in the third feature map, or the first C atom coordinates of the corresponding first C atom target in the fourth feature map; the first backbone atomic density map features are the density map features on the first 3D atomic density map that match the first backbone atomic coordinates.
[0087] Here, in this embodiment of the invention, when the 3D image recognition model performs backbone atomic density region fusion on the first 3D atomic density map, the second feature map, the third feature map, and the fourth feature map, it first adds two feature channels for marking atom types and atomic coordinates on the first 3D atomic density map and initializes them to default values, and uses the first 3D atomic density map with the newly added channels initialized as the initialized first backbone atomic density map; then, based on the first Cα atom coordinates of the first Cα atom target, the first N atom coordinates of the first N atom target, and the first C atom coordinates of the first C atom target in the second feature map, the third feature map, and the fourth feature map, it marks the first backbone atomic density map with Cα atoms, N atoms, and C atoms to obtain the corresponding second Cα atoms, first N atoms, and first C atoms, and adds the corresponding atom types and atomic coordinates to the two newly added feature channels of each corresponding point of the second Cα atom, first N atom, and first C atom; then Then, taking each second Cα atom, first N atom, and first C atom as the center, the density regions around each second Cα atom, first N atom, and first C atom are marked and confirmed according to the pre-set corresponding atomic radius, corresponding atom, and / or corresponding residue neighborhood density decay function. Other density regions on the first backbone atomic density map that have not been marked and confirmed are deleted to obtain a first backbone atomic density map that only retains the density regions of each second Cα atom, first N atom, and first C atom. Finally, the density regions of each second Cα atom, first N atom, and first C atom on the first backbone atomic density map may intersect with each other to form a connected overall density region. In this embodiment of the invention, this overall density region is regarded as a first backbone atomic density region. Therefore, each first backbone atomic density region may include one or more second Cα atoms, or one or more first N atoms, or one or more first C atoms.
[0088] It should also be noted that the 3D image recognition model in this embodiment of the invention is an intelligent model based on the 3D Unet model. Before using the 3D image recognition model, it needs to be trained. When training the 3D image recognition model, multiple sets of training data will be selected for training. Each set of training data will include at least the three-dimensional density map of the training input and the Cα atom density map and the backbone atom density map used for loss calculation. The loss function used for model training can be set according to specific implementation requirements. Options include: L1_loss, smooth_L1_loss, huber_loss, Lp_loss, L2_loss, L1+L2_loss, TV_loss, etc., which will not be elaborated here.
[0089] Step 3: Based on the first Cα atom density map and the first protein sequence, perform residue annotation fragment identification processing to generate multiple first annotation fragments;
[0090] The first labeled segment includes first segment feature data; the first segment feature data includes first segment number, first segment type sequence and first segment starting coordinates;
[0091] Specifically, this includes: Step 31, linking adjacent Cα atoms according to one or more first peptide bond directions of each first Cα atom in the first Cα atom density map to obtain multiple first Cα atom chains; and treating each first Cα atom chain as a corresponding first residue fragment;
[0092] The first Cα atom chain is composed of multiple first Cα atoms linked together. In the first Cα atom chain, each pair of linked first Cα atoms has a first peptide bond direction that overlaps with each other, and the linking relationship between the two is formed by the overlapping first peptide bond direction and the first Cα atom coordinates of the two first Cα atoms.
[0093] Here, as mentioned above, each first Cα atom in the first Cα atom density map corresponds to one or more first peptide bond directions. In this embodiment of the invention, when linking adjacent Cα atoms according to one or more first peptide bond directions of each first Cα atom in the first Cα atom density map: firstly, the usage state of all first Cα atoms in the first Cα atom density map is initialized to an unused state; then, first Cα atoms with only one first peptide bond direction are marked as atomic chain endpoints; then, any atomic chain endpoint in an unused state is taken as the current starting point, and a spherical region is drawn with the current starting point as the center and a preset atomic distance threshold as the radius. Other first Cα atoms in this spherical region that are in an unused state are recorded as neighboring atoms. The unique first peptide bond direction of the current starting point is matched and compared with any first peptide bond direction of each neighboring atom, and the neighboring atom corresponding to the other first peptide bond direction that matches the unique first peptide bond direction of the current starting point is identified. As a linking atom that matches the current starting point, a link relationship is established between the current starting point and the linking atom. The usage status of both the current starting point and the linking atom is changed to the used state. Then, the linking atom is used as the new current starting point, and the above steps of spherical region construction, neighboring atom labeling, peptide bond direction matching, linking atom positioning, linking and usage status modification are repeated to find the next linking atom and complete the linking until the latest current starting point has no matching linking atom. In this way, starting from one atomic chain endpoint, a first Cα atomic chain composed of multiple first Cα atoms linked sequentially can be obtained. After obtaining a first Cα atomic chain, another atomic chain endpoint in the first Cα atomic density map that is in an unused state is used as the current starting point and the linking is searched and linked in the above manner to obtain another corresponding first Cα atomic chain. The search can be stopped when the number of atomic chain endpoints in the first Cα atomic density map that are in an unused state is 0.
[0094] The above search can yield one or more isolated first Cα atom chains. By considering each first Cα atom chain as a residue fragment, i.e., a first residue fragment, multiple first residue fragments can be obtained.
[0095] Step 32: Statistically count the number of first residue type probabilities of the first Cα atoms on the first Cα atom density map to generate the corresponding total number of first residue types M; statistically count the residue type sequence lengths of the first protein sequence to generate the corresponding first sequence length L; and statistically count the number of first Cα atoms in each first residue fragment to generate the corresponding first fragment length L. x ;
[0096] Step 33: Based on the total number of first residue types M, the length of the first sequence L, and the length of the first fragment L of each first residue fragment... x The residue score and residue start position of each first residue fragment in the first protein sequence are predicted to generate the corresponding first residue fragment score and first residue fragment start position.
[0097] Specifically, this includes: Step 331, based on the total number of first residue types M and the first sequence length L, performing one-hot matrix encoding on the first protein sequence to obtain the corresponding first matrix vector F of shape M×L. * ;
[0098] Wherein, the first matrix vector F * Includes M*L vector data elements, the first matrix vector F * Each row corresponds to a residue type, and the first matrix vector F * The columns are sorted according to the residue type order of the first protein sequence; the first matrix vector F * In each column of M vector data elements, the vector data element whose code matches the residue type corresponding to the current column is set to 1, and the code of the remaining M-1 vector data elements is set to 0.
[0099] For example, suppose the first protein sequence is {QPJVRQ}, the total number of the first residue types is M = 5, the length of the first sequence is L = 6, and the first matrix vector is F. * The first row corresponds to residue type Q, the second row to residue type P, the third row to residue type J, the fourth row to residue type V, and the fifth row to residue type R. Therefore, the resulting first matrix vector F... * It should be:
[0100] The first residue type Q of the first protein sequence corresponding to the first column is due to the first matrix vector F. *The first row corresponds to residue type Q, so the encoding of the first vector data element in the first column is set to 1, and the encodings of the other four vector data elements are set to 0; the second column corresponds to the second residue type P of the first protein sequence, because the first matrix vector F * The second row corresponds to residue type P, so the encoding of the second vector data element in the second column is set to 1, and the encodings of the other four vector data elements are set to 0; the third column corresponds to the third residue type J of the first protein sequence, because the first matrix vector F * The third row corresponds to residue type J, so the code for the third vector data element in the third column is set to 1, and the codes for the other four vector data elements are set to 0; the fourth column corresponds to the fourth residue type V of the first protein sequence, because the first matrix vector F * The fourth row corresponds to residue type V, so the encoding of the fourth vector data element in the fourth column is set to 1, and the encoding of the other four vector data elements is set to 0; the fifth column corresponds to the fifth residue type R of the first protein sequence, because the first matrix vector F * The fifth row corresponds to residue type R, so the fifth vector data element in the fifth column is encoded as 1, and the other four vector data elements are encoded as 0; the sixth column corresponds to the sixth residue type Q of the first protein sequence, because the first matrix vector F * The first row corresponds to residue type Q, so the encoding of the first vector data element in the sixth column is set to 1, and the encoding of the other 4 vector data elements is set to 0.
[0101] Step 332, based on the total number M of the first residue type and the length L of the first fragment of the first residue fragment. x The first residue fragment is subjected to a two-dimensional matrix transformation to obtain the corresponding shape as M×L. x The second matrix vector G * ;
[0102] Wherein, the second matrix vector G * Including M*L x One vector data element, the second matrix vector G * Each row corresponds to a residue type, and the second matrix vector G * The columns are sorted according to the order of the first Cα atom of the first residue fragment; the second matrix vector G * The M vector data elements in each column represent the M first residue type probabilities of the first Cα atom corresponding to the current column;
[0103] For example, suppose the total number of first residue types is M = 5, and the 5 residue types include (type Q, type P, type J, type V, and type R); the length of the first fragment of the first residue segment is L. xIt is 3, composed of 3 first Cα atoms, namely Cα1, Cα2 and Cα3; the probabilities of Cα1 corresponding to the 5 first residue types are r Q1 r P1 r J1 r V1 and r R1 The probabilities of Cα2 corresponding to the five first residue types are r Q2 r P2 r J2 r V2 and r R2 The probabilities of Cα3 corresponding to the five first residue types are r Q3 r P3 r J3 r V3 and r R3 ; Second matrix vector G * The residue types corresponding to each row and the first matrix vector F * Each row corresponds to the same residue type: the first row corresponds to residue type Q, the second row to residue type P, the third row to residue type J, the fourth row to residue type V, and the fifth row to residue type R; the resulting second matrix vector G * It should be:
[0104]
[0105] Step 333: Let the sliding window's sliding step size be 1 column at a time, and let the sliding window's width be L. x The column; and based on the set sliding window step size and sliding window width, the first matrix vector F is... * Starting from the first column, a sliding window process is performed to obtain (LL) x +1) Shapes of M×L x The first sliding window matrix vector; and perform one-dimensional vector transformation on each first sliding window matrix vector to generate the corresponding length (M*L) x The first one-dimensional vector of ); and obtained from (LL x The shape formed by the +1) first one-dimensional vectors is (LL x +1)×(M*L x The third matrix vector F;
[0106] For example, given the first matrix vector F * for: The length of the first fragment of the first residue segment is L. x The length of the first sequence is L = 6, the sliding step is 1 column at a time, and the width of the sliding window is L. x =3 columns of the first matrix vector F * Starting from the first column, a sliding window process is performed to obtain (LL)x The four first sliding window matrix vectors with a shape of 5×3 (+1) are shown below:
[0107]
[0108] The four first one-dimensional vectors of length 15 obtained by performing one-dimensional vector transformation on the above four first sliding window matrix vectors are shown below:
[0109] The first one-dimensional vector {1,0,0,0,1,0,0,0,1,0,0,0,0,0,0}
[0110] The second first one-dimensional vector {0,0,0,1,0,0,0,1,0,0,0,1,0,0,0}
[0111] The third first one-dimensional vector {0,0,0,0,0,0,1,0,0,0,1,0,0,0,1}
[0112] The fourth first one-dimensional vector {0,0,1,0,0,0,0,0,0,1,0,0,0,1,0}
[0113] The third matrix vector F, consisting of the 1st, 2nd, 3rd, and 4th first one-dimensional vectors, has a shape of 4×15 and is:
[0114]
[0115] Step 334, for the second matrix vector G * Perform one-dimensional vector transformation to generate the corresponding vector of length (M*L) x The second one-dimensional vector is then subjected to a two-dimensional matrix upscaling process to obtain the corresponding shape (M*L). x The fourth matrix vector G is 1×1;
[0116] For example, given the second matrix vector G * for: For the second matrix vector G * After performing a one-dimensional vector transformation, the second one-dimensional vector with a length of 5*3=15 is obtained as: {r Q1 ,r P1 ,r J1 ,r V1 ,r R1 ,r Q2 ,r P2 ,r J2 ,r V2 ,r R2 ,r Q3 ,r P3 ,r J3 ,r V3 ,rR3 The fourth matrix vector G, with a shape of 15×1, is obtained by performing a two-dimensional matrix upscaling on the second one-dimensional vector:
[0117]
[0118] Step 335: Perform a cross product operation on the third matrix vector F and the fourth matrix vector G to generate a corresponding shape of (LL). x The fifth matrix vector S, S = F × G, is obtained by performing a one-dimensional vector reduction on the fifth matrix vector S to generate a corresponding vector of length (LL). x +1) the third one-dimensional vector; and from the third one-dimensional vector (LL) x The maximum value among the +1) vector data elements is selected as the corresponding first residue fragment score; and the sorting index of the first residue fragment score in the third one-dimensional vector is used as the starting position of the corresponding first residue fragment.
[0119] For example, let the fifth matrix vector S be as follows:
[0120]
[0121] Then, the third one-dimensional vector obtained by performing one-dimensional vector reduction on the fifth matrix vector is: {s1,s2,s3,s4}; let the maximum value be s3, then the score of the first residue fragment is s3, and the starting position of the first residue fragment is 3;
[0122] Here, a higher score for the first residue fragment indicates a greater likelihood that the first residue fragment contains a matching subsequence in the first protein sequence, and vice versa.
[0123] Step 34: Record the first residue fragment whose score exceeds the preset scoring threshold as the corresponding first pre-selected residue fragment; and based on the first protein sequence and the first fragment length L corresponding to each first pre-selected residue fragment... x The residue type of each first pre-selected residue fragment is corrected according to the starting position of the first residue fragment to generate the corresponding first fragment residue type sequence; and the average probability of each first pre-selected residue fragment is statistically analyzed to generate the corresponding first fragment average probability.
[0124] The first fragment residue type sequence includes multiple second residue types;
[0125] Specifically, this includes: step 341, recording the first residue fragment whose score exceeds a preset score threshold as the corresponding first pre-selected residue fragment;
[0126] Here, the scoring threshold is a pre-set score threshold. If the score of the first residue fragment exceeds the threshold, it means that the first residue fragment is likely to have a matching subsequence in the first protein sequence. It needs to be recorded as the first pre-selected residue fragment and further processed in subsequent steps.
[0127] Step 342, and based on the first protein sequence and the first fragment length L corresponding to each first preselected residue fragment. x The residue type of each first preselected residue fragment is corrected according to the starting position of the first residue fragment to generate the corresponding first fragment residue type sequence;
[0128] Specifically, this includes: step 3421, taking the start position of the first residue fragment as the extraction start position, and taking the length L of the first fragment as the extraction start position. x To extract the length, a subsequence is extracted from the first protein sequence as the corresponding first subsequence; the first subsequence includes L x Each residue type;
[0129] Step 3422, for the L of the first pre-selected residue fragment x Each of the first Cα atoms is traversed sequentially. During traversal, the sorting index of the currently traversed first Cα atom in the first pre-selected residue segment is used as the corresponding first index. The residue type with the highest probability among the M first residue type probabilities corresponding to the currently traversed first Cα atom is used as the corresponding second residue type. The residue type whose sorting index matches the first index in the first subsequence is used as the corresponding first target type. The matching between the second residue type and the first target type is checked. If they match, the traversal continues to the next first Cα atom until the last first Cα atom is traversed. If they do not match, the second residue type is reset using the first target type. After the reset, the traversal continues to the next first Cα atom until the last first Cα atom is traversed. At the end of the traversal, L is obtained. x The second residue type is sorted to form the corresponding first fragment residue type sequence;
[0130] Step 343, and statistically calculate the average probability of each first preselected residue fragment to generate the corresponding first fragment average probability;
[0131] Specifically, this includes: L of the first pre-selected residue fragment. x Each of the first Cα atoms is traversed sequentially; during the traversal, the maximum probability among the M first residue type probabilities corresponding to the currently traversed first Cα atom is taken as the corresponding first maximum probability; and at the end of the traversal, the obtained L... x The average probability of the first segment is generated by averaging the first maximum probability;
[0132] Step 35, set the length L of the first segment x The first pre-selected residue fragment that exceeds a preset length threshold and whose average probability of the first fragment exceeds a preset probability threshold is taken as the corresponding first labeled fragment; and the obtained multiple first labeled fragments are sorted according to the order of the starting positions of the corresponding first residue fragments and the sorting index is taken as the corresponding first fragment number; the first fragment residue type sequence corresponding to each first labeled fragment is taken as the corresponding first fragment type sequence; and the first Cα atom coordinate of the first Cα atom of each first labeled fragment is taken as the corresponding first fragment starting coordinate; and the first fragment number, the first fragment type sequence and the first fragment starting coordinate of each first labeled fragment are combined to form the corresponding first fragment feature data.
[0133] Here, the length threshold and the probability threshold are two pre-set threshold parameters. In this embodiment of the invention, the first pre-selected residue fragment with higher reliability (not short in length and with a relatively high average probability) is further screened from multiple first pre-selected residue fragments using these two threshold parameters as the fragment label, i.e., the first labeled fragment, required for subsequent processing steps.
[0134] Step 4: Based on all the first annotated fragments and the first backbone atomic density map, perform three-dimensional molecular structure optimization processing on the first 3D initial structure to generate the corresponding first optimized structure;
[0135] Specifically, this includes: Step 41, performing local optimization processing on the first 3D initial structure with all first annotated segments as the target to generate a new first 3D initial structure;
[0136] Specifically, this includes: marking the first Cα atoms of all first labeled fragments in the three-dimensional space of the first 3D initial structure as corresponding first target points; and recording the Cα atoms on the first 3D initial structure corresponding to each first target point as corresponding first initial points; iteratively optimizing the first 3D initial structure based on molecular dynamics simulation technology with all first target points as optimization targets, and calculating the point spacing between each first initial point and the corresponding first target point on the first process optimized structure obtained in each iteration to generate the corresponding first point spacing, and stopping the iterative optimization when all first point spacings are lower than the preset point spacing threshold, and outputting the latest first process optimized structure as the new first 3D initial structure;
[0137] Here, this embodiment of the invention targets multiple fragment tags, i.e., the first labeled fragments, and performs local structure optimization on a prior three-dimensional initial structure, i.e., the first 3D initial structure, based on molecular dynamics simulation technology. Molecular dynamics simulation technology is a conventional technique for molecular structure optimization, but if this technique is used for overall simulation, it is easy to encounter problems such as excessively long simulation time or failure to converge. Therefore, this embodiment of the invention uses molecular dynamics simulation technology to perform local structure optimization by targeting multiple fragment tags, which can improve optimization efficiency and shorten optimization time. During the simulation, this embodiment of the invention adds one or more external forces to the atoms corresponding to these fragment tags in the first 3D initial structure based on the published technical principles of molecular dynamics simulation and determines the force function of each external force. Based on the force function corresponding to each atom, the atomic motion state is simulated. The convergence condition of the simulation is whether the point spacing of each set of matching points enters within the preset point spacing threshold. Alternatively, the convergence condition can be set by the root mean square deviation (RMSD) method.
[0138] Step 42: Using the first backbone atomic density map as the target, perform global optimization processing on the first 3D initial structure according to the preset global optimization mode to generate a new first 3D initial structure;
[0139] The global optimization mode includes a first mode and a second mode;
[0140] Specifically, this includes: step 421, identifying the global optimization mode; when the global optimization mode is the first mode, proceeding to step 422; when the global optimization mode is the second mode, proceeding to step 423;
[0141] Step 422: Based on the selected first force field and first force field potential energy function, perform potential energy calculation on the first backbone atomic density map to generate the corresponding first potential energy; construct the corresponding first force field target potential energy function according to the first potential energy and first force field potential energy function; initialize the first iteration counter to 0; and based on molecular simulation technology, perform A iterations of optimization on the first 3D initial structure according to a preset iteration number threshold A, with the goal of minimizing the first force field target potential energy function. In each iteration, increment the first iteration counter by 1, and perform a second-process optimization on the newly obtained structure every preset iteration interval X, starting from when the count value of the first iteration counter equals the preset initial iteration number threshold B. Save the second process optimization structure of the first quantity Y at the end of the A iteration optimization, Y = int[(AB) / X] + 1, where int[] is the floor function; and record the second process optimization structure in the first quantity Y of the second process optimization structure whose structural potential energy is lower than the preset potential energy threshold as the corresponding third process optimization structure; and perform three-dimensional electron microscopy density map conversion on each third process optimization structure to generate the corresponding first electron microscopy density map; and calculate the correlation between each first electron microscopy density map and the first backbone atom density map to generate the corresponding first correlation, and take the third process optimization structure corresponding to the largest first correlation as the new first 3D initial structure output; go to step 43;
[0142] Step 423: Initialize the second iteration counter to 0; and based on molecular simulation technology, under the simulation conditions of a selected force field and a selected potential energy function, perform iterative optimization on the first 3D initial structure with the goal of maximizing the correlation between the 3D electron microscopy density map corresponding to the process-optimized structure and the first electron microscopy density map. In each iteration optimization, increment the second iteration counter by 1. Starting from when the count value of the second iteration counter equals the preset initial iteration number threshold E, perform a 3D electron microscopy density map conversion process on the latest obtained fourth process-optimized structure every preset iteration interval Z to generate the corresponding second electron microscopy density map. Calculate the correlation between the obtained second electron microscopy density map and the first backbone atom density map to generate the corresponding second correlation. Stop iterative optimization when the second correlation exceeds the preset correlation threshold and output the latest fourth process-optimized structure as the new first 3D initial structure.
[0143] Here, in step 42, this embodiment of the invention targets the first backbone atom density map, and performs global structural optimization on the first 3D initial structure optimized in step 41 based on molecular simulation technology. Molecular simulation technology is a conventional technique for molecular structure optimization, but if this technique is used to simulate the state of all atoms (backbone atoms and side chain atoms) in the structure, the simulation time will be too long. Therefore, this embodiment of the invention uses molecular simulation technology to perform global structural optimization targeting the first backbone atom density map, which is actually optimizing the backbone structure, which can improve optimization efficiency and shorten optimization time.
[0144] Step 43: The first 3D initial structure that has completed local and global structural optimization is output as the corresponding first optimized structure.
[0145] Here, in this embodiment of the invention, the first 3D initial structure is locally optimized using molecular dynamics simulation technology, and then the locally optimized first 3D initial structure is globally optimized using molecular simulation technology. Finally, the first 3D initial structure that has undergone both local and global optimization is output as the corresponding first optimized structure. It should be noted that, in another embodiment of the invention, the first 3D initial structure can be globally optimized using molecular simulation technology, and then the globally optimized first 3D initial structure can be locally optimized using molecular dynamics simulation technology. Finally, the first 3D initial structure that has undergone both global and local optimization is output as the corresponding first optimized structure. In yet another embodiment of the invention, a shared structure can be created to load the initial first 3D initial structure, and steps 41 and 42 can be processed simultaneously. Step 41 continuously iteratively optimizes the local structure of the shared structure, and step 42 continuously iteratively optimizes the global structure of the shared structure until both reach convergence, at which point the overall optimization ends and the final shared structure is output as the corresponding first optimized structure.
[0146] Furthermore, it should be noted that, since the molecular structures of RNA, DNA, and material macromolecules are similar to those of protein macromolecules, their 3D atomic density maps and base (group) type sequences (i.e., the corresponding RNA, DNA, and material macromolecule sequences) can be obtained using cryo-electron microscopy, and their prior three-dimensional initial structures can also be obtained. Therefore, after obtaining the (3D atomic density maps, molecular sequences, and three-dimensional initial structures) of RNA, DNA, and material macromolecules, structural optimization processing can be performed on the three-dimensional initial structures of RNA, DNA, and material macromolecules based on a processing method similar to steps 1-4 of the embodiments of the present invention.
[0147] Figure 2 This is a module structure diagram of a processing device for optimizing molecular structures based on three-dimensional atomic density maps, provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 2 As shown, the device includes: an acquisition module 201, an image recognition module 202, a fragment annotation module 203, and a structure optimization module 204.
[0148] The acquisition module 201 is used to acquire the first 3D atomic density map and the corresponding first protein sequence and first 3D initial structure.
[0149] The image recognition module 202 is used to perform target recognition processing on the first 3D atomic density map based on a preset 3D image recognition model to generate the corresponding first Cα atomic density map and first backbone atomic density map.
[0150] The fragment labeling module 203 is used to generate multiple first labeled fragments by performing residue labeling fragment identification processing based on the first Cα atom density map and the first protein sequence.
[0151] The structure optimization module 204 is used to perform three-dimensional molecular structure optimization processing on the first 3D initial structure based on all the first labeled fragments and the first backbone atomic density map to generate the corresponding first optimized structure.
[0152] The present invention provides a processing device for optimizing molecular structure based on three-dimensional atomic density map, which can execute the method steps in the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0153] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the acquisition module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and called and executed by a processing element of the device. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0154] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-on-a-Chip (SOC).
[0155] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0156] Figure 3 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be the aforementioned terminal device or server, or it can be a terminal device or server connected to the aforementioned terminal device or server that implements the method of the embodiments of the present invention. Figure 3 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing method embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.
[0157] exist Figure 3The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0158] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0159] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.
[0160] This invention also provides a chip for executing instructions, which is used to perform the processing steps described in the foregoing method embodiments.
[0161] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for optimizing molecular structures based on three-dimensional atomic density maps. A 3D image recognition model is used to perform semantic recognition on the molecular three-dimensional density map, resulting in two feature maps: a Cα atomic density map identifying all Cα atom features and a backbone atomic density map identifying the features of the protein's core atoms (C, Cα, N). Multiple residue-labeled fragments are extracted from the Cα atomic density map using the Cα atomic density map and a known protein sequence. Local structural optimization of a priori three-dimensional initial structure is performed using molecular dynamics simulation techniques, targeting these residue-labeled fragments. Global structural optimization of the same three-dimensional initial structure is then performed using molecular simulation techniques, targeting the backbone atomic density map. This invention, by using the priori three-dimensional initial structure as the optimization object and the Cα atomic density map and backbone atomic density map from the three-dimensional density map as targets for global and local optimization, avoids residue mismatch problems caused by unlabeled density maps, improves the accuracy of three-dimensional structure optimization, and enhances the optimization efficiency.
[0162] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0163] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0164] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for optimizing molecular structure based on three-dimensional atomic density maps, characterized in that, The method includes: Obtain the first 3D atomic density map and the corresponding first protein sequence and first 3D initial structure; Based on a preset 3D image recognition model, the first 3D atomic density map is processed to generate the corresponding first Cα atomic density map and first backbone atomic density map. Multiple first labeled fragments are generated by identifying residue labeled fragments based on the first Cα atom density map and the first protein sequence; Based on all the first labeled fragments and the first backbone atomic density map, the first 3D initial structure is subjected to three-dimensional molecular structure optimization processing to generate the corresponding first optimized structure; The first Cα atom density map includes multiple first Cα atoms; each first Cα atom includes first Cα atom coordinates, one or more first peptide bond orientations, and multiple first residue type probabilities; the number of first residue type probabilities is the same for each first Cα atom; the first backbone atom density map includes multiple first backbone atom density regions, wherein the backbone atoms include C atoms, Cα atoms, and N atoms; the first labeled fragment includes first fragment feature data; the first fragment feature data includes a first fragment number, a first fragment type sequence, and first fragment start coordinates; The step of generating a corresponding first Cα atomic density map and a first backbone atomic density map by performing target recognition processing on the first 3D atomic density map based on a preset 3D image recognition model specifically includes: performing residue feature recognition on the first 3D atomic density map based on the 3D image recognition model to obtain a corresponding first feature map; performing Cα atom feature recognition on the first 3D atomic density map to obtain a corresponding second feature map; performing N atom feature recognition on the first 3D atomic density map to obtain a corresponding third feature map; performing C atom feature recognition on the first 3D atomic density map to obtain a corresponding fourth feature map; fusing the first 3D atomic density map, the first feature map, and the second feature map with Cα atom features to generate a corresponding first Cα atomic density map; and fusing the first 3D atomic density map, the second feature map, the third feature map, and the fourth feature map with backbone atomic density regions to generate a corresponding first backbone atomic density map. The step of generating multiple first labeled fragments by identifying residue-labeled fragments based on the first Cα atom density map and the first protein sequence specifically includes: Based on one or more of the first peptide bond directions of each of the first Cα atoms in the first Cα atom density map, adjacent Cα atoms are linked to obtain multiple first Cα atom chains; and each first Cα atom chain is regarded as a corresponding first residue fragment; the first Cα atom chain is composed of multiple first Cα atoms linked together, and each pair of linked first Cα atoms in the first Cα atom chain has a first peptide bond direction that overlaps with each other, and the linking relationship between the two is formed by the pair of overlapping first peptide bond directions and the first Cα atom coordinates of the two first Cα atoms respectively; The number of first residue type probabilities of the first Cα atom on the first Cα atom density map is statistically analyzed to generate the corresponding total number of first residue types M; the residue type sequence lengths of the first protein sequence are statistically analyzed to generate the corresponding first sequence length L; and the number of first Cα atoms in each first residue fragment is statistically analyzed to generate the corresponding first fragment length L. x ; Based on the total number M of the first residue types, the length L of the first sequence, and the length L of each of the first residue fragments. x The residue score and residue start position of each first residue fragment in the first protein sequence are predicted to generate the corresponding first residue fragment score and first residue fragment start position. The first residue fragment whose score exceeds a preset score threshold is recorded as the corresponding first pre-selected residue fragment; and based on the first protein sequence and the first fragment length L corresponding to each first pre-selected residue fragment... x The fragment residue type of each of the first pre-selected residue fragments is corrected according to the starting position of the first residue fragment to generate a corresponding first fragment residue type sequence; and the average fragment probability of each of the first pre-selected residue fragments is statistically analyzed to generate a corresponding first fragment average probability; the first fragment residue type sequence includes multiple second residue types; The length L of the first segment x The first pre-selected residue fragments that exceed a preset length threshold and whose average probability exceeds a preset probability threshold are designated as the corresponding first labeled fragments; and the obtained multiple first labeled fragments are sorted according to the order of their corresponding first residue fragment start positions, with the sorting index serving as the corresponding first fragment number; the first fragment residue type sequence corresponding to each first labeled fragment is designated as the corresponding first fragment type sequence; and the first Cα atom coordinate of the first first Cα atom of each first labeled fragment is designated as the corresponding first fragment start coordinate; and the first fragment number, the first fragment type sequence, and the first fragment start coordinate of each first labeled fragment constitute the corresponding first fragment feature data; The step of performing three-dimensional molecular structure optimization processing on the first 3D initial structure based on all the first labeled fragments and the first backbone atomic density map to generate the corresponding first optimized structure specifically includes: performing local optimization processing on the first 3D initial structure with all the first labeled fragments as the target to generate a new first 3D initial structure; performing global optimization processing on the first 3D initial structure with the first backbone atomic density map as the target according to a preset global optimization mode to generate a new first 3D initial structure; the global optimization mode includes a first mode and a second mode; and outputting the first 3D initial structure that has completed local and global structure optimization as the corresponding first optimized structure.
2. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 1, characterized in that, The first 3D atomic density map is a three-dimensional atomic density map obtained by electron microscopy; the shape of the first 3D atomic density map is H0×W0×D0, where H0, W0 and D0 are the height, width and number of channels of the first 3D atomic density map, respectively. The first protein sequence is a residue type sequence of a protein molecule corresponding to the first 3D atomic density map; The first 3D initial structure is a standard three-dimensional protein molecular structure corresponding to the first protein sequence.
3. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 2, characterized in that, The 3D image recognition model is implemented based on the 3D Unet model.
4. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 2, characterized in that, The first feature map includes multiple first residue targets, each first residue target including one or more first peptide bond orientations and multiple first residue type probabilities; the second feature map includes multiple first Cα atom targets, each first Cα atom target including first Cα atom coordinates; the third feature map includes multiple first N atom targets, each first N atom target including first N atom coordinates; the fourth feature map includes multiple first C atom targets, each first C atom target including first C atom coordinates; the shapes of the first, second, third, and fourth feature maps are H1×W1×D1, H2×W2×D2, H3×W3×D3, and H4×W4×D4, respectively, where H1, H2, H3, and H4 are the heights of the corresponding feature maps, H1=H2=H3=H4=H0, W1, W2, W3, and W4 are the widths of the corresponding feature maps, W1=W2=W3=W4=W0, and D1, D2, D3, and D4 are the number of channels in the corresponding feature maps; The first Cα atom density map includes a plurality of first Cα atoms; each first Cα atom corresponds one-to-one with the first Cα atom target in the second feature map; each first Cα atom corresponds to a set of first Cα atom feature data, the first Cα atom feature data including the first Cα atom coordinates of the first Cα atom target corresponding to the first Cα atom target on the second feature map, the first density map feature matching the corresponding first Cα atom coordinates on the first 3D atom density map, one or more first peptide bond orientations and multiple first residue type probabilities of the first residue target matching the corresponding first Cα atom coordinates on the first feature map; The first backbone atom density map includes multiple first backbone atom density regions; each first backbone atom density region includes one or more second Cα atoms, or one or more first N atoms, or one or more first C atoms; each second Cα atom corresponds one-to-one with the first Cα atom target in the second feature map, each first N atom corresponds one-to-one with the first N atom target in the third feature map, and each first C atom corresponds one-to-one with the first C atom target in the fourth feature map; each second Cα atom, first N atom, and first C atom corresponds to a set of first backbone atom feature data; the first... The backbone atom feature data includes a first backbone atom type, first backbone atom coordinates, and a first backbone atom density map feature; the first backbone atom type includes Cα atom type, N atom type, and C atom type; the first backbone atom coordinates are the first Cα atom coordinates of the first Cα atom target corresponding to the second feature map, or the first N atom coordinates of the first N atom target corresponding to the third feature map, or the first C atom coordinates of the first C atom target corresponding to the fourth feature map; the first backbone atom density map feature is the density map feature on the first 3D atom density map that matches the first backbone atom coordinates.
5. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 1, characterized in that, The first residue type is determined based on the total number M, the first sequence length L, and the first fragment length L of each first residue fragment. x Predicting the residue score and residue start position of each first residue fragment in the first protein sequence to generate the corresponding first residue fragment score and first residue fragment start position, specifically including: Based on the total number M of the first residue types and the first sequence length L, the first protein sequence is subjected to one-hot matrix encoding to obtain a corresponding first matrix vector F of shape M×L. * The first matrix vector F * It includes M*L vector data elements, and the first matrix vector F * Each row corresponds to a residue type, and the first matrix vector F * The columns are sorted according to the residue type order of the first protein sequence; the first matrix vector F * In each column of M vector data elements, the vector data element whose code matches the residue type corresponding to the current column is set to 1, and the code of the remaining M-1 vector data elements is set to 0. Based on the total number M of the first residue type and the length L of the first fragment of the first residue fragment. x The first residue fragment is subjected to a two-dimensional matrix transformation to obtain a corresponding shape of M×L. x The second matrix vector G * The second matrix vector G * Including M*L x There are vector data elements, and the second matrix vector G * Each row corresponds to a residue type, and the second matrix vector G * The columns are sorted according to the order of the first Cα atom of the first residue fragment; the second matrix vector G * The M vector data elements in each column represent the M probabilities of the first residue type of the first Cα atom corresponding to the current column; Let the sliding window's sliding step size be 1 column at a time, and let the sliding window's width be L. x The column; and based on the set sliding window step size and sliding window width, the first matrix vector F is... * Starting from the first column, a sliding window process is performed to obtain (LL) x +1) Shapes of M×L x The first sliding window matrix vector; and perform one-dimensional vector transformation on each of the first sliding window matrix vectors to generate a corresponding length (M*L) x The first one-dimensional vector of ); and obtained from (LL x +1) of the first one-dimensional vectors form a shape corresponding to (LL) x +1)×(M*L x The third matrix vector F; For the second matrix vector G * Perform one-dimensional vector transformation to generate the corresponding vector of length (M*L) x The second one-dimensional vector is then subjected to a two-dimensional matrix upscaling process to obtain the corresponding shape (M*L). x The fourth matrix vector G is 1×1; Performing a cross product operation on the third matrix vector F and the fourth matrix vector G generates a corresponding shape of (LL). x The fifth matrix vector S, S=F×G, is obtained by performing a one-dimensional vector reduction process on the fifth matrix vector S to generate a corresponding vector of length (LL). x +1) the third one-dimensional vector; and from the (LL) of the third one-dimensional vector x The maximum value among the +1) vector data elements is selected as the corresponding score of the first residue fragment; and the sorting index of the first residue fragment score in the third one-dimensional vector is used as the starting position of the corresponding first residue fragment.
6. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 1, characterized in that, The first fragment length L corresponding to the first protein sequence and each of the first preselected residue fragments is... x The process of performing residue type correction on each of the first pre-selected residue fragments based on the start position of the first residue fragment to generate a corresponding first fragment residue type sequence specifically includes: Using the start position of the first residue fragment as the extraction start position, and the length L of the first fragment as the extraction start position... x To extract the length, a subsequence is extracted from the first protein sequence as the corresponding first subsequence; the first subsequence includes L x Each residue type; L of the first preselected residue fragment x Each of the first Cα atoms is traversed sequentially; during traversal, the sorting index of the currently traversed first Cα atom in the first pre-selected residue fragment is used as the corresponding first index; the residue type corresponding to the highest probability among the M first residue type probabilities corresponding to the currently traversed first Cα atom is used as the corresponding second residue type; the residue type whose sorting index matches the first index in the first sub-sequence is used as the corresponding first target type; and the matching between the second residue type and the first target type is identified; if they match, the traversal continues to the next first Cα atom until the last first Cα atom is traversed; if they do not match, the second residue type is reset using the first target type, and the traversal continues to the next first Cα atom until the last first Cα atom is traversed; and at the end of the traversal, L is obtained. x The second residue type is sorted to form the corresponding first fragment residue type sequence.
7. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 1, characterized in that, The step of statistically calculating the average probability of each of the first preselected residue fragments to generate the corresponding first fragment average probability specifically includes: L of the first preselected residue fragment x Each of the first Cα atoms is traversed sequentially; and during the traversal, the maximum probability among the M probabilities of the first residue type corresponding to the currently traversed first Cα atom is taken as the corresponding first maximum probability; and at the end of the traversal, the obtained L x The average probability of the first segment is generated by averaging the first maximum probability.
8. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 1, characterized in that, The step of performing local optimization processing on the first 3D initial structure targeting all the first annotated fragments to generate a new first 3D initial structure specifically includes: In the three-dimensional space of the first 3D initial structure, the first Cα atoms of all the first labeled fragments are marked as the corresponding first target points; and the Cα atoms on the first 3D initial structure corresponding to each of the first target points are recorded as the corresponding first initial points; and the first 3D initial structure is iteratively optimized based on molecular dynamics simulation technology with all the first target points as the optimization targets. During the iteration process, the point distance between each of the first initial points and the corresponding first target points on the first process optimized structure obtained in each iteration is calculated to generate the corresponding first point distance. When all the first point distances are lower than the preset point distance threshold, the iterative optimization is stopped and the latest first process optimized structure is output as the new first 3D initial structure.
9. The method for optimizing molecular structure based on three-dimensional atomic density maps according to claim 1, characterized in that, The step of performing global optimization processing on the first 3D initial structure according to a preset global optimization mode, with the first backbone atomic density map as the target, to generate a new first 3D initial structure, specifically includes: Identify the global optimization pattern; When the global optimization mode is the first mode, based on the selected first force field and the first force field potential energy function, the potential energy of the first backbone atom density map is calculated to generate the corresponding first potential energy; and a corresponding first force field target potential energy function is constructed according to the first potential energy and the first force field potential energy function; the first iteration counter is initialized to 0; and based on molecular simulation technology, the first 3D initial structure is optimized for A iterations according to a preset iteration number threshold A, with the goal of minimizing the first force field target potential energy function. The first iteration counter is incremented by 1 during each iteration optimization, and the newly obtained second process is optimized every preset iteration interval X starting when the count value of the first iteration counter equals the preset initial iteration number threshold B. The structure is saved once to obtain a first number Y of the second process optimized structures at the end of A iterations of optimization, where Y = int[(AB) / X] + 1, and int[] is the floor function; and the second process optimized structures whose structural potential energy is lower than a preset potential energy threshold among the first number Y of the second process optimized structures are recorded as the corresponding third process optimized structures; and each of the third process optimized structures is processed by 3D electron microscopy density map conversion to generate the corresponding first electron microscopy density map; and the correlation between each of the first electron microscopy density maps and the first backbone atom density map is calculated to generate the corresponding first correlation, and the third process optimized structure corresponding to the largest first correlation is output as the new first 3D initial structure; When the global optimization mode is the second mode, the second iteration counter is initialized to 0; and based on molecular simulation technology under the simulation conditions of a selected force field and a selected potential energy function, the first 3D initial structure is iteratively optimized with the goal of maximizing the correlation between the 3D electron microscopy density map corresponding to the process-optimized structure and the first electron microscopy density map. The second iteration counter is incremented by 1 during each iteration optimization. Starting from when the count value of the second iteration counter equals the preset initial iteration number threshold E, the latest obtained fourth process-optimized structure is converted into a 3D electron microscopy density map every preset iteration interval Z to generate the corresponding second electron microscopy density map. The correlation between the second electron microscopy density map obtained in this iteration and the first backbone atom density map is calculated to generate the corresponding second correlation. When the second correlation exceeds the preset correlation threshold, the iterative optimization is stopped and the latest fourth process-optimized structure is output as the new first 3D initial structure.
10. An apparatus for performing the processing method for optimizing molecular structure based on a three-dimensional atomic density map as described in any one of claims 1-9, characterized in that, The device includes: an acquisition module, an image recognition module, a fragment annotation module, and a structure optimization module; The acquisition module is used to acquire the first 3D atomic density map and the corresponding first protein sequence and first 3D initial structure; The image recognition module is used to perform target recognition processing on the first 3D atomic density map based on a preset 3D image recognition model to generate the corresponding first Cα atomic density map and first backbone atomic density map; The fragment labeling module is used to generate multiple first labeled fragments by performing residue labeling fragment identification processing based on the first Cα atom density map and the first protein sequence. The structure optimization module is used to perform three-dimensional molecular structure optimization processing on the first 3D initial structure based on all the first labeled fragments and the first backbone atomic density map to generate the corresponding first optimized structure.
11. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the steps of the method according to any one of claims 1-9; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the instructions of any one of claims 1-9.
Citation Information
Patent Citations
Energy-based multi-objective optimization fitting prediction method for atomic structures and electron density maps
CN111968707A
Protein structure modeling method and device, electronic equipment and storage medium
CN115035947A