Method for predicting protein structure, medium and electronic equipment

By employing a periodically restarting cosine annealing strategy and introducing a complex template library and multi-domain datasets, combined with ESMpair matrix and MSA features, the protein structure prediction model is optimized, addressing the insufficient accuracy issue in existing technologies and improving the accuracy and efficiency of protein structure prediction.

CN120853661APending Publication Date: 2025-10-28BIOMAP (BEIJING) INTELLIGENCE TECH LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510924624.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing AI-based protein structure prediction technologies suffer from low accuracy and a limited range of predictable protein sequence types.

Method used

A protein structure prediction model was trained using a periodically restarted cosine annealing strategy. A complex template library and a multi-domain dataset were introduced, and feature fusion was performed by combining the ESMpair matrix and MSA features to optimize the model training process.

Benefits of technology

It improves the accuracy and efficiency of protein structure prediction, especially in the absence of homologous templates, and can better predict the structure of protein complexes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853661A_ABST
    Figure CN120853661A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for predicting a protein structure, a medium and electronic equipment, and the method comprises the steps: searching a template matched with a to-be-analyzed protein sequence from a combined template library, and obtaining a target template, the combined template library comprising a single-chain template library and a compound template library; at least taking the target template as the input of a target protein structure prediction model, and predicting the structure of the to-be-analyzed protein sequence through the target protein structure prediction model, the target protein structure prediction model is obtained by training the protein structure prediction model based on a periodic restart cosine annealing strategy, and the periodic restart cosine annealing strategy comprises the steps of starting a periodic restart cosine annealing stage to enable a learning rate to be reduced to a minimum value according to a cosine function and then performing periodic resetting. According to the embodiment of the invention, the accuracy and efficiency of protein structure model prediction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computational biology, and more specifically, embodiments of this application relate to a method, medium, and electronic device for predicting protein structure. Background Technology

[0002] Protein structure prediction is a research area in computational biology and structural biology, aiming to predict the three-dimensional structure of proteins using computational methods. Since the function of a protein is determined by its three-dimensional structure, predicting protein structure is of great significance for fields such as new drug development and novel protein design.

[0003] Traditional experimental methods (such as X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryo-electron microscopy) can provide predicted protein structure information, but they are time-consuming, costly, and not applicable to the prediction of all protein structures. In recent years, related technologies have offered protein structure prediction based on artificial intelligence (AI), such as using AlphaFold to predict protein sequence structures. However, these AI-based prediction technologies all suffer from technical limitations, including low accuracy in structure prediction and a limited range of predictable protein sequence types. Summary of the Invention

[0004] The purpose of this application is to provide a method, medium, and electronic device for predicting protein structure. The embodiments of this application can improve the accuracy and efficiency of protein structure prediction by artificial intelligence network model (i.e., target protein structure prediction model).

[0005] In a first aspect, embodiments of this application provide a method for predicting protein structure. The method includes: retrieving a template matching a protein sequence to be analyzed from a combinatorial template library to obtain a target template, wherein the combinatorial template library includes a single-strand template library and a complex template library; using at least the target template as input to a target protein structure prediction model, and predicting the structure of the protein sequence to be analyzed through the target protein structure prediction model, wherein the target protein structure prediction model is obtained by training a protein structure prediction model based on a periodically restarted cosine annealing strategy, the periodically restarted cosine annealing strategy including: initiating a periodically restarted cosine annealing phase to reduce the learning rate to a minimum value according to a cosine function and then periodically resetting it, and the duration of each cosine annealing phase is the same or the duration of the subsequent cosine annealing phase is greater than the duration of the previous cosine annealing phase.

[0006] The embodiments of this application introduce a cosine annealing learning rate adjustment strategy for the first time during the training process of a protein structure model (this strategy includes: decreasing according to a cosine function, periodically restarting the learning rate, and dynamically adjusting the period length). That is, by periodically adjusting the learning rate to achieve periodic learning rate decay, it can more effectively optimize model parameters, accelerate the training process, and improve the model's convergence performance. Furthermore, the embodiments of this application introduce a complex template library during the model search phase, enabling the model to better learn the interactions between proteins, thereby improving the model's ability to predict the structure of various protein sequences.

[0007] In some embodiments of this application, before at least using the target template as input to the target protein structure prediction model, the method further includes: performing model initialization settings, wherein the initialization settings include: loading protein data, setting an upper limit of the learning rate, a lower limit of the learning rate, an initial period length, and a period growth factor; performing cyclic training according to periods, wherein the training process for the i-th period includes: resetting the learning rate to the upper limit of the learning rate; determining the learning rate of the i-th period based on the upper limit of the learning rate, the lower limit of the learning rate, the length of the i-th period, the number of steps in the i-th period, and a cosine function, and training the model based on the learning rate of the i-th period; ending the training of the i-th period when the number of steps in the i-th period is reached; dynamically adjusting the length of the i-th period based on the period growth factor, restarting, and entering the (i+1)-th training period; stopping training when the training termination condition is met, and saving the model parameters to obtain the target protein structure prediction model. Here, i is an integer greater than or equal to 1.

[0008] The inventors of this application discovered in their research that the adaptive learning rate training strategy commonly used in protein structure models such as alphafold may lead to local optima, and it is difficult to escape from local optima. Therefore, the inventors of this application designed a periodically restarting cosine annealing learning rate training strategy, which can solve the technical defects such as local optima caused by existing methods.

[0009] In some embodiments, the method further includes: searching for homologous sequences of the protein sequence to be analyzed from a homologous sequence database, and generating target multiple sequence alignment results based on the search results, wherein the target multiple sequence alignment results are used to characterize the evolutionary relationship between residues; wherein, at least using the target template as input to a target protein structure prediction model, and predicting the structure of the protein sequence to be analyzed by the target protein structure prediction model, includes: using the target multiple sequence alignment results and the target template as input to the target protein structure prediction model.

[0010] The embodiments of this application improve the accuracy of the model's prediction of protein sequence structure by using multiple sequence alignment results as input to the model.

[0011] In some embodiments, the step of searching for homologous sequences of the protein sequence to be analyzed from a homologous sequence database and generating a target multiple sequence alignment result based on the search results includes: identifying sequences with evolutionary relevance to the protein sequence to be analyzed from the homologous sequence database to obtain a target homologous sequence list; performing multiple sequence alignment on the protein sequence to be analyzed and the sequences in the target homologous sequence list to obtain an initial multiple sequence alignment result; performing feature pruning on the initial multiple sequence alignment result based on a sliding window, and retaining at least the target percentage of overlapping regions between adjacent windows to obtain an MSA feature matrix, and using the MSA feature matrix as the target multiple sequence alignment result, wherein the MSA feature matrix includes a residue frequency matrix and a co-evolutionary signal.

[0012] The embodiments of this application use a sliding window method for pruning to ensure that there is a certain overlap between each segment. This helps to preserve the contextual information of the sequence and avoids the loss of important evolutionary information due to pruning.

[0013] In some embodiments, the step of identifying sequences with evolutionary relevance to the protein sequence to be analyzed from the homology sequence database to obtain a target homology sequence list includes: preprocessing the protein sequence to be analyzed to obtain a standard protein sequence to be analyzed; segmenting the standard protein sequence to be analyzed into multiple short k-mer segments and constructing a hash-based k-mer index; obtaining multiple sub-blocks based on the homology sequence database, allocating different sub-blocks to different computing nodes or threads, and performing the following operations on each node or thread to obtain multiple first homology sequence lists: using a Bloom filter to filter homology sequences from the corresponding sub-blocks according to the hash k-mer index to obtain the first homology sequence list; if it is confirmed that the number of paired sequences in all first homology sequence lists is lower than a threshold, automatically performing cross-species data supplementation to obtain a second homology sequence list; filtering the second homology sequence list or the first homology sequence list according to a set retention parameter value to obtain the target homology sequence list, wherein the type of the retention parameter value includes: a first parameter value corresponding to the upper limit of the number of sequences retained in a single search and a second parameter value corresponding to the statistical significance of the sequence alignment results.

[0014] The embodiments of this application can significantly improve search speed and reduce computation time through segmentation and multithreading, thereby accelerating the entire MSA search process; the embodiments of this application use data from other species as a supplement when the number of pairs is insufficient. This supplementation mechanism ensures that even when the number of pairs is insufficient, the model can obtain sufficient co-evolution to improve the accuracy of the predicted structure; the embodiments of this application can improve the quality of sequences in the homologous sequence list and thus improve the accuracy of the predicted structure by setting the retention parameter value.

[0015] In some embodiments, the step of retrieving a template matching the sequence of the protein to be analyzed from the composite template library to obtain a target template includes: retrieving a template matching the sequence of the protein to be analyzed from the composite template library; if a matching template is found, the matching template in the composite template library is used as the target template; if no matching template is found, a search is performed in the single-stranded template library and the matching template retrieved in the single-stranded template library is used as the target template.

[0016] The embodiments of this application innovatively introduce complex templates to construct a complex template library, enabling the model to better learn the interactions between proteins, thereby improving the predictive ability of protein complex structures.

[0017] In some embodiments, the target protein structure prediction model includes a target Evoformer module and a target structure module. The step of predicting the structure of the protein sequence to be analyzed using the target protein structure prediction model includes: extracting the hidden state representation of each residue from the protein sequence to be analyzed; determining inter-residue mutual information for residue pairs based on the hidden state representation, wherein the inter-residue mutual information is used to characterize the covariation relationship of residue pairs during evolution; determining an inter-residue interaction information matrix based on the inter-residue mutual information, wherein the inter-residue interaction information matrix is ​​used to characterize the inter-residue mutual information and co-evolutionary signals, the inter-residue mutual information being used to characterize the covariation relationship of residue pairs during evolution, and the co-evolutionary signals being used to characterize the spatial proximity between hidden residues; and obtaining the structure at least based on the inter-residue interaction information matrix and the target Evoformer module and the target structure module, wherein the target Evoformer module and the target structure module are modules obtained after training.

[0018] Embodiments of this application use the ESMpair matrix, an inter-residue interaction information matrix, as one of the input features to help the model learn the interactions between residues. Specifically, embodiments of this application require the generation of an ESMpair matrix during the training and inference phases to assist in structure training and prediction. Adding ESMpair features to the protein structure model, and fusing them with MSA features, allows the ESMpair matrix to provide single-sequence-driven co-evolutionary signals. Some embodiments of this application fuse MSA features with the ESMpair matrix to obtain joint features, which are then input into the pre-trained Evoformer and structural modules.

[0019] In some embodiments, the method further includes: searching for homologous sequences of the protein sequence to be analyzed from a homologous sequence database, and generating a target multiple sequence alignment result based on the search results, wherein the target multiple sequence alignment result is characterized by an evolutionary matrix; wherein obtaining the structure at least based on the residue interaction information matrix and the target Evoformer module and the target structure module includes: concatenating or weighting the residue interaction information matrix and the evolutionary matrix to form a joint feature tensor; and inputting the joint feature tensor into the target Evoformer module and the target structure module to determine the structure.

[0020] In the feature construction stage, embodiments of this application input the joint feature obtained by fusing the ESMpair matrix with other features (such as MSA features) into the Evoformer module. For targets lacking homologous sequences (such as orphan proteins), the residue interaction information matrix ESMpair can provide single-sequence-driven co-evolutionary signals, thereby improving the accuracy of structural prediction for low homology protein sequences.

[0021] In some embodiments, before predicting the structure of the protein sequence to be analyzed using the target protein structure prediction model, the method further includes: obtaining each sample protein sequence from a target multi-domain combined dataset, wherein the target multi-domain combined dataset is obtained based on a natural complex dataset, a target pseudo-multimer dataset, and a predicted pseudo-multimer dataset; determining the hidden state of the sample protein sequence and determining a sample residue interaction information matrix based on the hidden state; obtaining a sample joint feature tensor based on the sample residue interaction information matrix and / or the sample evolution matrix, wherein the sample evolution matrix is ​​obtained based on the sample multi-sequence alignment results, which are obtained by feature pruning based on a sliding window; training the Evoformer module and the structure module for the i-th time based on the sample joint feature tensor to obtain the i-th prediction result, wherein i is an integer greater than or equal to 1; calculating a loss value based on the i-th prediction result, and updating parameters based on the loss value and the i-th target learning rate, wherein the i-th target learning rate is determined based on the periodically restarted cosine annealing strategy.

[0022] The embodiments of this application improve the training effect of the protein structure prediction model based on multi-domain datasets and cosine annealing algorithm, and ultimately improve the accuracy of structure prediction based on the trained model.

[0023] In some embodiments, before obtaining the protein sequences of each sample from the target multi-domain combined dataset, the method further includes: obtaining an original multi-domain dataset from multiple databases; performing conflict resolution, resolution filtering, and domain overlap removal on the data in the original multi-domain dataset to obtain a cleaned multi-domain dataset; performing structural domain splitting, pairwise single-domain structural combination, and complex verification on the cleaned multi-domain dataset to obtain an initial pseudo-multimer dataset; and sequentially performing multimer complex generation, sequence clustering, and structural deduplication on the initial pseudo-multimer dataset to obtain the target pseudo-multimer dataset.

[0024] Some embodiments of this application provide a method for obtaining pseudo-multiplex data (i.e., each data point in the target pseudo-multiplex dataset), thereby increasing the amount of data obtained in the multi-domain dataset and ultimately improving the effectiveness of model training based on the multi-domain dataset.

[0025] In some embodiments, obtaining the original multi-domain dataset from multiple databases includes: merging and deduplicating the data from the multiple databases to obtain the original multi-domain dataset, wherein the data in the original multi-domain dataset is obtained by deduplicating the merged dataset according to deduplication principles, the deduplication principles including: deduplication based on the number of domains and domain intervals, and the data in the original multi-domain dataset is a protein complex containing multiple structural domains; the step of performing conflict resolution, resolution filtering, and domain overlap removal on the data in the original multi-domain dataset to obtain a cleaned multi-domain dataset includes: when the same PDB entry is annotated by multiple databases, selection is made according to the priority set for the multiple databases; entries with the same number of domains and sequence overlap greater than a set percentage are retained; structures with atomic resolution less than a set threshold are retained; and entries with inter-domain sequence overlap are deleted.

[0026] The original multi-domain dataset obtained by some embodiments of this application has no duplicate data. This is because, during the merging process, duplicate data is removed according to the number of domains and the domain range, ensuring that each domain appears only once in the database and avoiding data redundancy. The original multi-domain dataset obtained by some embodiments of this application has reliable data quality. This is because, in the selection of data sources, priority (SCOP>CATH>ECOD) is used. When data conflicts occur, data from the database with higher priority is selected first, thereby ensuring the data quality of the merged database.

[0027] In some embodiments, the step of performing structural domain splitting, pairwise single-domain structure combination, and complex verification on the cleaned multi-domain dataset to obtain an initial pseudo-multimer dataset includes: segmenting the data of the cleaned multi-domain dataset according to the domain range to identify single-domain structures in the cleaned multi-domain dataset to obtain a single-domain dataset; and obtaining the initial pseudo-multimer dataset by combining two single-domain structures from the same protein database (PDB) in the single-domain dataset and confirming that the combined single-domain structures belong to a complex based on the distance between the two combined single-domain structures, wherein the data in the initial pseudo-multimer dataset is multi-domain data obtained by combining the single-domain structures.

[0028] Some embodiments of this application can generate pseudo-multimers through structural decomposition and verification, thereby improving the quantity and quality of training data and ultimately improving the accuracy of the trained model in predicting protein structures.

[0029] In some embodiments, before obtaining the individual sample protein sequences from the target multidomain combined dataset, the method further includes: performing domain prediction on multidomain protein sequences that are not annotated by a structural classification database to obtain the predicted pseudomultimer dataset.

[0030] Some embodiments of this application provide a method for obtaining a predicted pseudo-multiplex dataset, thereby increasing the amount of data obtained from the multi-domain dataset and ultimately improving the effectiveness of model training based on the multi-domain dataset.

[0031] In some embodiments, before obtaining the protein sequences of each sample from the target multi-domain combined dataset, the method further includes: performing experimental method screening, resolution filtering, and sequence multi-dimensional quality control on the multi-chain complex structures in the PDB (Protein Data Bank) database to obtain the natural complex dataset, wherein the dimensions of the sequence multi-dimensional quality control include: sequence composition and stereochemical rationality.

[0032] It should be noted that the PDB (Protein Data Bank) database is a database of three-dimensional structures of biological macromolecules, storing atomic-level three-dimensional structural data of biomolecules such as proteins, nucleic acids, and complexes that have been resolved through experimental methods.

[0033] Some embodiments of this application provide a method for obtaining a dataset of natural complexes, thereby increasing the amount of data obtained from the multi-domain dataset and ultimately improving the effectiveness of model training based on the multi-domain dataset.

[0034] In some embodiments, before obtaining the protein sequences of each sample from the target multidomain combined dataset, the method further includes: incorporating the predicted pseudomultimer dataset into the target pseudomultimer dataset to obtain an extended pseudomultimer dataset; and integrating the natural complex with the extended pseudomultimer dataset to obtain the target multidomain combined dataset.

[0035] The embodiments of this application introduce a large amount of high-quality multi-domain data (i.e., data from the target multi-domain combined dataset) into the model training based on the original protein structure model training data, enriching the model's training dataset, making the model learning process more comprehensive and improving its scalability, thereby enhancing the model's performance.

[0036] Secondly, some embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the method for predicting protein structure as described in any of the embodiments included in the first aspect.

[0037] Thirdly, some embodiments of this application provide an electronic device including a memory and a processor, the memory being configured to store a computer program and the processor being configured to read the program from the memory and execute it to implement the method for predicting protein structure as described in any of the embodiments included in the first aspect. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 One of the flowcharts for a method of predicting protein structure provided in the embodiments of this application;

[0040] Figure 2 A second flowchart illustrating the method for predicting protein structure provided in this application embodiment;

[0041] Figure 3 The third flowchart of the method for predicting protein structure provided in the embodiments of this application;

[0042] Figure 4 Flowchart four of the methods for predicting protein structure provided in the embodiments of this application;

[0043] Figure 5 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0044] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0045] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0046] Traditional protein structure prediction methods provided by related technologies mainly include homology modeling, free modeling, and fragment assembly techniques. Homology modeling relies on protein templates with known structures to predict the structure of the target protein through sequence alignment; free modeling predicts the structure de novo using physical energy functions and optimization algorithms. These methods have limitations in prediction accuracy and computational efficiency, especially in the absence of homology templates, where they cannot predict protein structures.

[0047] The limitations of deep learning-based protein structure prediction technologies include at least the following: poor generalization ability when dealing with complex protein systems, limited ability to predict dynamic structures and multi-molecular interactions, and poor performance in predicting complex proteins. In other words, while existing neural network-based protein structure prediction methods have made some progress, they still fall short in terms of prediction accuracy.

[0048] This application provides a method for predicting protein structure, which improves the accuracy and efficiency of protein structure prediction by a protein structure model. The method for predicting protein structure in this application introduces, for the first time, a complex template, a periodically restarted cosine annealing training strategy, innovatively introduces multi-domain data (i.e., a target multi-domain combined dataset) to enrich the training data of the protein structure model, applies the ESMPair (i.e., inter-residue interaction information matrix) strategy to optimize the training of the protein structure prediction model or improve the accuracy of the prediction results using the protein structure prediction model, and optimizes the MSA search process, significantly improving the accuracy of protein structure prediction.

[0049] Please refer to Figure 1 , Figure 1 A flowchart of a method for predicting protein structure provided in this application embodiment, the method comprising:

[0050] The first step is to perform multiple sequence alignment on the protein sequence to be analyzed to obtain the target multiple sequence alignment result; and then, to perform template search based on the protein sequence to be analyzed to obtain the target template.

[0051] The second step involves inputting the target multiple sequence alignment results, the target template, and the protein sequence to be analyzed into the target protein structure prediction model to obtain the target confidence assessment structure.

[0052] The third step is to output a PDF file, which provides the predicted structure obtained from the structural prediction of the protein sequence to be analyzed.

[0053] It should be noted that, Figure 1 The target protein structure prediction model shown can be an AlphaFold-based model, which achieves high-precision structure prediction by combining deep learning with multiple sequence alignment and attention mechanisms. As an example, this target protein structure prediction model includes a target feature construction module, a target Evoformer module (which processes multiple sequence alignment information and template information to extract interactions between residues), a target structure module (which generates atomic coordinates based on geometric constraints), and a target confidence evaluation model. These modules are all obtained through training. Figure 1This is just one implementation example. Those skilled in the art can adjust certain steps or use other deep network learning models (e.g., the RoseTTAFold model) as the target protein structure prediction model according to actual needs. For example, the structure file of the predicted protein sequence does not have to be a PDF file. It should be noted that in some embodiments of this application, structure prediction can be performed in the absence of target multiple sequence alignment results. For example, the structure of the protein to be analyzed can be determined by the residue interaction information matrix introduced later.

[0054] That is to say, in some embodiments of this application, such as Figure 2 As shown, embodiments of this application provide a method for predicting protein structure, the method comprising:

[0055] S110, retrieve a template that matches the sequence of the protein to be analyzed from the combined template library to obtain the target template, wherein the combined template library includes a single-stranded template library and a complex template library.

[0056] S120, at least the target template is used as input to the target protein structure prediction model, and the structure of the protein sequence to be analyzed is predicted by the target protein structure prediction model.

[0057] For example, in some embodiments of this application, S120 includes: using the target template as input to a target protein structure prediction model, and predicting the structure of the protein sequence to be analyzed through the target protein structure prediction model.

[0058] For example, in some other embodiments of this application, S120 includes: using the target multiple sequence alignment result (the specific acquisition process can be referred to in other parts of this document) and the target template as input to the target protein structure prediction model, and predicting the structure of the protein sequence to be analyzed through the target protein structure prediction model.

[0059] It should be noted that the target protein structure prediction model is obtained by training the protein structure prediction model based on a periodically restarted cosine annealing strategy. The periodically restarted cosine annealing strategy includes: starting a periodically restarted cosine annealing phase to reduce the learning rate to a minimum value according to the cosine function and then periodically resetting it, and the duration of each cosine annealing phase is the same or the duration of the next cosine annealing phase is longer than the duration of the previous cosine annealing phase.

[0060] The meaning of the periodically restarted cosine annealing strategy used in the embodiments of this application includes:

[0061] Cosine annealing means that the learning rate smoothly decreases from its initial value to its minimum value according to a cosine function.

[0062] Warm restarts mean that after each cycle, the learning rate jumps back to its initial value (or an adjusted initial value) and begins a new round of decline.

[0063] Dynamic period length means that it supports a fixed period length or a gradually increasing period (controlled by the T_mult parameter).

[0064] For example, in some embodiments of this application, before at least using the target template as input to the target protein structure prediction model, the method further includes: performing model initialization settings, wherein the initialization settings include: loading protein data (e.g., including template data of samples or MSA data of samples), setting an upper limit of learning rate, a lower limit of learning rate, an initial period length, and a period growth factor; performing cyclic training according to periods, wherein the training process for the i-th period includes: resetting the learning rate to the upper limit of learning rate; determining the learning rate of the i-th period according to the upper limit of learning rate, the lower limit of learning rate, the length of the i-th period, the number of steps in the i-th period, and a cosine function, and training the model according to the learning rate of the i-th period; ending the training of the i-th period when the number of steps in the i-th period is reached; dynamically adjusting the length of the i-th period according to the period growth factor, restarting, and entering the (i+1)-th training period; stopping training when the training termination condition is met, and saving the model parameters to obtain the target protein structure prediction model. Where i is an integer greater than or equal to 1.

[0065] It is easy to understand that the embodiments of this application introduce a cosine annealing learning rate adjustment strategy for the first time during the training process of the protein structure model. This involves periodically adjusting the learning rate, i.e., periodically decaying the learning rate, which can more effectively optimize model parameters, accelerate the training process, and improve the model's convergence performance. Furthermore, the embodiments of this application introduce a complex template library during the model search phase, enabling the model to better learn the interactions between proteins, thereby improving its ability to predict protein complex structures.

[0066] The following examples illustrate the process of obtaining target multiple sequence alignment results provided by some embodiments of this application.

[0067] In some embodiments of this application, the method for predicting protein structure further includes: searching for homologous sequences of the protein sequence to be analyzed from a homology sequence database (e.g., the UniRef database), and generating target multiple sequence alignment results based on the search results, wherein the target multiple sequence alignment results are used to characterize the evolutionary relationships between residues; the corresponding S120 exemplarily includes: using the protein sequence to be analyzed, the target multiple sequence alignment results, and the target template as input to a target protein structure prediction model, and predicting the structure of the protein sequence to be analyzed through the target protein structure prediction model. It is understood that the embodiments of this application, by using the multiple sequence alignment results as input to the model, can improve the accuracy of the model in predicting protein sequence structure.

[0068] For example, some embodiments of this application may include searching for homologous sequences of the protein sequence to be analyzed from a homologous sequence database (e.g., the UniRef database) by: obtaining evolutionary information through a homologous sequence database (such as UniRef), capturing conserved residues and co-evolutionary patterns, and extracting homologous sequences from the homologous sequence database using a tool (such as HHblits).

[0069] It should be noted that finding homologous sequences involves identifying sequences from relevant databases that have evolutionary relevance to the protein sequence being analyzed. These homologous sequences may contain conserved regions, which can help in understanding the function and structure of the target sequence. Multiple sequence alignment (MSA) results arrange multiple biological sequences (such as protein or nucleic acid sequences) according to their similarity to reveal their evolutionary relationships and common features. Understandably, in some embodiments of this application, after finding homologous sequences, multiple sequence alignment is performed on these sequences to determine their common structures and conserved regions. Exemplary alignment algorithms that can be used include Clustal Omega or MAFFT.

[0070] In some embodiments, the step of searching for homologous sequences of the protein sequence to be analyzed from a homologous sequence database and generating a target multiple sequence alignment result based on the search results includes: identifying sequences with evolutionary relevance to the protein sequence to be analyzed from the homologous sequence database to obtain a target homologous sequence list; performing multiple sequence alignment on the protein sequence to be analyzed and the sequences in the target homologous sequence list (i.e., aligning the searched homologous sequences to show their conserved regions and evolutionary relationships) to obtain an initial multiple sequence alignment result; performing feature trimming on the initial multiple sequence alignment result MSA based on a sliding window (the purpose of trimming is at least to obtain a matrix that fits the model input), and retaining at least the target percentage of overlapping regions between adjacent windows to obtain an MSA feature matrix (i.e., an evolutionary matrix), and using the MSA feature matrix as the target multiple sequence alignment result, wherein the MSA feature matrix includes a residue frequency matrix and a co-evolutionary signal.

[0071] It is understood that the embodiments of this application use a sliding window method for pruning to ensure a certain overlap between each segment. This helps to preserve the contextual information of the sequence and avoids the loss of important evolutionary information due to pruning. In other words, some embodiments of this application use a sliding window method for pruning when features are pruned to a fixed length, ensuring a certain overlap between each segment. This helps to preserve the contextual information of the sequence and avoids the loss of important evolutionary information due to pruning.

[0072] In some embodiments of this application, the step of identifying sequences with evolutionary relevance to the protein sequence to be analyzed from the homologous sequence database to obtain a target homologous sequence list includes: preprocessing the protein sequence to be analyzed to obtain a standard protein sequence to be analyzed; segmenting the standard protein sequence to be analyzed into multiple short fragment k-mer (converting long sequence alignment into short fragment matching to avoid the high time consumption of global alignment), and constructing a hash-based k-mer index; obtaining multiple sub-blocks based on the homologous sequence database, allocating different sub-blocks to different computing nodes or threads, and executing the following through each node or thread: The operation yields a first homologous sequence list: homologous sequences are filtered from the corresponding sub-blocks using a Bloom filter based on the k-mer index of the hash to obtain the first homologous sequence list; if it is confirmed that the number of paired sequences in all first homologous sequence lists is lower than a threshold, cross-species data supplementation is automatically performed to obtain a second homologous sequence list; the second homologous sequence list or the first homologous sequence list is filtered according to the set retention parameter value to obtain the target homologous sequence list, wherein the type of the retention parameter value includes: a first parameter value corresponding to the upper limit of the number of sequences retained in a single search and a second parameter value corresponding to the statistical significance of the sequence alignment results.

[0073] It should be noted that the aforementioned cross-species data supplementation is based on paired sequence supplementation using cross-species databases (such as EggNOG and OrthoDB), prioritizing species with closer evolutionary distances to the protein sequences being analyzed to ensure the biological validity of the co-evolutionary signal. Cross-species data supplementation ensures that the number of effective sequences in the comparative structure meets the quantity requirements, avoiding model underfitting due to data sparsity. For example, cross-species data supplementation can increase the coverage of the co-evolutionary signal by 25%-30%. The retention parameters set above include an upper limit on the number of sequences retained (i.e., the first parameter value) and an E-value (i.e., the second parameter value). The adjustment strategy is to appropriately increase the upper limit on the number of sequences retained and appropriately decrease the E-value threshold to reduce noise. For example, the specific strategy for increasing the upper limit on the number of sequences is to adjust the upper limit on the number of sequences retained in a single search from 1,000 to 2,000; the strategy for decreasing the E-value threshold is to adjust the E-value threshold for sequence screening from 1e-5 to 1e-10, eliminating low-confidence matches.

[0074] It is understood that the embodiments of this application can significantly improve search speed and reduce computation time through segmentation and multithreading, thereby accelerating the entire MSA search process. The embodiments of this application use data from other species as supplementary data when the number of pairings is insufficient. This supplementary mechanism ensures that the model can obtain sufficient co-evolutionary information even when the number of pairings is insufficient. The embodiments of this application can improve the quality of sequences in the homologous sequence list by setting retention parameter values.

[0075] In other words, some embodiments of this application optimize existing MSA strategies to more accurately capture evolutionary information of protein sequences, accelerate search efficiency, and thus improve the accuracy of structure prediction and reduce time consumption. The specific sub-strategies are as follows: (1) Using block, segmentation, and multi-threading strategies can significantly improve search speed and reduce computation time, thereby accelerating the entire search process. (2) When the number of pairs is insufficient, data from other species is used as a supplement. This supplementation mechanism ensures that the model can obtain sufficient co-evolution even when the number of pairs is insufficient. (3) Appropriately increase the upper limit of the number of sequences to be retained and appropriately decrease the E-value threshold to reduce noise.

[0076] The following example illustrates... Figure 1 The template search process, i.e. Figure 2 The implementation process of S110.

[0077] In some embodiments of this application, S110 includes, for example, retrieving a template from the complex template library that matches the sequence of the protein to be analyzed; if a template is found, using the matching template in the complex template library as the target template; if no template is found, searching the single-stranded template library and using the matching template retrieved in the single-stranded template library as the target template.

[0078] The embodiments of this application innovatively introduce complex templates to construct a complex template library, enabling the model to better learn the interactions between proteins, thereby improving the predictive ability of protein complex structures.

[0079] In other words, the embodiments of this application first optimize the template library during template search by introducing complex templates and constructing a complex template library. This enables the model to better learn the interactions between proteins, thereby improving its ability to predict the structure of protein complexes. For example, these sub-strategies include: innovatively introducing complex model logic into the model training and inference process to improve model accuracy. That is, in the case of the original single-stranded template, a new logic for using complex templates is added. The complex template information adds more useful features that help the model learn and applies them to the structural prediction of the target sequence to assist in predicting the three-dimensional structure of the target protein. The model iteratively updates this information to gradually optimize the prediction results. The updated strategy uses the logic of the original single-stranded template as a fallback. If the complex template logic does not find a matching template, the single-stranded template logic continues to be used.

[0080] like Figure 3 As shown, the template search process before optimization includes: Figure 3 The steps in the blue box on the left, namely, searching in the single-chain template library to obtain single-chain template search results, end the template search. The optimized template search process in this embodiment includes: performing a template search based on the constructed complex template library, determining whether there are search results in the complex template library, if so, using the template as the template search result; otherwise, continuing to search in the single-chain template library to obtain single-chain template search results and using those results as the template search result.

[0081] It should be noted that some embodiments of this application are based on the inter-residue interaction information matrix ESMpair, which is used to characterize inter-residue interactions or co-evolutionary information, as input. Figure 1 One of the features of the target Evoformer module is that it enables protein structure prediction models to learn the interactions between residues. During the inference phase, when using the target protein structure prediction model to predict the structure of the protein sequence being analyzed, it is also necessary to generate an inter-residue interaction information matrix to assist in the structure prediction.

[0082] For example, in some embodiments of this application, the target protein structure prediction model includes a target Evoformer module and a target structure module. S120, predicting the structure of the protein sequence to be analyzed using the target protein structure prediction model, exemplarily includes: extracting the hidden state representation of each residue from the protein sequence to be analyzed; determining inter-residue mutual information for residue pairs based on the hidden state representation, wherein the inter-residue mutual information is used to characterize the covariation relationship of residue pairs during evolution; determining an inter-residue interaction information matrix based on the inter-residue mutual information, wherein the inter-residue interaction information matrix is ​​used to characterize the inter-residue mutual information and co-evolutionary signals, the inter-residue mutual information being used to characterize the covariation relationship of residue pairs during evolution, and the co-evolutionary signals being used to characterize the spatial proximity between hidden residues; and obtaining the structure at least based on the inter-residue interaction information matrix and the target Evoformer module and the target structure module, wherein the target Evoformer module and the target structure module are modules obtained after training.

[0083] For example, in some embodiments of this application, the process of determining the inter-residue interaction information matrix based on the inter-residue mutual information includes: first, extracting the hidden state representation of each residue (e.g., the input sequence is forward-propagated through the ESM-1b model to extract the hidden state representation of each residue); second, mutual information calculation: performing matrix operations (such as covariance matrix or attention weights) on the hidden state representation to generate the MI value matrix of residue pairs (i.e., the inter-residue interaction information matrix ESMpair).

[0084] The embodiments of this application capture co-evolutionary signals between sequences through an inter-residue interaction information matrix, simulate conformational changes in protein complexes, and thus more accurately predict the structure of the complexes.

[0085] It should be noted that in some embodiments of this application, the information matrix of inter-residue interactions is obtained and the matrix is ​​fused with the evolutionary matrix (i.e., Figure 1 The process of (target multiple sequence alignment results) is by Figure 1 The target feature construction module performs this. In some embodiments of this application, if target multiple sequence alignment results exist, they are fused with ESMpair as... Figure 1 The target protein structure prediction module uses the input data to perform structure prediction; if there is no target multiple sequence alignment result, ESMpair is used directly as the input to the target protein structure prediction module for structure prediction.

[0086] For example, in some embodiments of this application, the method further includes: searching for homologous sequences of the protein sequence to be analyzed from a homologous sequence database, and generating target multiple sequence alignment results based on the search results, wherein the target multiple sequence alignment results are characterized by an evolutionary matrix; wherein obtaining the structure at least based on the inter-residue interaction information matrix and the target Evoformer module and the target structural module includes: splicing or weighting the inter-residue interaction information matrix and the evolutionary matrix (e.g., through...). Figure 1 The target feature construction module implements splicing and weighted fusion to form a joint feature tensor; the joint feature tensor is then input... Figure 1 The target Evoformer module (e.g., in some embodiments of this application, the module may add a processing layer for the composite template features) and the target structure module determine the structure.

[0087] The embodiments of this application, through feature fusion (i.e., if an evolutionary matrix exists, it is fused with the residue interaction information matrix ESMpair) and structure prediction (i.e., the target Evoformer module and the target structure module are trained by combining feature inputs to generate 3D coordinates), can provide a single sequence-driven co-evolutionary signal when performing structure prediction for targets lacking homologous sequences (such as orphan proteins) by replacing sparse MSA, thereby predicting the structure of such protein sequences.

[0088] In other words, some embodiments of this application have verified the significant advantages of this strategy in improving model performance by optimizing protein structure models through ESMpair. ESMpair captures co-evolutionary signals between sequences by jointly encoding protein sequence pairs and, combined with a dynamic learning mechanism, simulates conformational changes in protein complexes, thereby more accurately predicting complex structures and integrating ESMpair information into the training and inference process of the structure model.

[0089] like Figure 4 As shown in the figure, this process of obtaining joint feature tensors according to some embodiments of this application includes:

[0090] (1) ESM-1b model; the hidden state representation of each residue is extracted through this model.

[0091] (2) Inter-residue mutual information: The hidden state representation of residue pair (i,j) is calculated and the inter-residue mutual information (MI) is determined based on the hidden state representation. The formula is as follows:

[0092]

[0093] Where p(x,y) is the joint probability distribution of x and y, p(x) and p(y) are the marginal probability distributions, and x and y represent residue pairs (i,j), respectively.

[0094] (3) Co-evolution matrix: Convert the MI value into a symmetric matrix to obtain the co-evolution matrix, i.e., the ESMpair matrix.

[0095] (4) Joint feature fusion: The MSA feature (i.e., the evolutionary matrix corresponding to the evolutionary information) and ESMpair (i.e., the residue interaction information matrix corresponding to the co-evolutionary signal) are fused to obtain the joint feature tensor.

[0096] The following example illustrates how to train a protein structure prediction model. Figure 1 The process of predicting the target protein structure using a model.

[0097] It should be noted that some embodiments of this application introduce a periodically restarting cosine annealing strategy during the training process and train the protein structure prediction model by introducing multi-domain data (i.e., data from the target multi-domain combined dataset). Introducing multi-domain datasets can enrich the model's training dataset, making the model learning process more comprehensive and scalable, thereby improving model performance. The application of the periodically restarting cosine annealing learning rate training strategy in protein structure model training (the training of protein structure prediction models such as Alphafold usually adopts an adaptive learning rate training strategy) combines the characteristics of smooth descent of the cosine function and periodic learning rate reset. It aims to help the model escape local optima and accelerate convergence. That is, by periodically adjusting the learning rate—periodically decaying the learning rate—it can more effectively optimize model parameters, accelerate the training process, and improve the model's convergence performance.

[0098] In some embodiments of this application, the learning rate during the model training phase smoothly decreases from its initial value to its minimum value according to a cosine function. After each cycle, the learning rate jumps back to its initial value (or an adjusted initial value) and begins a new round of decrease. Embodiments of this application provide a dynamic cycle length, supporting either a fixed cycle length or a gradually increasing cycle length (controlled by the T_mult parameter). The update formula for the learning rate in embodiments of this application is:

[0099] The formula for updating the learning rate is:

[0100]

[0101] in

[0102] ·η t This is the current learning rate.

[0103] ·η min It is the minimum learning rate.

[0104] ηmax is the maximum learning rate (initial learning rate).

[0105] ·T cur It represents the number of iterations since the last reboot.

[0106] ·T i "This is the total number of iterations in the current period."

[0107] For example, in some embodiments of this application, before predicting the structure of the protein sequence to be analyzed using the target protein structure prediction model in step S120, the method further includes the following model training process:

[0108] The first step is to obtain the protein sequences of each sample from the target multi-domain combined dataset, wherein the target multi-domain combined dataset is obtained from the natural complex dataset, the target pseudo-multimer dataset, and the predicted pseudo-multimer dataset.

[0109] It should be noted that, in some embodiments of this application, the process of obtaining the target multi-domain combined data set includes, for example,:

[0110] 1. Extraction of interchain interactions

[0111] Interface residue identification: Using PDBePISA or PRODIGY tools, extract interchain contact residues ( Side chain non-hydrogen atom spacing ).

[0112] Hydrogen bond and hydrophobic interaction annotation: Hydrogen bonds and hydrophobic cores were analyzed using HBPLUS or LigPlot+.

[0113] 2. Feature information annotation

[0114] Evolutionary conservation score: Consurf score (1-9, 9 for high conservation) of interface residues was calculated by multiple sequence alignment (MSA).

[0115] Functional annotation: Indicate the biological function of the complex (such as "ATP binding", "signal transduction") and key residues (such as catalytic triplet).

[0116] Domain partitioning: Based on the SCOP or CATH database, the domain boundaries of each chain are marked.

[0117] 3. Template Library Construction

[0118] Data storage: Each complex template is stored as a separate entry, containing the following files:

[0119] Structure file: Standardized mmCIF format, containing atomic coordinates.

[0120] Comment file: JSON format, recording inter-chain interactions, conservatism score, and functional comments.

[0121] Metadata: PDB ID, resolution, experimental method, number of chains, etc.

[0122] The second step is to determine the hidden state of the sample protein sequence and determine the interaction information matrix between sample residues based on the hidden state.

[0123] The process of obtaining the inter-residue interaction information matrix can be referred to in the previous section on obtaining the inter-residue interaction information matrix. To avoid repetition, it will not be elaborated on here.

[0124] The third step involves obtaining a joint feature tensor for the samples based on the sample residue interaction information matrix and / or the sample evolution matrix. The sample evolution matrix is ​​obtained from the sample multiple sequence alignment results, which are obtained by feature pruning using a sliding window. The specific processing procedure for this step can be found in the previous section on obtaining the residue interaction information matrix; to avoid repetition, it will not be elaborated further here.

[0125] In some embodiments of this application, if a sample evolution matrix exists, a sample joint feature tensor is obtained based on the sample residue interaction information matrix and the sample evolution matrix; in some embodiments of this application, if an evolution matrix does not exist, the sample residue interaction information matrix is ​​used as the sample joint feature tensor.

[0126] It should be noted that the specific meanings of the parameters related to this step can be found in the descriptions above, and will not be elaborated upon here to avoid repetition.

[0127] The fourth step involves training the Evoformer module and the structure module for the i-th time based on the sample joint feature tensor, to obtain the i-th prediction result (which is the structure prediction result of the protein sequence of the sample to be predicted), where i is an integer greater than or equal to 1.

[0128] Fifth step: Calculate the loss value based on the i-th prediction result (i.e., determine the difference between the predicted structure and the actual structure), and update the parameters based on the loss value and the i-th target learning rate (the learning rate value is determined according to the periodically restarted cosine annealing strategy), wherein the i-th target learning rate is determined according to the periodically restarted cosine annealing strategy.

[0129] It should be noted that, in order to improve the accuracy of the model's prediction of complex structures, some embodiments of this application require the construction of a comprehensive and diverse training dataset, namely the target multi-domain combination dataset. This dataset includes three types of data: natural complexes, pseudo-complexes, and predicted complexes. Natural complexes (i.e., PDB multi-chain structures, as data in the natural complex dataset) are used to provide realistic interaction patterns. Training the model with natural complexes allows the trained model to learn real biophysical rules and binding interface features. Pseudo-complexes (pseudo-multimers, as pseudo-multimer data in the target pseudo-multimer dataset) expand the dataset coverage. In the absence of sufficient natural complexes, this generates more training samples by splitting multi-domain proteins, helping the model learn inter-domain interactions. Predicted complexes (predicted-pseudo-multimers, as data in the predicted pseudo-multimer dataset) supplement the experimental data with data generated by prediction tools, especially for newly discovered or low-homology protein complexes, enhancing the model's generalization ability.

[0130] Some embodiments of this application construct a multi-level, high-coverage complex training set by integrating three types of data: natural complexes, pseudo-complexes, and predicted complexes. This strategy not only compensates for the lack of experimental data but also expands the model's scope through computational means, enabling it to more accurately predict the structures of complex or emerging protein complexes, providing a powerful tool for drug design and functional research. It is easy to understand that the embodiments of this application improve the training effect of the protein structure prediction model based on multi-domain datasets and the cosine annealing algorithm, ultimately improving the accuracy of structure prediction based on the trained model.

[0131] The following examples illustrate the process of obtaining a target multi-domain combined data set provided by some embodiments of this application.

[0132] For example, in some embodiments of this application, before obtaining the individual sample protein sequences from the target multi-domain combined dataset, the method further includes the following process for obtaining the target pseudo-multimer dataset:

[0133] The first step is to obtain the original multi-domain dataset based on multiple data sources.

[0134] For example, in some embodiments of this application, this first step typically includes:

[0135] The data from the multiple databases (e.g., SCOP, CATH, ECOD) are merged and deduplicated to obtain the original multi-domain dataset. The data in the original multi-domain dataset is obtained by deduplicating the merged dataset according to deduplication principles, which include deduplication based on the number of domains and domain intervals. The data in the original multi-domain dataset is a protein complex containing multiple structural domains.

[0136] In the process of merging multiple data sets, the embodiments of this application remove duplicate data based on the number of fields and the field range, ensuring that each field appears only once in the database and avoiding data redundancy.

[0137] Understandably, the original multi-domain dataset provides the foundation for subsequent cleaning and generation.

[0138] The first step described above achieves domain database merging, a process designed to integrate data from different databases (SCOP, CATH, and ECOD) and avoid data duplication. The corresponding exemplary operation methods include: merging data from the three databases (SCOP, CATH, and ECOD); removing duplicate data based on the number of domains and domain ranges; and selecting data from different sources according to priority (SCOP > CATH > ECOD). The resulting original multi-domain dataset is a unified domain database that integrates all sources and eliminates duplicates.

[0139] The second step involves performing conflict resolution, resolution filtering, and domain overlap removal on the data in the original multi-domain dataset to obtain a cleaned multi-domain dataset.

[0140] In some embodiments of this application, conflict handling includes: when the same PDB entry is annotated by multiple databases, selecting entries with the same number of fields and ≥85% sequence overlap according to priority (e.g., SCOP>CATH>ECOD).

[0141] It should be noted that SCOP, CATH, and ECOD are three existing structural classification databases. The SCOP (Structural Classification of Proteins) database is a hierarchical classification driven by human experts, focusing on evolutionary and functional relationships; the CATH (Class, Architecture, Topology, Homology) database is a classification based on automated algorithms and human verification, emphasizing structural topology and homology; and the ECOD (Evolutionary Classification of Domains) database is a computationally driven evolutionary classification, revealing distant homology relationships.

[0142] In some embodiments of this application, resolution filtering includes: (As an example of setting a threshold) the structure ensures coordinate accuracy.

[0143] In some embodiments of this application, domain overlap removal includes deleting entries with overlapping sequences between domains (such as Domain1 residues 50-100, Domain2 residues 90-150).

[0144] The data in the cleaned multi-domain dataset is a high-quality multi-domain dataset after the above cleaning process.

[0145] The second step described above implements PDB filtering, which aims to select data suitable for subsequent analysis from the PDB database. A corresponding exemplary operation includes selecting data with a resolution less than or equal to... The PDB data obtained in this application ensures that there is no overlap between fields within the same chain. The high-resolution PDB data obtained in this embodiment provides a high-quality data foundation for subsequent analysis.

[0146] The third step involves splitting the cleaned multi-domain dataset into structural domains, combining pairs of single-domain structures, and verifying the complex to obtain an initial pseudo-multimer dataset. The data in the initial pseudo-multimer dataset is multi-domain data obtained by combining the single-domain structures.

[0147] For example, in some embodiments of this application, the third step is used to identify and segment the single-domain structure and obtain multi-domain data based on the single-domain structure. The corresponding process includes, for example, the following:

[0148] Separate individual domains from a clear multi-domain dataset based on the scope of the domain.

[0149] Identify discontinuous domain structures within chains and discontinuous structures across chains.

[0150] Randomly pair all single fields within the same PDB.

[0151] Based on the distance between calcium atoms being less than or equal to Heavy atom distance greater than or equal to and less than or equal to The standard is used to determine whether it is a complex.

[0152] The embodiments of this application, through this third step, can accurately classify PDB data into single-domain structures and synthesize multi-domain structures, preparing for subsequent analysis of different types of structures.

[0153] In some embodiments of this application, domain splitting includes:

[0154] Based on the database annotations, the atomic coordinates of each structural domain are extracted to generate independent PDB files.

[0155] Sequence splicing: For cross-chain or discontinuous structural domains, splice the sequence from the N end to the C end.

[0156] In some embodiments of this application, complex verification includes:

[0157] Main chain contact: (Spatial proximity).

[0158] Side chain contact: non-hydrogen heavy atom spacing (Reasonable interaction).

[0159] Multi-strand extension: Identify three- or four-strand complexes using NetworkX distance maps.

[0160] In other words, embodiments of this application split and verify the cleaned multi-domain data to generate pseudo-multimers.

[0161] The fourth step involves sequentially performing multimer complex generation, sequence clustering, and structure deduplication on the initial pseudo-multimer dataset to obtain the target pseudo-multimer dataset.

[0162] The fourth step, generating the multimeric complex, includes, for example:

[0163] Trimer and tetramer generation is a step aimed at generating possible trimer and tetramer structures. The generation process includes, for example, using NetWorkX tools to plot distance maps between dimer, trimer, and tetramer complexes, thereby generating these multimer structures and obtaining possible trimer and tetramer structure data to enrich the multi-domain structure database.

[0164] Dimeric domain complex generation is a step aimed at generating dimeric domain complexes. The corresponding implementation process, exemplarily, includes the same dimeric portion operations as in the trimer and tetramer generation steps, with the goal of generating dimeric domain complexes and obtaining dimeric domain complex data to provide data support for studying dimeric interactions.

[0165] The fourth step, sequence clustering, aims to cluster complexes according to their sequences, reducing data redundancy and improving data efficiency. An exemplary implementation process includes: splitting the complex into single strands; using the mmseqs2 tool to perform clustering at a sequence similarity threshold of 40%; and classifying the complexes into multiple categories based on the results of sequence clustering, ultimately obtaining complex data clustered by sequence similarity, making the data more representative and reducing redundant information.

[0166] The fourth step, structure deduplication, aims to remove structurally similar complexes, retain representative structures, and further improve data quality. A specific implementation example includes: removing complex structures with 100% sequence similarity; obtaining the cluster center structures of the complex clusters; comparing the remaining structures with the cluster centers; and retaining those with a RMSD (root mean square deviation) less than or equal to... The process of aligning the clusters requires chain alignment and sequence alignment to ensure the accuracy of structural comparison. For the elements of the complex clusters, the above method is used to perform deduplication again on the multi-domain complexes split in the same PDB to ensure the uniqueness of the data and finally obtain the deduplicated complex data, which retains the most representative structure and improves the overall quality of the database.

[0167] In some embodiments of this application, sequence clustering uses MMSeqs2 to cluster at 40% sequence similarity, merging evolution-related entries.

[0168] In some embodiments of this application, structural deduplication includes:

[0169] Sequence deduplication: Preserve cluster centers and delete entries that are 100% identical in sequence.

[0170] RMSD screening for quantitative structural differences: retention and cluster centers The structure (ensuring conformational diversity).

[0171] Performing this fourth step can remove the redundant pseudo-multiplex dataset. In other words, the embodiments of this application optimize the size and diversity of the dataset through clustering deduplication, preparing for subsequent integration.

[0172] In other words, the target pseudo-multiplex dataset involved in some embodiments of this application is obtained in the following way:

[0173] The first step is the collection of multi-domain data, which includes the following steps 2.1 and 2.2.

[0174] 2.1 Data Source

[0175] As more and more protein structures are resolved, structure-based protein classification libraries are gradually being improved. SCOP (Structural Classification of Proteins), CATH, and ECOD (Evolutionary Classification of Protein Domains) are three such libraries. They use structural data from PDB libraries as their research object, divide domains using different criteria, and use these domains as the basis for different levels of protein classification. Each of these three libraries has its own annotation files that label the sequence ranges corresponding to each domain in the PDB.

[0176] 2.2 Data Cleaning

[0177] 2.2.1 Database Priority Sort

[0178] Based on whether the domain annotations for a protein chain in the database are greater than one, some embodiments of this application can coarsely screen out multi-domain proteins. For a protein database (PDB), if its domain annotations appear in only one database, the annotation results of that database are directly retained. If a PDB is annotated by multiple databases, deduplication of the three databases is required, and the result of one of the databases is selected as the domain partitioning result. Based on whether the data has been manually annotated, the amount of data, and the frequency of database updates, the confidence priority of the three databases is SCOP > CATH > ECOD.

[0179] 2.2.2 Database Deduplication

[0180] When three databases have different annotations for the same PDB, some embodiments of this application first analyze the annotation results. If there are significant differences between the different databases, it is considered that the domain division of this structure is too controversial and is directly discarded. If the division between the different databases meets certain conditions, the information in the corresponding database is extracted according to the priority shown in 2.2.1. The criteria for determining whether to retain the division result for a certain PDB are as follows:

[0181] a. Multiple libraries are divided into the same number of structural domains.

[0182] b. Align the sequences of the domains such that the minimum overlap between sequences is greater than 85%.

[0183] 2.2.3 Data Refinement

[0184] Considering that subsequent determination of inter-domain interactions based on inter-atomic distance information is needed to filter for high-quality pseudo-multimers, some embodiments of this application employ near-atomic resolution to ensure the accuracy of atomic coordinates. The threshold was used to further filter the multi-domain data. At the same time, after inspection, it was found that there was sequence overlap between domains in some databases. Such cases were directly discarded.

[0185] The second step is the collection of data from the target pseudo-multimer dataset, which includes, for example, steps 3.1, 3.2, and 3.3 below.

[0186] 3.1 Generation of Sequence and Structure Data

[0187] a. Based on the annotation information in the multidomain library, extract the atomic coordinates corresponding to each domain of the PDB to generate a new PDB. The index of the PDB coordinate file may contain integers, integers plus characters, or be non-contiguous.

[0188] b. Extract the corresponding sequence information from the SEQRES of the original PDB based on the amino acid information of the new PDB. For domains that are discontinuous or cross-chain, splice the domain sequences directly from the N-terminus to the C-terminus according to the domain division results.

[0189] 3.2 Judgment of pseudo-multimer data

[0190] To ensure that the generated pseudo-multimer data are spatially close enough to resemble real complexes, some embodiments of this application require calculating the distance between pairwise domains to determine whether they form a complex. The criteria for forming a complex are as follows:

[0191] The distance between a.Cα is less than or equal to

[0192] b. The distance between non-hydrogen heavy atoms in the side chain is greater than or equal to... Less than or equal to

[0193] c. To further obtain triple-stranded and quadruple-stranded pseudo-multimers, the selected double-stranded pseudo-multimers were used to extract complex composition information by plotting distance maps using NetworkX.

[0194] 3.3 Data Clustering

[0195] The structure identified as pseudo-multimer in section 3.2 was split into single strands and clustered using MMSeqs2 based on a similarity of 40%.

[0196] 3.4 Data Deduplication

[0197] a. First, find the center of the sequence cluster and use it as the cluster center of the pseudo-multimer. Then, find other pseudo-multimers that have the exact same sequence clustering result as the pseudo-multimer, and remove pseudo-multimers with 100% sequence similarity.

[0198] b. Calculate the RMSD of other pseudo-multimer clusters with the cluster centers in the pseudo-multimer clustering results, and retain... The structure.

[0199] c. For pseudo-multimers split from the same PDB, further deduplication is performed according to the same criteria as a and b.

[0200] Some embodiments of this application provide a method for obtaining pseudo-multiplex datasets, thereby increasing the amount of data obtained from multi-domain datasets and ultimately improving the effectiveness of model training based on multi-domain datasets.

[0201] For example, in some embodiments of this application, before obtaining the individual sample protein sequences from the target multi-domain combined dataset, the method further includes the following process for obtaining a predicted pseudo-multimer dataset:

[0202] In some embodiments of this application, before obtaining the individual sample protein sequences from the target multidomain combined dataset, the method further includes: performing domain prediction on multidomain protein sequences that are not annotated by structural classification databases (e.g., these structural classification databases include: SCOP, CATH, ECOD, etc.) to obtain the predicted pseudomultimer dataset.

[0203] For example, some embodiments of this application generate complexes by predicting multi-domain protein structures or protein chain assembly patterns that have not been resolved experimentally using computational tools. This process exemplarily includes: multi-domain prediction: using AlphaFold to predict proteins containing multiple domains and simulating their interdomain interactions; chain assembly prediction: predicting the binding conformation of two independent protein chains using tools such as RosettaDock; construction steps: domain / chain prediction: generating potential structures using deep learning or molecular docking tools; contact screening: applying the same geometric criteria as the pseudo-complexes (e.g., ...). Verify the rationality of the interaction; Data integration: merge the prediction results with the experimental data to form a training set.

[0204] In other words, in some embodiments of this application, unannotated multidomain proteins from the SCOP / CATH / ECOD databases are used. These proteins can be analyzed using domain prediction methods, and data that meets the requirements can be selected according to standards to obtain a predicted pseudomultimer dataset.

[0205] The data in the predicted pseudomultimer dataset of this application embodiment can achieve the following technical effects: filling data gaps, that is, providing supplementary data for complexes lacking experimental structures (such as newly discovered proteins); enhancing model generalization ability: improving the model's robustness in predicting unknown complexes through training with diverse conformations; reliability verification: screening high-confidence prediction results through cross-validation (such as comparison with known complexes); dynamic update mechanism: continuously updating the prediction dataset as prediction tools (such as AlphaFold-Multimer) advance.

[0206] Some embodiments of this application provide a method for obtaining a predicted pseudo-multiplex dataset, thereby increasing the amount of data obtained from the multi-domain dataset and ultimately improving the effectiveness of model training based on the multi-domain dataset.

[0207] For example, in some embodiments of this application, before obtaining the protein sequences of each sample from the target multi-domain combined dataset, the method further includes the following process for obtaining a natural complex dataset: performing experimental method screening, resolution filtering, and sequence multi-dimensional quality control on the multi-chain complex structures in the PDB (Protein Data Bank) database to obtain the natural complex dataset, wherein the dimensions of the sequence multi-dimensional quality control include: sequence composition and stereochemical rationality.

[0208] In some embodiments of this application, the experimental method screening includes retaining only high-confidence data such as X-ray diffraction and cryo-electron microscopy.

[0209] In some embodiments of this application, resolution filtering includes: Ensure structural reliability.

[0210] In some embodiments of this application, sequence quality control includes:

[0211] The proportion of a single amino acid is less than 80% (to avoid interference from repetitive sequences).

[0212] For sequences longer than 20 residues, remove entries containing non-natural amino acids (such as X).

[0213] For example, in some embodiments of this application, the process of obtaining a dataset of natural complexes includes, exemplarily, the following:

[0214] For structures in the PDB database with more than one protein chain, those that have had their non-amino acid components removed from the sequence file and meet the following screening criteria can be directly identified as native complexes in the native complex dataset:

[0215] a.

[0216] b. The proportion of a single amino acid in the entire sequence is <80%.

[0217] c. Sequence length>20aa

[0218] d. 90% of the amino acids are non-natural amino acids or marked with an X and should be deleted directly.

[0219] e. Experimental methods: X-ray Diffraction, Electron Microscopy, Solution NMR, Solid-state NMR, Electron Crystallography.

[0220] Embodiments of this application integrate experimental validation data through a natural complex dataset to form a complete training set. Some embodiments of this application provide a method for obtaining a natural complex dataset, thereby increasing the amount of data obtained in the multi-domain dataset and ultimately improving the effectiveness of model training based on the multi-domain dataset.

[0221] It should be noted that, in some embodiments of this application, before obtaining the protein sequences of each sample from the target multi-domain combined dataset, the method further includes: merging the predicted pseudo-multimer dataset into the target pseudo-multimer dataset to obtain an extended pseudo-multimer dataset; and integrating the natural complex with the extended pseudo-multimer dataset to obtain the target multi-domain combined dataset. The embodiments of this application, based on the original protein structure model training data, introduce a large amount of high-quality multi-domain data (i.e., data from the target multi-domain combined dataset) to the model training, enriching the model's training dataset, making the model learning process more comprehensive and improving its scalability, thereby enhancing model performance.

[0222] In other words, in order to obtain the data in the target multi-domain combined data set, some embodiments of this application require merging data from different databases (such as CATH, PCB, etc.), and then performing filtering (such as PCB filter), error handling and network generation, deduplication, multidomain combination, sequence clustering, structure deduplication, and other steps.

[0223] In some embodiments of this application, the step of calculating the loss value based on the i-th prediction result includes: classifying the complexes in the natural complex dataset and the extended pseudo-multimer dataset according to function or structure; counting the number of samples in each class; calculating the weight based on the class frequency of the number of samples in each class; and multiplying the loss of each class of samples by the weight of the corresponding class to obtain the corresponding loss value.

[0224] Some embodiments of this application also optimize the loss function, thereby improving the accuracy of the trained model in predicting structures.

[0225] In summary, some embodiments of this application introduce complex templates into protein structure prediction for the first time, significantly improving the predictive ability of protein complex structures; embodiments of this application apply a periodically restarted cosine annealing learning rate training method to protein structure prediction models for the first time, improving training quality; embodiments of this application introduce multi-domain supplementary datasets to train protein structure models for the first time. Embodiments of this application apply the ESMPair strategy to the training and prediction of protein structure models for the first time, improving model prediction accuracy. Embodiments of this application optimize the MSA search process (i.e., first search in the complex template library, then search in the single-strand template library) to improve the accuracy of protein structure models.

[0226] Some embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the method for predicting protein structure as described in any of the above embodiments.

[0227] like Figure 5 As shown, some embodiments of this application provide an electronic device 400, which includes a memory 410 and a processor 420. The memory 410 is configured to store a computer program, and the processor 420 is configured to read the program from the memory 410 and execute it to implement the method for predicting protein structure as described in any of the above embodiments.

[0228] Processor 420 can process digital signals and may include various computing architectures. For example, it may be a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements multiple instruction set combinations. In some examples, processor 420 may be a microprocessor.

[0229] Memory 410 can be used to store instructions executed by processor 420 or data related to the execution of instructions. These instructions and / or data may include code used to implement some or all of the functions of one or more modules described in the embodiments of this application. The processor 420 of the embodiments of this disclosure can be used to execute the instructions in memory 410 to implement… Figures 1 to 4 The method shown. Memory 410 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memory well known to those skilled in the art.

[0230] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0231] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0232] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0233] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0234] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0235] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A method for predicting protein structure, characterized in that, The method includes: A target template is obtained by retrieving a template that matches the sequence of the protein to be analyzed from a combined template library, wherein the combined template library includes a single-stranded template library and a complex template library; The target template is used as input to the target protein structure prediction model. The target protein structure prediction model is used to predict the structure of the protein sequence to be analyzed. The target protein structure prediction model is obtained by training the protein structure prediction model based on a periodically restarted cosine annealing strategy. The periodically restarted cosine annealing strategy includes: starting a periodically restarted cosine annealing phase to reduce the learning rate to a minimum value according to the cosine function and then periodically resetting it. The duration of each cosine annealing phase is the same or the duration of the next cosine annealing phase is longer than the duration of the previous cosine annealing phase.

2. The method as described in claim 1, characterized in that, Before at least using the target template as input to the target protein structure prediction model, the method further includes: Perform model initialization settings, including: loading protein data, setting the upper limit of the learning rate, the lower limit of the learning rate, the initial period length, and the period growth factor; The training process for the i-th cycle, which is performed according to a cycle, includes: The learning rate is reset to the upper limit; the learning rate for the i-th cycle is determined based on the upper limit, the lower limit, the length of the i-th cycle, the number of steps in the i-th cycle, and the cosine function, and the model is trained based on the learning rate of the i-th cycle; when the number of steps in the i-th cycle is reached, the training of the i-th cycle ends; the length of the i-th cycle is dynamically adjusted based on the cycle growth factor, and after restarting, the training enters the (i+1)-th training cycle; If the training termination condition is met, training is stopped, and the model parameters are saved to obtain the target protein structure prediction model.

3. The method as described in claim 1, characterized in that, The method further includes: The homologous sequences of the protein sequence to be analyzed are searched from a homologous sequence database, and a target multiple sequence alignment result is generated based on the search results, wherein the target multiple sequence alignment result is used to characterize the evolutionary relationship between residues; in, The step of using at least the target template as input to a target protein structure prediction model, and predicting the structure of the protein sequence to be analyzed using the target protein structure prediction model, includes: The target multiple sequence alignment results and the target template are used as inputs to the target protein structure prediction model.

4. The method as described in claim 3, characterized in that, The step of searching for homologous sequences of the protein sequence to be analyzed from a homologous sequence database and generating target multiple sequence alignment results based on the search results includes: Sequences with evolutionary relevance to the protein sequence to be analyzed are identified from the homology sequence database to obtain a list of target homology sequences; Multiple sequence alignment is performed on the protein sequence to be analyzed and the sequences in the target homology sequence list to obtain initial multiple sequence alignment results; The initial multiple sequence alignment results are pruned using a sliding window, and the overlapping region between adjacent windows is retained to at least the target proportion, to obtain the MSA feature matrix. The MSA feature matrix is ​​then used as the target multiple sequence alignment result. The MSA feature matrix includes a residue frequency matrix and a co-evolutionary signal.

5. The method according to any one of claims 1-4, characterized in that, The step of retrieving a template from the combinatorial template library that matches the sequence of the protein to be analyzed to obtain the target template includes: A template matching the sequence of the protein to be analyzed is retrieved from the complex template library. If a matching template is found, it is used as the target template. If no matching template is found, the single-stranded template library is searched, and the matching template retrieved from the single-stranded template library is used as the target template.

6. The method as described in claim 2, characterized in that, The target protein structure prediction model includes a target Evoformer module and a target structure module. The target multiple sequence alignment results are represented using an evolutionary matrix. The step of predicting the structure of the protein sequence to be analyzed using the target protein structure prediction model includes: Extract the hidden state representation of each residue from the protein sequence to be analyzed; For residue pairs, mutual information between residues is determined based on the hidden state representation, wherein the mutual information between residues is used to characterize the covariation relationship of residue pairs during evolution; The residue interaction information matrix is ​​determined based on the residue mutual information, wherein the residue interaction information matrix is ​​used to characterize the residue mutual information and the co-evolution signal, the residue mutual information is used to characterize the covariation relationship of residue pairs during evolution, and the co-evolution signal is used to characterize the spatial proximity between hidden residues; The residue interaction information matrix is ​​concatenated or weighted and fused with the evolutionary matrix to form a joint feature tensor; The joint feature tensor is input into the target Evoformer module and the target structure module to determine the structure.

7. The method as described in claim 1 or 2, characterized in that, Before predicting the structure of the protein sequence to be analyzed using the target protein structure prediction model, the method further includes: Protein sequences of each sample are obtained from a target multi-domain combined dataset, wherein the target multi-domain combined dataset is obtained from a natural complex dataset, a target pseudo-multimer dataset, and a predicted pseudo-multimer dataset. Determine the hidden state of the sample protein sequence and determine the interaction information matrix between sample residues based on the hidden state; The sample joint feature tensor is obtained based on the sample residue interaction information matrix and / or sample evolution matrix, wherein the sample evolution matrix is ​​obtained based on the sample multiple sequence alignment results, and the sample multiple sequence alignment results are obtained based on feature pruning using a sliding window. The Evoformer module and the structure module are trained for the i-th time based on the sample joint feature tensor to obtain the i-th prediction result, where i is an integer greater than or equal to 1; The loss value is calculated based on the i-th prediction result, and the parameters are updated based on the loss value and the i-th target learning rate, wherein the i-th target learning rate is determined according to the periodically restarted cosine annealing strategy.

8. The method as described in claim 7, characterized in that, Prior to obtaining the protein sequences of each sample from the target multi-domain combined dataset, the method further includes: The original multi-domain dataset was obtained from multiple databases; The original multi-domain dataset is processed by conflict resolution, resolution filtering, and domain overlap removal to obtain a cleaned multi-domain dataset. The initial pseudo-multimer dataset is obtained by performing structural domain splitting, pairwise single-domain structural combination, and complex verification on the cleaned multi-domain dataset. The target pseudo-multimer dataset is obtained by sequentially performing multimer complex generation, sequence clustering, and structure deduplication on the initial pseudo-multimer dataset. The predicted pseudomultimer dataset is obtained by predicting the structural domains of multidomain protein sequences that are not annotated by the structural classification database. The natural complex dataset is obtained by screening experimental methods, filtering resolution, and performing multi-dimensional sequence quality control on the multi-chain complex structures in the PDB database. The dimensions of the multi-dimensional sequence quality control include sequence composition and stereochemical rationality. The predicted pseudo-multiplex dataset is incorporated into the target pseudo-multiplex dataset to obtain an extended pseudo-multiplex dataset; The natural complex is integrated with the extended pseudo-multimer dataset to obtain the target multi-domain combined dataset.

9. The method as described in claim 8, characterized in that, The process of obtaining the original multi-domain dataset from multiple databases includes: The data from the multiple databases are merged and deduplicated to obtain the original multi-domain dataset. The data in the original multi-domain dataset is obtained according to the deduplication principle, which includes deduplication based on the number of domains and domain intervals. The data in the original multi-domain dataset is a protein complex containing multiple structural domains. The process of performing conflict resolution, resolution filtering, and domain overlap removal on the data in the original multi-domain dataset to obtain a cleaned multi-domain dataset includes: When the same PDB entry is commented out by multiple databases, the selection is based on the priority set for those databases; Entries with the same number of reserved fields and whose sequence overlap exceeds the set percentage; Retain structures with atomic resolution less than a set threshold; Remove entries with overlapping sequences between domains.

10. The method as described in claim 8, characterized in that, The process of performing structural domain splitting, pairwise single-domain structural combination, and complex verification on the cleaned multi-domain dataset to obtain the initial pseudo-multimer dataset includes: The data in the cleaned multi-domain dataset is segmented according to the domain range in order to identify the single-domain structure in the cleaned multi-domain dataset and obtain a single-domain dataset. The initial pseudo-multiplex dataset is obtained by combining two single-domain structures from the same PDB in the single-domain dataset and confirming that the combined single-domain structures belong to a complex based on the distance between the two combined single-domain structures. The data in the initial pseudo-multiplex dataset is multi-domain data obtained by combining the single-domain structures.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for predicting protein structure as described in any one of claims 1-10.

12. An electronic device comprising a memory and a processor, the memory being configured to store a computer program, and the processor being configured to read the program from the memory and execute it to implement the method of predicting protein structure as claimed in any one of claims 1-10.

Citation Information

Patent Citations

  • Protein complex structure similar template searching method

    CN119694412A

  • Heterodimer interchain residue contact prediction method

    CN120072030A

  • Protein structure prediction

    CN120092293A

  • Machine-learning foundation model for generating biopolymer embeddings

    US20240404649A1