Protein sequence and structure co-design method and device, computer equipment and medium
By jointly generating protein sequences and structures in discrete space, this method utilizes a generative flow model and an asynchronous generation strategy to solve the problem of sequence and structure design separation in existing technologies. This improves the accuracy and efficiency of protein design, and the generated protein structures are closer to natural ones, making it suitable for drug development and industrial enzyme design.
Patent Information
- Application Number
- CN202511499266.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-04
- Filing Date
- 2025-10-20
- Publication Date
- 2026-01-13
AI Technical Summary
Existing protein design methods suffer from poor performance in practical applications due to the separation of sequence and structure design, inefficient continuous space modeling, and insufficient multimodal learning, especially in complex tasks that require simultaneous optimization of sequence and structure.
A generative flow model is used to jointly generate protein sequences and structures in discrete space. Through continuous-time Markov chain iterative generation, the target protein sequence and structure are gradually generated from the initial noise and restored to three-dimensional atomic coordinates. Combined with asynchronous generation strategy and time features, the coupling degree and consistency of sequence and structure are improved.
It significantly improves the accuracy and efficiency of protein design, generates protein structures that are closer to natural proteins, enhances the consistency between generated sequences and structures, reduces scRMSD values by about 8 times, and improves stability and design capabilities, making it suitable for drug development, industrial enzyme design, and basic biological research.
Smart Images

Figure CN121331218A_ABST
Abstract
Description
Technical Field
[0001] This invention mainly relates to the fields of artificial intelligence and computational biology, and in particular to a method, apparatus, computer equipment, and medium for co-designing protein sequences and structures. Background Technology
[0002] Protein design is a key area in bioengineering and drug development, aiming to achieve specific biological functions by designing novel protein sequences and structures. Traditional protein design methods are mainly divided into two categories: directed evolution and rational design. Directed evolution optimizes protein function step by step by simulating the natural selection process, but this method is time-consuming and costly. Rational design relies on energy functions and geometric constraints to optimize protein structure, but this method usually requires a large amount of computational resources and has limited effectiveness when designing complex proteins.
[0003] In recent years, with the development of artificial intelligence technology, generative models have been widely used in protein design. In particular, protein language models (such as ESM and ProtGPT2) and diffusion models (such as RFDiffusion and FrameDiff) have made significant progress in protein sequence and structure generation. These models can generate protein sequences or structures with specific functions, greatly improving the efficiency of protein design. However, existing methods still have some limitations: 1. Separation of Sequence and Structure Design: Most existing methods employ a two-stage design strategy, first designing the protein's backbone structure and then designing the corresponding sequence. This separate design approach fails to fully capture the complex relationship between sequence and structure, resulting in poor performance of the designed proteins in practical applications.
[0004] 2. Limitations of continuous space modeling: Generative methods such as diffusion models typically model protein structures in continuous three-dimensional space. Although they can generate high-quality structures, they are less efficient when dealing with discrete amino acid sequences and it is difficult to achieve joint optimization of sequence and structure.
[0005] These limitations restrict the application of existing methods in protein design, especially in complex tasks requiring simultaneous optimization of sequence and structure. Therefore, there is an urgent need for a design method that can effectively and jointly optimize protein sequence and structure to improve the accuracy and efficiency of protein design. Summary of the Invention
[0006] The purpose of this invention is to provide a method, apparatus, computer equipment, and medium for co-designing protein sequences and structures, aiming to solve problems such as the separation of sequence and structure design, low efficiency of continuous space modeling, and insufficient multimodal learning in existing protein design methods.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: On one hand, the present invention provides a method for co-designing protein sequences and structures, comprising: Input initial noise ,in and These represent discrete protein sequence tokens and structure tokens, respectively, and the protein sequence in the initial noise. or / and structure It contains a mask token, indicating that the protein sequence and / or structural information at the location corresponding to the mask token is hidden; Based on a generative flow model, a continuous-time Markov chain is used to iterate through T time steps to reduce initial noise. To the target protein sequence and structure The gradual generation, It does not contain a mask token, representing the target protein sequence and the protein sequences at all positions in the structure. and structure All information has been confirmed; Target protein sequence and structure The discrete structure is restored to three-dimensional atomic coordinates to generate the final protein structure.
[0008] Furthermore, a protein sequence and structure co-design device is provided, comprising: Input module, used to input initial noise ,in and These represent discrete protein sequence tokens and structure tokens, respectively, and the protein sequence in the initial noise. or / and structure It contains a mask token, indicating that the protein sequence and / or structural information at the location corresponding to the mask token is hidden; The generation module is used to generate data from initial noise using a generative flow model, through a continuous-time Markov chain iterating over T time steps. To the target protein sequence and structure The gradual generation, It does not contain a mask token, representing the target protein sequence and the protein sequences at all positions in the structure. and structure All information has been confirmed; The decoding module is used to connect the target protein sequence with its structure. The discrete structure is restored to three-dimensional atomic coordinates to generate the final protein structure.
[0009] On the other hand, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described protein sequence and structure co-design method.
[0010] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described protein sequence and structure co-design method.
[0011] On the other hand, the present invention provides a computer program product stored on a computer-readable storage medium and including computer instructions that, when executed by a processor, cause a computer device to implement the steps of the above-described protein sequence and structure co-design method.
[0012] Compared with the prior art, the technical effects of the present invention are as follows: By introducing a discrete generative flow model, this invention enables the joint generation of protein sequences and structures in a discrete space, significantly improving the accuracy, efficiency, and diversity of protein design.
[0013] This invention supports de novo design and conditional generation (such as specific fragment completion), and can be widely applied to drug development, industrial enzyme design, and basic biological research. By introducing an asynchronous generation strategy and temporal features, the consistency between the generated sequence and structure is significantly improved, with an average scRMSD value (a lower value indicates higher consistency) reaching 2.7 Å, approximately 8 times lower than the ESM3 method. Simultaneously, the generated protein structure is closer to the natural protein. Furthermore, it exhibits higher stability and design capability when processing long protein sequences (400 to 500 residues). Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram illustrating the principle of a protein sequence and structure co-design method provided in one embodiment; Figure 2 This is a schematic diagram of the architecture of a generative flow model provided in one embodiment; Figure 3 This paper presents a comparison of the performance of the protein sequence and structure co-design method provided by this invention and four baseline methods based on the ESM3 model in one embodiment. Figure 3(a) is a comparison chart of the scRMSD of the present invention and four baseline methods based on the ESM3 model; Figure 3 (b) pTM score distribution diagrams for the structures generated by the present invention and four baseline methods based on the ESM3 model; Figure 3 (c) is a diversity comparison chart of the present invention and four baseline methods based on the ESM3 model; Figure 4 This image shows a comparison between a protein structure generated under unconstrained conditions and a predicted structure using the protein sequence and structure co-design method provided by the present invention in one embodiment. Figure 5 This figure shows a comparison of the number of motif completion problems solved using the protein sequence and structure co-design method provided by the present invention and four baseline methods in one embodiment; Figure 6 The illustration shows a successful case of solving 20 motif completion problems using the protein sequence and structure co-design method provided by the present invention in one embodiment. Figure 7 This paper presents a comparison of the performance of the protein sequence and structure co-design method (CoFlow) provided by this invention and the ESM3 model in structure generation and sequence design tasks in one embodiment. Figure 7 (a) This paper presents a comparison of the mean squared errors of the protein sequence and structure co-design method (CoFlow) and the ESM3 model in the structure generation task, showing the results. Figure 7 (b) shows a comparison of the natural sequence recovery rate (NSR) of the protein sequence and structure co-design method (CoFlow) provided by this invention and the ESM3 model in sequence design tasks. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0017] Reference Figure 1 , Figure 1 This is a schematic diagram illustrating the principle of a protein sequence and structure co-design method provided in one embodiment. The present invention provides a protein sequence and structure co-design method, comprising: Input initial noise ,in and These represent discrete protein sequence tokens and structure tokens, respectively, and the protein sequence in the initial noise. or / and structure It contains a mask token, indicating that the protein sequence and / or structural information at the location corresponding to the mask token is hidden; Based on a generative flow model, a continuous-time Markov chain is used to iterate through T time steps to reduce initial noise. To the target protein sequence and structure The gradual generation, It does not contain a mask token, representing the target protein sequence and the protein sequences at all positions in the structure. and structure All information has been confirmed; Target protein sequence and structure The discrete structure is restored to three-dimensional atomic coordinates to generate the final protein structure.
[0018] The generative flow model predicts the true tokens corresponding to mask tokens in the protein sequences and structures within the initial noise. If the protein sequences and structures in the initial noise are all mask tokens, unconstrained generation is performed. Unconstrained generation automatically fills in the mask tokens in the input until all sequence and structure mask tokens are filled with non-mask tokens, but it does not change the number of tokens. Therefore, the only given condition for unconstrained generation is the length of the protein, i.e., the number of amino acids. Based on the protein sequence and structure co-design method proposed in this invention, the generated protein in the unconstrained generation task is a protein of a given length. In subsequent experimental verification for the unconstrained generation task, the effectiveness of this invention is evaluated based on consistency and diversity. For unconstrained generated proteins, which include both sequence and structure, consistency refers to whether the generated sequence can truly fold into the generated structure. The structure corresponding to the generated sequence is predicted using AlphaFold structure prediction software. If the error between the predicted structure and the generated structure is less than a threshold, the generated sequence and structure are considered consistent; otherwise, they are inconsistent. For example, if 100 proteins are generated unconstrainedly, and 80 of them have consistent sequences and structures, then the consistency ratio is 80%. Diversity refers to the structural similarity between generated and identical proteins. Higher similarity indicates lower diversity, and vice versa. For example, if 100 proteins are generated without constraints, and 80 of them have identical sequences and structures, then the structural similarity between each pair of these 80 proteins is calculated and averaged. If the average similarity is 0.6, then the diversity is 1 - 0.6 = 0.4.
[0019] If a portion of the protein sequence and / or structure in the initial noise is known (i.e., the initial noise contains a portion of unmasked tokens), then this portion of unmasked tokens in the initial noise is used as input. The generative flow model only predicts the masked tokens, while the given unmasked tokens (i.e., the given protein sequence tokens and structure tokens) remain unchanged. The initial noise can be a given partial sequence, and the proposed method is used to predict the structure and the remaining masked sequence based on the initial noise. Alternatively, the initial noise can be a given partial structure, and the proposed method is used to predict the sequence and the remaining masked structure based on the initial noise. Finally, the initial noise can be a given partial sequence and partial structure, and the proposed method is used to predict the remaining masked sequence and structure based on the initial noise.
[0020] Figure 1 The right side of the diagram represents the different tasks the model can perform. From top to bottom, it comprises four parts: The first part includes the protein's structure and sequence, representing the model's unconstrained generation of both the protein's skeletal structure and sequence. The sequence is represented by strings, with each letter representing an amino acid, and different letters representing different types of amino acids (e.g., M for methionine, K for lysine, W for tryptophan, E for glutamic acid, T for threonine, etc.). The second part only includes the protein's structure, representing the model generating the structure based on the sequence. The third part only includes the protein's sequence, representing the model generating the sequence based on the skeletal structure. Again, the sequence is represented by strings, with each letter representing an amino acid, and different letters representing different types of amino acids. The fourth part represents the model completing the protein from a given partial sequence and structure, where each letter represents an amino acid, and different letters represent different types of amino acids.
[0021] To standardize the processing of protein sequences and structures, this invention represents proteins as discrete protein sequences and structures. Specifically, a standard amino acid encoding method is used to represent the protein sequence as a discrete symbolic sequence, resulting in a protein sequence token. A vector quantization-variable autoencoder (VQ-VAE) is used to discretize the three-dimensional atomic coordinates of the protein into structure tokens, while simultaneously providing the ability to reconstruct the three-dimensional coordinates from the discrete tokens.
[0022] Reference Figure 2 , Figure 2 This is a schematic diagram of the architecture of a generative flow model provided in one embodiment. The present invention achieves the transformation from initial noise through a continuous-time Markov chain iterating over T time steps. To the target protein sequence and structure During the gradual generation process, the neural network at each time step is based on... Protein sequence and structural data at time points predict Protein sequence and structural data at time points , , Specifically, this includes: Will Protein sequence and structural data at time points protein sequences and structure Mapped to a vector representation; By using multi-frequency sine and cosine functions The time step corresponding to each moment is encoded as a high-dimensional vector to capture the periodic features of time information, as... Fourier time-series characteristics at any given moment; Get Vector representations of protein sequences and structures that integrate Fourier temporal features; Using neural networks based The vector representation of protein sequence and structure, which integrates Fourier temporal features, is used to predict the results. Protein sequence and structural data at time points ; The process continues iterating until the maximum number of iterations is reached, at which point the iteration stops, generating the final target protein sequence and structure. , The maximum number of iterations is determined based on actual circumstances or experience, and is generally set to 200, 400, or 500 steps.
[0023] One-hot encoding is used to map protein sequences and structures to vector representations. Proteins are composed of 20 amino acids, each with a unique symbol. In one-hot encoding, a 20-dimensional vector is assigned to each amino acid, with each vector corresponding to one amino acid. Only one position in each vector is 1, and the rest are 0. If a protein sequence has m amino acids, the entire protein sequence will be an m×20 matrix. Similarly, protein structures correspond to multiple structure types. In one-hot encoding, a unique integer index is assigned to each structure type. Then, a vector of all zeros with a length equal to the number of structure types is initialized. The corresponding position in the vector is set to 1 according to the structure type index, while the remaining positions remain 0. This encoding method converts discrete values into unique binary vectors, facilitating machine learning models in processing categorical data. After one-hot encoding, a linear transformation is used to convert the sparse one-hot vectors into dense vectors.
[0024] One embodiment proposes, Fourier time series characteristics of time As shown below:
[0025] in, It is a set of predefined frequencies used to generate sine and cosine components with different periods. , ,in This indicates the current iteration number. This encoding method effectively captures the continuity and periodicity of time steps, providing richer temporal information for the generative model.
[0026] Without loss of generality, frequency The expression is as follows:
[0027] in d The value is taken as Fourier temporal feature Embedding ( t Divide the vector dimension of ) by 2, and the Fourier temporal feature embedding ( t The vector dimension of the protein sequence and structure encoding is consistent with the vector dimension obtained from the protein sequence and structure encoding. , These are hyperparameters that are set without loss of generality. A value of 10 is generally acceptable. -2 , Possible value: 10 3 .
[0028] Fourier time series characteristics are expressed by multi-frequency sine and cosine functions representing time steps. Encoding as high-dimensional vectors effectively captures multi-scale temporal dependencies while providing a smooth and continuous temporal representation. This encoding method not only enhances the model's ability to model short-term and long-term dynamics but is also particularly suitable for handling periodic data. Furthermore, Fourier feature computation is efficient and requires no additional learnable parameters, significantly improving the efficiency of the diffusion model in utilizing temporal information, thereby improving generation quality and training stability.
[0029] The bidirectional Transformer is input by feature vectors of protein sequences and structures that incorporate Fourier temporal features. Leveraging its powerful self-attention mechanism, the bidirectional Transformer can capture global dependencies between positions within a protein sequence. For protein sequences, the bidirectional Transformer can effectively learn the interactions between amino acids and structural information.
[0030] This invention uses a bidirectional Transformer as the backbone network combined with Fourier temporal features to guide the model to make dynamic adjustments at each generation step, and outputs the class distribution of protein sequences and structures.
[0031] Figure 2 In the middle, "from Medium sampling This indicates the process of sampling from the conditional distribution. Generative flow models are based on the probability of the conditional distribution. ,current Protein sequence and structural data at time points (Including protein sequence and structural information) and model parameters Predicting the next moment Protein sequence and structural data at time points (Including protein sequence and structural information). This process can be formally represented as... ,in This indicates that the model parameters are The class distribution predicted by the neural network, This represents the one-hot vector of the mask token. That is, in At any given moment, if a position in the protein sequence or structure is a mask token, then at... At that moment, there will be The probability is sampled from the class distribution predicted by the model, and the sampled data is used as the data at the mask token position in the next time step. The probability remains unchanged with the mask Token, thus obtaining the next time step. Protein sequence and structure data at time points .
[0032] Based on the vector representation of protein sequence and structure incorporating Fourier temporal features, predictions were obtained. Protein sequence and structural data at time points During the process, an asynchronous generation strategy is adopted, prioritizing the generation of protein sequences in each generation step, and then using the generated protein sequences to guide the generation of structures. This strategy effectively improves the coupling between sequences and structures and the consistency of generation.
[0033] Asynchronous generation strategy refers to generating sequence data in two modalities—protein sequence and structure—in an alternating order. Specifically, the model first generates the first element of the protein sequence, then the first element of the structure, then the second element of the protein sequence, and so on. This strategy allows the model to dynamically and alternately consider the dependencies between the two modalities during the generation process, thereby better capturing the correlations between modalities.
[0034] The generative flow model in this invention achieves the stepwise generation of target sequences and structural markers from initial noise using a continuous-time Markov chain. The generation process of protein sequences and structures is dynamically coupled through a joint flow model. Specifically, during generation, protein sequences and structures are mapped to a common latent space via a neural network to interact, ensuring mutual influence and information sharing between them. An asynchronous generation strategy is employed, prioritizing the generation of protein sequences in each generation step, and then using the generated protein sequences to guide the generation of protein structures. This strategy effectively improves the coupling between sequences and structures and the consistency of generation.
[0035] This invention effectively addresses the problems of low sequence-structure matching, low design efficiency, and insufficient long-sequence protein generation capability in existing technologies by constructing a protein sequence structure co-design system based on a generative flow model. The co-flow mechanism couples the protein sequence and structure generation processes together and dynamically alternates between sequence and structural markers through an asynchronous generation strategy, ensuring a close correlation between the two and significantly improving the consistency and accuracy of the generated sequences. Experiments show that this invention outperforms existing methods by approximately 8 times in the consistency index (scRMSD) of unconditional generation tasks. The generated long-sequence protein structures are stable and highly natural, and their template modeling score (pTM) distribution is closer to that of natural proteins.
[0036] Furthermore, this invention offers significant advantages in supporting multifunctional design, enabling de novo design, conditional completion, and function-driven generation based on diverse task requirements. By optimizing the discretization representation and generation strategy of the VQ-VAE model, this invention demonstrates a high success rate in specific fragment completion tasks, while simultaneously improving generation efficiency and resource utilization. Compared to traditional methods, this invention significantly reduces computation time while maintaining generation quality, providing an efficient and reliable technical means for protein design applications in drug development, industrial enzyme design, and other fields.
[0037] This invention constructs an efficient and flexible protein design system based on a generative flow model. The system includes a data preprocessing module, a generation module, and a decoding module. It couples sequence and structural design together through joint flow technology, achieving high-quality generation of the target protein from initial conditions. Specifically, it includes: Input module, used to input initial noise ,in and Representing discrete protein sequence tokens and structural tokens, respectively, the initial noise protein sequence... or / and structure It contains a mask token, indicating that the protein sequence and / or structural information at the location corresponding to the mask token is hidden; The generation module is used to generate data from initial noise using a generative flow model, through a continuous-time Markov chain iterating over T time steps. To the target protein sequence and structure The gradual generation, The absence of a mask token indicates that the protein sequence and structural information at all positions in the target protein sequence and structure have been determined. The decoding module is used to connect the target protein sequence with its structure. The discrete structure is restored to three-dimensional atomic coordinates to generate the final protein structure.
[0038] The input module supports inputting functional requirements, sequence fragments, or structural fragments of the target protein, providing flexible conditional generation capabilities. The generation module uses a bidirectional Transformer-based backbone network, combined with Fourier temporal features, to guide the model to dynamically adjust at each generation step. The decoding module uses a structure decoder to restore discrete structural markers to three-dimensional coordinates, generating the final protein structure.
[0039] This generative model offers flexible conditional generation capabilities, enabling it to accomplish diverse tasks. Specifically, if the initial noise... If the initial noise only contains a mask token (i.e., the protein sequence and structure are unknown), the model will generate it unconditionally; if the initial noise is... protein sequence Known structure If the initial noise is unknown, the model makes structural predictions; conversely, if the initial noise is unknown, the model makes structural predictions. protein structure Known protein sequence If the initial noise is unknown, the model will design the protein sequence; finally, if the initial noise... protein structure and sequence Since all parts are known, the model will be designed to complete it.
[0040] The model training process is divided into two stages: pre-training and fine-tuning. The pre-training stage utilizes a large-scale protein database (such as MGnify30) to ensure the model has broad generative capabilities. The fine-tuning stage further trains the model on a high-quality PDB dataset to optimize the ability of the generated results to approximate natural structures. The inference stage employs multi-step iterative sampling, alternately generating sequences and structural markers to ensure high consistency and accuracy of the generated proteins.
[0041] To demonstrate the effectiveness of this invention, the protein sequence and structure co-design method (CoFLow) provided by this invention is applied in... Figure 1The invention was compared with baseline methods on four types of tasks described in the paper (unconstrained generation task, sequence-to-structure generation task, sequence-to-backbone structure generation task, and complete protein generation task based on given partial sequences and structures) to demonstrate its effectiveness. The invention significantly improves the consistency between the generated sequences and structures, with an average scRMSD value (a lower value indicates higher consistency) of 2.7 Å, approximately 8 times lower than the ESM3 method. Simultaneously, the generated protein structures are closer to the natural proteins. Neither the protein sequence and structure co-design method provided by this invention nor the baseline methods used for comparison are limited to the design of a specific type / class of protein sequences and structures; they can be used for the design of various types of protein sequences and structures.
[0042] The first type of experimental task is the unconstrained generation of protein sequence structures (corresponding to...). Figure 1 The first part of the task): The initial input sequence and structure are both mask tokens. The sequence and structure are represented by two arrays of equal length. The elements of the arrays are 32 and 4096, respectively, representing the sequence mask and the structure mask. The length of the array is the number of amino acids in the protein.
[0043] In the first type of task, the protein sequence and structure co-design method (CoFLow) provided by this invention was compared with four baseline methods based on the ESM3 model (Thomas Hayes et al., Simulating 500 million years of evolution with a language model. Science 387, 850-858 (2025). DOI: 10.1126 / science.ads0018). The baseline methods based on the ESM3 model used the same input pattern as the protein sequence and structure co-design method (CoFLow) provided by this invention. The four baseline methods based on the ESM3 model include: 1. ESM3: s→r: Use ESM3 to generate sequences and structures sequentially. 2. ESM3: r→s: Use ESM3 to generate structures and sequences sequentially.
[0044] 3. ESM3: ss→s→r: First, the ESM3 model is used to generate the protein secondary structure (alpha helix, beta fold, coil random coil), then the sequence is generated based on the secondary structure, and finally the three-dimensional structure is generated.
[0045] 4. ESM3: ss→r→s: Uses the ESM3 model to generate secondary structure, three-dimensional structure and sequence in sequence.
[0046] For the first type of task, the performance of this invention and the four baseline methods based on the ESM3 model mentioned above were evaluated through the following experiments: 500 protein samples of varying lengths were randomly generated for each method, and the generated proteins were evaluated from two dimensions: consistency and diversity. Consistency: For the generated sequence and structure, the protein structure prediction model AlphaFold2 is first used to predict the new structure based on the sequence. Then, the root mean square deviation (scRMSD) between the generated structure and the predicted structure is calculated. The smaller the scRMSD, the higher the consistency between the generated sequence and the structure.
[0047] Diversity: For generated proteins that meet the consistency condition (scRMSD < 2Å), the average template modeling score (TMScore) is calculated for each pair of proteins, and diversity is defined as 1 minus this average. Lower values indicate higher structural similarity of the generated proteins but lower diversity; conversely, higher values indicate higher diversity.
[0048] In the first category of tasks, the comparison results between the protein sequence and structure co-design method provided by this invention and four baseline methods based on the ESM3 model are as follows: Figure 3 As shown, where Figure 3 (a) is a comparison of scRMSD between the present invention and four baseline methods based on the ESM3 model. The lower the scRMSD, the higher the consistency. Figure 3(b) is a distribution of pTM scores of the generated structures of the present invention and four baseline methods based on the ESM3 model. The higher the pTM, the higher the quality of the generated structure. Figure 3 (c) This is a diversity comparison chart of the present invention and four baseline methods based on the ESM3 model; a higher score indicates higher diversity. (Refer to...) Figure 3 (a) All four baseline methods based on ESM3 exhibit high mean squared error (scRMSD), indicating poor consistency between the generated sequences and corresponding structures. Furthermore, referring to... Figure 3 (b) The predicted template modeling score (pTM) distributions of the four baseline methods based on the ESM3 model exhibit large variance, suggesting disorder issues in many generated structures. This reflects the significant limitations of the four ESM3 models in unconstrained structure generation. In contrast, the protein sequence and structure co-design method (CoFlow) provided in this invention demonstrates superior performance compared to the four baseline methods based on the ESM3 model, exhibiting higher consistency and pTM scores, indicating the generation of more rational protein structures. The average scRMSD of this invention (a lower scRMSD indicates higher consistency) can reach over 2.7 Å, significantly improving the consistency of generated sequences and structures, and reducing it by approximately 8 times compared to the ESM3 method. (See reference...) Figure 3(c) In terms of diversity metrics, the present invention has a slight advantage over the four baseline methods based on the ESM3 model.
[0049] Figure 4 This image shows a comparison between the protein structure generated and the predicted structure under unconstrained conditions using the protein sequence and structure co-design method provided by this invention. It can be seen that the protein generated by the method of this invention exhibits good consistency. (Refer to...) Figure 4 The generated structure is compared with the predicted structure, where red represents the generated structure and green represents the structure predicted using AlphaFold2 based on the generated sequence. Length represents the length of the protein; scTM and scRMSD represent the template modeling score and the mean squared error of the structure, respectively.
[0050] The second type of task: The input sequence and structure are partially known, and the rest are masked tokens, representing the task of completing the protein based on the given partial sequence and structure (corresponding to...). Figure 1 The fourth task in the paper). In this task, the test was conducted on 24 motif completion problems proposed in the relevant literature (Yim J, et al. Improved motif-scaffolding with SE (3) flowmatching[J]. Transactions on Machine Learning Research.). Each problem gives a partial fragment of a protein and requires designing a complete protein.
[0051] In the second type of task, the protein sequence and structure co-design method provided by this invention is compared with four other methods. These four methods include ESM3:s→r and ESM3:r→s, as described in the first task, and two methods based on the ESM3 model, namely: ESM3→ESMFold: Sequences were generated using ESM3 and structures were predicted using ESMFold (Zeming Lin et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023). DOI:10.1126 / science.ade2574); ESM3→ProteinMPNN: Generating structures using ESM3 and predicting sequences using ProteinMPNN (J. Dauparas et al., Robust deep learning-based protein sequence design using ProteinMPNN. Science 378, 49-56 (2022). DOI: 10.1126 / science.add2187); For each problem in the second type of task, 100 candidate proteins, including sequences and structures, are generated based on a given partial sequence and structure using five methods: ESM3:s→r, ESM3:r→s, ESM3→ESMFold, ESM3→ProteinMPNN, and the protein sequence and structure co-design method (CoFlow) provided in this invention. Then, the structure is predicted based on the generated sequences using AlphaFold2. The problem is considered successfully solved if the generated protein meets two criteria: (1) the template modeling score (TMscore) between the generated structure and the predicted structure exceeds 0.8; (2) the mean squared error (RMSD) between the predicted structure and the native structure on a given motif is less than 1 Å.
[0052] For the second type of task, Figure 5 This chart shows a comparison of the number of motif completion problems solved by the protein sequence and structure co-design method provided by this invention and four baseline methods. The method of this invention successfully solved 20 out of 24 problems, outperforming all baseline models. Figure 6 The figure further illustrates 20 successful cases of motif completion problems solved using the protein sequence and structure co-design method provided in this invention. Figure 6 The red portion represents the given motif fragment, and the green portion represents the content generated based on the motif. The first line of text indicates the problem name, the second line indicates the length of the generated protein, the third line indicates the template modeling score for the predicted motif structure and the generated structure, and the fourth line indicates the mean squared error of the predicted motif structure and the generated structure.
[0053] The third task is structure generation: the input sequence is given and known, meaning the tokens in the sequence array do not contain mask tokens, while the structure consists entirely of mask tokens, representing the model generating the structure based on the sequence (corresponding to this application). Figure 1 The second part of the task is to generate a structure based on the sequence.
[0054] The fourth task is sequence design: the input structure is given and known, meaning the tokens in the structure array do not contain mask tokens, while the sequence consists entirely of mask tokens, representing the model generating the sequence based on the structure (corresponding to this application). Figure 1 The third task in the process is to generate a sequence based on the structure.
[0055] For the third and fourth tasks, a test set containing 438 proteins, constructed using relevant literature (Campbell A, et al. Generative flows on discrete state-spaces: enabling multimodal flows with applications to protein co-design[C] / / Proceedings of the 41st International Conference on Machine Learning. 2024: 5453-5512.), was used for comparison with the ESM3 model. The evaluation metric for the third task was the mean squared error (RMSD) between the generated and actual structures; a lower metric indicates better structure generation. The evaluation metric for the fourth task was the natural sequence recovery rate (NSR); a higher metric indicates better sequence generation. (See reference...) Figure 7 The figure shows a comparison of the performance of the protein sequence and structure co-design method (CoFlow) provided by this invention and the ESM3 model in structure generation and sequence design tasks. Figure 7 (a) This paper presents a comparison of the mean squared errors of the protein sequence and structure co-design method (CoFlow) and the ESM3 model in the structure generation task, showing the results. Figure 7 (b) A comparison chart showing the natural sequence recovery rate (NSR) of the protein sequence and structure co-design method (CoFlow) provided by this invention and the ESM3 model in sequence design tasks; from Figure 7 (a) shows that the RMSD index of the method of this invention (CoFlow) is significantly lower than that of the ESM3 model, indicating that it can generate structures from sequences more accurately than the baseline method. From Figure 7 (b) It can be seen that the average NSR of the method of the present invention (CoFlow) reaches 0.56, while that of ESM3 is 0.5, indicating that the method of the present invention can generate more accurate sequences based on a given structure.
[0056] On the other hand, the present invention provides a computer device including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the protein sequence and structure co-design method provided in any of the above embodiments. The computer device may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device stores sample data. The network interface of the computer device is used for communication with external terminals via a network connection.
[0057] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the protein sequence and structure co-design method provided in any of the above embodiments.
[0058] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0059] Matters not covered in this invention are common knowledge.
[0060] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0061] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application.
[0062] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for co-designing protein sequences and structures, characterized in that, include: Input initial noise ,in and Representing discrete protein sequence tokens and structure tokens, respectively, the protein sequence in the initial noise... or / and structure It contains a mask token, indicating that the protein sequence and / or structural information at the location corresponding to the mask token is hidden; Based on a generative flow model, a continuous-time Markov chain is used to iterate through T time steps to reduce initial noise. To the target protein sequence and structure The gradual generation, It does not contain a mask token, representing the target protein sequence and the protein sequences at all positions in the structure. and structure All information has been confirmed; Target protein sequence and structure The discrete structure is restored to three-dimensional atomic coordinates to generate the final protein structure.
2. The protein sequence and structure co-design method according to claim 1, characterized in that, The initial noise is reduced by iterating through a continuous-time Markov chain over T time steps. To the target protein sequence and structure During the gradual generation process, the neural network at each time step is based on... Protein sequence and structural data at time points predict Protein sequence and structural data at time points .
3. The protein sequence and structure co-design method according to claim 2, characterized in that, according to Protein sequence and structural data at time points predict Protein sequence and structural data at time points ,include: Will Protein sequence and structural data at time points protein sequences in and structure Mapped to a vector representation; By using multi-frequency sine and cosine functions The time step corresponding to each moment is encoded as a high-dimensional vector to capture the periodic features of time information, as... Fourier time-series characteristics at any given moment; Get Vector representations of protein sequences and structures that integrate Fourier temporal features; Using neural networks based The vector representation of protein sequence and structure, which integrates Fourier temporal features, is used to predict the results. Protein sequence and structural data at time points ; The process continues iterating until the maximum number of iterations is reached, at which point the iteration stops, generating the final target protein sequence and structure. , .
4. The protein sequence and structure co-design method according to claim 3, characterized in that, Fourier time series characteristics of time As shown below: in, It is a set of predefined frequencies used to generate sine and cosine components with different periods. , ,in Indicates the current iteration number.
5. The protein sequence and structure co-design method according to claim 4, characterized in that, frequency The expression is as follows: in d The value is Fourier temporal feature Embedding ( t Divide the vector dimension of ) by 2, and the Fourier temporal feature embedding ( t The vector dimension of the protein sequence and structure encoding is consistent with the vector dimension obtained from the protein sequence and structure encoding. , These are the hyperparameters that are set.
6. The protein sequence and structure co-design method according to claim 5, characterized in that, The value is 10 -2 , The value is 10 3 .
7. The protein sequence and structure co-design method according to claim 3, 4, 5, or 6, characterized in that, Using neural networks based The vector representation of protein sequence and structure, which integrates Fourier temporal features, is used to predict the results. Protein sequence and structural data at time points This process is represented as ,in This indicates that the model parameters are The class distribution predicted by the neural network, The one-hot vector representing the mask token, in At any given moment, if a position in the protein sequence or structure is a mask token, then at... At that moment, there will be The probability is sampled from the class distribution predicted by the model, and the sampled data is used as the data at the mask token position in the next time step. The probability remains unchanged with the mask Token, thus obtaining the next time step. Protein sequence and structural data at time points .
8. A device for co-designing protein sequences and structures, characterized in that, include: Input module, used to input initial noise ,in and Representing discrete protein sequence tokens and structure tokens, respectively, the protein sequence in the initial noise... or / and structure It contains a mask token, indicating that the protein sequence and / or structural information at the location corresponding to the mask token is hidden; The generation module is used to generate data from initial noise using a generative flow model, through a continuous-time Markov chain iterating over T time steps. To the target protein sequence and structure The gradual generation, It does not contain a mask token, representing the target protein sequence and the protein sequences at all positions in the structure. and structure All information has been confirmed; The decoding module is used to connect the target protein sequence with its structure. The discrete structure is restored to three-dimensional atomic coordinates to generate the final protein structure.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes a computer program, it implements the steps of the protein sequence and structure co-design method as described in claim 1.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the protein sequence and structure co-design method as described in claim 1.