A method for multi-domain protein assembly guided by flexible residues
By combining a flexible residue prediction network with a multi-objective optimization model, various distance maps are generated, which solves the problem of inter-domain conformational uncertainty in multi-domain protein structure prediction, realizes high-precision multi-domain protein structure modeling, and supports in-depth research in structural biology and protein engineering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV OF TECH
- Filing Date
- 2026-04-21
- Publication Date
- 2026-05-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for predicting multi-domain protein structures suffer from problems such as relative conformational uncertainty between domains, multiple conformational states caused by flexible connecting regions, and missing template information, which limit high-precision modeling.
By using a flexible residue prediction network to decouple co-evolutionary information in multiple sequence alignment, various distance maps are generated and introduced into a multi-objective optimization model as an energy function constraint to guide the conformation sampling process, ultimately obtaining a high-precision multi-domain protein structure model.
It achieves high-precision prediction of multi-domain protein structures, which can accurately reflect their conformational space and function, providing a reliable modeling method for structural biology, protein engineering and drug design.
Smart Images

Figure CN122090939A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of bioinformatics and computational intelligence, and in particular relates to a method for multi-domain protein assembly based on flexible residue guidance. Background Technology
[0002] Over 70% of proteins in nature consist of multiple domains. These domains often represent relatively independent and stable folding units within a protein, each undertaking a specific biological function, such as substrate recognition, catalytic reactions, signal transduction, or molecular binding. Different domains do not exist in isolation but rather form complex, synergistic systems through flexible linkers or interfacial interactions, enabling proteins to coordinate spatial movement and achieve coupled functional regulation. This modular structural feature endows proteins with high functional diversity and adaptability, forming the basis for complex life activities in organisms. For example, multi-domain enzymes achieve efficient substrate transport and cascade catalysis through interdomain synergy, while receptor proteins rely on conformational transitions between different domains to complete signal recognition and transmission. However, precisely because of the highly dynamic nature of the flexible linkers connecting domains and the diverse relative orientations between domains during functional transitions, the structural prediction of multi-domain proteins has become a key technological bottleneck that urgently needs to be overcome in the field of computational structural biology.
[0003] In recent years, the rapid development of artificial intelligence technology has driven breakthroughs in the field of protein structure prediction. Advanced models such as AlphaFold2 (based on attention-based isovariant Transformer networks), AlphaFold3 (incorporating a diffusion model), RoseTTAFold (using multi-trajectory information fusion and graph neural networks), and ESMFold (based on protein language models) have enabled high-precision prediction of single-domain protein structures. These models can not only accurately predict the main chain and side chain conformations of proteins, but also effectively capture long-range interactions between residues through multiple sequence alignment (MSA) and structural template information, providing a reliable structural basis for understanding protein function. However, despite their excellent performance in single-domain proteins, these methods still face significant challenges in multi-domain protein structure prediction, such as the uncertainty of relative conformations between domains, the multiple conformational states caused by flexible connecting regions, and the lack of template information. These factors limit their direct application to complex multi-domain proteins. Therefore, with the maturity of single-domain structure prediction, the research frontier and core issues have shifted to high-precision modeling of multi-domain proteins. The key lies in how to integrate geometric deep learning methods to further analyze the conformation space based on accurate prediction of static inter-domain orientation, thereby achieving high-precision modeling that can truly reflect the structure and function of multi-domain proteins. Summary of the Invention
[0004] To overcome the limitations of existing models in predicting multi-domain protein structures, this invention proposes a flexible residue-guided multi-domain protein structure modeling method. By deeply mining the dynamic information hidden in the protein static database (PDB) and combining it with a large language model to analyze the multiple conformational features contained in the protein amino acid sequence, this method can achieve accurate modeling of protein structural diversity. This method can significantly improve the prediction accuracy of multi-domain proteins and provide strong support for revealing the functional mechanisms of proteins.
[0005] The technical solution adopted by this invention to solve its technical problem is: A method for multi-domain protein assembly based on flexible residue guidance first utilizes a flexible residue prediction network to decouple co-evolutionary information in multiple sequence alignment (MSA) to generate multiple distance maps that reflect conformational heterogeneity. Subsequently, the distance maps are introduced into a multi-objective optimization model as an energy function to constrain the conformational sampling process, guiding it to search for the optimal solution in a broader conformational space, and finally obtaining a high-precision multi-domain protein structure model.
[0006] The technical concept of this invention is as follows: First, key features of proteins at the atomic, residue, and overall topological structures are extracted and input into a diffusion model based on graph neural networks (GNNs) for learning. During the denoising process, the Modeller algorithm is introduced to physically correct the generated conformations to ensure their rationality and biological credibility. Subsequently, through iterative optimization, a protein conformation ensemble that conforms to physical constraints is continuously generated, thereby accurately capturing the dynamic behavior and functional conformational changes of proteins, providing a more reliable modeling method for structural biology, protein engineering, and drug design.
[0007] The beneficial effects of this invention are as follows: by using a flexible residue prediction network to decouple the co-evolutionary information in MSA, a distance map that can characterize multiple heterogeneous conformations of proteins is generated; multiple distance maps are introduced into a multi-objective optimization algorithm to achieve efficient and large-scale exploration of the protein conformation space; finally, by combining a high-precision model quality assessment module, candidate conformations are screened to obtain a more accurate and conformationally reasonable three-dimensional protein structure. Attached Figure Description
[0008] Figure 1 This is a flowchart of a multi-domain protein assembly method based on flexible residue guidance.
[0009] Figure 2 This is the three-dimensional structure of the protein 2rcsH predicted based on a flexible residue-guided multi-domain protein assembly method. Detailed Implementation
[0010] The present invention will now be further described with reference to the accompanying drawings.
[0011] Reference Figure 1 and Figure 2 A flexible residue-guided multi-domain protein assembly method is proposed. This method first utilizes a flexible residue prediction network to decouple the co-evolutionary information in multiple sequence alignment (MSA) in a targeted manner, thereby generating multiple distance maps that can reflect conformational heterogeneity. Subsequently, the distance maps are introduced into a multi-objective optimization model as an energy function to constrain the conformational sampling process, guiding it to search for the optimal solution in a broader conformational space, and finally obtaining a high-precision multi-domain protein structure model.
[0012] The method for assembling multi-domain proteins based on flexible residue guidance in this embodiment includes the following steps: 1) Given the amino acid sequence of a protein; 2) Obtain the structural information of the protein using the AlphaFold3 method based on the sequence; 3) Input the protein structure into the flexible residue prediction network MoReFold to decouple the co-evolutionary information in multiple sequence alignment (MSA) and obtain multiple distance maps containing heterogeneous conformational states; 4) Construct an energy function based on multiple distance maps : (1); in, The length of the protein's amino acid sequence. , They represent the first and second digits in the sequence, respectively. and the One residue, Indicating the first in the prediction structure The residue and the first residues Distance between Indicates the first in the reference structure The residue and the first residues Distance between; 5) Construct a multi-objective optimization model, using the energy function as a constraint, and obtain the optimal solution under multiple energy functions. (2); (3); in, , Representing the An energy function, The total number of energy functions. For decision-making space The decision vector in the conformation space is a protein structure in the conformation space. Two structures in protein conformation space , When in any energy function middle, The values are not inferior to ,and Superior to at least one energy function Then it is called To find an optimal solution, we can group all Pareto optimal solutions into a set to obtain the Pareto solution set. 6) Set parameters, population size is The population size is The number of iterations is Initial iteration count ; 7) Population Initialization: The overall structure of the target protein is predicted using AlphaFold3 and then broken down into multiple domains using DomainParser. The spatial orientation of each domain is then randomly perturbed and combined to generate multiple possible multi-domain conformations, thereby constructing a population containing... The initial population of individuals; 8) Parental selection: Individuals are evaluated based on an energy function constructed from a distance map containing heterogeneous conformations. A fitness-based selection strategy is adopted to select superior conformations into the mating pool to form the parental population. 9) Crossover operation: Based on the preset crossover probability To determine whether the parent individuals have undergone crossover operations to generate new recombinant conformation individuals; 10) Mutation operation: based on the mutation probability =0.01, randomly select parent individuals for local perturbation and conformational fragment transformation to enhance population diversity and global search capability; 11) Individuals that have undergone crossover and mutation are evaluated and screened using a multi-objective optimization model. Individuals in the Pareto optimal solution set are retained with a high probability (0.95) for the next generation of population update, while non-Pareto solutions are accepted with a low probability (0.05) to maintain population diversity and prevent premature convergence; 12) Loop: Repeat the above steps until... Break out of the loop; 13) For the multiple protein conformations in the final output, the DeepUMQA-X method, a protein complex model quality assessment software based on a combination of single-model methods and consensus strategies, is used to select the high-precision protein structures.
[0013] This embodiment uses the protein 2rcsH with a sequence length of 217 as an example, and successfully predicted the three-dimensional structure of this protein using the above method. The prediction accuracy surpasses AlphaFold3, with a corresponding template matching score (TM-score) of 0.95. The results are shown in the figure below. Figure 2 As shown.
[0014] The above describes the results of an example provided by the present invention. Obviously, the present invention is not only suitable for the above embodiments, but can also be implemented in various ways without departing from the basic spirit of the present invention and without exceeding the content involved in the substantive content of the present invention.
Claims
1. A method for assembling multi-domain proteins based on flexible residue guidance, characterized in that, First, a flexible residue prediction network is used to decouple the co-evolutionary information in multiple sequence alignment (MSA) in a targeted manner, thereby generating multiple distance maps that can reflect conformational heterogeneity. Then, the distance maps are introduced into a multi-objective optimization model as an energy function to constrain the conformational sampling process, guiding it to search for the optimal solution in a broader conformational space, and finally obtaining a high-precision multi-domain protein structure model.
2. The method for multi-domain protein assembly based on flexible residue guidance as described in claim 1, characterized in that, The multi-domain protein assembly method includes the following steps: 1) Given the amino acid sequence of a protein; 2) Obtain the structural information of the protein using the AlphaFold3 method based on the sequence; 3) Input the protein structure into the flexible residue prediction network MoReFold to decouple the co-evolutionary information in multiple sequence alignment (MSA) and obtain multiple distance maps containing heterogeneous conformational states.
3. The method for multi-domain protein assembly based on flexible residue guidance as described in claim 2, characterized in that, The method further includes the following steps: 4) Construct an energy function based on multiple distance maps : (1); in, The length of the protein's amino acid sequence. , They represent the first and second digits in the sequence, respectively. and the One residue, Indicating the first in the prediction structure The residue and the first residues Distance between Indicates the first in the reference structure The residue and the first residues The distance between them.
4. The method for multi-domain protein assembly based on flexible residue guidance as described in claim 3, characterized in that, The method further includes the following steps: 5) Construct a multi-objective optimization model, using the energy function as a constraint, and obtain the optimal solution under multiple energy functions. (2); (3); in, , Representing the An energy function, The total number of energy functions. For decision-making space The decision vector in the conformation space is a protein structure in the conformation space. Two structures in protein conformation space , When in any energy function middle, The values are not inferior to ,and Superior to at least one energy function Then it is called To find an optimal solution, we can group all the optimal Pareto solutions into a set to obtain the Pareto solution set.
5. The method for multi-domain protein assembly based on flexible residue guidance as described in claim 4, characterized in that, The method further includes the following steps: 6) Set parameters, population size is The population size is The number of iterations is Initial iteration count ; 7) Population Initialization: The overall structure of the target protein is predicted using AlphaFold3, and then split into multiple domains using DomainParser. The spatial orientation of each domain is then randomly perturbed and combined to generate multiple possible multi-domain conformations, thereby constructing a population containing... The initial population of individuals.
6. The method for multi-domain protein assembly based on flexible residue guidance as described in claim 5, characterized in that, The method further includes the following steps: 8) Parental selection: Individuals are evaluated based on an energy function constructed from a distance map containing heterogeneous conformations. A fitness-based selection strategy is adopted to select superior conformations into the mating pool to form the parental population. 9) Crossover operation: Based on the preset crossover probability To determine whether the parent individuals have undergone crossover operations to generate new recombinant conformation individuals; 10) Mutation operation: based on the mutation probability Randomly select parent individuals for local perturbation and conformational segment transformation to enhance population diversity and global search capabilities.
7. The method for multi-domain protein assembly based on flexible residue guidance as described in claim 6, characterized in that, The method further includes the following steps: 11) Multi-objective screening: For individuals that have undergone crossover and mutation, a multi-objective optimization model is used for evaluation and screening. Individuals in the Pareto optimal solution set are retained with a higher probability for the next generation of population update, while non-Pareto solutions are accepted with a lower probability, in order to maintain population diversity and prevent premature convergence. 12) Loop: Repeat the above steps until... Break out of the loop; 13) For the multiple protein conformations in the final output, the DeepUMQA-X method, a protein complex model quality assessment software based on a combination of single-model methods and consensus strategies, is used to select the high-precision protein structures.