Oligomeric protein de novo design method, system and equipment and storage medium
By extracting functional characteristics and design rules from natural protein databases, generating a candidate mutant sequence library, and using protein structure prediction models for screening and iterative optimization, the problem of low efficiency of multi-subunit interaction interfaces in oligomeric protein design was solved, achieving efficient and accurate oligomeric protein design.
Patent Information
- Application Number
- CN202511474901.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies have low efficiency in predicting and designing multi-subunit interaction interfaces when designing oligomeric proteins, making it difficult to quickly obtain oligomeric proteins that meet specific needs, resulting in long design cycles, high costs, and low success rates.
By extracting functional characteristics from natural protein databases, determining design rules, generating a candidate mutant sequence library, and using protein structure prediction models for screening and iterative optimization, combined with deep learning models to optimize sequences, the designed oligomeric protein sequences are ensured to meet the preset requirements.
It improves the efficiency of oligomeric protein design, shortens the design cycle, reduces the cost of experimental validation, and enhances the accuracy and efficiency of the design.
Smart Images

Figure 40EAE0A4-438A-47DB-867E-48709E31D3DC 
Figure 90B76F49-D811-4455-9971-DC40D8D4DDA3 
Figure FA875D8E-9D1A-4547-AC68-F51108F3760A
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of protein design, in particular to an oligomeric protein de novo design method, system, terminal device and computer readable storage medium. BACKGROUND
[0002] Proteins, as the most important functional molecules in life systems, play a key role in metabolic regulation, material transport, energy conversion and signal transmission. Among them, oligomeric proteins are assembled by multiple subunits, have higher stability and more complex functions, and occupy an important position in biological processes such as catalytic reaction, electron transfer and molecular recognition.
[0003] Protein design methods mainly rely on random mutation and high-throughput screening, or rational design based on structural templates. These methods have problems such as long design cycle, high experimental cost and low hit rate, making it difficult to quickly obtain oligomeric proteins that meet specific needs. In recent years, with the development of deep learning and artificial intelligence technology, related technical personnel have been able to predict the three-dimensional structure of proteins using tools such as AlphaFold2, and achieve sequence design of monomeric proteins through generation models such as RFdiffusion.
[0004] However, the prediction and design of multi-subunit interaction interfaces are still difficult, therefore, there is an urgent need for a method combining artificial intelligence and structure prediction to efficiently and accurately design and optimize oligomeric proteins, thereby expanding the application potential of proteins in the fields of energy, materials and medicine. SUMMARY
[0005] The embodiments of the present application provide an oligomeric protein de novo design method, system, terminal device and computer readable storage medium, which solve the problem of low efficiency in predicting and designing multi-subunit interaction interfaces when designing oligomeric proteins, and improve the design efficiency of oligomeric proteins.
[0006] The embodiments of the present application provide an oligomeric protein de novo design method, which comprises: obtaining the functional characteristics of a target protein and determining a first design rule based on the functional characteristics; extracting a first protein sequence and a second protein sequence from a natural protein database based on the first design rule, and determining a second design rule; generating a candidate mutant sequence library according to the first design rule and the second design rule; performing structure prediction and stability calculation on the sequences in the candidate mutant sequence library, and screening out optimized sequences that meet the preset requirements; Simulate the function characteristics of the optimized sequence, and output the final target protein sequence if the simulation is passed; otherwise, feed back the verification result to the protein sequence optimization model for iterative optimization design.
[0007] Optionally, the step of extracting the first protein sequence and the second protein sequence from the natural protein database based on the first design rule and determining the second design rule comprises: Performing multiple sequence alignment on the first protein sequence and the second protein sequence to identify super secondary structure elements in the sequence; Based on the super secondary structure elements, determine the amino acid residue sites directly related to the function characteristics, and construct the function related region; According to the conserved region in the first protein sequence and the second protein sequence and the function related region, determine the second design rule.
[0008] Optionally, the step of generating a candidate mutation sequence library according to the first design rule and the second design rule comprises: According to the first design rule, determine the main chain template sequence corresponding to the function characteristics; According to the second design rule, generate sequence constraint conditions for the main chain template sequence; wherein the sequence constraint conditions include fixed constraints and variable constraints, the fixed constraints specify that the amino acid residue sites corresponding to the conserved region and their residue types cannot be changed; the variable constraint is that for each mutation target point in the function related region, at least one candidate residue type set that allows substitution is specified; Based on the fixed constraint and the variable constraint, generate an amino acid sequence that meets the constraint condition by a sequence space sampling algorithm, forming the candidate mutation sequence library.
[0009] Optionally, the step of performing structure prediction and stability calculation on the sequences in the candidate mutation sequence library to screen out the optimized sequence meeting the preset requirements comprises: Input the candidate mutation sequence into the trained protein structure prediction model, and output the predicted three-dimensional structure; Calculate the fit degree of the predicted three-dimensional structure and the preset structure characteristics determined based on the design requirements, and calculate the folding free energy change of the predicted three-dimensional structure; wherein the preset structure characteristics include at least one of target folding type, relative arrangement of secondary structure elements, geometric shape of interaction interface, and spatial conformation of functional site; Determine that the sequence with the fit degree lower than the first threshold value and the folding free energy change lower than the second threshold value meets the preset requirements.
[0010] Optionally, after the step of generating a candidate mutant sequence library according to the first design rule and the second design rule, the following steps are included: The candidate mutant sequence library is input into a protein sequence generation model, which optimizes the sequences based on the functional characteristics and outputs an optimized candidate mutant sequence library.
[0011] Optionally, the step of feeding back the validation results to the protein sequence optimization model for iterative optimization design includes: If the simulation verification fails, analyze the key structural areas that cause the functional defects; Based on the key structural regions, the set of mutation targets or residue types in the protein sequence optimization model is adjusted to generate a new round of candidate mutation sequence libraries.
[0012] Optionally, before the step of obtaining the functional properties of the target protein and determining the first design rule based on the functional properties, the method includes: Design requirements for receiving the target protein; The design requirements are analyzed to determine the functional characteristics of the target protein.
[0013] Furthermore, to achieve the above objectives, embodiments of the present invention also provide a de novo design system for oligomeric proteins, the system comprising: The rule determination module is used to acquire the functional characteristics of the target protein and determine a first design rule based on the functional characteristics, and to extract protein sequences from a natural protein database based on the first design rule and determine a second design rule. A sequence generation module is used to generate a candidate mutation sequence library according to the first design rule and the second design rule; The screening module is used to perform structure prediction and stability calculation on the sequences in the candidate mutation sequence library, and screen out the optimized sequences that meet the preset requirements. The verification and iteration module is used to simulate and verify the functional characteristics of the optimized sequence. If the verification is successful, the final target protein sequence is output; otherwise, the verification result is fed back to the sequence generation module for iterative optimization design.
[0014] In addition, to achieve the above objectives, embodiments of the present invention also provide a terminal device, including a memory, a processor, and an oligomeric protein de novo design program stored in the memory and executable on the processor. When the processor executes the oligomeric protein de novo design program, it implements the method described above.
[0015] In addition, to achieve the above objectives, embodiments of the present invention also provide a computer-readable storage medium storing an oligoprotein de novo design program, which, when executed by a processor, implements the method described above.
[0016] One or more technical solutions provided in the embodiments of this application have at least the following technical effects: By extracting and refining protein sequence evolution rules from natural protein databases, a reliable physicochemical and structural biological basis is provided for design, ensuring the rationality of the initial design direction and significantly narrowing the sequence search space. Then, a protein structure prediction model is used to perform rapid and large-scale virtual screening of candidate sequences, eliminating structurally unstable or conformationally inconsistent schemes in advance, reducing the cost and time required for later experimental verification. Finally, an iterative feedback mechanism is introduced to progressively design oligomeric protein sequences that meet the design requirements, thereby improving the efficiency of oligomeric protein design. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating Example 1 of the de novo design method for oligomeric proteins in this application; Figure 2 This is a flowchart illustrating Example 2 of the de novo design method for oligomeric proteins in this application; Figure 3 This is a schematic diagram of the terminal structure of the hardware operating environment involved in one embodiment of this application. Detailed Implementation
[0018] To address the low efficiency issue caused by the instability of multi-subunit interaction interfaces in de novo oligoprotein design, this application proposes a de novo oligoprotein design method. First, structures and design rules satisfying the functional characteristics of the target protein are extracted from natural protein databases. Then, a candidate sequence library is constructed using constrained sequence generation technology. The candidate sequences are screened for stability and functionality using protein structure prediction and physical computation models. Finally, through a feedback loop, the simulation validation results are continuously iteratively optimized to obtain the target protein sequence that meets the design requirements, thus improving the efficiency of de novo oligoprotein design.
[0019] To better understand the above technical solutions, exemplary embodiments of this application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0020] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0021] Example 1 In this embodiment, a de novo design method for oligomeric proteins is provided.
[0022] Reference Figure 1 The de novo design method for oligomeric proteins in this embodiment includes the following steps: Step S100: Obtain the functional characteristics of the target protein and determine the first design rule based on the functional characteristics; In this embodiment, functional properties refer to the biochemical, physical, or structural attributes of a protein that can be measured and designed, derived from design requirements. Functional properties include at least one of the following: binding properties, catalytic properties, physicochemical properties, and environmental response properties. Binding properties include binding affinity, specificity, and target molecule class. Catalytic properties include enzyme activity, substrate specificity, and transformation number; physicochemical properties include conductivity, mechanical strength, fluorescence, and self-assembly ability; environmental response properties include sensitivity to pH, temperature, light, or specific molecules, and conformational changes. The first design rule refers to an oligomeric protein template determined based on the functional properties.
[0023] As an alternative implementation, before obtaining the functional characteristics of the target protein, the design requirements of the target protein are received first, the design requirements are analyzed, and then the functional characteristics of the target protein are determined.
[0024] For example, suppose the received design requirement is "design a highly conductive bio-nanowire". Analyzing this requirement reveals that the functional characteristics of the target protein are the ability to carry charge transport by having continuous π-π stacks or ion channels to achieve electron / hole transport, the ability to spontaneously form extended and ordered supramolecular structures through self-assembly, and the ability to maintain structural stability under specific conditions.
[0025] As another optional implementation, when determining the first design rule based on functional characteristics, at least one protein structure classification database is queried, and then a list of protein domains with relevant molecular functions or structural features is retrieved based on functional characteristics. Protein folding types that are compatible with the preset oligomeric state of the target protein are selected from the list, and at least one selected protein domain is determined as the first design rule.
[0026] For example, when the functional characteristic of a target protein is defined as "forming a cation channel on the cell membrane that is sensitive to specific signaling molecules," a preset protein structure classification database is first queried. Based on the core function of "cation channel," keywords including "ion channel," "cation selectivity," and "transmembrane pore" are searched to obtain an initial list containing domains such as voltage-gated potassium channels, ligand-gated ion channels, and mechanosensitive ion channels. Subsequently, preset screening rules are applied to select protein folding types whose natural oligomeric state is "homotetramer" from the initial list. Finally, since the monomeric domain of a voltage-gated potassium channel with a "chandelier-like pore domain" fold can construct a precise ion-selective filter and pore when forming a tetramer, the monomeric domain of a voltage-gated potassium channel with a "chandelier-like pore domain" fold is used as the first design rule.
[0027] It should be noted that protein function is largely determined by its three-dimensional structure, and oligomerization is a fundamental and crucial property of this structure; specific functions require specific oligomerization states to be realized. Upon receiving design requirements, it is possible to reasonably determine the target oligomerization state based on known biological knowledge.
[0028] Step S200: Based on the first design rule, extract the first protein sequence and the second protein sequence from the natural protein database, and determine the second design rule; In this embodiment, the first protein sequence refers to the amino acid sequence of a natural protein or its homologous protein that possesses the aforementioned functional characteristics. This sequence serves as a positive template for design, and the structure of its core functional region is to be inherited and imitated in the design. The second protein sequence refers to the amino acid sequence of a natural protein that is homologous to the first protein sequence but lacks or significantly weakens the aforementioned functional characteristics. This sequence serves as a negative control template for design; by comparing it with the first protein sequence, the key amino acid residues and structural features necessary to achieve the target function can be highlighted.
[0029] As an optional implementation, multiple sequence alignment is performed on the first and second protein sequences to identify supersecondary structural elements. Based on these supersecondary structural elements, amino acid residue sites directly related to the functional characteristics are determined, forming a function-related region. Then, a second design rule is determined based on the conserved regions in the first and second protein sequences and the function-related region.
[0030] For example, assuming the goal is to design highly conductive protein nanowires, the PilA protein sequences of various conductive bacteria such as Geobacterium and Shewanella are selected as the first protein sequence, and homologous sequences of PilA from non-conductive bacteria such as Pseudomonas aeruginosa are selected as the second protein sequence. Through multiple sequence alignment of these two sets of sequences, a supersecondary structural element spanning all PilA sequences was successfully identified: an extended α-helix composed of an N-terminal region. Based on this supersecondary structural element, amino acid residue sites directly related to conductivity were further identified. Comparison revealed that aromatic amino acids are regularly present at specific axial positions of this α-helix in the PilA of all conductive bacteria, while the corresponding sites in non-conductive bacteria are occupied by aliphatic or polar residues. Therefore, these sites were determined to constitute the functionally relevant regions responsible for π-π stacking for charge transport. Finally, a second design rule was determined based on the conserved regions in the sequence (i.e., the core hydrophobic residues maintaining the α-helix structure and protein oligomerization) and the identified functionally relevant regions (i.e., the axially aligned aromatic residue sites). The second design rule is to keep the conserved region residues unchanged in subsequent mutation design, and focus on these specific axial sites in the functionally relevant regions, replacing or optimizing them with aromatic amino acids with stronger π orbital overlap capabilities to directionally enhance the protein's conductivity.
[0031] Step S300: Generate a candidate mutation sequence library according to the first design rule and the second design rule; In this embodiment, multiple candidate mutant sequences can be generated according to the first design rule and the second design rule, so as to conduct subsequent functional and stability tests and iteratively select the optimal target protein. The collection of multiple candidate mutant sequences generates a candidate mutant sequence library.
[0032] As an optional implementation, the main strand template sequence corresponding to the functional characteristics is first determined according to the first design rule, and then sequence constraints on the main strand template sequence are generated according to the second design rule. The sequence constraints include fixed constraints and variable constraints. Under the joint constraints of fixed constraints and variable constraints, an amino acid sequence that meets the constraints is generated through a sequence space sampling algorithm to form the candidate mutant sequence library.
[0033] For example, based on the first design rule, the α-helix structure of the Geobacterium PilA protein is determined as the main strand template sequence, i.e., the first design rule. Then, according to the second design rule, sequence constraints on this main strand template are generated. The fixed constraint specifies that the core residue sites maintaining the rigidity of the α-helix and the oligomerization interface remain unchanged, while the variable constraint defines the aromatic residue sites located along the helix axis, which are confirmed by alignment analysis to be key to conductivity. The set of allowed mutation types for these sites is limited to tryptophan, tyrosine, and phenylalanine to optimize the π-π stacking effect. Under the joint constraints of this fixed and variable constraint, the Monte Carlo sequence spatial sampling algorithm is used to generate all amino acid sequence variants that meet the above constraints, thereby constructing a structurally stable candidate mutant sequence library focused on enhancing charge transport capabilities, laying the design foundation for subsequent high-throughput screening.
[0034] Step S400: Perform structure prediction and stability calculation on the sequences in the candidate mutant sequence library, and select optimized sequences that meet the preset requirements; In this embodiment, the protein sequences in the candidate mutant sequence library are subjected to structure prediction and stability analysis, and protein sequences whose structure prediction results do not meet the design requirements and those that are unstable are eliminated.
[0035] As an optional implementation, candidate mutant sequences are input into a trained protein structure prediction model, which outputs the predicted three-dimensional structure. The fit between the predicted three-dimensional structure and the preset structural features determined based on design requirements is calculated, and the fold free energy change of the predicted three-dimensional structure is calculated. Sequences with a fit below a first threshold and a fold free energy change below a second threshold are determined to meet the preset requirements, and these candidate mutant sequences are selected as optimized sequences. The preset structural features include at least one of the following: target fold type, relative arrangement of secondary structure elements, geometry of interaction interfaces, and spatial conformation of functional sites.
[0036] For example, RFdiffusion's sequence redesign capabilities can be used to perform local structural optimization on each candidate sequence, particularly targeting its π-π stacking interface, to eliminate potential side-chain conflicts and optimize charge transport paths, generating a more energy-efficient 3D structural model. This optimized structure is then used as input for high-precision 3D structure prediction via AlphaFold2, outputting the final 3D conformation. The degree of fit between the predicted structure and the preset structural features determined based on the design requirement of "forming a linear, extended α-helical fiber with continuous aromatic ring stacks" is calculated. Specifically, the overall α-helix content, the centroid distance of its axial aromatic residues, and the face-to-face parallelism are calculated to quantify its matching degree with the geometry of the "ideal conductive interface." Finally, the folding free energy change of this conformation is calculated using a physical force field. Sequences with a structural fit score below the first threshold (i.e., good structural feature matching; the fit score uses negative logic, with lower values indicating higher matching) and a folding free energy change below the second threshold (i.e., the structure itself is highly stable) are identified as stable conformations that meet the preset requirements. These candidate sequences are then selected as optimized sequences for the next round of functional verification.
[0037] As another alternative implementation, after performing structural prediction and stability analysis on candidate mutation sequences, a deep learning model can be used to optimize the sequences and generate new optimized sequences.
[0038] For example, after screening preliminary optimized sequences through structure prediction and stability analysis, the optimized sequences and their corresponding three-dimensional structures, predicted to have good stability, are used as conditional inputs to a deep learning model. This model is configured to prioritize retaining fixed constraint regions defined by the second design rule, while simultaneously proposing new residue substitution schemes for function-related regions limited by variable constraints, based on the learned sequence-structure-function mapping relationship. This generates novel sequence variants that are more rational in chemical-physical space. The new candidate mutant sequences generated by the deep learning model then proceed to subsequent functional verification.
[0039] Step S500: Perform functional characteristic simulation verification on the optimized sequence. If the verification is successful, the final target protein sequence is output; otherwise, the verification result is fed back to the protein sequence optimization model for iterative optimization design.
[0040] In this embodiment, the optimized sequence undergoes functional validation to test whether the candidate mutant sequence meets the design requirements. If the validation fails, the validation result is fed back to the protein sequence optimization model to continue mutation-structure prediction and stability analysis of the candidate mutant sequence. After multiple rounds of iterative optimization, the target protein that meets the design requirements is finally output.
[0041] As an optional implementation, appropriate functional verification tools can be selected based on the functional characteristics of the protein. For example, for protein stability verification, molecular dynamics simulation software can be used to simulate conformational changes of the protein under different conditions; for protein-ligand binding ability, molecular docking software can be used to predict the binding mode and affinity between the protein and ligand. During verification, appropriate simulation parameters are first set. If the verification result meets the preset functional threshold, the verification is considered successful, and the verified target protein sequence is output. If the verification result does not meet the preset functional threshold, the key structural regions leading to functional defects are analyzed. Based on the key structural regions, the set of mutation targets or residue types in the protein sequence optimization model is adjusted, a new round of candidate mutation sequence library is regenerated, and the steps of performing structure prediction and stability calculation on the sequences in the candidate mutation sequence library are continued to screen out optimized sequences that meet the preset requirements.
[0042] For example, based on the design requirement of "designing a highly conductive bio-nanowire," conductivity verification is performed. This involves combining molecular dynamics simulations with quantum mechanical calculations to evaluate the charge transport efficiency of the optimized sequence in its predicted three-dimensional structure. If the simulation verification shows that its conductivity is higher than a preset functional threshold, the sequence is deemed to have passed verification and is output as the final target protein sequence. If verification fails, the structural root causes of poor conductivity are analyzed in depth, and the analysis results, such as "excessive π-π stacking distance" or "poor orbital overlap due to side chain orientation," are used as qualitative performance feedback. Simultaneously, the sequences that fail verification are fed back into the protein sequence optimization model to continue performing structural prediction and stability calculations on sequences in the candidate mutant sequence library, ultimately selecting optimized sequences that meet the preset requirements.
[0043] In this embodiment, the functional characteristics of the target protein are obtained through design requirement analysis, and a first design rule is determined based on these characteristics. A first protein sequence and a second protein sequence are extracted based on the first design rule, and then a second design rule is determined, namely, the conserved and functional regions of the protein sequence. The possible amino acid positions for mutation are designed, and the candidate mutant sequences are subjected to structure prediction and stability calculations. Optimized sequences that meet preset requirements are selected, and functional simulations are performed on the selected optimized sequences for verification. If the verification is successful, the final target protein sequence is output; if the verification fails, the reasons for the failure are analyzed, and the verification results are fed back to the protein sequence optimization model for iterative optimization design. This completes the entire process from design requirements to target protein generation, improving the efficiency of de novo oligoprotein design.
[0044] Example 2 Based on Embodiment 1, another embodiment of this application is proposed, with reference to...Figure 2 After the step of generating a candidate mutant sequence library according to the first design rule and the second design rule, the following steps are included: Step S310: Input the candidate mutant sequence library into the protein sequence generation model, and the protein sequence generation model optimizes the sequences based on the functional characteristics and outputs the optimized candidate mutant sequence library.
[0045] In this embodiment, in order to improve the efficiency of structural testing and stability analysis, after initially generating a candidate mutant sequence library, the protein sequences in the candidate sequence library can be optimized and mutated based on functional characteristics, and then the optimized and mutated candidate mutant sequences can be updated to the candidate mutant sequence library.
[0046] As an alternative implementation, functional characteristics can be used as direct guiding signals to control the protein sequence generation model. That is, each sequence in the candidate sequence library, along with its corresponding physicochemical property indicators obtained from structure prediction, is used as an input pair and input into a protein sequence generation model based on a conditional diffusion model.
[0047] For example, suppose the candidate PilA mutant sequence is paired with two physicochemical indicators: the predicted average centroid distance between aromatic residues and the hydrophobic core stacking score. The conditional diffusion model is set to optimize "shortening the centroid distance" and "increasing the hydrophobic core score." The protein sequence generation model, through directed exploration in the sequence space, outputs sequence variants that produce tighter π-π stacking and a denser hydrophobic core while maintaining structural stability. For example, replacing a phenylalanine residue at a certain site with a larger tryptophan residue enhances stacking, or replacing an excessively large residue in the side chain with isoleucine eliminates spatial conflicts, thereby optimizing core packaging. The output sequence variants are then updated in the candidate mutant sequence library.
[0048] As an alternative implementation, functional characteristics can be transformed into energy functions, and sequence optimization can be driven by optimizing these energy values. Specifically, a library of candidate mutant sequences is input into an energy-guided protein sequence generation model, which incorporates one or more custom energy functions associated with the functional characteristics. By sampling in the sequence space and scoring the sampled points based on the energy functions, sequence mutations that reduce negative energies (such as steric hindrance energy and electrostatic repulsion energy) and enhance positive energies (such as π-π stacking interaction energy and hydrogen bond energy) are preferentially selected, thereby achieving directed sequence evolution.
[0049] For example, taking the design of conductive proteins as an example, three key energy functions are defined for the protein sequence generation model: π-orbital overlap energy, steric hindrance energy, and oligomerization interface binding energy. When optimizing sequences in the candidate library, various point mutations and combined mutations are tried, and the aforementioned energy changes of the mutated sequences are quickly calculated. For example, mutating a non-aromatic residue to tyrosine will be retained if this mutation significantly reduces the π-orbital overlap energy (enhancing conductivity) without significantly increasing the steric hindrance energy and interface binding energy (maintaining stability). In this way, the model can automatically "discover" complex mutation combinations and output an optimized sequence library that is superior in terms of energy.
[0050] In this embodiment, after generating the candidate sequence library, the mutant sequences in the sequence library are first optimized based on their functional characteristics, and candidate mutant sequences that are more functionally compatible with the design requirements are output. This makes it more efficient for the candidate mutant sequences output after structural analysis and stability analysis to pass functional verification, thereby improving the efficiency of oligomeric protein design.
[0051] Example 3 In this application embodiment, an oligomeric protein de novo design device is proposed.
[0052] Reference Figure 3 , Figure 3 This is a schematic diagram of the terminal structure of the hardware operating environment involved in one embodiment of this application.
[0053] like Figure 3 As shown, the control terminal may include: a processor 1001, such as a CPU, a network interface 1003, a memory 1004, and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The network interface 1003 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1004 may be high-speed RAM or stable non-volatile memory, such as disk storage. Alternatively, the memory 1004 may be a storage device independent of the aforementioned processor 1001.
[0054] Those skilled in the art will understand that Figure 3 The terminal structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0055] like Figure 3 As shown, the memory 1004, which serves as a computer storage medium, may include an operating system, a network communication module, and an oligomeric protein de novo design program.
[0056] exist Figure 3In the hardware structure of the oligoprotein de novo design device shown, the processor 1001 can call the oligoprotein de novo design program stored in the memory 1004 and perform the following operations: Obtain the functional properties of the target protein and determine a first design rule based on the functional properties; Based on the first design rule, a first protein sequence and a second protein sequence are extracted from a natural protein database, and a second design rule is determined. A candidate mutation sequence library is generated according to the first design rule and the second design rule; The sequences in the candidate mutant sequence library are subjected to structure prediction and stability calculation, and optimized sequences that meet the preset requirements are selected. The optimized sequence is subjected to functional characteristic simulation verification. If the verification is successful, the final target protein sequence is output; otherwise, the verification result is fed back to the protein sequence optimization model for iterative optimization design.
[0057] Optionally, the processor 1001 may call the oligoprotein de novo design program stored in the memory 1004 and also perform the following operations: Multiple sequence alignment was performed on the first protein sequence and the second protein sequence to identify supersecondary structural elements in the sequences; Based on the supersecondary structural elements, amino acid residue sites directly related to the functional characteristics are identified to form functionally relevant regions; The second design rule is determined based on the conserved regions and functionally relevant regions in the first and second protein sequences.
[0058] Optionally, the processor 1001 may call the oligoprotein de novo design program stored in the memory 1004 and also perform the following operations: Determine the main chain template sequence corresponding to the functional characteristics according to the first design rule; The sequence constraints on the main chain template sequence are generated according to the second design rule; wherein, the sequence constraints include fixed constraints and variable constraints, the fixed constraints specify that the amino acid residue sites and their residue types corresponding to the conserved regions cannot be changed; the variable constraints specify at least one set of candidate residue types that can be allowed to be replaced for each mutation target in the functionally relevant regions. Based on the fixed constraints and the variable constraints, an amino acid sequence that meets the constraints is generated using a sequence space sampling algorithm, forming the candidate mutant sequence library.
[0059] Optionally, the processor 1001 may call the oligoprotein de novo design program stored in the memory 1004 and also perform the following operations: The candidate mutation sequence is input into a trained protein structure prediction model, which outputs the predicted three-dimensional structure. The degree of fit between the predicted three-dimensional structure and the preset structural features determined based on design requirements is calculated, and the folding free energy variation of the predicted three-dimensional structure is calculated; wherein, the preset structural features include at least one of the target folding type, the relative arrangement of secondary structural elements, the geometry of the interaction interface, and the spatial configuration of the functional sites. A sequence whose fit is lower than a first threshold and whose folding free energy change is lower than a second threshold is determined to meet the preset requirements.
[0060] Optionally, the processor 1001 may call the oligoprotein de novo design program stored in the memory 1004 and also perform the following operations: The candidate mutant sequence library is input into a protein sequence generation model, which optimizes the sequences based on the functional characteristics and outputs an optimized candidate mutant sequence library.
[0061] Optionally, the processor 1001 may call the oligoprotein de novo design program stored in the memory 1004 and also perform the following operations: If the simulation verification fails, analyze the key structural areas that cause the functional defects; Based on the key structural regions, the set of mutation targets or residue types in the protein sequence optimization model is adjusted to generate a new round of candidate mutation sequence libraries.
[0062] Optionally, the processor 1001 may call the oligoprotein de novo design program stored in the memory 1004 and also perform the following operations: Design requirements for receiving the target protein; The design requirements are analyzed to determine the functional characteristics of the target protein.
[0063] Furthermore, to achieve the above objectives, embodiments of the present invention also provide a de novo design system for oligomeric proteins, comprising: The rule determination module is used to acquire the functional characteristics of the target protein and determine a first design rule based on the functional characteristics, and to extract protein sequences from a natural protein database based on the first design rule and determine a second design rule. A sequence generation module is used to generate a candidate mutation sequence library according to the first design rule and the second design rule; The screening module is used to perform structure prediction and stability calculation on the sequences in the candidate mutation sequence library, and screen out the optimized sequences that meet the preset requirements. The verification and iteration module is used to simulate and verify the functional characteristics of the optimized sequence. If the verification is successful, the final target protein sequence is output; otherwise, the verification result is fed back to the sequence generation module for iterative optimization design.
[0064] In addition, to achieve the above objectives, embodiments of the present invention also provide a terminal device, including a memory, a processor, and an oligomeric protein de novo design program stored in the memory and executable on the processor. When the processor executes the oligomeric protein de novo design program, it implements the oligomeric protein de novo design method as described above.
[0065] In addition, to achieve the above objectives, embodiments of the present invention also provide a computer-readable storage medium storing an oligoprotein de novo design program, which, when executed by a processor, implements the oligoprotein de novo design method as described above.
[0066] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0067] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0068] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0069] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0070] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. This application can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, third, etc., does not indicate any order. These words can be interpreted as names.
[0071] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0072] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of the invention. Therefore, if these modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
Claims
1. A method for de novo design of oligomeric proteins, characterized in that, The method includes: Obtain the functional properties of the target protein and determine a first design rule based on the functional properties; Based on the first design rule, a first protein sequence and a second protein sequence are extracted from a natural protein database, and a second design rule is determined. A candidate mutation sequence library is generated according to the first design rule and the second design rule; The sequences in the candidate mutant sequence library are subjected to structure prediction and stability calculation, and optimized sequences that meet the preset requirements are selected. The optimized sequence is subjected to functional characteristic simulation verification. If the verification is successful, the final target protein sequence is output; otherwise, the verification result is fed back to the protein sequence optimization model for iterative optimization design.
2. The de novo design method for oligomeric proteins as described in claim 1, characterized in that, The steps of extracting a first protein sequence and a second protein sequence from a natural protein database based on the first design rule, and determining the second design rule, include: Multiple sequence alignment was performed on the first protein sequence and the second protein sequence to identify supersecondary structural elements in the sequences; Based on the supersecondary structural elements, amino acid residue sites directly related to the functional characteristics are identified to form functionally relevant regions; The second design rule is determined based on the conserved regions and functionally relevant regions in the first and second protein sequences.
3. The de novo design method for oligomeric proteins as described in claim 1, characterized in that, The step of generating a candidate mutant sequence library according to the first design rule and the second design rule includes: Determine the main chain template sequence corresponding to the functional characteristics according to the first design rule; The sequence constraints on the main chain template sequence are generated according to the second design rule; wherein, the sequence constraints include fixed constraints and variable constraints, the fixed constraints specify that the amino acid residue sites and their residue types corresponding to the conserved regions cannot be changed; the variable constraints specify at least one set of candidate residue types that can be allowed to be replaced for each mutation target in the functionally relevant regions. Based on the fixed constraints and the variable constraints, an amino acid sequence that meets the constraints is generated using a sequence space sampling algorithm, forming the candidate mutant sequence library.
4. The de novo design method for oligomeric proteins as described in claim 1, characterized in that, The step of performing structure prediction and stability calculation on sequences in the candidate mutant sequence library to select optimized sequences that meet preset requirements includes: The candidate mutation sequence is input into a trained protein structure prediction model, which outputs the predicted three-dimensional structure. The degree of fit between the predicted three-dimensional structure and the preset structural features determined based on design requirements is calculated, and the folding free energy variation of the predicted three-dimensional structure is calculated; wherein, the preset structural features include at least one of the target folding type, the relative arrangement of secondary structural elements, the geometry of the interaction interface, and the spatial configuration of the functional sites. A sequence whose fit is lower than a first threshold and whose folding free energy change is lower than a second threshold is determined to meet the preset requirements.
5. The de novo design method for oligomeric proteins as described in claim 1, characterized in that, After the step of generating a candidate mutant sequence library according to the first design rule and the second design rule, the following is included: The candidate mutant sequence library is input into a protein sequence generation model, which optimizes the sequences based on the functional characteristics and outputs an optimized candidate mutant sequence library.
6. The de novo design method for oligomeric proteins as described in claim 1, characterized in that, The step of feeding the validation results back to the protein sequence optimization model for iterative optimization design includes: If the simulation verification fails, analyze the key structural regions that cause the functional defects; Based on the key structural regions, the set of mutation targets or residue types in the protein sequence optimization model is adjusted to generate a new round of candidate mutation sequence libraries.
7. The de novo design method for oligomeric proteins as described in claim 1, characterized in that, Before the step of obtaining the functional properties of the target protein and determining the first design rule based on the functional properties, the following steps are included: Design requirements for receiving the target protein; The design requirements are analyzed to determine the functional characteristics of the target protein.
8. A de novo design system for oligomeric proteins, characterized in that, The system includes: The rule determination module is used to acquire the functional characteristics of the target protein and determine a first design rule based on the functional characteristics, and to extract protein sequences from a natural protein database based on the first design rule and determine a second design rule. A sequence generation module is used to generate a candidate mutation sequence library according to the first design rule and the second design rule; The screening module is used to perform structure prediction and stability calculation on the sequences in the candidate mutation sequence library, and screen out the optimized sequences that meet the preset requirements. The verification and iteration module is used to simulate and verify the functional characteristics of the optimized sequence. If the verification is successful, the final target protein sequence is output; otherwise, the verification result is fed back to the sequence generation module for iterative optimization design.
9. A terminal device, characterized in that, The method includes a memory, a processor, and an oligomeric protein de novo design program stored in the memory and executable on the processor, wherein when the processor executes the oligomeric protein de novo design program, it implements the method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an oligoprotein de novo design program, which, when executed by a processor, implements the method described in any one of claims 1-7.