Multi-dimensional screening method for protein design based on lexicographical order optimization strategy
Through a multi-dimensional screening method of dictionary sequence optimization strategy, combined with deep learning algorithms to generate and screen protein sequences, the time-consuming problems in traditional methods are solved, and efficient and accurate protein design is achieved, which is suitable for drug development, industrial enzyme engineering and gene therapy.
Patent Information
- Application Number
- CN202510775259.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to quickly and accurately screen out proteins that meet the expected structural, stability and functional requirements among thousands of protein candidate sequences, and traditional methods are time-consuming and costly.
A multi-dimensional screening method based on dictionary order optimization strategy is adopted to generate candidate amino acid sequences through deep learning algorithms, and a multi-dimensional screening index of structural accuracy, thermodynamic stability, thermal stability and physical and chemical properties is combined to optimize the candidate sequence to generate the final candidate sequence.
It improves the efficiency and accuracy of protein design, reduces the number of experimental verifications, can handle complex protein structure and functional relationships, and provides strong support for the fields of drug design, enzyme engineering and gene therapy.
Smart Images

Figure CN120280002A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of computational biology and synthetic biology, and specifically relates to a multi - dimensional screening method for protein design based on the lexicographic optimization strategy, which is applicable to the efficient design of highly stable and highly functional proteins in scenarios such as drug development, industrial enzyme engineering, and gene therapy. Background Art
[0002] Proteins are the main executors of life activities and are involved in almost all physiological processes in organisms, including catalyzing biochemical reactions, transmitting signals, providing structural support, and participating in immune defense. The function of a protein is determined by its three - dimensional structure, which in turn is determined by its amino acid sequence. With the development of structural biology and computational biology, scientists have been able to resolve the structures of a large number of proteins and understand their functional mechanisms. However, the functions of natural proteins are often restricted by their evolutionary history and cannot fully meet the needs of humans in the fields of medicine, industry, and agriculture. Therefore, designing and modifying proteins to achieve specific functions has become an important research direction in the field of biotechnology. Traditional protein design methods mainly rely on experimental screening and optimization. For example, directed evolution is a commonly used protein engineering strategy that, by simulating the process of natural selection, mutates and screens proteins in the laboratory for multiple rounds to obtain proteins with improved functions. Although directed evolution has been successful in many applications, its process is time - consuming and costly, and relies on large - scale experimental screening. In addition, due to the huge size of the protein sequence space (the sequence space of a protein composed of 100 amino acids is 20^100), it is unrealistic to rely solely on experimental methods to explore all possible sequences.
[0003] In recent years, with the breakthroughs in protein structure prediction methods (such as AlphaFold, ESMFold) and the development of deep - learning - based sequence generation algorithms (such as ProteinMPNN), computer - aided protein design has gradually become a research hotspot. However, due to the extremely large protein sequence space and diverse targets, it remains a difficult problem to quickly and accurately screen out the best candidates that meet the expected structural, stability, and functional requirements from thousands or even tens of thousands of candidate sequences. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a multi - dimensional screening method for protein design based on the lexicographic optimization strategy.
[0005] The purpose of the present invention is achieved by the following technical solutions: A multi - dimensional screening method for protein design based on the lexicographic optimization strategy, comprising the following steps: Predict the three - dimensional structure of the complex of the target protein and the ligand to obtain the interaction information between the target protein and the ligand; Construct a three-dimensional backbone model by extracting the three-dimensional structure of the complex of the target protein and ligand. Use a deep learning algorithm to generate candidate amino acid sequences that match the three-dimensional backbone model, and optimize the binding affinity and structural stability of the candidate amino acid sequences by dynamically learning the interaction information between the target protein and ligand. Use multi-dimensional screening metrics to hierarchically screen the candidate amino acid sequences. The screening metrics include: structural accuracy screening metrics, thermodynamic stability screening metrics, thermal stability screening metrics, and physicochemical property screening metrics. Based on the lexicographical optimization strategy, optimize the screening metrics in order of priority to generate the final candidate sequences. The priority order is structural accuracy > thermodynamic stability > thermal stability > physicochemical properties.
[0006] Furthermore, the three-dimensional structure of the complex includes complexes of proteins, nucleic acids, small molecule ligands, or metal ions.
[0007] Furthermore, the step of optimizing the binding affinity and structural stability of the candidate amino acid sequences by dynamically learning the interaction information between the target protein and ligand includes: dynamically learning the physicochemical patterns between the target protein and ligand, including hydrogen bonds and hydrophobic interactions, and automatically optimizing the key residues in the binding pocket, thereby generating sequences with both high affinity and structural stability.
[0008] Furthermore, the screening metrics are specifically: Structural accuracy screening metrics, the structural accuracy metrics pLDDT and pTM evaluated by structural prediction methods, are used to screen three-dimensional folding conformations with high confidence; among them, pLDDT is used to evaluate local structural accuracy, and pTM is used to evaluate the global structure of the protein. Thermodynamic stability screening metrics, the free energy change metrics evaluated by thermodynamic calculation methods, are used to screen sequences with stable folding states. Thermal stability screening metrics, the metrics evaluated by thermal stability prediction methods, are used to screen stable sequences adapted to specific temperature environments. Physicochemical property screening metrics, the metrics calculated by physicochemical property analysis methods, are used to optimize the water solubility, purification feasibility, and functional adaptability of the sequences.
[0009] Furthermore, the threshold of the structural accuracy screening metrics is pLDDT > 70, and the candidate sequences with the highest pLDDT score are preferentially retained.
[0010] Furthermore, the thermodynamic stability screening threshold is ΔΔG < 0, and the free energy ΔG of the wild-type protein WTScreening is carried out based on a benchmark, where ΔΔG represents the free energy change.
[0011] Furthermore, the thermal stability screening threshold is set according to the target application scenario, including the stability score threshold in a high-temperature environment.
[0012] Furthermore, the physicochemical properties include hydrophobicity, isoelectric point or aromaticity, and the screening threshold is GRAVY < 0, where GRAVY represents the average hydrophilicity.
[0013] Furthermore, the lexicographic optimization strategy specifically includes: Retaining sequences that meet the structural accuracy threshold in the first-level screening; In the second-level screening, screening sequences that meet the thermodynamic stability threshold from the results of the first-level screening; In the third-level screening, screening sequences that meet the thermal stability threshold from the results of the second-level screening; In the fourth-level screening, screening sequences that meet the physicochemical property threshold from the results of the third-level screening.
[0014] The present invention also provides a multi-dimensional screening device for protein design based on the lexicographic optimization strategy, including one or more processors for implementing the multi-dimensional screening method for protein design based on the lexicographic optimization strategy.
[0015] The beneficial effects of the present invention are as follows: By combining the deep learning method of protein design, the present invention provides an efficient, accurate and fully automated protein redesign process. This method can design new proteins while retaining their core functions based on known functional proteins, greatly improving the design efficiency and accuracy. Compared with traditional protein redesign methods, the present invention can significantly reduce the number of experimental verifications and can handle complex protein structure and function relationships, providing strong support for fields such as drug design, enzyme engineering and gene therapy. In the method of sequence screening, the present invention adopts lexicographic optimization, which ensures that key objectives (such as structural accuracy, stability) are preferentially met and avoids interference of secondary indicators in traditional multi-objective optimization with core performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic flow chart of de novo design of proteins according to the present invention; Figure 2 It is a schematic flow chart of screening protein sequences by lexicographic optimization according to the present invention; Figure 3 It is a spatial structure diagram of a protein-small molecule complex predicted by RFAA in Example S1 of the present invention; Figure 4 It is a schematic diagram of the change in the number of sequences during the screening process by lexicographic optimization according to the present invention. Detailed Implementation Modes
[0017] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation modes described in the following exemplary embodiments do not represent all implementation modes consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0018] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0019] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0020] As Figure 1 shown, an embodiment of the present invention provides a multi-dimensional screening method for protein design based on a lexicographical order optimization strategy, including the following steps: Step 1: Prediction of the structure of a protein complex The prediction of the structure of a protein complex is crucial for subsequent protein design. Precise prediction of the protein complex structure requires the use of an all-atom modeling method, which can finely present the backbone and side-chain information of the protein and can accurately describe the microscopic structural features inside the protein and at the complex interface. To predict the three-dimensional structure of complex complexes containing proteins, nucleic acids, small molecules, metal ions, and covalent modifiers, etc., the present invention uses the RosettaFold All-Atom (RFAA) method for complex structure prediction. This method comprehensively utilizes deep learning and energy optimization algorithms, and through a large amount of training data and advanced computational models, significantly improves the accuracy and stability of protein complex structure prediction, effectively reduces the computational cost and time consumption, and provides a solid foundation for protein function research and drug screening.
[0021] Step 2: Construction of the protein backbone In the process of protein design, constructing the three-dimensional backbone of a protein is a crucial step. Traditional methods usually rely on experimental data to obtain the structural information of the target protein, which requires a large amount of experimental time and resources. Modern computational methods, through deep learning techniques, can efficiently obtain backbone information from the structures of known proteins. The present invention uses the RFdiffusion method to construct the backbone of the target protein. This method can be trained in a large-scale protein structure database to identify the core constituent elements of the protein backbone, thereby generating the backbone of the target protein. Specifically, by processing the structural data of known protein-ligand complexes using the RFdiffusion method, local and global geometric features and structural patterns are captured, and a new protein backbone is generated based on the geometric features and structural patterns. Without directly relying on specific amino acid sequences, a backbone model that meets the target function and structural requirements is created, greatly reducing the manual intervention and experimental cycle in the process of protein backbone construction.
[0022] Step 3: Generation of protein sequences After the backbone construction is completed, the next step is to generate the corresponding protein sequence. Traditional protein design methods usually generate protein sequences through random mutation or directed evolution, but this method cannot efficiently ensure the functionality and structural stability of the sequence. Modern methods use deep learning algorithms such as message passing neural networks (MPNNs) to generate protein sequences with specific functions starting from the backbone information. The present invention uses the LigandMPNN method to generate sequences for the backbone of the target protein. This method is a protein sequence generation model based on message passing neural networks, and its core advantage lies in explicitly modeling ligand-protein interactions: by taking ligand atoms as nodes in the input graph, the model can dynamically learn physical and chemical patterns such as hydrogen bonds and hydrophobic interactions between the two, and automatically optimize the key residues in the binding pocket, thereby generating sequences with both high affinity and structural stability. In addition, LigandMPNN significantly improves the computational efficiency, can complete sequence sampling within seconds, avoids the bottleneck of traditional Monte Carlo simulations that take several days, and supports batch generation of diverse candidate sequences, effectively breaking through the local optimum solution limit. At the application level, this model co-optimizes structural stability and binding free energy through a multi-task loss function, and the experimental verification success rate can reach 30-50% (traditional methods are usually below 20%).
[0023] Step 4: Screening indicators and screening methods for protein sequences Screening the generated protein sequences is crucial because this process can ensure that the candidate sequences not only have the expected structure and function theoretically, but also possess sufficient stability and biological activity in practical applications. By comprehensively considering multi-dimensional indicators such as structure prediction, stability assessment, free energy calculation, and physicochemical property analysis, sequences that do not meet the experimental requirements can be effectively filtered, thereby reducing the time and cost risks of subsequent experimental verification and providing more accurate and efficient candidate molecules for fields such as protein engineering and drug development.
[0024] 4.1 Screening Metrics The present invention uses ESMFold to perform three-dimensional structure prediction on the generated protein sequences and obtains the corresponding pLDDT and pTM scores to evaluate the accuracy and reliability of the predicted structure; uses FoldX to calculate the free energy of the protein sequence as an evaluation index of thermodynamic stability; uses TemStaPro to predict the thermal stability of the protein sequence to determine its feasibility in biological systems; and combines Biopython to comprehensively analyze the physicochemical properties of the protein sequence, including but not limited to the determination of parameters such as aromaticity, isoelectric point, and hydrophobicity, to achieve multi-dimensional and high-precision screening of the protein sequence.
[0025] 4.1.1 Structure Accuracy Screening Metrics Evaluating the structural quality of the protein sequences generated by LigandMPNN is a crucial step. Protein structure prediction methods can model the protein folding rules through deep learning models and directly predict the three-dimensional structure of proteins from the amino acid sequences. The present invention uses the ESMFold method to perform structure prediction and quality assessment on the generated sequences. ESMFold is a deep learning-based protein structure prediction method that can quickly predict the three-dimensional structure of proteins through sequence information. Its core innovation lies in combining a large-scale protein language model (ESM-2) with a folding algorithm to achieve efficient three-dimensional structure modeling with single-sequence input. After using ESMFold to predict the protein structure, each structure will generate corresponding pLDDT (Predicted Local Distance Difference Test) and pTM (Predicted Template Modelingscore) scores. pLDDT is an index to measure the precision of each amino acid residue in the predicted structure. The higher the pLDDT value, the more accurate the predicted residue position. The pTM value is used to evaluate the global and local accuracy of the predicted protein structure, especially the degree of matching with the reference template. By setting the thresholds of the pLDDT and pTM values, proteins with higher structural quality and higher prediction accuracy are screened out.
[0026] 4.1.2 Thermodynamic Stability Screening Metrics Folding thermodynamics analysis is used to evaluate the energy changes required for a protein to transition from a disordered state to a stable folded state, which is crucial for determining the stability of proteins in biological systems. Through thermodynamic calculations, the folding stability of proteins under different environmental conditions (such as temperature, pH, ionic strength, etc.) can be understood. In this invention, the FoldX method is used to perform folding thermodynamics calculations on proteins. By means of energy calculations and free energy analysis, the stability of protein structures can be efficiently evaluated. This method is based on an empirical force field model, quantifying the stability of proteins as the value of free energy (ΔG). By decomposing energy terms such as van der Waals forces, hydrogen bonds, solvation effects, backbone tension, and entropy loss, the energy state of the protein in the folded state can be accurately simulated. Its advantages include fast calculation speed (only a few seconds for single mutation analysis), low operation threshold (compatible with PDB format input), and high consistency with experimental data. Finally, by calculating the free energy change (ΔΔG) of the wild-type protein and the newly designed protein, the folding stability of the protein is evaluated.
[0027] 4.1.3 Thermal Stability Screening Index Thermal stability is a key parameter for proteins to stably maintain their functions in different environments. Especially in enzyme catalysis, biomedicine, and industrial applications, the thermal stability of proteins is crucial for their practical applications. By using deep learning methods that have flourished in recent years for sequence analysis, it can help solve this prediction problem. Such methods can not only support the training of larger-scale data but also may provide possibilities for the development of universal therapeutic factors adapted to multiple temperature ranges. In this invention, the TemStaPro method is used to evaluate the thermal stability of proteins. This method applies the principle of transfer learning, based on the sequence embeddings generated by a protein language model, to predict the thermal stability of the input protein sequence. Based on large protein language models pre-trained on hundreds of millions of known sequences, the embedding features generated by these models use more than one million sequences collected from organisms with annotated growth temperatures to efficiently train and validate a high-performance prediction method. In this invention, the generated protein sequences are screened for thermal stability using TemStaPro to evaluate the stability of proteins at different temperatures. By evaluating the structural maintenance of proteins in a high-temperature environment, proteins that can maintain stable folding and continue to function at high temperatures are screened out.
[0028] 4.1.4 Physicochemical Property Screening Index Biopython is a Python library for bioinformatics that provides comprehensive computational modules in the screening of the physicochemical properties of protein sequences. Biopython uses some built-in functions and modules to evaluate various properties of protein sequences. For example, the built-in GRAVY (Grand Average of Hydropathy), isoelectric_point, and aromaticity modules can calculate key indicators such as hydrophobicity, isoelectric point, and aromaticity to evaluate the water solubility, purification method, and stability of the tertiary structure of proteins. The present invention calculates and predicts the hydrophobicity, isoelectric point, and aromaticity of proteins through Biopython. The GRAVY module integrates classical hydrophobicity scales such as Kyte-Doolittle and Eisenberg, calculates the local hydrophobicity score through the sliding window method (default window size is 9), generates a hydrophobicity map, and is used to identify transmembrane regions or hydrophobic cores to assist in the optimal design of protein solubility. The isoelectric point calculation uses an iterative algorithm based on the Bjellqvist method to accurately solve the pH value when the net charge of the protein is zero, which is crucial for predicting the solubility of proteins in specific buffers and chromatography purification strategies (such as ion exchange chromatography). Aromaticity analysis evaluates the tendency of the rigid structure of proteins or the potential of π-π interactions by statistically analyzing the proportion and distribution of aromatic amino acids such as phenylalanine (F), tyrosine (Y), and tryptophan (W), providing a basis for the design of enzyme active sites or the modification of antibody CDR regions.
[0029] 4.2 Screening Method It is crucial to screen the generated protein sequences because this process can ensure that the candidate sequences not only have the expected structure and function theoretically but also have sufficient stability and biological activity in practical applications. By comprehensively considering multi-dimensional indicators such as structure prediction, stability evaluation, free energy calculation, and physicochemical property analysis, sequences that do not meet the experimental requirements can be effectively filtered out, thereby reducing the time and cost risks of subsequent experimental verification. The present invention screens the indicators of protein sequences based on lexicographic optimization.
[0030] 4.2.1 Concept and Mathematical Model of Lexicographic Optimization Lexicographic optimization is a multi-objective optimization method. Its core idea is to optimize each objective in turn according to the priority order of the objective functions, and the optimization at each stage needs to be carried out on the premise of satisfying the optimal solution of the previous objective. Specifically, lexicographic optimization first optimizes the objective with the highest priority until it reaches the optimum or meets the threshold. Then, on the premise of ensuring that the objective with the highest priority is not damaged, it optimizes the objective with the second highest priority, and so on until all objectives are processed. In protein sequence screening, this method can effectively integrate multi-dimensional indicators such as structure, stability, function, and physicochemical properties, and improve the design efficiency and success rate through hierarchical screening. The mathematical model of lexicographic optimization is as follows. Suppose we need to maximize or minimize n objective functions and assume that the priorities of the objective functions are sorted according to , that is is the most important objective. Then lexicographic optimization can be described as the following recursive problem: A. Solve the first-level optimization problem: or .
[0031] Obtain the optimal value and its optimal solution set .
[0032] B. Solve the second-level optimization problem on the set : or .
[0033] Obtain the optimal value and its optimal solution set .
[0034] C. And so on. For the th objective, solve on : or .
[0035] Obtain the optimal solution set .
[0036] Finally, after n-level optimization, obtain the solution set that satisfies all objective functions.
[0037] 4.2.2 Application of Lexicographic Optimization in Protein Sequence Screening Compared with multi-objective methods such as Pareto optimization and weighted sum method, lexicographic optimization is more suitable for situations where there is an obvious priority ranking among objectives. In protein sequence screening, some key indicators (such as structural reliability and thermal stability) are usually considered much more important than other indicators. Therefore, using lexicographic optimization can ensure that while meeting the core requirements, other indicators are optimized, resulting in candidate sequences that are more balanced in all aspects.
[0038] By forcing the priority order, lexicographic optimization avoids the interference of low-priority objectives on high-priority objectives in traditional methods (for example, sacrificing structural accuracy in pursuit of higher thermal stability). Usually, structure prediction and thermal stability, as core objectives, should be ranked first in lexicographic optimization, while secondary indicators such as physicochemical properties and conservativeness are placed later. Based on this, the priority ranking of screening indicators for different protein sequences is as follows: a. Structural accuracy (such as pLDDT / pTM): The function of a protein is determined by its three-dimensional structure. Incorrect structure prediction will directly lead to functional failure, so it has the highest priority.
[0039] b. Thermodynamic stability (such as ΔΔG): A stable folded state is a prerequisite for a protein to maintain its function.
[0040] c. Thermal stability: Additional requirements for specific application scenarios (such as high-temperature enzymes).
[0041] d. Physicochemical properties (such as isoelectric point and hydrophobicity): Affect experimental operability and industrial production.
[0042] According to the priority, the present invention constructs corresponding objective functions for each screening indicator and substitutes the screening indicators into the mathematical model of lexicographic optimization, that is: represents structural accuracy, represents thermodynamic stability, represents thermal stability, represents physicochemical properties. That is, the priority ranking of the objective functions is, .
[0043] 4.2.3 Design process of lexicographic optimization Assume that a set of candidate protein sequences (generated by the previous sequence generation module) has been obtained, and each sequence can calculate the values of each objective function. Next, perform screening and sorting on according to the steps of lexicographic optimization: First-level optimization: Structural accuracy ( ) Use ESMFold to perform structure prediction on all generated protein sequences and finally obtain the pLDDT and pTM scores. For the interpretation of pLDDT and pTM, according to the AlphaFold official website guide, a pLDDT value greater than 70 usually indicates that the structure prediction of this region is very reliable and worthy of being used as a high-quality candidate structure. Therefore, lower pLDDT and pTM mean a higher probability of misfolding of the protein structure, so protein sequences with low scores can be preferentially excluded. Since pLDDT and pTM are positively correlated, that is, when pLDDT is larger, pTM is also larger; when pLDDT is smaller, pTM is also smaller. Therefore, usually only the value of pLDDT needs to be concerned about among these two scores. The mathematical expression is as follows: 。
[0044] By sorting the pLDDT scores of all candidate sequences, find the sequences with pLDDT scores greater than 70 as the optimal solutions and its set of optimal solutions 。This step ensures that the selected candidate sequences all have reliable protein structures.
[0045] Second-level optimization: Thermodynamic stability ( ) The stability of a protein can be judged by calculating its free energy (ΔG). The smaller the free energy, the more stable the protein. Calculate the free energies of the wild-type protein and the newly designed protein using FoldX to determine whether the stability of the newly designed protein is increased or decreased compared to the wild-type protein. The free energy change (ΔΔG) is the difference between the free energies of the newly designed protein and the wild-type protein. Calculating ΔΔG can intuitively compare whether the protein stability is increased or decreased, that is, when ΔΔG is negative, it means that the stability of the newly designed protein is increased compared to the wild-type protein; when ΔΔG is positive, it means that the stability of the newly designed protein is decreased compared to the wild-type protein. The following is the formula for calculating ΔΔG: ΔΔG = ΔG i – ΔG WT 。
[0046] where ΔΔG is the free energy change, ΔG i is the free energy of the redesigned protein, and ΔG WT is the free energy of the wild-type protein. It is hoped that the newly designed protein has higher stability, that is, retain the proteins with negative ΔΔG. Therefore, continue to solve based on the set : 。
[0047] Obtain the optimal value and its set of optimal solutions 。
[0048] Third - level optimization: Thermal stability ( ) TemStaPro can predict the thermal stability of protein sequences within a temperature range. According to requirements, given a temperature range as the threshold for thermal stability, continue to solve within the set : .
[0049] Obtain the optimal subset , so that the candidate sequences have better thermal stability that meets the experimental requirements while ensuring the optimality of the first two - level objectives.
[0050] Fourth - level optimization: Physicochemical properties ( ) The physicochemical properties of proteins affect the ease of protein purification and production. By predicting properties such as the hydrophobicity of protein sequences through Biopython, the number of experiments in the later wet - experiment verification can be effectively reduced. Therefore, during the optimization in , according to requirements, select a relatively important physicochemical property of the protein for optimization. For example, if more attention is paid to whether the protein is easily soluble in water, let represent water solubility. The more negative the GRAVY value predicted by Biopython, the more hydrophilic it is. Therefore, take GRAVY = 0 as the threshold to screen protein sequences with negative GRAVY values. Thus, optimization matching can be carried out in : .
[0051] Obtain the optimal subset , where each sequence further improves the functional matching while meeting the previous objectives. This step is applicable to other physical properties of proteins that are concerned, such as aromaticity, isoelectric point, etc.
[0052] Finally, optimize according to lexicographical order to obtain a set of candidate sequence sets that strictly meet or are close to the optimal conditions of each level . At this time, sort according to the subtle differences of each sequence in the last - level objective, or select the first several sequences according to actual needs to enter the experimental verification stage. Example
[0053] This example takes the de novo design of synthase Y as an example to demonstrate the full process from protein backbone construction, sequence generation to multi - dimensional screening.
[0054] S1. Prediction of protein complex structure and generation of protein backbone and sequence First, obtain the sequence of the synthase and the small molecule structural formula, and use RFAA to predict the structures of the enzyme and the small molecule. Then use RFdiffusion to perform de novo design of the synthase. Before using RFdiffusion, fix the amino acid positions of the key active sites of the known synthase and its adjacent regions to ensure that the newly generated protein backbone has an innovative overall structure while retaining the original catalytic function. Use this method to generate 1000 new protein backbones, each of which contains a fixed functional region. After obtaining the new backbones, use LigandMPNN to automatically generate the corresponding enzyme sequences for each backbone. Specifically, each backbone generates 4 candidate sequences through a deep learning algorithm, thus obtaining a total of 4000 brand-new enzyme sequences containing key active sites. During this process, while maintaining the integrity of the backbone structure, the generated sequences are optimized by algorithms to be more likely to meet the design requirements in computational prediction.
[0055] S2. Screening of protein sequences Screen the above-obtained 4000 protein sequences using lexicographic optimization to obtain the optimal sequence set. The screening indicators used in the present invention are pLDDT and pTM predicted by ESMFold, hydrophobicity, aromaticity and isoelectric point predicted by Biopython, free energy change predicted by FoldX, and thermal stability predicted by TemStaPro. The present invention predicts all the above screening indicators for the 4000 selected sequences. As shown in Table 1 and Table 2, the predicted values of some protein sequences are presented. After obtaining all the predicted data, screen the protein sequences according to the lexicographic optimization screening process.
[0056] Table 1: Predicted data 1 of candidate sequences
[0057] Table 2: Predicted data 2 of candidate sequences
[0058] S2.1. First-level optimization: Structural accuracy ( ) Since pLDDT and pTM are positively correlated, only the value of pLDDT is concerned here. A pLDDT value greater than 70 indicates that the protein sequence is relatively reliable and can be used as a high-quality candidate structure. Therefore, 70 is used as the threshold for the pLDDT score to obtain a set containing the optimal solution , the set It contains all protein sequences with a pLDDT score greater than 70. Table 3 shows some of the protein sequences above the threshold. The first row, WT, is the score of the wild-type enzyme, and the others are the scores of the de novo designed protein structures. Since a pLDDT value greater than 70 indicates a relatively reliable structure prediction, in this round of screening, proteins with a pLDDT greater than 70 were retained from 4000 designed proteins, that is, 2575 proteins were retained.
[0059] Table 3: pLDDT scores predicted by ESMFold
[0060] S2.2, Second-level optimization: Thermodynamic stability ( ) The free energy (total energy, or ΔG) of all proteins was calculated using FoldX, and the free energy (ΔG WT ) of the wild-type protein was used as the threshold to filter and screen the proteins below this threshold. That is, in the set , protein sequences with a negative free energy change (ΔΔG = ΔG i –ΔG WT ) were retained, and sequences with a positive free energy change were excluded, generating a new set . Table 4 shows the free energy changes of some proteins. Excluding 415 protein sequences with a positive free energy change, the set contains 2160 proteins with a pLDDT greater than 70 and a negative free energy change.
[0061] Table 4: Partial free energy changes
[0062] S2.3, Third-level optimization: Thermal stability ( ) TemStaPro predicts the thermal stability of proteins in a certain temperature range and can determine whether the proteins are stable in this temperature range. In this example, more attention is paid to the thermal stability of proteins at 65 degrees, so the threshold is set to 0.9, that is, protein sequences with a score greater than 0.9 are retained. Table 5 shows the TemStaPro prediction scores of proteins at 65 °C. Therefore, based on the set, through thermal stability prediction and the setting of temperature difference and threshold, the optimized sequences are 1796, generating a new set.
[0063] Table 5: TemStaPro prediction scores of thermal stability at 65 °C
[0064] S2.4, Fourth-level optimization: Physicochemical properties ( ) The optimization of physicochemical properties can be regarded as application-oriented optimization. Since hydrophobicity, isoelectric point, and aromaticity can be used to evaluate the solubility of proteins in water, the ease of separation and purification, and the stability of the tertiary structure using Biopython. Therefore, when optimizing this step, it is necessary to determine which properties of the protein are of the greatest concern. For example, in this instance, more attention is paid to whether the de novo designed enzyme is easily soluble in water. Therefore, the GRAVY values of the protein sequences calculated by Biopython are sorted, as shown in Table 6. Finally, 233 optimal protein sequences are selected according to the requirements for wet experiment verification.
[0065] Table 6: Average coefficient of hydrophilicity predicted by Biopython
[0066] Through the above specific implementation manners, the present invention provides a method for de novo design of synthetic enzymes based entirely on computer-aided design and simulation evaluation. The entire process does not rely on experimental verification and functional testing, but through high-precision structure prediction, free energy calculation, and physicochemical parameter analysis, realizes multi-level screening and iterative optimization of candidate proteins. This method not only greatly shortens the design cycle, but also can efficiently generate proteins with excellent stability and expected functions in a fully automated environment, laying a solid foundation for further research and application on the computational simulation platform.
[0067] The embodiment of the present invention also provides a device for de novo design of proteins based on deep learning and lexicographical order optimization, including one or more processors for implementing the method for de novo design of proteins based on deep learning and lexicographical order optimization.
[0068] The embodiment of the device for de novo design of proteins based on deep learning and lexicographical order optimization of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.
[0069] The implementation processes of the functions and roles of each unit in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0070] For the device embodiments, since they basically correspond to the method embodiments, reference may be made to the relevant descriptions in the method embodiments for details. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0071] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.
Claims
1. A multi-dimensional screening method for protein design based on a lexicographic order optimization strategy, characterized in that, It includes the following steps: Predict the three-dimensional structure of the complex of the target protein and ligand to obtain the interaction information between the target protein and the ligand; Construct a three-dimensional backbone model by extracting the three-dimensional structure of the complex of the target protein and ligand; Use a deep learning algorithm to generate a candidate amino acid sequence that matches the three-dimensional backbone model, and optimize the binding affinity and structural stability of the candidate amino acid sequence by dynamically learning the interaction information between the target protein and the ligand; Use multi-dimensional screening metrics to hierarchically screen the candidate amino acid sequences, and the screening metrics include: structural accuracy screening metrics, thermodynamic stability screening metrics, thermal stability screening metrics, and physicochemical property screening metrics; Based on the lexicographic optimization strategy, optimize the screening metrics in order of priority to generate the final candidate sequence, and the priority order is structural accuracy > thermodynamic stability > thermal stability > physicochemical properties.
2. The method according to claim 1, characterized in that The three-dimensional structure of the complex includes complexes of proteins, nucleic acids, small molecule ligands, or metal ions.
3. The method according to claim 1, wherein The step of optimizing the binding affinity and structural stability of the candidate amino acid sequence by dynamically learning the interaction information between the target protein and the ligand includes: dynamically learning the physicochemical patterns between the target protein and the ligand, including hydrogen bonds and hydrophobic interactions, and automatically optimizing the key residues in the binding pocket, thereby generating a sequence with both high affinity and structural stability.
4. The method according to claim 1, wherein The screening metrics are specifically: Structural accuracy screening metrics, the structural accuracy metrics pLDDT and pTM evaluated by structure prediction methods, which are used to screen three-dimensional folding conformations with high confidence; among them, pLDDT is used to evaluate local structural accuracy, and pTM is used to evaluate the global structure of the protein; Thermodynamic stability screening metrics, the free energy change metrics evaluated by thermodynamic calculation methods, which are used to screen sequences with stable folding states; Thermal stability screening metrics, the metrics evaluated by thermal stability prediction methods, which are used to screen stable sequences adapted to specific temperature environments; Physicochemical property screening metrics, the metrics calculated by physicochemical property analysis methods, which are used to optimize the water solubility, purification feasibility, and functional suitability of the sequence.
5. The method according to claim 4, characterized in that The threshold of the structural accuracy screening metric is pLDDT > 70, and the candidate sequence with the highest pLDDT score is preferentially retained.
6. The method according to claim 4, characterized in that The thermodynamic stability screening threshold is ΔΔG < 0, and screening is carried out based on the free energy ΔG of the wild-type protein WT as a reference, where ΔΔG represents the free energy change.
7. The method according to claim 4, characterized in that, The thermal stability screening threshold is set according to the target application scenario, including the stability score threshold in high-temperature environments.
8. The method according to claim 1, wherein The physicochemical properties include hydrophobicity, isoelectric point, or aromaticity, and the screening threshold is GRAVY < 0, where GRAVY represents the average hydrophilicity.
9. The method according to claim 1, wherein The lexicographic optimization strategy specifically includes: Retain the sequences that meet the structural accuracy threshold in the first-level screening; In the second-level screening, screen the sequences that meet the thermodynamic stability threshold from the results of the first-level screening; In the third-level screening, screen the sequences that meet the thermal stability threshold from the results of the second-level screening; In the fourth-level screening, screen the sequences that meet the physicochemical property threshold from the results of the third-level screening.
10. A multi-dimensional screening device for protein design based on a lexicographical order optimization strategy, characterized in that, Comprising one or more processors for implementing a multi-dimensional screening method for protein design based on a lexicographical order optimization strategy as described in any one of claims 1-9.
Citation Information
Patent Citations
Enzyme thermal stability mutant prediction method and device, electronic equipment and storage medium
CN118197408A
High-precision drug screening method, system and equipment based on hierarchical multi-modal representation learning and readable storage medium
CN119724333A
Cited By
Multi-modal deep learning intelligent design method based on protein sequence and structure
CN121213645A
Multimodal Deep Learning Intelligent Design Method Based on Protein Sequence and Structure
CN121213645B
Method and system for efficiently screening high-activity isoenzyme
CN122201503A