Method, apparatus and related equipment for screening mutant proteins

By simulating single mutation, language model prediction and folding energy calculation methods, the target mutant protein was screened out, which solved the problems of high difficulty in designing and transformation of mutant proteins and poor screening effect in the prior art, and achieved the effect of reducing calculation amount and cost.

CN119296640BActive Publication Date: 2025-06-10CIXI INST OF BIOMEDICAL ENG NINGBO INST OF IND TECH CHINESE ACAD OF SCI NINGBO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411836877.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-06-10
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

In the prior art, mutant protein design and transformation is difficult, poor screening effect, large calculation amount, long test cycle and high cost.

Method used

By simulated single mutations of multiple amino acid sites in the protein sequence of the target protein, adaptive predictions are performed based on the target protein language model, structural stability values ​​are calculated in combination with the protein folding energy calculation tool, normalized processing and multi-index sorting, and the target mutant proteins are screened out.

Benefits of technology

It reduces the difficulty of protein design and transformation, improves the screening effect of mutant proteins, reduces the calculation amount and test cycle, and reduces the cost of consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119296640B_ABST
    Figure CN119296640B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and related equipment for screening mutant proteins, which relates to the technical field of material determination and analysis. The method includes the steps of: respectively performing simulated single mutations on multiple amino acid sites in the protein sequence of the target protein to obtain multiple simulated mutant proteins, and determining the adaptability evaluation value F of the multiple simulated mutant proteins based on the target protein language model; determining the structural stability value S of the multiple simulated mutant proteins based on the protein folding energy calculation tool; respectively performing normalization processing on the adaptability evaluation value F and the structural stability value S, and respectively performing single-index sorting and comprehensive-index sorting on the normalization processing results; screening the multiple simulated mutant proteins according to the target screening quantity, the single-index sorting result and the comprehensive-index sorting result to obtain the target mutant protein. The present invention can reduce the difficulty of protein design and modification, improve the screening effect of mutant proteins, reduce the computational amount required for protein modification and screening, and reduce the cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of material determination and analysis, and particularly to a method and device for screening mutant proteins and related equipment thereof. Background Art

[0002] Proteins are one of the most diverse and important macromolecules in organisms and play crucial roles in various fields such as the environment, industrial production, medicine, and materials. However, in actual applications, natural proteins often cannot meet complex application requirements. Therefore, some methods and means are needed to perform specific mutation modification designs on proteins. Given many problems such as high experimental design costs, long experimental cycles, and great difficulties in physicochemical property analysis, the method of optimizing and modifying protein mutations designed through simulation calculations is considered to have great potential in the field of protein engineering.

[0003] However, the number of mutant proteins obtained through simulation calculations is large. How to screen out mutant proteins with potential improvement characteristics is an urgent problem to be solved currently. The mutant protein screening schemes in related technologies have problems such as high difficulty in designing and modifying mutant proteins, poor screening effects of mutant proteins, large amounts of calculations required for protein modification and screening, long experimental cycles, and high costs.

[0004] In view of the above problems in related technologies, no effective solutions have been proposed yet. Summary of the Invention

[0005] A method for screening mutant proteins and related equipment provided by an embodiment of the present invention at least solves the problems of high difficulty in designing and modifying mutant proteins, poor screening effects of mutant proteins, large amounts of calculations required for protein modification and screening, long experimental cycles, and high costs in related technologies.

[0006] To solve the above problems, one aspect of an embodiment of the present invention provides a method for screening mutant proteins, including:

[0007] Performing simulated single mutations on multiple amino acid sites in the protein sequence of a target protein to obtain multiple simulated mutant proteins, and performing adaptability prediction on the sequences of the simulated single mutants corresponding to the multiple simulated mutant proteins based on a target protein language model to obtain the adaptability evaluation value F of each simulated single mutant sequence; wherein, any sequence of a simulated single mutant includes the sequence of a single mutant protein corresponding to the corresponding amino acid site;

[0008] Determining the structures of the simulated single mutants corresponding to the multiple simulated mutant proteins respectively, calculating the folding energy change values before and after protein mutation of the multiple simulated single mutant structures respectively based on a protein folding energy calculation tool, and determining the structural stability value S of each simulated single mutant structure based on the folding energy change values;

[0009] The adaptability evaluation value F and the structural stability value S are respectively normalized, and the single-index sorting and the comprehensive-index sorting are respectively performed on the results of the normalization process;

[0010] Multiple simulated mutant proteins are screened according to the target screening quantity, the single-index sorting result, and the comprehensive-index sorting result to obtain the target mutant protein.

[0011] In some of these embodiments, the steps of respectively normalizing the adaptability evaluation value F and the structural stability value S, and respectively performing single-index sorting and comprehensive-index sorting on the results of the normalization process include:

[0012] Based on the min-max normalization method, the adaptability evaluation value F and the structural stability value S are respectively normalized to obtain the first adaptability evaluation normalization value F' and the first structural stability normalization value S';

[0013] The first adaptability evaluation normalization value F' and the first structural stability normalization value S' are respectively sorted to obtain a single-index sorting result including the first adaptability sorting value RF and the first structural stability sorting value RS;

[0014] The sorting average values of the first adaptability sorting value RF and the first structural stability sorting value RS corresponding to each simulated mutant protein are calculated, and the sorting average values of the multiple simulated mutant proteins are sorted to obtain the first comprehensive-index sorting result R.

[0015] In some of these embodiments, the method further includes:

[0016] Based on the Z-Score normalization method, the adaptability evaluation value F and the structural stability value S are respectively normalized to obtain the second adaptability evaluation normalization value F'' and the second structural stability normalization value S'';

[0017] According to the second adaptability evaluation normalization value F'', the second structural stability normalization value S'', the adaptability index weight coefficient, and the structural stability index weight coefficient, the comprehensive-index numerical values of each simulated mutant protein are calculated;

[0018] The comprehensive-index numerical values are sorted to obtain the second comprehensive-index sorting result K.

[0019] In some of these embodiments, the number of protein folding energy calculation tools is N, and the number of structural stability values S calculated for the simulated mutant proteins based on the protein folding energy calculation tools is N; the step of normalizing the structural stability values S corresponding to the multiple simulated mutant proteins includes:

[0020] The N groups of structural stability values Sn respectively calculated by the N protein folding energy calculation tools are respectively normalized, where n ∈ [1, N];

[0021] For the N structural stability normalization values Sn' corresponding to any simulated mutant protein, an average value is calculated to obtain the structural stability normalization result corresponding to the simulated mutant protein.

[0022] In some of these embodiments, the target screening quantity includes the single-index screening quantity and the comprehensive-index screening quantity; the step of screening multiple simulated mutant proteins according to the target screening quantity, the single-index sorting result, and the comprehensive-index sorting result to obtain the target mutant protein includes:

[0023] Determine the first target mutant protein according to the single-index screening quantity, the first adaptability sorting value RF, and the first structural stability sorting value RS;

[0024] After removing the first target mutant protein from the comprehensive-index sorting result, determine the second target mutant protein according to the comprehensive-index screening quantity and the comprehensive-index sorting result. The set of the first target mutant protein and the second target mutant protein is the target mutant protein; wherein, the comprehensive-index sorting result includes the first comprehensive-index sorting result R and / or the second comprehensive-index sorting result K.

[0025] In some of these embodiments, before the step of determining the target mutant protein, the method further includes:

[0026] Remove the simulated mutant proteins in the comprehensive-index sorting result with an adaptability evaluation value F less than 0 and a structural stability value S greater than 0.

[0027] In some of these embodiments, the method further includes the steps of building and parameter tuning of a target protein language model:

[0028] Obtain a basic protein language model, replace the positional encoding layer of the basic protein language model with a learnable embedding encoding, and construct a fully connected neural network layer in the basic protein language model to obtain the target protein language model; wherein, the learnable embedding encoding is used to capture the distance information in the amino acid sequence; the fully connected neural network layer is used to perform weighted calculation on the type and corresponding site of each amino acid;

[0029] Perform unsupervised training on the target protein language model to achieve model parameter tuning processing of the target protein model.

[0030] To solve the above problems, in one aspect of the embodiments of the present invention, a mutant protein screening device is provided, including:

[0031] An adaptability evaluation value calculation module is used to perform simulated single mutations on multiple amino acid sites in the protein sequence of the target protein respectively to obtain multiple simulated mutant proteins, and perform adaptability prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins based on the target protein language model to obtain the adaptability evaluation value F of each simulated single mutant sequence; wherein, any simulated single mutant sequence includes the single mutant protein sequence of the corresponding amino acid site.

[0032] A structure stability value calculation module is used to determine the simulated single mutant structures corresponding to the multiple simulated mutant proteins respectively, calculate the folding energy change values of the multiple simulated single mutant structures before and after protein mutation respectively based on the protein folding energy calculation tool, and determine the structure stability value S of each simulated single mutant structure based on the folding energy change values.

[0033] A data processing module is used to perform normalization processing on the adaptability evaluation value F and the structure stability value S respectively, and perform single-index sorting and comprehensive-index sorting on the normalization processing results respectively.

[0034] A screening module is used to screen the multiple simulated mutant proteins according to the target screening quantity, the single-index sorting result and the comprehensive-index sorting result to obtain the target mutant protein.

[0035] To solve the above problems, in one aspect of the embodiments of the present invention, an electronic device is provided, including: a processor, and a memory storing a program, the program includes instructions, and the instructions, when executed by the processor, cause the processor to execute any one of the above mutant protein screening methods.

[0036] To solve the above problems, in one aspect of the embodiments of the present invention, a non-transitory machine-readable medium storing computer instructions is provided, and the computer instructions are used to cause a computer to execute any one of the above mutant protein screening methods.

[0037] Advantages of the embodiments of the present invention: By performing simulated single mutations on multiple amino acid sites in the protein sequence of the target protein to obtain multiple simulated mutant proteins, and based on the target protein language model, performing adaptability prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins to obtain the adaptability evaluation value F of each simulated single mutant sequence; wherein, any one of the simulated single mutant sequences includes the single mutant protein sequence corresponding to the amino acid site; determining the simulated single mutant structures corresponding to the multiple simulated mutant proteins respectively, calculating the folding energy change values of the multiple simulated single mutant structures before and after protein mutation respectively based on the protein folding energy calculation tool, and determining the structural stability value S of each simulated single mutant structure based on the folding energy change value; performing normalization processing on the adaptability evaluation value F and the structural stability value S respectively, and performing single-index sorting and comprehensive-index sorting on the normalization processing results respectively; screening the multiple simulated mutant proteins according to the target screening quantity, the single-index sorting result and the comprehensive-index sorting result to obtain the target mutant protein. By determining the adaptability evaluation value of the simulated protein sequence based on the protein language model, determining the structural stability value of the simulated mutant protein structure based on the protein folding energy calculation tool, and then screening out the mutant proteins with potential functional improvement by combining the protein multi-scale evaluation index normalization screening method, the technical effects of reducing the difficulty of protein design and modification, improving the screening effect of mutant proteins, reducing the calculation amount required for protein modification and screening, shortening the test period and reducing the cost are achieved.

[0038] Details of one or more embodiments of the present invention are set forth in the following drawings and description to make other features, objects, and advantages of the present invention more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings without creative efforts.

[0040] Figure 1 It is a schematic diagram of the main process of the mutant protein screening method according to an embodiment of the embodiments of the present invention;

[0041] Figure 2 It is a schematic diagram of the main process of the mutant protein screening method according to another embodiment of the embodiments of the present invention;

[0042] Figure 3 It is a schematic diagram of performing data processing based on the min-max normalization method according to an embodiment of the embodiments of the present invention;

[0043] Figure 4 It is a schematic diagram of data processing based on the Z-Score normalization method in an embodiment of the present invention;

[0044] Figure 5 It is a schematic diagram of screening for target mutant proteins based on the determined data in an embodiment of the present invention;

[0045] Figure 6 It is a schematic diagram of the main framework of a mutant protein screening device in an embodiment of the present invention;

[0046] Figure 7 It is a schematic diagram of the structure of an electronic device of the present invention. Detailed implementation manners

[0047] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0048] Related research shows that the function of a protein is determined by its structure, and the protein structure is in turn formed by the folding of its amino acid sequence. Therefore, in protein mutation improvement design, it is often designed based on the protein structure or based on the protein sequence. However, in related technologies, if the mutation modification design is carried out only from the perspective of the protein structure or only from the perspective of the protein sequence, the complete protein properties are often not fully considered. For example, if the mutation modification design is carried out only based on the protein structure, due to its complete structure simulation and optimization process, the computational cost is high and the time consumption is long. Therefore, most studies can only consider some key local regions. And if the mutation modification design is carried out only based on the protein sequence, although it can quickly consider the characteristics of all amino acid sites within the global scope, it lacks the key structural information that leads to the final function of the protein to a certain extent. And if both the protein structure information and the protein sequence information are taken into consideration, the design factors to be considered are very complex. Most of the solutions provided by related technologies require designers to have rich prior knowledge of the physical and chemical properties of the protein to be designed, and at the same time, they also need to have rich practical experience in protein design methods. Moreover, due to the variety of proteins designed, it is relatively difficult to fully understand the properties of each designed protein. Therefore, how to combine the computationally obtained key and important protein structure feature information with the protein sequence design method that can consider the global scope has very important research significance for computer-aided protein engineering design and modification.

[0049] To solve the above problems, an embodiment of the present invention provides a method for screening mutant proteins, as Figure 1 shown. The method for screening mutant proteins mainly includes:

[0050] Step S101, perform simulated single mutations on multiple amino acid sites in the protein sequence of the target protein respectively to obtain multiple simulated mutant proteins, and perform adaptability prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins based on the target protein language model to obtain the adaptability evaluation value F of each simulated single mutant sequence; wherein, any simulated single mutant sequence includes the single mutant protein sequence of the corresponding amino acid site.

[0051] The present invention first regards the protein performance optimization problem (screening out suitable mutant proteins) as a protein adaptability search problem, that is, assuming that the protein mutants with high adaptability (i.e., the above-mentioned mutant proteins) are mutants with potential functional improvements.

[0052] By performing simulated single mutations on multiple amino acid sites in the protein sequence of the target protein (the target protein is the object of protein engineering research, and the simulated mutant protein is the mutant that may appear during the evolution of the target protein), all possible single mutants can be systematically generated and evaluated, and then the effects of mutations at each amino acid site on protein function and stability can be comprehensively understood. Then, through the protein language model to perform adaptability prediction on the simulated single mutant sequence, the adaptability evaluation value S is obtained. Based on this, the mutations that are more likely to be retained during evolution can be identified (the higher the value of the adaptability evaluation value S, the higher the possibility that the simulated mutant is retained during evolution), thus providing valuable candidate mutant proteins for protein engineering.

[0053] Among them, the target protein provided by the embodiment of the present invention is DNA polymerase, preferably Phi29 DNA polymerase or any other DNA polymerase with similar structure and function.

[0054] This is because the functional optimization of DNA polymerase is an important goal in biotechnology and molecular biology research. Through single-site mutation, mutants with higher activity, higher fidelity or longer read length can be screened. At the same time, DNA polymerase also has high stability under extreme conditions such as high temperature and high salt concentration. Single-site mutation can help screen mutants that are more stable under these conditions, thereby expanding its application range.

[0055] According to a specific embodiment of the present invention, Phi29 DNA polymerase is known for its high fidelity and long read length (which means that Phi29 DNA polymerase can continuously synthesize longer DNA fragments during DNA synthesis without frequent termination or errors). It has a very high fidelity during DNA replication and can continuously synthesize thousands of bases without errors. This property makes it very useful in applications such as genome sequencing, cloning, and amplification. Through single-site mutation, precise changes can be made at specific amino acid sites to study the effects of these changes on enzyme activity, stability, and other functions. The precise mutation method provided by the embodiments of the present invention helps to deeply understand the structure-function relationship of the enzyme.

[0056] Applying the mutant protein screening method provided by the embodiments of the present invention to the screening of mutants of DNA polymerase, especially to the screening of mutants of Phi29 DNA polymerase or DNA polymerase with similar structure and function, can screen out Phi29 DNA polymerase mutants with higher fidelity, thereby reducing the error rate during sequencing and cloning. It can also screen out mutants that can still maintain high activity and stability at high temperatures based on the stability of DNA polymerase under high-temperature conditions; based on the high fidelity and long read length characteristics of Phi29 DNA polymerase, Phi29 DNA polymerase mutants with longer read lengths can be screened out to improve the coverage and accuracy of genome sequencing.

[0057] According to a specific embodiment of the present invention, single substitution mutations of the remaining 19 natural amino acids other than the original amino acid can be performed on all amino acid sites of the target protein, and 19*L (L is the length of the target protein sequence) single mutant protein sequences can be obtained as the input for subsequent prediction. The original amino acid refers to the amino acid originally present at a specific position in the target protein sequence. Excluding the original amino acid helps to avoid redundancy, reduce the amount of simulation calculation, and simplify subsequent data analysis. It should be noted that there are many types of amino acids in proteins. The number of amino acids involved in the mutation in the above specific embodiment is not a limitation of the present invention, and the types of amino acids required for simulated mutation can be adjusted according to the actual mutation design and modification requirements of protein engineering. For example, according to a specific embodiment of the present invention, single-site mutations of all amino acid sites of the target protein are simulated to consider the mutation effects of all sites from a global perspective and provide comprehensive adaptability prediction.

[0058] In some of these embodiments, the above method further includes the steps of building and parameter tuning of the target protein language model: obtaining a basic protein language model, replacing the positional encoding layer of the basic protein language model with a learnable embedding encoding, and constructing a fully connected neural network layer in the basic protein language model to obtain the target protein language model; wherein, the learnable embedding encoding is used to capture distance information in the amino acid sequence; the fully connected neural network layer is used to perform weighted calculation on the type and corresponding site of each amino acid; and the target protein language model is trained unsupervised to implement model parameter tuning processing of the target protein model.

[0059] According to a specific embodiment of the present invention, the Tranception model is selected as the basic protein language model. This basic protein language model is based on the Transformer architecture. Compared with other large-scale protein language models, the Tranception model focuses on protein adaptability prediction tasks, which helps to improve the accuracy of the calculated adaptability values. At the same time, in order to further improve the adaptability prediction effect of the model, in the construction of the target protein language model: in addition to using the Tranception model itself, a learnable embedding encoding is used to replace the "Grouped ALiBi" mechanism of the basic model in the position encoding layer of the basic model. This learnable embedding encoding can better capture the distance information in the amino acid sequence. In order to fully consider the distance information of all amino acid sites, a fully connected neural network is also constructed to perform weighted calculations on the type encoding of each amino acid and the corresponding site (specific steps: input the type encoding of each amino acid (for example, the one-hot encoding of 20 natural amino acids) and its position information in the sequence; perform weighted calculations on these inputs through the fully connected layer to generate a new embedding vector; the generated embedding vector will be used as the input of the Transformer model), further enhancing the expression ability of the model. For the parameter tuning step: by obtaining a basic data set, which includes the functional labels of multiple proteins (suitable for evaluating the performance of the protein language model), the target protein language model is trained and optimized based on the zero-shot data set and the partial-shot data set, where different weight coefficients are configured for the prediction results of the zero-shot data set and the partial-shot data set. For example, "ProteinGym" can be selected as the benchmark data set. In multiple zero-shot cases (in the zero-shot case, the model does not see any data of specific proteins and only relies on pre-trained knowledge. A relatively large weight coefficient can be set to maximize the prediction effect under unsupervised conditions) and partial-shot cases (in the partial-shot case, the model can access a small amount of data of specific proteins for fine-tuning the model. A relatively small weight coefficient can be set to supplement the prediction results in the zero-shot case), the target protein language model is used for training and prediction, and the optimal parameters are obtained according to the prediction results. Among them, the weight of the prediction result obtained in the zero-shot case can be set to 0.8 in order to maximize the prediction effect of the target protein language model under unsupervised conditions.

[0060] Step S102: Determine the simulated single mutant structures corresponding to multiple simulated mutant proteins, calculate the folding energy change values of the multiple simulated single mutant structures before and after protein mutation respectively based on the protein folding energy calculation tool, and determine the structural stability value S of each simulated single mutant structure based on the folding energy change values.

[0061] By using protein folding energy calculation tools (such as molecular modeling software Discovery Studio and FoldX) to calculate the change values of folding energy of the simulated single mutant structures before and after mutation, the impact of each mutation on the protein structure stability can be quantitatively evaluated. The structure stability value S can be used to screen out those mutants that still maintain or improve the structure stability after mutation, which helps to quickly screen out candidate mutants with stable structures, reduce the workload of experimental verification, and improve the efficiency and accuracy of mutant protein screening.

[0062] According to a specific embodiment of the present invention, the virtual amino acid mutation function of Discovery Studio can be used to specify single mutations at all amino acid sites, and then based on the calculation formula designed within the Discovery Studio software, the structure stability value S (i.e., ) of each simulated single mutant structure is calculated. The calculation formula is as follows:

[0063]

[0064] Wherein,

[0065]

[0066] In the above formula is the free energy in the protein folding state, is the free energy when the protein unfolds (or the folding is disrupted), wild type represents the wild-type structure of the target protein, and mut represents the mutant structure of the simulated mutant protein. Among them, is used as the structure stability value S. A negative value indicates that the mutant is more stable than the wild type, that is, the mutation has a stabilizing effect; a positive value indicates that the mutant is more unstable than the wild type, that is, the mutation has an unstable effect.

[0067] In some of these embodiments, the steps to determine the simulated single mutant structures corresponding to multiple simulated mutant proteins are: using the molecular modeling software Discovery Studio to perform simulated single mutations at all amino acid sites of the target protein to generate multiple simulated single mutant structures.

[0068] In some of these embodiments, before the step of determining the simulated single mutant structures corresponding to the multiple simulated mutant proteins, the above method further includes: obtaining a protein structure file of the wild-type of the target protein, and performing a preliminary screening on the protein structure file based on structure evaluation factors to obtain a target protein structure file; wherein, the protein structure file of the wild-type includes the three-dimensional structure of the target protein in its natural state (it is also possible to first obtain the three-dimensional structure of the target protein through existing experimental methods (such as X-ray crystallography or NMR), and if the structure observed by the experimental method is not available, then use a high-quality protein structure prediction model (such as AlphaFold) to generate the protein structure file of the wild-type); performing preprocessing on the target protein structure file; wherein, the preprocessing includes simulating the removal of crystallization water and simulating the assignment of force field parameters.

[0069] According to a specific embodiment of the embodiments of the present invention, when performing a preliminary screening on the protein structure file based on structure evaluation factors, the following several structure evaluation factors are mainly considered: resolution (resolution is an important indicator for measuring the accuracy of a protein structure, and a high resolution means more accurate atomic position information), RMSD value (RMSD (root mean square deviation) is used to measure the similarity between two structures. A lower RMSD value indicates a smaller difference between the two structures, and usually a structure with a lower RMSD is selected to ensure the consistency and reliability of the structure), presence or absence of a ligand (the function of some proteins is closely related to the ligand they bind. If the modification target involves the binding of a specific ligand, then it is necessary to select a structure file containing the correct ligand), and ligand quality assessment (even if the structure contains a ligand, it is necessary to evaluate whether the position and conformation of the ligand are reasonable. An incorrect position or conformation of the ligand may affect the subsequent calculation and analysis results), etc. After determining a suitable protein structure file (i.e., the above-mentioned target protein structure file), Discovery Studio can be used to perform preprocessing on the target protein file, including the removal of crystallization water and the assignment of the CHARMm force field (which means parameterizing the protein structure using the CHARMm force field to ensure the accuracy of subsequent calculations), etc.

[0070] In some of these embodiments, due to different calculation methods for the folding energy of the simulated mutant proteins, the structure stability values S calculated by using different protein folding energy calculation tools are slightly different. To further ensure the accuracy of the calculated structure stability values, in a specific embodiment of the embodiments of the present invention, another protein folding energy calculation tool FolfX is also selected to calculate the structure stability value of the mutant (i.e., ) of the following formula:

[0071]

[0072]

[0073] Among them, : The sum of the contributions of the van der Waals interactions of all atoms relative to the solvent; is the energy term The corresponding weight coefficient. and : The differences in solvation energy of non-polar and polar groups from the unfolded state to the folded state, respectively; and are the energy terms and The corresponding weight coefficients. : Refers to the additional stabilizing free energy provided by a single water molecule forming multiple hydrogen bonds (called water bridges) with the protein. : Refers to the free energy difference between the formation of intramolecular hydrogen bonds and the formation of intermolecular hydrogen bonds (with the solvent). : Refers to the electrostatic contribution of charged groups, including the helical dipole moment. : Refers to the entropy cost of fixing the main chain in the folded state, is the energy term The corresponding weight coefficient. : Refers to the entropy cost of fixing the side chain in a specific conformation, is the energy term The corresponding weight coefficient.

[0074] Using Discovery Studio and FoldX software, these two professional molecular modeling softwares are used to perform simulation calculations on the structural stability of the target protein, thus providing high-precision structural information, which helps to more accurately evaluate the mutation effect. Specifically, based on the above molecular modeling software, the wild-type structure of the target protein is first changed to the mutant type, and the change in the folding free energy of the protein structure before and after the mutation is calculated to determine the change in the stability of the mutant, and the structural stability value S is obtained (used to represent the structural stability effect of the mutant and serve as a reference basis for screening mutant sites with improved stability). Through structural simulation, the impact of the mutation on the three-dimensional structure and stability of the protein can be more accurately evaluated.

[0075] It can be understood that in order to further improve the structural stability value S of the calculated simulated mutant, other protein folding energy calculation tools with higher precision can also be used, or the number of protein folding energy calculation tools can be increased.

[0076] Step S103, normalize the adaptability evaluation value F and the structural stability value S respectively, and perform single-index sorting and comprehensive-index sorting on the normalization results respectively.

[0077] The adaptive evaluation value F and the structural stability value S usually have different dimensions and numerical ranges. Through normalization, these indicators of different scales can be unified into the same interval (such as 0 to 1), making the weights between different indicators more balanced, avoiding a certain indicator having too much influence on the comprehensive score due to its large value, and thus facilitating comparison and comprehensive evaluation. By performing single-index sorting and comprehensive-index sorting on the results of normalization respectively, among which single-index sorting can help researchers quickly identify mutants that perform excellently in a specific indicator, while comprehensive-index sorting combines information from multiple aspects and provides a more comprehensive evaluation criterion, which helps to screen out mutants that perform well in multiple aspects.

[0078] At the same time, the normalization process and sorting method can automatically process a large amount of data, reduce the workload of manual analysis, and improve the screening efficiency. Through a systematic sorting method, the optimal mutants can be quickly screened out, accelerating the process of protein engineering projects. Combining multiple indicators for comprehensive evaluation can reduce the bias caused by a single indicator, improve the robustness and reliability of the screening results. Comprehensive evaluation of multiple indicators can better reflect the performance of mutants in practical applications and improve the practicality and credibility of the screening results.

[0079] In some of these embodiments, the steps of respectively normalizing the adaptive evaluation value F and the structural stability value S and respectively performing single-index sorting and comprehensive-index sorting on the results of normalization include: normalizing the adaptive evaluation value F and the structural stability value S respectively based on the min-max normalization method to obtain the first adaptive evaluation normalized value F' and the first structural stability normalized value S'; respectively sorting the first adaptive evaluation normalized value F' and the first structural stability normalized value S' to obtain a single-index sorting result including the first adaptive sorting value RF and the first structural stability sorting value RS; calculating the sorting average value of the first adaptive sorting value RF and the first structural stability sorting value RS corresponding to each simulated mutant protein, and sorting the sorting average values of multiple simulated mutant proteins to obtain the first comprehensive-index sorting result R.

[0080] For the adaptive evaluation values F corresponding to multiple simulated mutant proteins under the adaptive indicator, or the structural stability values S corresponding to multiple simulated mutant proteins under the structural stability indicator. Taking X (referring to the adaptive indicator or the structural stability indicator) as an example, when applying the min-max normalization method, first find the minimum value Xmin and the maximum value Xmax of this indicator, and then use the following formula for normalization:

[0081]

[0082] where is the value after normalization.

[0083] The above steps are a specific implementation manner of the embodiment of the present invention for normalizing the adaptability evaluation index and the structural stability evaluation index, and respectively performing single-index sorting and comprehensive-index sorting on the results of the normalization process. Among them, the maximum-minimum normalization method is a data preprocessing method that scales the data to between [0, 1] by using the maximum and minimum values in the data column.

[0084] Through the minimum-maximum normalization method, the adaptability evaluation value F and the structural stability value S are normalized to between 0 and 1, so that the numerical ranges of different indexes (i.e., the sequence adaptability evaluation index and the structural stability index) are the same, facilitating the comparison of the performances of different indexes on the same scale. By respectively sorting the first adaptability evaluation value F' and the first structural stability value S' obtained after the normalization process, the single-index sorting results (i.e., the first adaptability sorting value RF and the first structural stability sorting value RS) can be obtained, which can intuitively display the performances of each mutant in terms of adaptability evaluation and structural stability. Further, by re-sorting the sorting average values of the multiple single-index sorting results corresponding to each simulated mutant protein, the first comprehensive-index sorting result R is obtained. This method combines the information of multiple indexes, provides a more comprehensive evaluation criterion, reduces the deviation caused by a single index, can better reflect the performance of the mutant in practical applications, improves the practicality and credibility of the screening results, and improves the robustness and reliability of the screening results, which can help researchers screen out mutants that perform well in both adaptability and stability.

[0085] In some of these embodiments, the above method further includes: respectively normalizing the adaptability evaluation value F and the structural stability value S based on the Z-Score normalization method to obtain the second adaptability evaluation normalized value F'' and the second structural stability normalized value S''; calculating the comprehensive-index numerical values of each simulated mutant protein according to the second adaptability evaluation normalized value F'', the second structural stability normalized value S'', the adaptability index weight coefficient, and the structural stability index weight coefficient; and sorting the comprehensive-index numerical values to obtain the second comprehensive-index sorting result K.

[0086] According to a specific implementation manner of the embodiment of the present invention, the specific formula for Z-Score normalization calculation is as follows:

[0087]

[0088] Among them, is the mean value of each index respectively, is the standard deviation of each index respectively, and X is the index value of each simulated mutant.

[0089] The above steps are another specific implementation manner of the embodiments of the present invention for normalizing the adaptability evaluation index and the structural stability evaluation index, and sorting the comprehensive index for the results of the normalization process. Among them, the Z-Score normalization method is also a data normalization method, which converts two or more groups of data into unitless Z-Score scores, making the data standard unified and improving the comparability of the data.

[0090] By performing Z-Score normalization on the adaptability evaluation value F and the structural stability value S, calculating the comprehensive index value based on the normalized values, the adaptability index weight coefficient, and the structural stability index weight coefficient, and then sorting the comprehensive index value to obtain the second comprehensive index sorting result K, the comparability between different indicators can be improved, and a comprehensive evaluation can be carried out. Among them, Z-Score normalization not only unifies the scale but also considers the distribution characteristics of the data. The normalized value reflects the relative position of each mutant in its respective index, which helps to more accurately evaluate its performance. This method not only helps to screen out mutants with potential functional improvement and structural stability but also provides an important reference basis for subsequent experimental design, improving the efficiency and success rate of protein engineering research.

[0091] In some of the embodiments, the number of the above protein folding energy calculation tools is N, and the number of the structural stability values S calculated for the above simulated mutant proteins based on the protein folding energy calculation tools is N; the step of normalizing the structural stability values S corresponding to multiple simulated mutant proteins includes: respectively normalizing the N groups of structural stability values Sn calculated by N protein folding energy calculation tools, where n ∈ [1, N]; for any one of the N normalized structural stability values Sn' corresponding to the simulated mutant protein, taking the average value to obtain the structural stability normalization result corresponding to the simulated mutant protein.

[0092] Different protein folding energy calculation tools may be based on different algorithms and models, so they may have different advantages and limitations in calculating structural stability. By using multiple tools, their respective deficiencies can be complemented, the robustness and reliability of the calculation results can be improved, and thus a more accurate stability evaluation can be provided. Moreover, the structural stability value of each simulated mutant protein is calculated by N tools respectively, which is equivalent to multiple verifications, helping to identify mutants that perform consistently in different tools, thereby improving the credibility of the results.

[0093] Taking the average value of the N normalized structural stability values Sn' of each simulated mutant protein to obtain the final structural stability normalization result realizes the integration of the calculation results of multiple tools, can reduce the random noise in the calculation of a single tool, and improve the stability and accuracy of the evaluation results.

[0094] Step S104: Screen multiple simulated mutant proteins according to the target screening quantity, single-index sorting results, and comprehensive-index sorting results to obtain target mutant proteins.

[0095] By setting a specific target screening quantity, it is possible to ensure that the number of mutants finally screened meets the research requirements, which helps control the scale of the experiment and avoid waste of resources. Based on the single-index sorting results, mutants that perform excellently in a specific aspect can be quickly identified. Combining the comprehensive sorting results of multiple indicators can comprehensively evaluate the overall performance of each mutant and screen out mutants that perform well in multiple aspects.

[0096] In some of these embodiments, the above-mentioned target screening quantity includes a single-index screening quantity and a comprehensive-index screening quantity; the step of screening multiple simulated mutant proteins according to the target screening quantity, single-index sorting results, and comprehensive-index sorting results to obtain target mutant proteins includes: determining the first target mutant protein according to the single-index screening quantity, the first adaptability sorting value RF, and the first structural stability sorting value RS; after removing the first target mutant protein from the comprehensive-index sorting results, determining the second target mutant protein according to the comprehensive-index screening quantity and the comprehensive-index sorting results. The set of the first target mutant protein and the second target mutant protein is the target mutant protein; wherein, the comprehensive-index sorting results include the first comprehensive-index sorting result R and / or the second comprehensive-index sorting result K.

[0097] The above steps provide a specific embodiment of screening target mutant proteins based on the target screening quantity, single-index sorting results, and comprehensive-index sorting results. Through these two screening methods, namely single-index screening and comprehensive-index screening, the mutants are evaluated from different perspectives, improving the accuracy and precision of the screening. Mutants that perform excellently in a single index and mutants that perform well in comprehensive indexes can be obtained simultaneously, thus achieving diversified screening. Specifically, single-index screening may miss mutants that perform well in comprehensive performance, while comprehensive-index screening may ignore mutants that are particularly outstanding in specific attributes. Through two-step screening, the possibility of omission can be reduced.

[0098] On the other hand, due to different selections of data processing methods (normalization processing method, sorting method), multiple comprehensive sorting results (such as the above-mentioned first comprehensive-index sorting result R and second comprehensive-index sorting result K) may be obtained finally. For this comprehensive-index sorting result, after single-index screening, re-screening can be carried out respectively for each comprehensive-index sorting result, and only the same target mutant proteins need to be de-duplicated after the screening is completed.

[0099] In some of these embodiments, before the step of determining the target mutant protein, the above method further includes: eliminating the simulated mutant proteins with an adaptability evaluation value F less than 0 and a structural stability value S greater than 0 in the comprehensive index sorting result.

[0100] The adaptability evaluation value F obtained based on the prediction of the protein language model represents the probability that a mutation can be retained during the evolution of the protein. The higher the value, the more it means that the mutation makes the protein more adaptable to the biological environment. An adaptability evaluation value F less than 0 usually indicates that the mutant is disadvantageous during the evolution process and may reduce the function or stability of the protein. The structural stability value S obtained based on the simulation calculation of the protein folding energy effect represents the impact of the mutation on the structural stability of the protein. Generally, a negative value indicates that the mutation makes the structure more stable, that is, the energy required for folding is less than the energy required for deconstruction. The smaller this value, the less energy is required for protein folding, that is, theoretically the protein structure is more stable, and a structural stability value S greater than 0 indicates that the structure of the protein becomes unstable after the mutation. Therefore, it can be determined that the simulated mutant proteins with an adaptability evaluation value F less than 0 and a structural stability value S greater than 0 are defective mutants. Research shows that mutants that show defects in any index are likely to have poor characterization results in actual experiments.

[0101] Therefore, by eliminating the simulated mutant proteins with positive values in the protein stability calculation results and the mutant proteins with negative values in the adaptability prediction, it helps to improve the screening quality (eliminating bad mutants), improves the reliability of the screening results, optimizes resource allocation, improves the guidance of experimental design, and increases the success rate.

[0102] The embodiments of the present invention also achieve the following technical effects based on the above steps: 1. A method for evaluating the effect of protein mutation is proposed, which combines the calculation of protein folding energy effect (the stability value index of protein structure) and protein adaptability prediction (the adaptability index of protein sequence) simultaneously, providing an effective tool for judging the possible effects brought by protein mutation and a basis for basic research on protein mutation; 2. A method for screening protein mutations by using protein stability and protein adaptability is proposed. By integrating protein structure information, protein sequence and evolutionary information, the screening of protein mutants is realized, solving the problem that the existing methods cannot make full use of the existing information of the target protein; 3. A method for normalizing and screening protein multi-scale evaluation indexes is proposed. By means of the maximum-minimum normalization method, the Z-score standardization method and the means of selecting the longest first and then excluding, the comprehensive screening of protein mutants with potential functional improvement is realized under multiple dimensions and multiple systems. 4. Through multiple protein stability calculation tools, the problem of insufficient robustness of a single method at the present stage is solved. By using different calculation methods, the problem of insufficient consideration of folding energy terms is made up, and the screening accuracy is improved.

[0103] For the above-mentioned mutant protein screening method provided by the embodiments of the present invention, by respectively performing simulated single mutations on multiple amino acid sites in the protein sequence of the target protein to obtain multiple simulated mutant proteins, and based on the target protein language model, performing adaptability prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins to obtain the adaptability evaluation value F of each simulated single mutant sequence; wherein, any simulated single mutant sequence includes the single mutant protein sequence corresponding to the amino acid site; determining the simulated single mutant structures corresponding to the multiple simulated mutant proteins respectively, calculating the folding energy change values of the multiple simulated single mutant structures before and after protein mutation respectively based on the protein folding energy calculation tool, and determining the structural stability value S of each simulated single mutant structure based on the folding energy change value; respectively performing normalization processing on the adaptability evaluation value F and the structural stability value S, and respectively performing single-index sorting and comprehensive-index sorting on the normalization processing results; screening the multiple simulated mutant proteins according to the target screening quantity, the single-index sorting result and the comprehensive-index sorting result to obtain the target mutant protein. By determining the adaptability evaluation value of the simulated protein sequence based on the protein language model, determining the structural stability value of the simulated mutant protein structure based on the protein folding energy calculation tool, and then combining the protein multi-scale evaluation index normalization screening method to screen out the mutant protein with potential functional improvement, the technical effects of reducing the difficulty of protein design and transformation, improving the screening effect of mutant proteins, reducing the calculation amount required for protein transformation and screening, shortening the test period and reducing the cost are achieved.

[0104] The embodiments of the present invention also provide a mutant protein screening method, as Figure 2As shown, the mutant protein screening mainly includes:

[0105] First, perform simulated single mutations on all amino acid sites in the protein sequence of the target protein. After obtaining the protein sequences and protein structures corresponding to multiple simulated mutant proteins respectively, for the protein sequences of the simulated mutant proteins, an unsupervised protein language model is used for adaptive prediction to obtain an adaptive evaluation value F. For the protein sequences of the simulated mutant proteins, the molecular modeling software Discovery Studio and FoldX are used to perform stability calculations of the protein folding energy effect respectively to obtain structure stability values S1 and S2. Then, two numerical normalization algorithms (min-max normalization method and Z-Score standardization method) are used for numerical processing. Finally, determine the required quantity (target screening quantity). First, select the extreme value (i.e., screen the sorting results for a single index) to obtain the first target mutant protein; then screen the sorting results of the comprehensive index to obtain the second target mutant protein. The first target mutant protein and the second target mutant protein constitute the finally screened mutant library.

[0106] According to a specific implementation manner of an embodiment of the present invention, the data processing steps performed based on the min-max normalization method are as Figure 3 shown. First, the adaptive evaluation value F (also called the adaptive prediction value F) obtained based on the protein language model, the structure stability value S1 (also called the stability calculation value S1) calculated based on Discovery Studio, and the structure stability value S2 (also called the stability calculation value S2) calculated based on FoldX are respectively normalized by the min-max normalization method to obtain the first adaptive evaluation normalized value F', the first structure stability normalized values S1' and S2'. Among them, since both S1' and S2' are structure stability indicators, the arithmetic mean of these two values is calculated to obtain S'. The first adaptive evaluation normalized value F' and the first structure stability normalized value S' are respectively sorted to obtain a single-index sorting result including the first adaptive sorting value RF and the first structure stability sorting value RS. Then, calculate the sorting average values of the first adaptive sorting value RF and the first structure stability sorting value RS corresponding to each simulated mutant protein, and sort the sorting average values of multiple simulated mutant proteins to obtain the first comprehensive index sorting result R.

[0107] According to a specific implementation manner of an embodiment of the present invention, the data processing steps performed based on the Z-Score standardization method are as Figure 4As shown below. First, the adaptability evaluation value F (also called the adaptability prediction value F) obtained based on the protein language model, the structural stability value S1 calculated based on Discovery Studio (also called the stability value calculation value S1), and the structural stability value S2 calculated based on FoldX (also called the stability value calculation value S2) are respectively normalized by the Z-Score normalization method to obtain the second adaptability evaluation normalized value F'', the second structural stability normalized value S1'', and S2''. Then, according to the second adaptability evaluation normalized value F'', the second structural stability normalized value S'', the adaptability index weight coefficient, and the structural stability index weight coefficient, the comprehensive index value FS of each simulated mutant protein is calculated, and the second comprehensive index sorting result K is obtained by sorting the comprehensive index values. Among them, the expression for calculating the comprehensive index value FS is as follows:

[0108]

[0109] Among them, the adaptability index weight coefficient is set to 0.6, and the structural index is set to 0.2. It should be noted that the specific numerical setting of the weight coefficient is not a limitation of the present invention and can be adjusted adaptively according to actual experimental conditions and protein modification requirements.

[0110] As Figure 5 shown, the present invention also provides a specific method for screening the target mutant protein, including: first, determining the target screening quantity according to the number of simulated mutant proteins in the simulated mutant library (as shown in the above example, assuming that all amino acid sites of the target protein are individually replaced with the other 19 natural amino acids except the original amino acid, 19*L (L is the length of the target protein sequence) single mutant protein sequences can be obtained), that is Figure 5The total demand S), and then first screen from the single-index sorting results, that is, screen from the special features. Specifically, mutants with the highest ranking in each index and a small index gap with the mutant of the next rank can be regarded as mutants with "special features". This part of the mutants will only consider a single index, and about 20% of the total screening quantity S will be selected. Then screen from the comprehensive sorting index. For example, 30% of the total screening quantity S can be screened from the first comprehensive sorting result R, and 30% of the total screening quantity S can be screened from the second comprehensive sorting result K. At the same time, considering that there will be duplicates in the target mutant proteins screened from the first comprehensive sorting result R and the second comprehensive sorting result K, after deduplication, the total number of the screened target mutant proteins does not meet the target screening quantity (that is, does not meet the total demand). At this time, simulated mutant proteins with an adaptability evaluation value F less than 0 and a structural stability value S greater than 0 among the remaining simulated mutant proteins can be excluded, and then about 10% of the total screening quantity S can be screened again from the first comprehensive sorting result R and the second comprehensive sorting result K. It should be noted that the above values are only examples and do not limit the present invention.

[0111] Based on the above mutant protein screening method provided by the embodiments of the present invention, the embodiments of the present invention also provide a mutant protein screening device, as Figure 6 shown. The mutant protein screening device 600 includes:

[0112] An adaptability evaluation value calculation module 601, configured to perform simulated single mutations on multiple amino acid sites in the protein sequence of the target protein respectively to obtain multiple simulated mutant proteins, and perform adaptability prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins based on the target protein language model to obtain the adaptability evaluation value F of each simulated single mutant sequence; wherein, any simulated single mutant sequence includes the single mutant protein sequence corresponding to the amino acid site.

[0113] By performing simulated single mutations on multiple amino acid sites in the protein sequence of the target protein (the target protein is the object of protein engineering research, and the simulated mutant protein is a mutant that may occur during the evolution of the target protein), all possible single mutants can be systematically generated and evaluated, and then the impact of mutations at each amino acid site on protein function and stability can be comprehensively understood. Then, the adaptability prediction is performed on the simulated single mutant sequence through the protein language model to obtain the adaptability evaluation value S. Based on this, mutations that are more likely to be retained during evolution can be identified (the higher the value of the adaptability evaluation value S, the higher the possibility that the simulated mutant is retained during evolution), thereby providing valuable candidate mutant proteins for protein engineering.

[0114] In some of these embodiments, the above-mentioned mutant protein screening device 600 further includes a model building and parameter tuning module, which is used to: obtain a basic protein language model, replace the positional encoding layer of the basic protein language model with a learnable embedding encoding, and construct a fully connected neural network layer in the basic protein language model to obtain a target protein language model; wherein, the learnable embedding encoding is used to capture the distance information in the amino acid sequence; the fully connected neural network layer is used to perform weighted calculation on the type and corresponding site of each amino acid; perform unsupervised training on the target protein language model to realize the model parameter tuning process of the target protein model.

[0115] The structural stability value calculation module 602 is used to determine the simulated single mutant structures corresponding to multiple simulated mutant proteins, calculate the folding energy change values of the multiple simulated single mutant structures before and after protein mutation based on a protein folding energy calculation tool, and determine the structural stability value S of each simulated single mutant structure based on the folding energy change values.

[0116] Calculating the folding energy change values of the simulated single mutant structures before and after mutation through a protein folding energy calculation tool (such as molecular modeling software Discovery Studio and FoldX) can quantitatively evaluate the impact of each mutation on the structural stability of the protein. The structural stability value S can be used to screen out those mutants that still maintain or improve the structural stability after mutation, which helps to quickly screen out candidate mutants with stable structures, reduce the workload of experimental verification, and improve the efficiency and accuracy of mutant protein screening.

[0117] It can be understood that, in order to further improve the calculated structural stability value S of the simulated mutant, other protein folding energy calculation tools with higher precision can also be used, or the number of protein folding energy calculation tools can be increased.

[0118] The data processing module 603 is used to perform normalization processing on the adaptability evaluation value F and the structural stability value S respectively, and perform single-index sorting and comprehensive-index sorting on the normalization processing results respectively.

[0119] The adaptive evaluation value F and the structural stability value S usually have different dimensions and numerical ranges. Through normalization, these indicators of different scales can be unified into the same interval (such as 0 to 1), making the weights between different indicators more balanced, avoiding an excessive impact of a certain indicator with a large value on the comprehensive score, and thus facilitating comparison and comprehensive evaluation. By separately performing single-index sorting and comprehensive-index sorting on the results of normalization, among which, single-index sorting can help researchers quickly identify mutants that perform excellently in a specific indicator, while comprehensive-index sorting combines information from multiple aspects, provides a more comprehensive evaluation criterion, and helps to screen out mutants that perform well in multiple aspects.

[0120] Among them, the normalization method provided by the embodiments of the present invention may include the min-max normalization method, the Z-Score standardization method, and other methods for unifying data scales.

[0121] In some of these embodiments, the above data processing module 603 is configured to: respectively perform normalization processing on the adaptive evaluation value F and the structural stability value S based on the min-max normalization method to obtain a first adaptive evaluation normalized value F' and a first structural stability normalized value S'; respectively sort the first adaptive evaluation normalized value F' and the first structural stability normalized value S' to obtain a single-index sorting result including a first adaptive sorting value RF and a first structural stability sorting value RS; calculate the sorting average value of the first adaptive sorting value RF and the first structural stability sorting value RS corresponding to each simulated mutant protein, and sort the sorting average values of multiple simulated mutant proteins to obtain a first comprehensive-index sorting result R.

[0122] Through the min-max normalization method, the adaptive evaluation value F and the structural stability value S are normalized between 0 and 1, making the numerical ranges of different indicators (i.e., the sequence adaptive evaluation indicator and the structural stability indicator) consistent, so as to facilitate comparing the performances of different indicators on the same scale. By respectively sorting the first adaptive evaluation value F' and the first structural stability value S' obtained after the normalization processing, a single-index sorting result (i.e., the first adaptive sorting value RF and the first structural stability sorting value RS) can be obtained, which can intuitively display the performances of each mutant in terms of adaptive evaluation and structural stability. Further, by re-sorting the sorting average values of multiple single-index sorting results corresponding to each simulated mutant protein, a first comprehensive-index sorting result R is obtained. This method combines information from multiple indicators, provides a more comprehensive evaluation criterion, reduces the deviation caused by a single indicator, can better reflect the performance of mutants in practical applications, improves the practicality and credibility of the screening results, and improves the robustness and reliability of the screening results, and can help researchers screen out mutants that perform well in terms of adaptability and stability.

[0123] In some of these embodiments, the above data processing module 603 is further configured to: perform normalization processing on the adaptability evaluation value F and the structural stability value S respectively based on the Z-Score normalization method to obtain a second adaptability evaluation normalized value F'' and a second structural stability normalized value S''; calculate the comprehensive index values of each simulated mutant protein according to the second adaptability evaluation normalized value F'', the second structural stability normalized value S'', the adaptability index weight coefficient, and the structural stability index weight coefficient; and sort the comprehensive index values to obtain a second comprehensive index sorting result K.

[0124] By performing Z-Score normalization processing on the adaptability evaluation value F and the structural stability value S, calculating the comprehensive index values based on the normalized values, the adaptability index weight coefficient, and the structural stability index weight coefficient, and then sorting the comprehensive index values to obtain the second comprehensive index sorting result K, the comparability between different indicators can be improved, and a comprehensive evaluation can be carried out. Among them, Z-Score normalization not only unifies the scale but also considers the distribution characteristics of the data. The normalized values reflect the relative positions of each mutant in their respective indicators, which helps to more accurately evaluate their performance. This method not only helps to screen out mutants with potential functional improvements and structural stability but also provides an important reference basis for subsequent experimental design, improving the efficiency and success rate of protein engineering research.

[0125] The screening module 604 is configured to screen a plurality of simulated mutant proteins according to the target screening quantity, the single-index sorting result, and the comprehensive-index sorting result to obtain the target mutant protein.

[0126] By setting a specific target screening quantity, it can be ensured that the number of finally screened mutants meets the research requirements, which helps to control the experimental scale and avoid waste of resources. Based on the single-index sorting result, mutants that perform excellently in a certain specific aspect can be quickly identified. Combining the comprehensive sorting results of multiple indicators, the overall performance of each mutant can be comprehensively evaluated, and mutants that perform well in multiple aspects can be screened out.

[0127] In some of these embodiments, the above screening module 604 is further configured to: determine a first target mutant protein according to the single-index screening quantity, the first adaptability sorting value RF, and the first structural stability sorting value RS; after removing the first target mutant protein from the comprehensive-index sorting result, determine a second target mutant protein according to the comprehensive-index screening quantity and the comprehensive-index sorting result. The set of the first target mutant protein and the second target mutant protein is the target mutant protein; wherein, the comprehensive-index sorting result includes the first comprehensive-index sorting result R and / or the second comprehensive-index sorting result K.

[0128] Through these two screening methods, namely single-index screening and comprehensive-index screening, the evaluation of mutants from different perspectives is realized, improving the precision and accuracy of screening. Mutants with excellent performance in a single index and mutants with good performance in comprehensive indexes can be obtained simultaneously, thus achieving diversified screening. Specifically, single-index screening may miss mutants with good comprehensive performance, while comprehensive-index screening may overlook some mutants that are particularly outstanding in specific attributes. Through two-step screening, the possibility of omission can be reduced.

[0129] In some of these embodiments, the above mutant protein screening device further includes an elimination module. Before the step of determining the target mutant protein, the elimination module is used to: eliminate the simulated mutant proteins with an adaptability evaluation value F less than 0 and a structural stability value S greater than 0 in the comprehensive-index sorting result.

[0130] Eliminating the simulated mutants with positive values in the protein stability calculation results and the mutants with negative values in the adaptability prediction helps improve the screening quality (eliminating bad mutants), enhances the reliability of the screening results, optimizes resource allocation, improves the guidance for experimental design, and increases the success rate.

[0131] The above-mentioned mutant protein screening device provided by the embodiments of the present invention, due to the adoption of an adaptability evaluation value calculation module, is used to perform simulated single mutations on multiple amino acid sites in the protein sequence of the target protein respectively to obtain multiple simulated mutant proteins, and perform adaptability prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins based on the target protein language model to obtain the adaptability evaluation value F of each simulated single mutant sequence; wherein, any simulated single mutant sequence includes the single mutant protein sequence of the corresponding amino acid site; a structure stability value calculation module, which is used to determine the simulated single mutant structures corresponding to the multiple simulated mutant proteins respectively, calculate the folding energy change values of the multiple simulated single mutant structures before and after protein mutation respectively based on the protein folding energy calculation tool, and determine the structure stability value S of each simulated single mutant structure based on the folding energy change value; a data processing module, which is used to perform normalization processing on the adaptability evaluation value F and the structure stability value S respectively, and perform single-index sorting and comprehensive-index sorting on the normalization processing results respectively; a screening module, which is used to screen the multiple simulated mutant proteins according to the target screening quantity, the single-index sorting result and the comprehensive-index sorting result to obtain the target mutant protein. It realizes determining the adaptability evaluation value of the simulated protein sequence based on the protein language model, determining the structure stability value of the simulated mutant protein structure based on the protein folding energy calculation tool, and then screening out the mutant proteins with potential functional improvement by combining the protein multi-scale evaluation index normalization screening method, achieving the technical effects of reducing the difficulty of protein design and transformation, improving the screening effect of mutant proteins, reducing the calculation amount required for protein transformation and screening, shortening the test cycle and reducing the cost.

[0132] The embodiments of the present invention also provide a non-transitory machine-readable medium storing a computer program, wherein the above-mentioned computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention.

[0133] The embodiments of the present invention also provide a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention. Among them, the computer program product should be understood as a software product that mainly realizes the above-mentioned method of the present invention through the computer program.

[0134] The embodiments of the present invention also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The above-mentioned memory stores a computer program that can be executed by the at least one processor, and the above-mentioned computer program, when executed by the at least one processor, is used to cause the electronic device to execute the method of the embodiments of the present invention.

[0135] Reference Figure 7, the structural block diagram of an electronic device such as a server or a client that can be an embodiment of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0136] As Figure 7 shown, the electronic device includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0137] Multiple components in the electronic device are connected to the I / O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709. The input unit 706 can be any type of device that can input information into the electronic device. The input unit 706 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 707 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 708 can include but is not limited to a magnetic disk, an optical disk. The communication unit 709 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0138] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a CPU, a graphics processing unit (GPU), various special artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 702 and / or the communication unit 709. In some embodiments, the computing unit 701 can be configured to execute the above methods in any other suitable way (for example, by means of firmware).

[0139] The computer program for implementing the method of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of the embodiments of the present invention, the machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable signal medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium will include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0141] It should be noted that the term "including" and its variants used in the embodiments of the present invention are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0142] The information and data involved in the embodiments of the present invention (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or reject.

[0143] The various steps described in the method embodiments provided by the embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The protection scope of the present invention is not limited in this regard.

[0144] The term "embodiment" in this specification means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments are referred to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiments.

[0145] The above-described embodiments only represent several implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be understood as a limitation of the protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A method for screening mutant proteins, characterized in that: Includes steps: Simulating single mutations at multiple amino acid sites in the protein sequence of the target protein to obtain multiple simulated mutant proteins, and performing adaptability prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins based on the target protein language model to obtain the adaptability evaluation value F of each of the simulated single mutant sequences; wherein any of the simulated single mutant sequences includes a single mutant protein sequence at the corresponding amino acid site; Determine the simulated single mutant structures corresponding to the multiple simulated mutant proteins respectively, calculate the folding energy change values ​​of the multiple simulated single mutant structures before and after the protein mutation based on the protein folding energy calculation tool, and determine the structural stability value S of each simulated single mutant structure based on the folding energy change value; The adaptability evaluation value F and the structural stability value S are respectively normalized, and the normalized results are respectively sorted by single index and comprehensive index; The plurality of simulated mutant proteins are screened according to the target screening quantity, the single index ranking results and the comprehensive index ranking results to obtain the target mutant protein; It also includes the steps of building the target protein language model and tuning parameters: Obtaining a basic protein language model, replacing the position encoding layer of the basic protein language model with a learnable embedding code, and constructing a fully connected neural network layer in the basic protein language model to obtain the target protein language model; wherein the learnable embedding code is used to capture distance information in the amino acid sequence; and the fully connected neural network layer is used to perform weighted calculation on the type and corresponding site of each amino acid; The target protein language model is subjected to unsupervised training to achieve model parameter tuning processing for the target protein model.

2. The screening method according to claim 1, characterized in that The steps of normalizing the adaptability evaluation value F and the structural stability value S respectively, and performing single index sorting and comprehensive index sorting on the normalized processing results respectively include: Based on the minimum-maximum normalization method, the adaptability evaluation value F and the structural stability value S are respectively normalized to obtain a first adaptability evaluation normalized value F' and a first structural stability normalized value S'; The first adaptability evaluation normalized value F' and the first structural stability normalized value S' are sorted respectively to obtain a single index sorting result including a first adaptability sorting value RF and a first structural stability sorting value RS; The ranking average values ​​of the first adaptability ranking values ​​RF and the first structural stability ranking values ​​RS corresponding to each of the simulated mutant proteins are calculated, and the ranking average values ​​of the multiple simulated mutant proteins are ranked to obtain a first comprehensive index ranking result R.

3. The screening method according to claim 2, characterized in that Also includes the steps: Based on the Z-Score standardization method, the adaptability evaluation value F and the structural stability value S are respectively normalized to obtain a second adaptability evaluation normalized value F'' and a second structural stability normalized value S''; According to the second adaptability evaluation normalized value F'', the second structural stability normalized value S'', the adaptability index weight coefficient, and the structural stability index weight coefficient, the comprehensive index value of each simulated mutant protein is calculated; The comprehensive indicator values ​​are sorted to obtain a second comprehensive indicator sorting result K.

4. The screening method according to claim 2, characterized in that The number of the protein folding energy calculation tools is N, and the number of structural stability values ​​S calculated by the simulated mutant proteins based on the protein folding energy calculation tools is N; The step of normalizing the structural stability values ​​S corresponding to the plurality of simulated mutant proteins comprises: The N groups of structural stability values ​​Sn calculated by N protein folding energy calculation tools are normalized, n∈[1,N]; The N structural stability normalized values ​​Sn' corresponding to any simulated mutant protein are averaged to obtain the structural stability normalized processing result corresponding to the simulated mutant protein.

5. The screening method according to claim 3, characterized in that The target screening number includes a single index screening number and a comprehensive index screening number; the step of screening the multiple simulated mutant proteins according to the target screening number, the single index ranking result and the comprehensive index ranking result to obtain the target mutant protein includes: Determine the first target mutant protein according to the single index screening quantity, the first adaptability ranking value RF, and the first structural stability ranking value RS; After the first target mutant protein is eliminated from the comprehensive index ranking result, the second target mutant protein is determined according to the comprehensive index screening quantity and the comprehensive index ranking result, and the set of the first target mutant protein and the second target mutant protein is the target mutant protein; wherein the comprehensive index ranking result includes the first comprehensive index ranking result R and / or the second comprehensive index ranking result K.

6. The screening method according to claim 5, characterized in that Before the step of determining the target mutant protein, the following steps are also included: In the comprehensive index ranking results, simulated mutant proteins with adaptability evaluation values ​​F less than 0 and structural stability values ​​S greater than 0 are eliminated.

7. A mutant protein screening device, characterized in that: include: A fitness evaluation value calculation module is used to simulate single mutations at multiple amino acid sites in the protein sequence of the target protein to obtain multiple simulated mutant proteins, and to perform fitness prediction on the simulated single mutant sequences corresponding to the multiple simulated mutant proteins based on the target protein language model to obtain the fitness evaluation value F of each of the simulated single mutant sequences; wherein any of the simulated single mutant sequences includes a single mutant protein sequence at the corresponding amino acid site; A structural stability value calculation module, used to determine the simulated single mutant structures corresponding to the multiple simulated mutant proteins, respectively calculate the folding energy change values ​​of the multiple simulated single mutant structures before and after the protein mutation based on the protein folding energy calculation tool, and determine the structural stability value S of each simulated single mutant structure based on the folding energy change value; A data processing module, used to normalize the adaptability evaluation value F and the structural stability value S, and to sort the normalized results by single index and comprehensive index; A screening module, used to screen the multiple simulated mutant proteins according to the target screening quantity, the single index ranking results and the comprehensive index ranking results to obtain the target mutant protein; The steps of building the target protein language model and tuning parameters include: Obtaining a basic protein language model, replacing the position encoding layer of the basic protein language model with a learnable embedding code, and constructing a fully connected neural network layer in the basic protein language model to obtain the target protein language model; wherein the learnable embedding code is used to capture distance information in the amino acid sequence; and the fully connected neural network layer is used to perform weighted calculation on the type and corresponding site of each amino acid; The target protein language model is subjected to unsupervised training to achieve model parameter tuning processing for the target protein model.

8. An electronic device comprising: A processor and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the screening method according to any one of claims 1 to 6.

9. A non-transitory machine-readable medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the screening method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Protein fitness prediction method based on deep learning

    CN115472221A

  • Enzyme thermal stability mutant prediction method and device, electronic equipment and storage medium

    CN118197408A