Multi-modal protein design method, device and system and storage medium thereof
Through the multimodal protein design method, the zero-sample scoring model and cluster dimensionality reduction technology are used to generate efficient and multi-objective optimized protein sequences, solving the problems of inefficient and insufficient generalization capabilities in the existing technology, and achieving more efficient protein design and optimization.
Patent Information
- Application Number
- CN202510413097.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing protein design methods are inefficient, insufficient generalization capabilities, limited experimental throughput, low degree of automation and lack of dynamic optimization capabilities, making it difficult to meet the needs of efficient and widely applicable.
The multimodal protein design method is used to obtain the original sequence information of the target protein, determine the mutable candidate regions, and use the trained zero-sample scoring model to generate the candidate protein sequence, perform cluster dimensionality reduction processing and sampling screening to obtain the target mutant protein sequence.
It improves the efficiency and automation of protein design, enhances the adaptability and optimization capabilities of the model, and can quickly generate high-quality candidate sequences with limited experimental data, achieving the improvement of multi-objective optimization and generalization capabilities.
Smart Images

Figure CN119943206A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of protein design and machine learning technology, and in particular to a multimodal protein design method, device, system and storage medium thereof. Background Art
[0002] As key molecules in organisms, proteins are widely used in biomedicine, industrial catalysis, materials science and other fields. Their unique structure and function make them a hot topic of research, especially in the development of new treatment methods, improving industrial catalytic efficiency and materials science. However, natural proteins often cannot meet specific application requirements, especially in terms of efficiency, stability and functional diversity. Therefore, protein design and optimization have become key ways to solve these problems. Optimizing the design of proteins through artificial synthesis or computer simulation methods can improve their performance to a certain extent and make them more suitable for specific usage environments.
[0003] Traditional protein design methods mainly include random mutation and semi-rational design. Although these methods have been widely used, they face many challenges in actual operation. Random mutation usually relies on screening a large number of mutants. Such methods are inefficient, and the optimization of specific functions often relies on the experience of experts. Although semi-rational design introduces certain structural knowledge, it also faces the problem of difficulty in balancing multi-objective optimization, especially when multiple functions need to be optimized at the same time, it is often impossible to obtain the best design results. In recent years, with the rapid development of artificial intelligence technology, protein design methods based on pre-trained models have gradually emerged. This type of method improves design efficiency in a data-driven way, especially in predicting the structure and function of proteins. However, these methods still have some technical problems that need to be solved, such as insufficient generalization ability, limited experimental throughput, low degree of automation, and lack of dynamic optimization capabilities. To achieve more efficient and more widely applicable protein design, it is urgent to break through these bottlenecks, improve the adaptability and optimization ability of the model, and especially improve design efficiency when experimental data is limited. Summary of the invention
[0004] In a first aspect, the present invention provides a multimodal protein design method, comprising: Obtaining original sequence information of the target protein; wherein the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence; Determining candidate regions of the target protein that can be mutated based on the original sequence information; Based on the candidate mutable region and the original sequence information, a recommended set of candidate protein sequences is obtained using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes; The candidate protein sequence recommendation set is subjected to clustering and dimensionality reduction processing according to the zero-sample scoring model, and is sampled and screened according to latent space representation to obtain a target mutant protein sequence.
[0005] In an optional embodiment, determining the candidate region of the target protein that can be mutated based on the original sequence information includes: Using a protein stability model, single-point saturation mutation scoring is performed on the original sequence information to obtain a stability scoring result; The candidate regions that can be mutated under preset conditions are screened out according to the stability scoring results.
[0006] In an optional embodiment, the calculation method of the stability score result includes: ; where i represents the site index of the i-th amino acid in the protein sequence; S i represents the stability score result of the ith site; μ i represents the average stability score when the i-th site is mutated into 20 amino acids; σ i represents the standard deviation of the stability score when the i-th site is mutated to 20 amino acids; represents the average structural stability score of the ith site within the sliding window; M represents the size of the sliding window; Represents the sum of the structural stability scores from the iM / 2th site to the i+M / 2th site.
[0007] In an optional embodiment, the step of obtaining a candidate protein sequence recommendation set based on the candidate mutable region and the original sequence information using a trained zero-sample scoring model includes: Performing mutations in the candidate mutable region based on the original sequence information to generate different mutant sequences and obtain a mutant sequence library; Using the zero-sample scoring model to score and evaluate each mutant sequence in the mutant sequence library to obtain a comprehensive evaluation result; Based on the comprehensive evaluation results, each of the mutant sequences is optimized using a multi-objective optimization algorithm to obtain an optimized recommended set of candidate protein sequences.
[0008] In an optional embodiment, the scoring and evaluation of each mutant sequence in the mutant sequence library using the zero-sample scoring model to obtain a comprehensive evaluation result includes: Using the zero-sample scoring model to evaluate each mutant sequence in the mutant sequence library, obtaining a scoring evaluation result corresponding to each scoring algorithm, and integrating all the scoring evaluation results to obtain a comprehensive evaluation result; Wherein, the scoring algorithm in the zero-sample scoring model includes at least one of a protein sequence naturalness algorithm, a protein structure naturalness algorithm, a protein stability algorithm, a protein affinity algorithm and a functional activity algorithm; The scoring evaluation results include at least one of sequence naturalness, sequence rationality, structural naturalness, structural stability, protein affinity, functional activity and mutation rationality.
[0009] In an optional embodiment, after obtaining a candidate protein sequence recommendation set based on the candidate mutable region and the original sequence information using a trained zero-sample scoring model, the method further includes: Acquiring experimental data of a verification experiment of the mutant sequences in the mutant sequence library based on the comprehensive evaluation results; Constructing a multimodal target attribute scoring model, and using the experimental data to train the multimodal target attribute scoring model; The multimodal target attribute scoring model is iteratively optimized based on the verification experiment to obtain an updated multimodal target attribute scoring model, and an optimized candidate protein sequence recommendation set is obtained based on the updated multimodal target attribute scoring model.
[0010] In an optional implementation, the iterative optimization includes: Evaluating the candidate protein sequence recommendation set according to the multimodal target attribute scoring model and the zero-sample scoring model to obtain a scoring evaluation result; According to the scoring and evaluation results, candidate sequences are screened and a candidate sequence set is constructed; Acquire experimental data of a verification experiment corresponding to each candidate sequence in the candidate sequence set, and update the multimodal target attribute scoring model using the experimental data; Using the updated multimodal target attribute scoring model and the zero-sample scoring model, the candidate protein sequence recommendation set is evaluated to obtain a new scoring evaluation result; Using the new scoring evaluation results to screen out target sequences, construct a new candidate protein sequence recommendation set, and obtain new experimental data corresponding to the new candidate protein sequence recommendation set; Based on the new experimental data, the iterative optimization of the multimodal target attribute scoring model is evaluated to facilitate completing the iterative optimization.
[0011] In an optional embodiment, the evaluating the iterative optimization of the multimodal target attribute scoring model based on the new experimental data to complete the iterative optimization includes: Determining whether the multimodal target attribute scoring model reaches a preset optimization goal based on the new experimental data; If yes, it is determined that the iterative optimization of the multimodal target attribute scoring model is completed; If not, it is determined that the iterative optimization of the multimodal target attribute scoring model is not completed, and the evaluation of the candidate protein sequence recommendation set according to the multimodal target attribute scoring model and the zero-sample scoring model is returned to obtain a scoring evaluation result.
[0012] Determining whether the multimodal target attribute scoring model reaches a preset optimization goal based on the new experimental data; If yes, it is determined that the iterative optimization of the multimodal target attribute scoring model is completed; If not, it is determined that the iterative optimization of the multimodal target attribute scoring model is not completed, and the evaluation of the candidate protein sequence recommendation set according to the multimodal target attribute scoring model and the zero-sample scoring model is returned to obtain a scoring evaluation result.
[0013] In an optional embodiment, the candidate protein sequence recommendation set is subjected to clustering dimensionality reduction processing according to the zero-sample scoring model, and sampling and screening are performed according to latent space representation to obtain a target mutant protein sequence, including: Embedding the protein sequences in the candidate protein sequence recommendation set by the zero-shot scoring model or the multimodal target attribute scoring model to obtain high-dimensional representation information in a latent space; Performing clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space to obtain low-dimensional representation information; In the low-dimensional space, sampling is performed according to the distribution of the low-dimensional evidence information to generate a sampling screening result; wherein the sampling screening result includes the screened candidate mutation sequences; The target mutant protein is determined according to the sampling and screening results.
[0014] In an optional implementation, the clustering and dimensionality reduction processing of the high-dimensional representation information in the latent space includes: Using a clustering algorithm to perform clustering processing on the high-dimensional representation information in the latent space, and grouping similar sequences into different clusters; The high-dimensional representation information after clustering is mapped to a low-dimensional space by using a dimensionality reduction method to obtain the low-dimensional representation information; wherein the dimensionality reduction method includes at least one of principal component analysis, t-random neighbor embedding and UMAP nonlinear dimensionality reduction technology.
[0015] In a second aspect, the present invention provides a multimodal protein design device, comprising: A data acquisition module, used to acquire the original sequence information of the target protein; wherein the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence; A candidate region module, used to determine the candidate regions of the target protein that can be mutated based on the original sequence information; A sequence recommendation module, for obtaining a candidate protein sequence recommendation set based on the candidate mutable region and the original sequence information using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes; The sampling generation module performs clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-sample scoring model, and performs sampling screening according to the latent space representation to obtain the target mutant protein sequence.
[0016] In a third aspect, the present invention provides a computer device, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the multimodal protein design method described in any one of the aforementioned embodiments.
[0017] In a fourth aspect, the present invention provides a computer storage medium storing a computer program, which, when executed on a processor, implements the multimodal protein design method according to any one of the aforementioned embodiments.
[0018] The present invention provides a multimodal protein design method, device, system and storage medium thereof, wherein the multimodal protein design method realizes high efficiency, automation and multi-objective optimization capability of protein design by integrating multiple advanced technical means.
[0019] This method realizes the full-process automated design from obtaining the original sequence information of the target protein to obtaining the target mutant protein sequence, without the need for human intervention. It greatly reduces the time and effort required for traditional protein design that relies on expert experience and manual screening, significantly improves design efficiency, and can complete multiple rounds of optimization in a short period of time, accelerating the improvement of protein function.
[0020] Through the different target attribute scoring algorithms included in the zero-sample scoring model, this method can simultaneously evaluate multiple key properties of proteins, such as sequence rationality, structural stability, functional activity, etc., avoiding the problem of deterioration of other properties that may be caused by single-target optimization in traditional methods, and achieving balanced optimization among multiple objectives. In addition, by using clustering dimensionality reduction processing and sampling screening, the Pareto optimal solution set can be searched in the multi-dimensional target space, providing a more comprehensive and reliable optimization solution for practical applications.
[0021] In the absence of prior knowledge or experimental data, the zero-shot scoring model can provide reasonable mutation suggestions based on large-scale unsupervised learning, thereby achieving a rapid start of protein design in the cold start phase. With the accumulation of verification experimental data, a self-learning module is introduced to use the accumulated experimental data to construct a protein target attribute scoring model based on a multimodal pre-training model, thereby guiding protein optimization towards higher performance goals. Based on a multimodal pre-training model and an efficient parameter fine-tuning architecture, this model can stably predict target attributes even with only a small amount of experimental data, significantly improving the prediction effect and iteration efficiency of the model. At the same time, this method exhibits strong generalization capabilities by integrating the sequence, structure, and evolutionary information of proteins, and can adapt to different protein families and mutation regions. It can also show good performance even in unseen mutation regions or new protein families, providing a wider range of applications for protein design.
[0022] The target attribute scoring model can continuously optimize itself based on its self-learning ability as experimental data accumulates, and automatically adjust model parameters to better adapt to new data and task requirements. This self-learning mechanism not only improves the prediction accuracy of the model, but also enhances its adaptability and stability under different experimental conditions. Through continuous learning, the model can continuously absorb new knowledge, thereby providing more accurate and valuable suggestions in subsequent protein design, further promoting the efficiency and quality of protein design.
[0023] Through clustering and dimensionality reduction, the high-dimensional protein sequence space is reduced to a low-dimensional space and uniformly sampled. This method not only improves computational efficiency, but also ensures the diversity of the mutant sequences in function and structure, avoiding the local optimal problem that may occur in traditional methods. In addition, the generated candidate mutant sequences can cover different functional and structural characteristics, providing a wealth of choices for experiments, making full use of limited experimental resources, and improving the efficiency and success rate of experiments.
[0024] The modules of this method are designed to be pluggable and can be flexibly adjusted according to different design requirements and experimental conditions, which has strong flexibility and scalability. Whether it is high-throughput or low-throughput experimental conditions, this method can efficiently perform protein design and optimization and adapt to a variety of protein design scenarios.
[0025] Through multiple rounds of iterative optimization, this method can significantly improve the functional performance of proteins and optimize the overall performance of proteins, making them more reliable and efficient in practical applications. For example, in the fields of industrial catalysis and biomedicine, optimized proteins can show higher activity, stability and specificity, which is better than traditional artificial design paths.
[0026] In summary, the multimodal protein design method provided in this application shows significant advantages in high efficiency, multi-objective optimization capability, small sample adaptability, generalization ability, optimization efficiency, flexibility and scalability, and can be widely used in protein design and optimization in biomedicine, industrial catalysis, materials science and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present application and should not be regarded as limiting the scope of protection of the present application. For ordinary technicians in this field, other related drawings can also be obtained based on these drawings without creative work.
[0028] Figure 1 A schematic diagram of the structure of the hardware operating environment involved in the embodiment of the multimodal protein design method of the present invention; Figure 2 Schematic diagram of the process of Example 1 of the multimodal protein design method of the present invention; Figure 3 This is a schematic diagram of a detailed process of step S200 in Example 2 of the multimodal protein design method of the present invention; Figure 4 Schematic diagram of the overall process including the refinement of step S300 in Example 3 of the multimodal protein design method of the present invention; Figure 5 Schematic diagram of the process of Example 4 of the multimodal protein design method of the present invention; Figure 6 This is a schematic diagram of a detailed process of step S700 in Example 4 of the multimodal protein design method of the present invention; Figure 7 FIG. 5 is a schematic diagram of a detailed process of step S400 in Example 5 of the multimodal protein design method of the present invention; Figure 8 This is a schematic diagram of a detailed process of step S420 in Example 5 of the multimodal protein design method of the present invention; Fig. 9 This is a schematic diagram of the model performance evaluation results in Example 6 of the multimodal protein design method of the present invention; Fig.10 Schematic diagram of the module connections of the multimodal protein design device of the present invention. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.
[0030] The components of the embodiments of the present application generally described and shown in the drawings herein may be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0031] Hereinafter, the terms "including", "having" and their cognates, which may be used in various embodiments of the present application, are intended only to indicate specific features, numbers, steps, operations, elements, components, or combinations of the foregoing items, and should not be understood as first excluding the existence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations of the foregoing items.
[0032] Furthermore, the terms “first”, “second”, “third”, etc. are merely used for distinguishing descriptions and are not to be understood as indicating or implying relative importance.
[0033] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meanings as those generally understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meanings as the contextual meanings in the relevant technical field and will not be interpreted as having idealized meanings or overly formal meanings unless clearly defined in the various embodiments of the present application.
[0034] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0035] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0036] like Figure 1 , which is a schematic diagram of the structure of the hardware operating environment of the terminal involved in the embodiment of the present invention.
[0037] The multimodal protein design system of the embodiment of the present invention can be a PC, or a mobile terminal device such as a smart phone, a tablet computer or a portable computer. The multimodal protein design system may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005 and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen, an input unit such as a keyboard, and a remote control. The user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory, such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001. Optionally, the multimodal protein design system may also include an RF (Radio Frequency) circuit, an audio circuit, a WiFi module, and the like. In addition, the multimodal protein design system can also be configured with other sensors such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., which will not be described in detail here.
[0038] Those skilled in the art will understand that Figure 1 The multimodal protein design system shown in the figure does not constitute a limitation thereof, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Figure 1 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a data interface control program, a network connection program, and a multimodal protein design program.
[0039] In summary, the method provided by the present invention improves the accuracy of fall prediction by real-time monitoring of key status data, and can quickly respond to high-risk fall states and perform protective actions. This method reduces robot structural damage, reduces safety risks to surrounding personnel, improves the environmental adaptability and economy of the robot, and enhances its stability and reliability in a changing environment.
[0040] Embodiment 1: Reference Figure 2 This embodiment provides a multimodal protein design method, comprising: Step S100, obtaining original sequence information of the target protein; wherein the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence.
[0041] As mentioned above, the target protein refers to a specific protein that researchers hope to design, modify or optimize during the protein design or optimization process. It can be a known natural protein or a protein with specific functional requirements.
[0042] As mentioned above, raw sequence information refers to the basic information of the target protein, including its amino acid sequence and corresponding PDB structure information. This information is the starting point for protein design and optimization.
[0043] The amino acid sequence provides the basic composition information of proteins. The amino acid sequence refers to the linear arrangement order of amino acids in a protein. Proteins are polypeptide chains composed of amino acids connected by peptide bonds. The amino acid sequence determines the primary structure of the protein. The amino acid sequence is the basis of protein design and determines the chemical properties and possible structures of the protein. By modifying the amino acid sequence, the function, stability and other characteristics of the protein can be changed.
[0044] PDB (Protein Data Bank) structural information refers to the three-dimensional spatial structural information of proteins, which is usually stored in the PDB file format. PDB files contain the coordinate information of each atom in the protein and describe the spatial conformation of the protein. In the PDF structural information, each line represents an atom in the protein, including information such as the atom name, residue name, chain identifier, residue number, atomic coordinates, etc. PDB structural information helps to understand the spatial conformation of proteins, which is crucial for predicting the function, stability and interaction of proteins with other molecules. In protein design, PDB structural information can be used to evaluate the impact of mutations on protein structure and ensure that the designed protein remains stable and functional in three-dimensional space.
[0045] In this step, basic information of the target protein is obtained, including its amino acid sequence and corresponding PDB structure information.
[0046] Specifically, the amino acid sequence of the target protein can be obtained by database query or experimental determination. If the target protein already has PDB structure information, it can be directly obtained; if not, a structure prediction algorithm (such as xTrimoStructure) can be called to obtain the predicted structure PDB. Thus, the amino acid sequence and PDB structure information of the target protein are obtained, which provides a basis for the subsequent determination and design of the mutation region.
[0047] This step ensures that the starting point of the design is accurate and complete, and provides a reliable data basis for subsequent multimodal analysis and optimization. Relevant information can be obtained from public protein databases (such as UniProt, PDB), or the protein structure can be determined using experimental techniques (such as X-ray crystallography, nuclear magnetic resonance, etc.).
[0048] Step S200, determining the candidate mutable regions of the target protein according to the original sequence information.
[0049] Based on the amino acid sequence and PDB structure information of the target protein, determine which regions of amino acids can be mutated.
[0050] For example, a protein stability prediction model can be used to score single-point saturation mutations in a sequence, and the structural stability score can be calculated for all sites based on the model. The mutation effect of each site is calculated using a formula, and regions with less mutation effect on stability and better overall stability are considered. At the same time, a sliding window smoothing process is performed on local regions to enhance the reliability of the score, thereby obtaining candidate regions that can be mutated. At the same time, users can also customize candidate mutation regions according to their needs.
[0051] This step can accurately locate the area suitable for mutation, avoiding loss of protein function or structural instability caused by blind mutation. Implementation method: Use existing protein stability prediction models, or train new models based on machine learning methods to evaluate the impact of mutations on protein stability.
[0052] Step S300, based on the candidate mutable region and the original sequence information, a recommended set of candidate protein sequences is obtained using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes.
[0053] As mentioned above, in protein design and optimization, the target attributes are multifaceted, including but not limited to functional activity, sequence naturalness, structural naturalness, protein stability, affinity, etc. These attributes jointly determine the performance and applicability of the protein and are key factors that need to be considered and optimized during the protein design process.
[0054] As mentioned above, scoring algorithms are computational methods used to evaluate the potential performance of protein sequences on specific target properties. These algorithms predict the performance of proteins through mathematical models or machine learning models based on the sequence information, structural information or evolutionary information of proteins. For example, the sequence rationality scoring algorithm evaluates whether the protein sequence conforms to known biochemical laws or evolutionary conservation. For example, multiple sequence alignment (MSA) information is used to determine whether a mutation is evolutionarily reasonable; the structural stability scoring algorithm is used to evaluate the stability of the protein structure; the effect of mutations on stability is predicted by calculating the change in protein free energy (ΔG) before and after mutation; the functional activity scoring algorithm predicts the functional activity of proteins; machine learning-based models are used to predict the catalytic efficiency of enzymes or the binding affinity of antibodies, and so on.
[0055] The Zero-Shot Scoring Model mentioned above is a special machine learning model that can directly evaluate and score protein sequences without experimental data.
[0056] This model is based on a pre-trained protein language model (such as ESM-2, PGLM) or a structural model (such as RDE), and uses the knowledge learned from large-scale protein data to evaluate new protein sequences.
[0057] Zero-Shot analysis, scoring and design is a highly innovative method. It can directly use the pre-trained model to analyze and score the original sequence information of the target protein without any experimental data, thereby generating a recommended set of candidate protein sequences. This method breaks through the reliance of traditional design methods on a large amount of experimental data, allowing protein design to start quickly in the cold start phase without prior knowledge, greatly improving design efficiency. In this way, researchers can evaluate and screen a large number of potential protein sequences in a very short time, quickly lock in mutant protein sequences with potential functions, and provide a solid foundation for subsequent experimental verification and optimization. This zero-shot design capability not only accelerates the process of protein design, but also opens up new ways to explore new protein functions and applications.
[0058] The zero-shot scoring model used in this step does not require experimental data and can provide reasonable mutation suggestions in the cold start phase (i.e., without experimental data); it can evaluate a large number of candidate sequences in a short time and improve design efficiency. Specifically, it can encode protein sequences based on pre-trained protein language models (such as ESM-2) or structural models (such as RDE), combine multiple sequence alignment (MSA) information, and use models such as MSA Transformer to capture evolutionary information, and score protein sequences on different target attributes by fine-tuning or directly using the zero-shot capabilities of pre-trained models.
[0059] In this step, a zero-shot scoring model is used to generate a recommended set of candidate protein sequences based on the candidate regions that can be mutated and the original sequence information. The zero-shot scoring model can evaluate the potential mutations in the mutable regions based on the pre-trained protein language model and structure model, as well as the evolutionary information model of multiple sequence alignment. The model contains scoring algorithms for different target attributes, such as sequence naturalness, structural stability, functional activity, etc. These algorithms are used to score the mutated sequences to obtain a set of recommended candidate protein sequences. Each sequence corresponds to a score, reflecting its potential performance on different target attributes. The zero-shot scoring model can directly evaluate the mutated sequences without experimental data, which greatly improves the design efficiency and is especially suitable for the cold start stage. Implementation method: Existing pre-trained protein language models and structure models can be used, and fine-tuned on this basis or directly use their zero-shot capabilities.
[0060] Step S400, clustering and dimensionality reduction processing is performed on the candidate protein sequence recommendation set according to the zero-sample scoring model, and sampling and screening are performed according to the latent space representation to obtain the target mutant protein sequence.
[0061] It should be noted that protein sequence space is usually high-dimensional, because each amino acid position can have 20 different amino acid choices, which leads to a very high dimensionality of the sequence space. Direct optimization and screening in high-dimensional space will face: the computational cost in high-dimensional space is extremely high, especially when generating and evaluating a large number of candidate sequences; data sparsity, the data points in high-dimensional space are sparse, and it is difficult to find representative areas; and high-dimensional data is also difficult to intuitively understand and visualize.
[0062] The purpose of clustering and dimensionality reduction is to map the high-dimensional protein sequence space to a low-dimensional space while retaining key information. This makes calculation and optimization in low-dimensional space more efficient; data points in low-dimensional space are denser, making it easier to find representative areas; and low-dimensional space (such as two-dimensional or three-dimensional) is easier to visualize and understand.
[0063] As mentioned above, latent space is the representation after mapping high-dimensional data to low-dimensional space through dimensionality reduction techniques (such as PCA, t-SNE, Autoencoder, etc.). In protein design, latent space is usually a low-dimensional embedding space generated by a pre-trained protein language model or structural model. These models compress complex high-dimensional data into low-dimensional space by learning the intrinsic laws of protein sequences or structures. Latent space representation refers to the representation of protein sequences in latent space. It is a low-dimensional vector that contains the key features and information of protein sequences.
[0064] In this step, in order to adapt to different experimental fluxes and improve design efficiency, the recommended set of candidate protein sequences is further screened and optimized, including clustering and dimensionality reduction, so as to obtain the final target mutant protein sequence. Specifically, the latent space representation of the zero-shot scoring model can be used to cluster and reduce the dimensionality of the candidate protein sequences. Through clustering, similar sequences are grouped together to reduce duplication and redundancy; then uniform sampling is performed in the latent space after dimensionality reduction to generate a set of uniformly distributed candidate mutant sequences; finally, the optimal mutant protein sequence is selected based on the latent space representation and scoring results.
[0065] Through the processing method in this step, optimized target mutant protein sequences can be obtained, which perform well in multiple target attributes and have high diversity and innovation. This step can effectively reduce the amount of calculation and improve the screening efficiency through clustering dimensionality reduction and uniform sampling, while ensuring the quality and diversity of the mutant sequences.
[0066] Through the above steps, this multimodal protein design method can efficiently and automatically generate target mutant protein sequences with excellent performance, which is suitable for a variety of application scenarios.
[0067] In summary, in this embodiment, the multimodal protein design method starts with obtaining the original sequence information of the target protein, and uses a zero-sample scoring model and a scoring algorithm for multiple target attributes to quickly generate a recommended set of candidate protein sequences in the absence of experimental data. Through sampling and screening of clustering dimensionality reduction and latent space representation, this method can balance and optimize key properties such as sequence rationality, structural stability, and functional activity of proteins. At the same time, the integration of sequence, structure, and evolutionary information enables it to have a strong generalization ability and adapt to different protein families and mutation regions. In addition, the method efficiently utilizes experimental resources to ensure the diversity and global optimization ability of mutant sequences. The entire process is automated and flexible and scalable, which can significantly improve protein performance and optimize the overall design effect, making it more reliable and efficient in practical applications.
[0068] Embodiment 2: Reference Figure 3 This embodiment provides a multimodal protein design method. Based on the above embodiment 1, the step S200 determines the candidate regions of the target protein that can be mutated according to the original sequence information, including: Step S210, using a protein stability model, single-point saturation mutation scoring is performed on the original sequence information to obtain a stability scoring result.
[0069] As mentioned above, the protein stability model is a computational tool used to evaluate the ability of a protein to maintain stability in its three-dimensional structure and function after a change in its amino acid sequence (such as a mutation). The model predicts the effect of mutations on protein stability by analyzing the sequence information, structural information, and possible physical and chemical properties of the protein, thereby helping researchers select appropriate mutation sites during protein design and optimization.
[0070] In the multimodal protein design method, the role of the protein stability model is to determine the candidate regions of the target protein that can be mutated. By evaluating the effect of each amino acid site on protein stability after mutation, the model can help screen out those regions that can still maintain high stability after mutation, thereby improving the success rate and efficiency of protein design.
[0071] Furthermore, the calculation method of the stability score result includes: ; where i represents the site index of the i-th amino acid in the protein sequence; S i represents the stability score result of the ith site; μ i represents the average stability score when the i-th site is mutated into 20 amino acids; σ i represents the standard deviation of the stability score when the i-th site is mutated to 20 amino acids; represents the average structural stability score of the ith site within the sliding window; M represents the size of the sliding window; Represents the sum of the structural stability scores from the iM / 2th site to the i+M / 2th site.
[0072] Step S220, screening out candidate regions that can be mutated under preset conditions according to the stability scoring result.
[0073] The "preset conditions" mentioned above refer to the conditions used to screen out amino acid sites suitable for mutation when determining candidate regions for mutation. These conditions are usually thresholds or ranges set based on stability scoring results to determine whether a site is suitable for mutation.
[0074] The preset conditions are one or more screening criteria set based on the stability score results calculated by the protein stability model to determine which amino acid sites can be used as candidate regions for mutation. These conditions are usually intended to ensure that the mutated protein can achieve the expected performance optimization while maintaining its functional and structural stability.
[0075] In this step, single-point saturation mutation is performed on each amino acid site of the target protein, that is, each site is mutated to the other 19 amino acids. A protein stability model (such as xTrimoStability) is used to score the stability of each mutated sequence to obtain the stability score results when each site is mutated to different amino acids.
[0076] Based on the stability score results and preset conditions (such as the stability score is higher than a certain threshold or within a certain range), candidate regions that can be mutated are screened out to obtain amino acid site regions in the target protein that are suitable for mutation. These regions have less impact on the overall stability of the protein after mutation and are more likely to produce mutants with optimized performance.
[0077] In this step, the candidate regions for mutation screened out by stability scoring are more targeted and reduce the risk of blind mutation. In addition, the scope of mutation is narrowed, which reduces the computational cost and experimental workload of subsequent design and screening. Existing protein stability prediction models are used, such as the xTrimoStability model based on deep learning. Single-point saturation mutation simulation is performed on each amino acid site in the target protein sequence and input into the stability model for scoring. The candidate regions for mutation are screened out based on the scoring results and preset conditions (such as thresholds).
[0078] Embodiment 3: Reference Figure 4 This embodiment provides a multimodal protein design method. Based on the above embodiment 1, the step S300, based on the candidate regions that can be mutated and the original sequence information, uses a trained zero-sample scoring model to obtain a candidate protein sequence recommendation set, including: Step S310, performing mutations in the candidate mutable region based on the original sequence information to generate different mutant sequences and obtain a mutant sequence library.
[0079] In this step, the amino acid sequence of the target protein is mutated within the determined candidate regions to generate a series of different mutant sequences, forming a mutant sequence library.
[0080] For the amino acid sites in the target protein that have been determined to be suitable for mutation, mutation operations are performed on each site. The mutation operation can be a single-point mutation, multiple-point mutation, or random mutation, etc. For example, the amino acid at a certain site is replaced with any of the other 19 amino acids.
[0081] Through the above mutation operation, a large number of different mutation sequences are generated to form a mutation sequence library. As a result, a mutation sequence library containing a variety of mutation sequences is obtained, which is the basis for subsequent evaluation and optimization.
[0082] This step can generate a large number of different mutation sequences, increasing the possibility of finding the optimized sequence; mutations are only performed within the candidate regions that can be mutated, reducing unnecessary mutation attempts and improving efficiency.
[0083] Step S320, using the zero-sample scoring model to score and evaluate each mutant sequence in the mutant sequence library to obtain a comprehensive evaluation result.
[0084] In this step, the trained zero-shot scoring model is used to score and evaluate each mutant sequence in the mutant sequence library, evaluate the performance of these sequences on different target attributes, and obtain a comprehensive evaluation result.
[0085] First, input the mutation sequence, and input each mutation sequence in the mutation sequence library into the zero-shot scoring model; then, the zero-shot scoring model scores each mutation sequence on different target attributes (such as sequence naturalness, structural stability, functional activity, etc.) according to its internal scoring algorithm; finally, integrate the scoring results of each mutation sequence on different target attributes to obtain a comprehensive evaluation result of each sequence. As a result, the scoring evaluation results of each mutation sequence on different target attributes and the comprehensive evaluation results are obtained.
[0086] In this step, multiple target attributes are comprehensively considered to comprehensively evaluate the performance of the mutant sequence; the zero-sample scoring model can quickly evaluate without experimental data, thereby improving design efficiency.
[0087] Specifically, a pre-trained model can be used: a pre-trained protein language model (such as ESM-2) or a structural model (such as RDE) is used as the basis for the zero-shot scoring model. Multiple scoring algorithms are integrated into the model to score different target attributes separately. The scoring results of different target attributes can be integrated into a comprehensive evaluation result using methods such as weighted average and voting mechanism.
[0088] Step S330: Based on the comprehensive evaluation result, each of the mutant sequences is optimized using a multi-objective optimization algorithm to obtain an optimized recommended set of candidate protein sequences.
[0089] In this step, based on the comprehensive evaluation results, a multi-objective optimization algorithm is used to optimize the mutant sequence, screen out the mutant sequence with the best performance, and form a recommended set of candidate protein sequences.
[0090] Multi-Objective Optimization Algorithm (MOO) is an algorithm used to find the best balance between multiple conflicting objectives. It aims to find a set of "Pareto Optimal Solutions" that achieve the best trade-off between different objectives, that is, no solution is better than other solutions in all objectives.
[0091] The Pareto Optimal Solution mentioned above means that if a solution improves on a certain objective, it will inevitably sacrifice other objectives, then this solution is the Pareto Optimal Solution. The set of all Pareto Optimal Solutions is called the Pareto Front.
[0092] In this step, through the multi-objective optimization algorithm, it is possible to balance multiple target attributes and find the mutation sequence with the best performance; it can be flexibly adjusted according to different design requirements and constraints to adapt to various application scenarios.
[0093] Specifically, a multi-objective optimization algorithm such as the non-dominated sorting genetic algorithm (NSGA-III) can be used. The scoring evaluation results in the comprehensive evaluation results are used as the objective function, and the weights and constraints are designed according to actual needs.
[0094] Through the above steps, we finally obtain an optimized set of candidate protein sequence recommendations. These sequences perform well in multiple target attributes and are candidates for subsequent experimental verification and application.
[0095] In some embodiments, the step S320, using the zero-sample scoring model to score and evaluate each mutant sequence in the mutant sequence library to obtain a comprehensive evaluation result, includes: Step S321, using the zero-sample scoring model to evaluate each mutant sequence in the mutant sequence library, obtaining a scoring evaluation result corresponding to each scoring algorithm, and integrating all the scoring evaluation results to obtain a comprehensive evaluation result.
[0096] The scoring algorithm in the zero-sample scoring model includes at least one of a protein sequence naturalness algorithm, a protein structure naturalness algorithm, a protein stability algorithm, a protein affinity algorithm and a functional activity algorithm.
[0097] The scoring evaluation results include at least one of sequence naturalness, sequence rationality, structural naturalness, structural stability, protein affinity, functional activity and mutation rationality.
[0098] It should be noted that in the multimodal protein design method, the scoring algorithm and evaluation results of the zero-shot scoring model can comprehensively consider multiple aspects to comprehensively evaluate the characteristics of protein sequence and structure. The scoring algorithm may include but is not limited to protein sequence naturalness algorithm, structural naturalness algorithm, stability algorithm, affinity algorithm, functional activity algorithm, and additional sequence diversity algorithm, sequence conservation algorithm, structure prediction algorithm, protein-protein interaction algorithm, protein-ligand binding algorithm, thermodynamic stability algorithm and kinetic stability algorithm. These algorithms can evaluate the sequence rationality, structural rationality, structural stability, protein affinity, functional activity, mutation rationality of proteins, as well as sequence variability, structural variability, binding affinity, thermodynamic stability, kinetic stability and functional prediction score. By integrating these multimodal information, the scoring and evaluation results can not only provide an evaluation of the naturalness and rationality of protein sequence and structure, but also predict the potential impact of mutations on protein function and structure, thereby providing important guidance and decision support in the process of protein design and optimization.
[0099] In this step, the zero-shot scoring model is used to evaluate each mutant sequence in the mutant sequence library to determine its performance on different target attributes. Each mutant sequence in the mutant sequence library is input into the zero-shot scoring model. The zero-shot scoring model evaluates each mutant sequence based on its internal multiple scoring algorithms.
[0100] Then, the evaluation results of each scoring algorithm on the mutant sequence are extracted from the zero-shot scoring model. The zero-shot scoring model may include multiple scoring algorithms, such as protein sequence naturalness algorithm, protein structure naturalness algorithm, protein stability algorithm, protein affinity algorithm and functional activity algorithm. For each mutant sequence, the evaluation results of these scoring algorithms are extracted respectively.
[0101] For example, when designing a zero-shot scoring model, ensure that the output of each scoring algorithm can be extracted separately. For each mutant sequence, record the output value of each scoring algorithm separately.
[0102] Finally, the scoring and evaluation results of each mutant sequence on different target attributes are integrated to obtain a comprehensive evaluation result. For each mutant sequence, the scoring and evaluation results on different target attributes are weighted averaged, voted, or other integrated methods are used. Determine the weight of each target attribute and adjust the weight distribution according to actual needs. Result: The comprehensive evaluation result of each mutant sequence is obtained, which reflects its overall performance on multiple target attributes. This step comprehensively considers multiple target attributes and provides a more comprehensive evaluation; the weight distribution can be adjusted according to actual needs to adapt to different design goals.
[0103] Specifically, the zero-shot scoring model includes a variety of scoring algorithms to evaluate the performance of mutant sequences on different target attributes. Among them, the protein sequence naturalness algorithm can evaluate whether the mutant sequence is evolutionarily reasonable. The protein structure naturalness algorithm can evaluate whether the three-dimensional structure of the mutant sequence is stable. The protein stability algorithm can evaluate the thermal stability and chemical stability of the mutant sequence. The protein affinity algorithm can evaluate the binding affinity of the mutant sequence with other molecules (such as substrates, ligands, antigens, etc.). The functional activity algorithm can evaluate the functional activity of the mutant sequence, such as the catalytic efficiency of the enzyme, the binding affinity of the antibody, etc. Correspondingly, the scoring evaluation results include the evaluation results of multiple target attributes, which are used to comprehensively evaluate the performance of the mutant sequence.
[0104] Furthermore, the step S330, based on the comprehensive evaluation result, optimizes each of the mutant sequences using a multi-objective optimization algorithm to obtain the optimized candidate protein sequence recommendation set, including: Step S331, obtaining the scoring evaluation result corresponding to each mutant sequence in the comprehensive evaluation result as the objective function; and determining the constraint conditions and decision variables, and using the objective function, the constraint conditions and the decision variables as the actual demand target; the constraint conditions include the candidate amino acid type that can be mutated, the limit on the number of mutations, the structural stability requirement and the functional activity requirement; the decision variables include the mutation site and the amino acid type; In the above, the scoring evaluation results obtained after the zero-shot scoring model evaluates each mutant sequence in the mutant sequence library are used as the objective function of the multi-objective optimization algorithm. Specifically, the scoring evaluation results of each mutant sequence on different target attributes are extracted from the comprehensive evaluation results of the zero-shot scoring model. These results will be used as the objective function of the optimization algorithm to obtain the objective function value of each mutant sequence, which will be used in the subsequent optimization process. The objective function is the basis of the optimization algorithm. A clear objective function helps the optimization algorithm to find the optimal solution more accurately. Specifically, the scoring evaluation results can be directly extracted from the output of the zero-shot scoring model and used as the value of the objective function.
[0105] In this step, the constraints and decision variables in the multi-objective optimization problem are defined, and the objective function, constraints and decision variables are integrated into the actual demand target.
[0106] Among them, constraints define the conditions that need to be met during the optimization process, such as the types of candidate amino acids that can be mutated, the limit on the number of mutations, structural stability requirements, and functional activity requirements.
[0107] Among them, decision variables define the variables that need to be optimized, such as mutation sites and amino acid types.
[0108] The above-mentioned actual demand goal is to integrate the objective function, constraints and decision variables into the actual demand goal and clarify the specific form of the optimization problem.
[0109] In this step, the specific form of the optimization problem is clarified to provide clear guidance for the subsequent optimization algorithm, which helps the optimization algorithm to search the solution space more efficiently and find the optimal solution that meets actual needs.
[0110] Step S332, randomly generating initial mutation sequences as an initial population according to the objective function, and calculating the objective function values of the sequences in the initial population; In this step, a set of initial mutation sequences are randomly generated as the initial population of the genetic algorithm, and the objective function values of these sequences are calculated. The initial population is randomly generated, specifically, mutation sites and amino acid types are randomly selected in the candidate region for mutation, to generate a set of initial mutation sequences. For each mutation sequence in the initial population, its objective function value is calculated, so as to obtain the initial population and its objective function value, which provides a starting point for the subsequent optimization process.
[0111] The diversity of the initial population helps the optimization algorithm to search the solution space more comprehensively and avoid falling into the local optimal solution. A random number generator can be used to randomly select mutation sites and amino acid types in the mutable region to generate the initial population. Then the objective function is called to calculate the value of each sequence.
[0112] Step S333, performing non-dominated sorting processing on the initial mutation sequence in the initial population based on the objective function value, so as to divide the initial mutation sequence into a plurality of non-dominated levels; In this step, the mutant sequences in the initial population are non-dominatedly sorted according to the objective function value, and the sequences are divided into multiple non-dominated levels.
[0113] As mentioned above, non-dominated sorting is to compare the objective function values of each sequence in the initial population to determine which sequences are non-dominated (i.e., no worse than other sequences in one objective, and better than other sequences in at least one objective). The non-dominated sequences are divided into the first level, and then the non-dominated sorting is continued in the remaining sequences, divided into the second level, and so on, so as to obtain the non-dominated level of the initial population, which provides a basis for subsequent selection operations. Non-dominated sorting helps the optimization algorithm find a balance between multiple objectives and avoids over-optimizing a certain objective while ignoring other objectives.
[0114] Specifically, the initial population is sorted using a non-dominated sorting algorithm. For example, for two sequences A and B, if A is not worse than B in all objectives and is better than B in at least one objective, then A dominates B.
[0115] Step S334, taking the dimension of the objective function and the set of objective function values as the target space, and determining reference points in the target space; the number of the reference points is determined according to the dimension of the objective function and the preset target number.
[0116] Step S335 , performing reference point association processing on each of the initial mutation sequences and the nearest reference point, and calculating the distance between the individual objective function value and the reference point.
[0117] In this step, a reference point is determined in the target space, and each mutant sequence is associated with the nearest reference point, and the distance between the individual objective function value and the reference point is calculated.
[0118] According to the dimension of the objective function and the preset number of targets, a set of reference points are uniformly distributed in the target space. The distance between the objective function value of each mutation sequence and each reference point is calculated, and each sequence is associated with the nearest reference point to obtain the association relationship between each mutation sequence and the reference point, as well as the distance between the individual objective function value and the reference point.
[0119] The introduction of reference points helps the optimization algorithm to distribute solutions more evenly in the target space and improve the diversity of solutions. Specifically, the distance calculation formula (such as Euclidean distance) can be used to calculate the distance between the objective function value of each sequence and the reference point.
[0120] Step S336, selecting individuals according to the non-dominated hierarchy and the distance between the individual objective function value and the reference point, performing crossover and mutation operations, generating a selected mutation sequence, and forming a progeny population.
[0121] As mentioned above, individuals with higher non-dominated levels are preferentially selected according to the non-dominated levels. For individuals at the same level, individuals with larger distance are selected according to the distance between the individual objective function value and the reference point.
[0122] Among them, a crossover operation is used, such as selecting two parent individuals and generating one or more offspring individuals by exchanging some genetic information. Also, a mutation operation is used to randomly mutate the parent individuals, introduce new genetic information, increase the diversity of the population, and thus generate offspring populations, providing new candidate solutions for the subsequent optimization process.
[0123] Selection, crossover, and mutation operations help the optimization algorithm to gradually search for better solutions while maintaining the diversity of solutions. For example, roulette selection, tournament selection, and other methods can be used for selection operations; single-point crossover, multi-point crossover, and other methods can be used for crossover operations; and random mutation and other methods can be used for mutation operations.
[0124] Step S337, taking the initial population as the parent population; and merging the parent population with the child population to obtain a merged population.
[0125] In the above, the initial population is used as the parent population, and the parent population is merged with the child population to form a new population. The population can be merged, that is, all individuals in the parent population and the child population are merged into one population, so as to obtain a merged population, which provides a wider range of candidate solutions for the subsequent optimization process. Merging populations helps the optimization algorithm to search the solution space more comprehensively and avoid converging to the local optimal solution too early. Specifically, the individuals in the parent population and the child population can be simply merged into a list.
[0126] Step S338, repeatedly performing the non-dominated sorting process and the reference point association process on the merged population until a preset stop condition is reached, thereby obtaining an optimized recommended sequence set; This step is an iterative process of repeated optimization of the recommended sequence set. The non-dominated sorting and reference point association processing are repeated on the merged population until the preset stop condition is met; the non-dominated sorting and reference point association processing are repeated on the merged population, and individuals are selected for crossover and mutation operations to generate a new offspring population; it is determined whether the preset stop condition is met, such as reaching the maximum number of iterations, the diversity of the population is lower than a certain threshold, etc., so as to obtain the optimized recommended sequence set.
[0127] Repeating the optimization process helps the optimization algorithm gradually approach the global optimal solution and improve the quality of the solution.
[0128] A specific implementation method may be to set a loop, and perform non-dominated sorting, reference point association, selection, crossover and mutation operations once in each loop until a stop condition is met.
[0129] Step S339, determining the Pareto optimal solution in the recommended sequence set; and, based on the actual demand target, screening sequences from the Pareto optimal solution to obtain the candidate protein sequence recommendation set.
[0130] In this step, the Pareto optimal solution is first determined from the optimized recommended sequence set. The Pareto front is extracted, that is, non-dominated individuals are extracted from the recommended sequence set to form a Pareto front, thereby obtaining a Pareto optimal solution set, which achieves the best balance between multiple objectives. The Pareto optimal solution set provides a variety of options for practical applications, and the optimal solution can be selected from it according to specific needs. Specifically, the recommended sequence set can be non-dominated sorted, and the individuals of the first level can be extracted as the Pareto optimal solution. Then, according to the actual demand objectives, sequences that meet specific conditions can be screened out from the Pareto optimal solution to form a candidate protein sequence recommendation set. Specifically, according to the actual demand objectives, such as structural stability requirements, functional activity requirements, etc., sequences that meet the conditions can be screened out from the Pareto optimal solution.
[0131] Embodiment 4: Reference Figure 5 This embodiment provides a multimodal protein design method. Based on the above embodiment 3, the step S300, after obtaining a candidate protein sequence recommendation set based on the candidate mutable region and the original sequence information using a trained zero-sample scoring model, further includes: Step S500, obtaining experimental data of a verification experiment of a mutant sequence in the mutant sequence library based on the comprehensive evaluation result.
[0132] In protein design, although zero-shot scoring models can quickly generate a set of candidate sequence recommendations without experimental data, their prediction results may be affected by the limitations of the model itself, such as inaccurate modeling of certain complex biochemical processes.
[0133] In this embodiment, by introducing experimental data to iteratively optimize the multimodal target attribute scoring model, the actual experimental results can be fed back to the model, and the model parameters can be further adjusted and optimized, thereby improving the model's prediction accuracy for protein performance. This iterative optimization not only enhances the model's understanding of complex biochemical processes, but also enables the model to more comprehensively evaluate the performance of protein sequences, taking into account a variety of factors such as sequence naturalness, structural stability, and functional activity. In addition, the multimodal target attribute scoring model can better adapt to new design requirements and experimental conditions through iterative optimization, and improve the generalization ability of the model. Ultimately, this process helps to reduce the number of experimental verifications, reduce experimental costs and time, and improve the overall efficiency of protein design.
[0134] In this step, the sequences in the mutant sequence library obtained by the comprehensive evaluation of the zero-sample scoring model are experimentally verified to obtain experimental data. Some or all mutant sequences are selected from the mutant sequence library for experimental verification. Experimental verification may include but is not limited to functional activity testing, stability testing, affinity testing, etc., to obtain the actual performance data of the sequence, thereby obtaining the performance data of each mutant sequence in the actual experiment, which will be used for subsequent model training and optimization.
[0135] By verifying the prediction results of the model through experimental data from the validation experiment, the accuracy and reliability of the model can be improved. The experimental data provides real feedback for the iterative optimization of the model, which helps to further improve the model performance.
[0136] Specifically, you can design an experimental plan to test the functional activity and stability of the mutant sequence. Collect experimental data, including quantitative activity values, stability scores, etc.
[0137] Step S600: construct a multimodal target attribute scoring model, and use the experimental data to train the multimodal target attribute scoring model.
[0138] In this step, a multimodal target attribute scoring model is constructed and trained based on the experimental data. This model can comprehensively consider multiple target attributes to score the mutation sequence.
[0139] First, a multimodal target attribute scoring model is constructed, which can be a model based on machine learning, such as a neural network, a support vector machine, etc.; then, the model is trained using experimental data, and the model parameters are adjusted to minimize the prediction error, thereby obtaining a trained multimodal target attribute scoring model that can more accurately score mutation sequences.
[0140] In this step, a multimodal target attribute scoring model is constructed and trained using experimental data, taking into account multiple target attributes to provide a more comprehensive evaluation; the model is trained based on experimental data to improve the accuracy and reliability of the model. Specifically, a machine learning framework (such as TensorFlow, PyTorch) can be used to build the model. The experimental data is used as the training set, and the model is trained using an optimization algorithm (such as gradient descent).
[0141] Step S700, iteratively optimizing the multimodal target attribute scoring model based on the verification experiment to obtain an updated multimodal target attribute scoring model, and obtaining an optimized candidate protein sequence recommendation set based on the updated multimodal target attribute scoring model.
[0142] The experimental data from the validation experiment is used to iteratively optimize the multimodal target attribute scoring model, update the model parameters, and use the updated model to re-evaluate the candidate protein sequence recommendation set.
[0143] In this step, the multimodal target attribute scoring model is evaluated using experimental data, and the prediction error of the model is calculated; the model parameters are adjusted according to the prediction error, and iterative optimization is performed; then, the updated model is used to re-score and sort the candidate protein sequence recommendation set to obtain an optimized recommendation set, thereby obtaining an optimized multimodal target attribute scoring model and a candidate protein sequence recommendation set optimized based on the model.
[0144] The prediction accuracy of the model is improved through iterative optimization; a higher quality candidate protein sequence recommendation set is obtained based on the optimized model.
[0145] Specifically, methods such as cross-validation can be used to evaluate model performance and calculate prediction errors; optimization algorithms (such as gradient descent) can be used to adjust model parameters and perform iterative optimization; and the updated model can be used to re-score and re-rank the candidate protein sequence recommendation set.
[0146] In summary, through the above steps in this embodiment, the performance of the multimodal target attribute scoring model can be gradually improved, so that it can play a greater role in protein design. The optimized multimodal target attribute scoring model can more accurately predict the performance of the protein, thereby improving the design success rate and reducing the risk of failure; the experimental data provides direct feedback for model optimization, so that the model can continue to learn and improve, and further improve the prediction accuracy; the multimodal target attribute scoring model can comprehensively consider multiple target attributes, provide a more comprehensive evaluation, and avoid the problems that may be caused by single target optimization; in addition, through iterative optimization, the model can flexibly adapt to different design requirements and experimental conditions, and has strong scalability.
[0147] For further reference, Figure 6 , the iterative optimization comprises: Step S710, evaluating the candidate protein sequence recommendation set according to the multimodal target attribute scoring model and the zero-sample scoring model to obtain a scoring evaluation result.
[0148] In this step, the multimodal target attribute scoring model and the zero-shot scoring model are combined to evaluate the candidate protein sequence recommendation set to obtain the scoring evaluation results of each sequence.
[0149] First, a multimodal target attribute scoring model can be used to score each sequence in the candidate protein sequence recommendation set; the zero-shot scoring model can be used to score the same sequence; the scoring results of the two models can be combined to obtain a comprehensive evaluation result for each sequence, thereby obtaining a comprehensive scoring evaluation result for each candidate protein sequence on multiple target attributes.
[0150] This step combines the advantages of both models to provide a more comprehensive evaluation; it evaluates sequence performance from multiple perspectives and improves the accuracy of the evaluation.
[0151] As mentioned above, each sequence can be scored using a multimodal target attribute scoring model and a zero-shot scoring model respectively; in addition, weights can be set to perform a weighted average on the scoring results of the two models to obtain a comprehensive evaluation result.
[0152] Step S720: Screen candidate sequences according to the scoring and evaluation results to construct a candidate sequence set.
[0153] In this step, based on the comprehensive scoring evaluation results, candidate sequences with better performance are screened out to construct a candidate sequence set.
[0154] First, based on a set threshold or ranking, sequences with better performance can be selected from the candidate protein sequence recommendation set; then a candidate sequence set containing these sequences is constructed for subsequent experimental verification and model optimization, thereby obtaining a candidate sequence set containing sequences with better performance.
[0155] Screening can obtain sequences that are more likely to meet design requirements, reduce the number of sequences that need to be experimentally verified, and improve efficiency.
[0156] Specifically, you can set screening criteria, such as sequences with scores higher than a certain threshold or ranked in the top N; select sequences from the recommended set based on the criteria to build a candidate sequence set.
[0157] Step S730, obtaining experimental data of a verification experiment corresponding to each candidate sequence in the candidate sequence set, and using the experimental data to update the multimodal target attribute scoring model.
[0158] As described above, each sequence in the candidate sequence set is experimentally verified to obtain experimental data, and the multimodal target attribute scoring model is updated using the data.
[0159] It should be noted that the experimental data in this step is the experimental data obtained by the current verification experiment on the sequence. If in the subsequent iteration process, it is necessary to re-perform the verification experiment on the screened sequence each time to obtain the corresponding new experimental data.
[0160] This step experimentally verifies each sequence in the candidate sequence set to obtain its performance data in actual experiments; and uses these experimental data to update the multimodal target attribute scoring model and adjust the model parameters to improve the prediction accuracy, thereby obtaining an updated multimodal target attribute scoring model that can more accurately predict sequence performance.
[0161] The model is updated through experimental data to improve its accuracy and reliability; moreover, the experimental data provides real feedback for model optimization, which helps to further improve model performance.
[0162] Candidate sequences can be tested for functional activity, stability, etc., and experimental data can be collected and used to update the multimodal target attribute scoring model.
[0163] Step S740: Using the updated multimodal target attribute scoring model and the zero-sample scoring model, the candidate protein sequence recommendation set is evaluated to obtain a new scoring evaluation result.
[0164] In this step, the updated multimodal target attribute scoring model and zero-shot scoring model are used to re-evaluate the candidate protein sequence recommendation set to obtain new scoring evaluation results.
[0165] Specifically, the updated multimodal target attribute scoring model can be used to score each sequence in the candidate protein sequence recommendation set. Then, the zero-shot scoring model is used to score the same sequence, and the scoring results of the two models are combined to obtain a new comprehensive evaluation result for each sequence, thereby obtaining a new comprehensive scoring evaluation result for each candidate protein sequence on multiple target attributes.
[0166] This step improves the accuracy of the evaluation by updating the model; combines the advantages of the two models to provide a more comprehensive evaluation; scores each sequence using the updated multimodal target attribute scoring model and the zero-sample scoring model respectively; sets weights and performs a weighted average on the scoring results of the two models to obtain a new comprehensive evaluation result.
[0167] Step S750, using the new scoring evaluation results to screen out target sequences, construct a new candidate protein sequence recommendation set, and obtain new experimental data corresponding to the new candidate protein sequence recommendation set; In this step, based on the new comprehensive scoring evaluation results, target sequences with better performance are screened out, a new candidate protein sequence recommendation set is constructed, and new experimental data for these sequences are obtained.
[0168] First, based on the set threshold or ranking, sequences with better performance can be selected from the candidate protein sequence recommendation set; then a new candidate protein sequence recommendation set containing these sequences can be constructed; each sequence in the new candidate sequence set can be experimentally verified to obtain its performance data in actual experiments, thereby obtaining a new candidate protein sequence recommendation set and its corresponding experimental data.
[0169] Through screening, sequences that are more likely to meet design requirements are obtained; the number of sequences that need to be experimentally verified is reduced to improve efficiency; in addition, the model is further optimized through experimental data to improve model performance.
[0170] Specifically, you can set screening criteria, such as sequences with scores above a certain threshold or the top N rankings. Select sequences from the recommended set according to the criteria to build a new set of candidate sequences; test the functional activity and stability of the new candidate sequences; and collect experimental data for subsequent model optimization.
[0171] Step S760, based on the new experimental data, evaluating the iterative optimization of the multimodal target attribute scoring model, so as to complete the iterative optimization.
[0172] In this step, the performance of the updated multimodal target attribute scoring model is evaluated using new experimental data. This may include, but is not limited to, checking the model's prediction accuracy, generalization ability, and stability for new data. By comparing the model's performance on new data with previous performance indicators, it is determined whether the model has improved. After checking the convergence conditions, if the model's performance has reached the target, or has not improved significantly in consecutive rounds of iterations, the iterative optimization is considered complete. At this point, the iteration process can be stopped, and the final multimodal target attribute scoring model and the optimized candidate protein sequence recommendation set can be output.
[0173] Through the above steps, it is ensured that the multimodal target attribute scoring model can be effectively updated and optimized according to new experimental data in each iteration, and finally achieve the preset optimization goal, thereby improving the accuracy and efficiency of protein design.
[0174] Furthermore, the step S760, based on the new experimental data, evaluates the iterative optimization of the multimodal target attribute scoring model to complete the iterative optimization, including: Step S761, judging whether the multimodal target attribute scoring model reaches a preset optimization target based on the new experimental data.
[0175] In this step, based on the newly acquired experimental data, it is evaluated whether the multimodal target attribute scoring model has achieved the preset optimization goal to decide whether to continue optimizing the model.
[0176] In this step, the prediction results of the multimodal target attribute scoring model can first be evaluated using new experimental data to calculate the prediction error or performance index of the model; then, the evaluation results are compared with the preset optimization goals to determine whether the model meets the requirements, so as to further determine whether the multimodal target attribute scoring model has achieved the optimization goal, thereby deciding whether to continue iterative optimization.
[0177] This step sets the optimization goal and clarifies the end point of model optimization to avoid endless optimization; based on whether the model reaches the optimization goal, the optimization strategy is dynamically adjusted to improve optimization efficiency.
[0178] Specifically, you can set optimization goals, such as prediction error below a certain threshold or model performance improvement exceeding a certain ratio. Use new experimental data to evaluate the model, calculate the prediction error or performance index; compare the evaluation results with the optimization goals to determine whether the conditions are met.
[0179] It should be noted that in the process of protein design and optimization, the preset optimization goal refers to the standard set during the iterative optimization process of the model to determine whether the model has achieved the expected performance. These goals are usually based on experimental data and design requirements, and are used to evaluate the prediction accuracy, stability and generalization ability of the model. The preset optimization goal is a specific standard used during the optimization process to determine whether the model meets the design requirements. These goals usually include the model's prediction accuracy, performance improvement, error range, etc., which are used to decide whether to continue iterative optimization.
[0180] Step S770: If yes, it is determined that the iterative optimization of the multimodal target attribute scoring model is completed.
[0181] As mentioned above, if the multimodal target attribute scoring model reaches the preset optimization goal, the iterative optimization process of the model is considered to be completed; when the evaluation result shows that the model has reached the optimization goal, the iterative optimization process is stopped; the optimized model parameters are saved for subsequent use, and finally an optimized multimodal target attribute scoring model is obtained, which can more accurately predict the performance of protein sequences.
[0182] Through judgment, unnecessary optimization steps are avoided, computing resources and time are saved; the optimized model has better stability and reliability under the preset goals; when the evaluation results meet the optimization goals, the optimization cycle is stopped; the final parameters of the model are saved for subsequent use.
[0183] Step S780, if not, it is determined that the iterative optimization of the multimodal target attribute scoring model is not completed, and returns to the step S720, and the candidate protein sequence recommendation set is evaluated according to the multimodal target attribute scoring model and the zero-sample scoring model to obtain a scoring evaluation result.
[0184] As mentioned above, if the multimodal target attribute scoring model does not reach the preset optimization goal, it is considered that the iterative optimization of the model is not completed and needs to be continued.
[0185] In this step, when the evaluation result shows that the model has not reached the optimization goal, iterative optimization can be continued. In this case, it is necessary to return to step S720 to evaluate the candidate protein sequence recommendation set using the multimodal target attribute scoring model and the zero-sample scoring model, re-perform the scoring evaluation, and continue the optimization process until the model reaches the preset optimization goal.
[0186] Through the above continuous optimization, the performance of the model is gradually improved until it meets the design requirements; the optimization strategy is dynamically adjusted according to the performance of the model to ensure that the model can adapt to different design requirements.
[0187] Specifically, when the evaluation result does not meet the optimization goal, the optimization cycle can be continued; the candidate protein sequence recommendation set can be re-evaluated using the updated model to obtain a new scoring evaluation result.
[0188] For example, the preset optimization goal is to increase the enzyme activity by at least 3 times. First, the multimodal target attribute scoring model and the zero-sample scoring model are used to evaluate the candidate protein sequence recommendation set. There is currently a set of candidate sequences, each of which is scored by these two models, and the scoring evaluation results obtained include scores in multiple dimensions such as sequence naturalness, structural stability, and functional activity. Based on these scoring evaluation results, candidate sequences with better performance are screened out to construct a candidate sequence set, and 20 sequences with higher comprehensive scores are screened out from 100 candidate sequences as the candidate sequence set.
[0189] The 20 candidate sequences were experimentally verified to obtain experimental data. The experiment mainly tested the activity of the enzymes corresponding to these sequences and compared them with the activity of the seed enzyme. The experimental results showed that the enzyme activity of 5 sequences increased to a certain extent, but none of them reached the target of more than 3 times. The highest sequence enzyme activity increased by 2.5 times. The experimental data of these 20 candidate sequences (including performance indicators such as enzyme activity) were used to update the multimodal target attribute scoring model. The model parameters were adjusted through optimization algorithms (such as gradient descent) so that the model can better predict target attributes such as enzyme activity.
[0190] The updated multimodal target attribute scoring model and zero-shot scoring model are used to evaluate the candidate protein sequence recommendation set (here, it can be a newly generated candidate sequence set or a re-evaluation of the previously unscreened sequences) to obtain new scoring evaluation results. According to the new scoring evaluation results, the target sequence is screened out again to construct a new candidate protein sequence recommendation set. This time, 15 sequences are screened out from the new candidate sequences. These 15 new candidate sequences are experimentally verified to obtain new experimental data. The experimental results show that the enzyme activity of 3 sequences has been further improved, and the enzyme activity of one sequence has reached more than 3 times the seed enzyme activity, meeting the preset optimization goal. Based on these new experimental data, the iterative optimization of the multimodal target attribute scoring model is evaluated. Since one sequence has met the preset optimization goal, it can be judged that the iterative optimization of the model has achieved the expected effect and the optimization process is completed. From the sequences that meet the optimization goal, other factors (such as structural stability, etc.) are comprehensively considered to determine the final target mutant protein sequence. The sequence with more than 3 times the enzyme activity and good structural stability is selected as the final result.
[0191] Through the above process, protein sequence optimization with the goal of increasing enzyme activity by at least 3 times was achieved. At the same time, the application and effect of the multimodal target attribute scoring model in the iterative optimization process was also demonstrated.
[0192] Embodiment 5: Reference Figure 7 This embodiment provides a multimodal protein design method. Based on the above embodiment 4, the step S400 performs clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-sample scoring model, and performs sampling and screening according to the latent space representation to obtain the target mutant protein sequence, including: Step S410: embed the zero-shot scoring model or the multimodal target attribute scoring model into the protein sequences in the candidate protein sequence recommendation set to obtain high-dimensional representation information in the latent space.
[0193] In this step, the optimized protein sequence is embedded through a zero-shot scoring model or a multimodal target attribute scoring model to generate high-dimensional representation information in the latent space.
[0194] Among them, the use of "the zero-sample scoring model or the multimodal target attribute scoring model" for embedding processing is mainly to provide a more flexible and comprehensive protein sequence characterization method. These two models each have their own characteristics and applicable scenarios. Choosing one of them or combining them can better meet different design requirements and optimization goals.
[0195] The zero-shot scoring model is characterized by the absence of experimental data: the zero-shot scoring model can directly evaluate and score protein sequences without experimental data. This is very useful for quickly generating candidate sequences in the early stages of design, especially in the absence of experimental verification. Rapid generation of candidate sequences: the zero-shot scoring model can quickly process a large number of sequences and generate a recommended set of candidate sequences, which is suitable for large-scale preliminary screening.
[0196] The multimodal target attribute scoring model integrates multiple target attributes. The multimodal target attribute scoring model can comprehensively consider multiple target attributes (such as sequence naturalness, structural stability, functional activity, etc.) to provide a more comprehensive evaluation. This is critical for optimizing the comprehensive performance of proteins. Using experimental data: The multimodal target attribute scoring model can be trained and optimized through experimental data to further improve the accuracy and reliability of the model.
[0197] The optimized protein sequence is input into a zero-shot scoring model or a multimodal target attribute scoring model; the model maps each protein sequence into a high-dimensional latent space through its internal neural network or other embedding mechanism to generate high-dimensional representation information, thereby obtaining high-dimensional representation information of each protein sequence in the latent space, which contains the key features and attributes of the sequence.
[0198] High-dimensional representation information can capture the complex features and properties of protein sequences; compress complex protein sequence information into a high-dimensional space for subsequent processing; use pre-trained neural network models (such as ESM-2, ProtBERT, etc.) for embedding; and map protein sequences to high-dimensional latent space through the forward propagation of the model.
[0199] Step S420, clustering and dimensionality reduction processing is performed on the high-dimensional representation information in the latent space to obtain low-dimensional representation information.
[0200] In this step, the high-dimensional representation information in the latent space is clustered and dimensionally reduced to generate low-dimensional representation information, which is convenient for subsequent sampling and screening.
[0201] Clustering algorithms (such as K-means, DBSCAN, etc.) can be used to cluster high-dimensional representation information and group similar sequences into different clusters; and dimensionality reduction methods (such as PCA, t-SNE, UMAP, etc.) can be used to map high-dimensional representation information to low-dimensional space to obtain low-dimensional representation information, thereby obtaining the representation information of each protein sequence in the low-dimensional space. This information retains the key features of the sequence while reducing the dimension of the data.
[0202] Low-dimensional representation information simplifies the data structure and facilitates subsequent processing; the computational cost in low-dimensional space is lower, which improves processing efficiency.
[0203] For example, the K-means algorithm is used to cluster the high-dimensional representation information and PCA or t-SNE is used to reduce the dimension of the high-dimensional representation information.
[0204] Step S430, in the low-dimensional space, sampling is performed according to the distribution of the low-dimensional evidence information to generate a sampling screening result; wherein the sampling screening result includes the screened candidate mutation sequences; In this step, sampling is performed in the low-dimensional space according to the distribution of the low-dimensional representation information to generate sampling screening results and screen out representative candidate mutation sequences.
[0205] Specifically, the distribution of low-dimensional representation information is analyzed to determine the sampling strategy; uniform sampling or distribution-based sampling is performed in the low-dimensional space to generate a set of candidate mutation sequences; based on the sampling results, representative candidate mutation sequences are screened out, which are representative in the low-dimensional space.
[0206] As mentioned above, the selected candidate mutation sequences are representative in the low-dimensional space and can cover different feature areas; through the sampling strategy, the generated candidate sequences are ensured to be diverse.
[0207] A uniform sampling or a distribution-based sampling method may be used; based on the sampling results, representative candidate mutation sequences are screened out.
[0208] Step S440, determining the target mutant protein according to the sampling and screening results.
[0209] As described above, the final target mutant protein sequence is determined based on the sampling and screening results.
[0210] In this step, the best performing candidate mutation sequence is selected from the sampling and screening results; the candidate mutation sequence is further evaluated and verified according to the actual demand objectives; the final target mutant protein sequence is determined to obtain optimized and screened target mutant protein sequences, which perform well in multiple target attributes.
[0211] The mutant protein sequence finally determined performed well in multiple target properties and met the design requirements; through sampling and screening, the target mutant protein sequence was efficiently determined.
[0212] Specifically, the sampling and screening results can be evaluated to select the best performing candidate mutant sequence. Experimental verification can be performed to ensure that the final mutant protein sequence meets the design requirements.
[0213] In some embodiments, reference Figure 8 The step S420, performing clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space, includes: Step S421, clustering the high-dimensional representation information in the latent space using a clustering algorithm to group similar sequences into different clusters; A clustering algorithm is used to cluster the high-dimensional representation information in the latent space and group similar protein sequences into different clusters.
[0214] In this step, a suitable clustering algorithm (such as K-means, DBSCAN, etc.) is selected, and the high-dimensional representation information is input into the clustering algorithm for clustering processing. Similar sequences are grouped into different clusters, and each cluster contains similar protein sequences.
[0215] Step S422, using a dimensionality reduction method, maps the high-dimensional representation information after clustering processing to a low-dimensional space to obtain the low-dimensional representation information; wherein the dimensionality reduction method includes at least one of principal component analysis, t-random neighbor embedding and UMAP nonlinear dimensionality reduction technology.
[0216] In this step, a dimensionality reduction method is used to map the high-dimensional representation information after clustering processing into a low-dimensional space to generate low-dimensional representation information.
[0217] You can choose a suitable dimensionality reduction method (such as PCA, t-SNE, UMAP, etc.). Input the high-dimensional representation information after clustering into the dimensionality reduction method for dimensionality reduction processing; obtain low-dimensional representation information, and thus obtain the representation information of each protein sequence in the low-dimensional space. This information retains the key features of the sequence while reducing the dimension of the data.
[0218] Low-dimensional representation information simplifies the data structure and facilitates subsequent processing; the computational cost in low-dimensional space is lower, which improves processing efficiency.
[0219] Specifically, PCA or t-SNE can be used for dimensionality reduction; the high-dimensional representation information after clustering is input into PCA or t-SNE for dimensionality reduction.
[0220] Embodiment 6: In order to better illustrate the multimodal protein design method provided in Examples 1 to 6, in this example, a multimodal pre-trained target attribute scoring model is constructed to encode the protein sequence, and then prediction is performed through a fusion network.
[0221] 1. Experimental design: (1) Dataset preparation: Collect a set of protein sequences with known functions and their corresponding structural information (PDB format). Ensure that the dataset contains sufficient diversity to cover different protein families and functional types.
[0222] (2) Model preparation: Prepare a multimodal pre-trained target attribute scoring model (referred to as the multimodal model in the figure), including a structural encoder, an MSA encoder, a sequence encoder, and a fusion network. Prepare other baseline models, including: ESM2_3B (PLM model), RDE (structural pre-trained model), PGLM_3B (PLM model), MSA_Transformer (a hybrid PLM model containing MSA data), etc.
[0223] (3) Experimental setup: Divide the dataset into training set, validation set, and test set. Set evaluation indicators such as Spearman rank correlation, mean square error (MSE), etc.
[0224] 2. Experimental process: All models are trained on the same training conditions and data sets, and the performance of each model on the training set and validation set is monitored, and the optimal model is selected by the minimum loss. In each iteration, the loss value and evaluation index during the training process are recorded, and the model is tested after the training is completed.
[0225] 3. Experimental results analysis: reference Fig. 9 , shows the performance comparison of different models in the protein sequence prediction task, using Spearman's rank correlation coefficient as the evaluation indicator. Spearman's rank correlation coefficient is a non-parametric statistic that measures the correlation between two rank variables, and its value ranges from -1 (complete negative correlation) to 1 (complete positive correlation). The closer the value is to 1, the better the correlation between the model prediction result and the experimental data.
[0226] The multimodal pre-trained target attribute scoring model performed best among all models, with a Spearman rank correlation coefficient of 0.851, which was significantly higher than other models. This shows that multimodal models have significant advantages in integrating protein sequence, structure, and evolutionary information, and can more accurately predict protein properties such as function and stability. The MSA_Transformer model performed second, with a Spearman rank correlation coefficient of 0.835, which is lower than the multimodal model, but still better than the ESM2_3B and RDE models. The PGLM_3B model performed moderately, with a Spearman rank correlation coefficient of 0.675. The ESM2_3B and RDE models performed relatively poorly, with Spearman rank correlation coefficients of 0.589 and 0.594, respectively, indicating that these models have limited prediction accuracy when dealing with protein sequence and structure information.
[0227] It can be seen that the excellent performance of the multimodal pre-trained target attribute scoring model verifies the effectiveness of the multimodal integration architecture proposed in the present invention. This architecture combines the protein sequence pre-training model, the structure pre-training model and the MSA pre-training model based on evolutionary information, and uses the LORA method for efficient pre-training parameter fine-tuning, which can efficiently optimize the target attributes with minimal experimental data. Compared with other models, the multimodal model not only performs well in integrating sequence information, structural information and evolutionary information between sequences in large-scale pre-training, but also maintains high prediction accuracy when experimental data is limited, which makes the model more advantageous in practical applications.
[0228] The efficient performance of the multimodal model indicates that it can quickly and accurately predict the target properties of proteins when there is limited experimental data. This is particularly important for protein design and optimization because it reduces the need for experimental verification and improves the efficiency and cost-effectiveness of research.
[0229] The high Spearman rank correlation coefficient of the multimodal model also shows that the model has good generalization ability and can adapt to different protein families and mutation regions, and can show good performance even in unseen mutation regions or new protein families.
[0230] In conclusion, Fig. 9 The experimental results support the effectiveness of the multimodal protein design method of the present invention. The multimodal model significantly improves the efficiency and accuracy of protein design and optimization by integrating multiple pre-trained models and fine-tuning with a small amount of experimental data.
[0231] In addition, reference Fig.10 In an embodiment of the present application, a multimodal protein design device is also provided, comprising: a data acquisition module 10, for acquiring the original sequence information of the target protein; wherein the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence; a candidate region module 20, for determining the mutable candidate region of the target protein according to the original sequence information; a sequence recommendation module 30, for obtaining a candidate protein sequence recommendation set based on the mutable candidate region and the original sequence information using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes; a sampling generation module 40, for performing clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model, and performing sampling screening according to the latent space representation to obtain a target mutant protein sequence.
[0232] The present application also provides a computer device. Exemplarily, the computer device includes a processor and a memory, wherein the memory stores a computer program, and the processor runs the computer program to enable the computer device to execute the functions of each module in the above-mentioned multimodal protein design method or the above-mentioned multimodal protein design device.
[0233] Among them, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or at least one of other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application.
[0234] The memory may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electric erasable programmable read-only memory (EEPROM), etc. The memory is used to store a computer program, and the processor may execute the computer program accordingly after receiving an execution instruction.
[0235] The present application also provides a computer storage medium for storing the computer program used in the above-mentioned computer device. The computer storage medium may be a readable storage medium, or a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium may include but is not limited to: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0236] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or the flow diagram, and the combination of boxes in the structure diagram and / or the flow diagram, can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0237] In addition, the functional modules or units in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0238] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0239] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A multimodal protein design method, characterized in that: include: Obtaining original sequence information of the target protein; wherein the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence; Determining candidate regions of the target protein that can be mutated based on the original sequence information; Based on the candidate mutable region and the original sequence information, a recommended set of candidate protein sequences is obtained using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes; The candidate protein sequence recommendation set is subjected to clustering and dimensionality reduction processing according to the zero-sample scoring model, and is sampled and screened according to latent space representation to obtain a target mutant protein sequence.
2. The multimodal protein design method according to claim 1, characterized in that: The step of determining the candidate regions of the target protein that can be mutated based on the original sequence information includes: Using a protein stability model, single-point saturation mutation scoring is performed on the original sequence information to obtain a stability scoring result; The candidate regions that can be mutated under preset conditions are screened out according to the stability scoring results.
3. The multimodal protein design method according to claim 2, characterized in that: The calculation method of the stability score result includes: ; Where i represents the site index of the i-th amino acid in the protein sequence; S i represents the stability score result of the ith site; μ i represents the average stability score when the i-th site is mutated into 20 amino acids; σ i represents the standard deviation of the stability score when the i-th site is mutated to 20 amino acids; represents the average structural stability score of the i-th site within the sliding window; M represents the size of the sliding window; Represents the sum of the structural stability scores from the iM / 2th site to the i+M / 2th site.
4. The multimodal protein design method according to claim 1, characterized in that: The method of obtaining a candidate protein sequence recommendation set based on the candidate mutable region and the original sequence information using a trained zero-sample scoring model includes: Performing mutations in the candidate mutable region based on the original sequence information to generate different mutant sequences and obtain a mutant sequence library; Using the zero-sample scoring model to score and evaluate each mutant sequence in the mutant sequence library to obtain a comprehensive evaluation result; Based on the comprehensive evaluation results, each of the mutant sequences is optimized using a multi-objective optimization algorithm to obtain an optimized recommended set of candidate protein sequences.
5. The multimodal protein design method according to claim 4, characterized in that: The step of scoring and evaluating each mutant sequence in the mutant sequence library using the zero-sample scoring model to obtain a comprehensive evaluation result includes: Using the zero-sample scoring model to evaluate each mutant sequence in the mutant sequence library, obtaining a scoring evaluation result corresponding to each scoring algorithm, and integrating all the scoring evaluation results to obtain a comprehensive evaluation result; Wherein, the scoring algorithm in the zero-sample scoring model includes at least one of a protein sequence naturalness algorithm, a protein structure naturalness algorithm, a protein stability algorithm, a protein affinity algorithm and a functional activity algorithm; The scoring evaluation results include at least one of sequence naturalness, sequence rationality, structural naturalness, structural stability, protein affinity, functional activity and mutation rationality.
6. The multimodal protein design method according to claim 4, characterized in that: After obtaining a candidate protein sequence recommendation set based on the candidate mutable region and the original sequence information using a trained zero-sample scoring model, the method further includes: Acquiring experimental data of a verification experiment of the mutant sequences in the mutant sequence library based on the comprehensive evaluation results; Constructing a multimodal target attribute scoring model, and using the experimental data to train the multimodal target attribute scoring model; The multimodal target attribute scoring model is iteratively optimized based on the verification experiment to obtain an updated multimodal target attribute scoring model, and an optimized candidate protein sequence recommendation set is obtained based on the updated multimodal target attribute scoring model.
7. The multimodal protein design method according to claim 6, characterized in that: The iterative optimization comprises: Evaluating the candidate protein sequence recommendation set according to the multimodal target attribute scoring model and the zero-sample scoring model to obtain a scoring evaluation result; According to the scoring and evaluation results, candidate sequences are screened and a candidate sequence set is constructed; Acquire experimental data of a verification experiment corresponding to each candidate sequence in the candidate sequence set, and update the multimodal target attribute scoring model using the experimental data; Using the updated multimodal target attribute scoring model and the zero-sample scoring model, the candidate protein sequence recommendation set is evaluated to obtain a new scoring evaluation result; Using the new scoring evaluation results to screen out target sequences, construct a new candidate protein sequence recommendation set, and obtain new experimental data corresponding to the new candidate protein sequence recommendation set; Based on the new experimental data, the iterative optimization of the multimodal target attribute scoring model is evaluated to facilitate completing the iterative optimization.
8. The multimodal protein design method according to claim 7, characterized in that: The evaluating the iterative optimization of the multimodal target attribute scoring model based on the new experimental data to complete the iterative optimization includes: Determining whether the multimodal target attribute scoring model reaches a preset optimization goal based on the new experimental data; If yes, it is determined that the iterative optimization of the multimodal target attribute scoring model is completed; If not, it is determined that the iterative optimization of the multimodal target attribute scoring model is not completed, and the evaluation of the candidate protein sequence recommendation set according to the multimodal target attribute scoring model and the zero-sample scoring model is returned to obtain a scoring evaluation result.
9. The multimodal protein design method according to claim 6, characterized in that: The candidate protein sequence recommendation set is clustered and dimensionally reduced according to the zero-sample scoring model, and sampling and screening are performed according to latent space representation to obtain a target mutant protein sequence, including: Embedding the protein sequences in the candidate protein sequence recommendation set by the zero-shot scoring model or the multimodal target attribute scoring model to obtain high-dimensional representation information in a latent space; Performing clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space to obtain low-dimensional representation information; In the low-dimensional space, sampling is performed according to the distribution of the low-dimensional evidence information to generate a sampling screening result; wherein the sampling screening result includes the screened candidate mutation sequences; The target mutant protein is determined according to the sampling and screening results.
10. The multimodal protein design method according to claim 9, characterized in that: The clustering and dimensionality reduction processing of the high-dimensional representation information in the latent space includes: Using a clustering algorithm to perform clustering processing on the high-dimensional representation information in the latent space, and grouping similar sequences into different clusters; The high-dimensional representation information after clustering is mapped to a low-dimensional space by using a dimensionality reduction method to obtain the low-dimensional representation information; wherein the dimensionality reduction method includes at least one of principal component analysis, t-random neighbor embedding and UMAP nonlinear dimensionality reduction technology.
11. A multimodal protein design device, characterized in that: include: A data acquisition module, used to acquire the original sequence information of the target protein; wherein the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence; A candidate region module, used to determine the candidate regions of the target protein that can be mutated based on the original sequence information; A sequence recommendation module, for obtaining a candidate protein sequence recommendation set based on the candidate mutable region and the original sequence information using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes; The sampling generation module performs clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-sample scoring model, and performs sampling screening according to the latent space representation to obtain the target mutant protein sequence.
12. A computer device, characterized in that: The computer device comprises a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the multimodal protein design method according to any one of claims 1 to 10.
13. A computer storage medium, characterized in that: The device stores a computer program, which, when executed on a processor, implements the multimodal protein design method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Protein sequence design implementation method based on multi-objective optimization
CN111554346A
Method and apparatus for evolutionary data driven design of protein and other sequence defined biomolecules using machine learning
CN114651064A
Novel glucose oxidase with high thermal stability based on saturation mutation and composite evaluation design
CN115181734A
Design method of pepsin mutant with high thermal stability and high activity
CN115197925A
Deep learning model of integrated sequence and structural features based on protein engineering and prediction method
CN115954050A
Cited By
Multi-modal deep learning intelligent design method based on protein sequence and structure
CN121213645A
Multimodal Deep Learning Intelligent Design Method Based on Protein Sequence and Structure
CN121213645B
Method and device for designing protease substrate oligopeptide based on EVE model
CN122245415A