A multi-modal protein design method, device, system and storage medium thereof

Through the multimodal protein design method, the zero-sample scoring model and cluster dimensionality reduction technology are used to generate efficient and multi-objective optimized protein mutant sequences, solving the problems of inefficiency and multi-objective optimization in traditional methods, and achieving the automation and generalization capabilities of protein design.

CN119943206BActive Publication Date: 2025-07-22BIOMAP (BEIJING) INTELLIGENCE TECH LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510413097.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing protein design methods have shortcomings in terms of efficiency, stability and functional diversity. The traditional methods are inefficient and difficult to take into account multi-objective optimization. They rely on experimental data and expert experience and lack generalization capabilities.

Method used

The multimodal protein design method is adopted to obtain the original sequence information of the target protein, and the mutable candidate regions are determined using the zero-sample scoring model. Combined with cluster dimensionality reduction and cryptospace characterization, the target mutant protein sequence is generated, and the scoring algorithm of multiple target attributes is integrated for optimization.

Benefits of technology

It realizes efficient automation of protein design, can quickly generate optimized candidate sequences in the absence of experimental data, improve design efficiency, adapt to different protein families, have strong generalization ability and flexibility, and optimize protein performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943206B_ABST
    Figure CN119943206B_ABST
Patent Text Reader

Abstract

The present application provides a multimodal protein design method, apparatus, system and storage medium, relating to the technical fields of protein design and machine learning. The multimodal protein design method includes: obtaining the original sequence information of a target protein; determining the mutable candidate regions of the target protein according to the original sequence information; obtaining a candidate protein sequence recommendation set based on the mutable candidate regions and the original sequence information by using a zero-shot scoring model; performing clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model and performing sampling and screening according to the latent space representation to obtain the target mutant protein sequence. The multimodal protein design method provided by the present application exhibits significant advantages in terms of high efficiency, multi-objective optimization ability, small sample adaptability, generalization ability, optimization efficiency, flexibility and scalability, and can be widely applied to protein design and optimization in the fields of biomedicine, industrial catalysis, material science, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of protein design and machine learning, and particularly to a multi-modal protein design method, apparatus, system, and its storage medium. Background Art

[0002] As a key molecule in organisms, proteins are widely used in fields such as biomedicine, industrial catalysis, and materials science. Their unique structure and function have made them a research hotspot, especially in the development of new treatment methods, improvement of industrial catalysis efficiency, and applications in materials science. However, natural proteins often cannot meet specific application requirements, especially in terms of high efficiency, stability, and functional diversity. Therefore, the design and optimization of proteins have become a key approach to solving these problems. By optimizing the design of proteins through artificial synthesis or computer simulation methods, their performance can be improved to a certain extent, making them more adaptable to specific usage environments.

[0003] Traditional protein design methods mainly include random mutagenesis, semi-rational design, etc. Although these methods have been widely used, they face many challenges in actual operation. Random mutagenesis usually relies on screening a large number of mutants, which is inefficient, and the optimization of specific functions often depends on the experience of experts. Although semi-rational design introduces certain structural knowledge, it also faces the problem of being difficult to balance multi-objective optimization. Especially when multiple functions need to be optimized simultaneously, it is often impossible to obtain the best design results. In recent years, with the rapid development of artificial intelligence technology, protein design methods based on pre-trained models have gradually emerged. These methods have improved the design efficiency through a data-driven approach, especially having significant advantages in predicting the structure and function of proteins. However, these methods still have some technical problems to be solved, such as insufficient generalization ability, limited experimental throughput, low automation level, and lack of dynamic optimization ability. To achieve more efficient and widely applicable protein design, it is urgent to break through these bottlenecks and improve the adaptability and optimization ability of the model, especially to improve the design efficiency in the case of limited experimental data. Summary of the Invention

[0004] In a first aspect, the present invention provides a multi-modal protein design method, including:

[0005] Obtain the original sequence information of the target protein; wherein, the original sequence information includes the amino acid sequence and the PDB structure information corresponding to the amino acid sequence;

[0006] Determine the mutable candidate regions of the target protein according to the original sequence information;

[0007] Based on the mutable candidate regions and the original sequence information, a candidate protein sequence recommendation set is obtained by using a trained zero-shot scoring model; different scoring algorithms for target attributes are included in the zero-shot scoring model;

[0008] The candidate protein sequence recommendation set is subjected to clustering and dimensionality reduction processing according to the zero-shot scoring model, and sampling and screening are performed according to the latent space representation to obtain the target mutant protein sequence.

[0009] In an alternative embodiment, determining the mutable candidate regions of the target protein according to the original sequence information includes:

[0010] Using a protein stability model, performing single-site saturation mutation scoring on the original sequence information to obtain a stability scoring result;

[0011] Filtering out the mutable candidate regions under preset conditions according to the stability scoring result.

[0012] In an alternative embodiment, the calculation method of the stability scoring result includes:

[0013] ; where i represents the site index of the i-th amino acid in the protein sequence; S i represents the stability scoring result of the i-th site; μ i represents the average stability score when the i-th site mutates into 20 kinds of amino acids; σ i represents the standard deviation of the stability scores when the i-th site mutates into 20 kinds of amino acids; represents the average structural stability score of the i-th site within the sliding window range; M represents the size of the sliding window; represents the sum of the structural stability scores from the (i - M / 2)-th site to the (i + M / 2)-th site.

[0014] In an alternative embodiment, obtaining a candidate protein sequence recommendation set by using a trained zero-shot scoring model based on the mutable candidate regions and the original sequence information includes:

[0015] Mutating based on the original sequence information within the mutable candidate regions to generate different mutant sequences, obtaining a mutant sequence library;

[0016] Performing scoring and evaluation on each mutant sequence in the mutant sequence library by using the zero-shot scoring model to obtain a comprehensive evaluation result;

[0017] Based on the comprehensive evaluation result, using a multi-objective optimization algorithm to optimize each mutant sequence to obtain the optimized candidate protein sequence recommendation set.

[0018] In an alternative embodiment, each mutation sequence in the mutation sequence library is scored and evaluated using the zero-shot scoring model to obtain a comprehensive evaluation result, including:

[0019] Each mutation sequence in the mutation sequence library is evaluated using the zero-shot scoring model to obtain a scoring evaluation result corresponding to each scoring algorithm, and all the scoring evaluation results are integrated to obtain a comprehensive evaluation result;

[0020] Among them, the scoring algorithms in the zero-shot scoring model include at least one of a protein sequence naturalness algorithm, a protein structure naturalness algorithm, a protein stability algorithm, a protein affinity algorithm, and a functional activity algorithm;

[0021] The scoring evaluation result includes at least one of sequence naturalness, sequence rationality, structure naturalness, structure stability, protein affinity, functional activity, and mutation rationality.

[0022] In an alternative embodiment, after obtaining the candidate protein sequence recommendation set using the trained zero-shot scoring model based on the mutable candidate region and the original sequence information, it further includes:

[0023] Obtain experimental data of the verification experiment of the mutation sequences in the mutation sequence library based on the comprehensive evaluation result;

[0024] Construct a multi-modal target attribute scoring model and train the multi-modal target attribute scoring model using the experimental data;

[0025] Iteratively optimize the multi-modal target attribute scoring model based on the verification experiment to obtain the updated multi-modal target attribute scoring model, and obtain an optimized candidate protein sequence recommendation set based on the updated multi-modal target attribute scoring model.

[0026] In an alternative embodiment, the iterative optimization includes:

[0027] Evaluate the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain a scoring evaluation result;

[0028] According to the scoring evaluation result, screen to obtain candidate sequences and construct a candidate sequence set;

[0029] Obtain experimental data of the verification experiment corresponding to each candidate sequence in the candidate sequence set and update the multi-modal target attribute scoring model using the experimental data;

[0030] Using the updated multi-modal target attribute scoring model and the zero-shot scoring model, evaluate the candidate protein sequence recommendation set to obtain a new scoring evaluation result;

[0031] Using the new scoring evaluation result, screen out the target sequences, construct a new candidate protein sequence recommendation set, and obtain new experimental data corresponding to the new candidate protein sequence recommendation set;

[0032] Based on the new experimental data, evaluate the iterative optimization of the multi-modal target attribute scoring model to facilitate the completion of the iterative optimization.

[0033] In an alternative embodiment, the evaluating the iterative optimization of the multi-modal target attribute scoring model based on the new experimental data to facilitate the completion of the iterative optimization includes:

[0034] Based on the new experimental data, determine whether the multi-modal target attribute scoring model reaches a preset optimization goal;

[0035] If so, determine that the iterative optimization of the multi-modal target attribute scoring model is completed;

[0036] If not, determine that the iterative optimization of the multi-modal target attribute scoring model is not completed, and return to evaluate the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain a scoring evaluation result.

[0037] Based on the new experimental data, determine whether the multi-modal target attribute scoring model reaches a preset optimization goal;

[0038] If so, determine that the iterative optimization of the multi-modal target attribute scoring model is completed;

[0039] If not, determine that the iterative optimization of the multi-modal target attribute scoring model is not completed, and return to evaluate the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain a scoring evaluation result.

[0040] In an alternative embodiment, the clustering and dimensionality reduction processing of the candidate protein sequence recommendation set according to the zero-shot scoring model and the sampling and screening according to the latent space representation to obtain the target mutant protein sequence includes:

[0041] Embed the protein sequences in the candidate protein sequence recommendation set using the zero-shot scoring model or the multi-modal target attribute scoring model to obtain high-dimensional representation information in the latent space;

[0042] Perform clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space to obtain low-dimensional representation information;

[0043] In the low-dimensional space, sample according to the distribution of the low-dimensional representation information to generate a sampling and screening result; wherein, the sampling and screening result includes the selected candidate mutant sequences;

[0044] Determine the target mutant protein according to the sampling and screening result.

[0045] In an alternative embodiment, the performing clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space includes:

[0046] Use a clustering algorithm to perform clustering processing on the high-dimensional representation information in the latent space, and group similar sequences into different clusters;

[0047] Utilize a dimensionality reduction method to map the high-dimensional representation information after clustering processing into a low-dimensional space to obtain the low-dimensional representation information; wherein, the dimensionality reduction method includes at least one of principal component analysis, t-stochastic neighborhood embedding, and UMAP non-linear dimensionality reduction technology.

[0048] In a second aspect, the present invention provides a multi-modal protein design device, including:

[0049] A data acquisition module for acquiring the original sequence information of a target protein; wherein, the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence;

[0050] A candidate region module for determining the mutatable candidate region of the target protein according to the original sequence information;

[0051] A sequence recommendation module for obtaining a candidate protein sequence recommendation set based on the mutatable candidate region and the original sequence information by using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes;

[0052] A sampling and generation module for performing clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model, and performing sampling and screening according to the latent space representation to obtain a target mutant protein sequence.

[0053] In a third aspect, the present invention provides a computer device, the computer device includes a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the multi-modal protein design method according to any one of the foregoing embodiments.

[0054] In a fourth aspect, the present invention provides a computer storage medium storing a computer program which, when executed on a processor, implements the multi-modal protein design method according to any one of the foregoing embodiments.

[0055] The present invention provides a multi-modal protein design method, apparatus, system and storage medium thereof. The multi-modal protein design method realizes the efficiency, automation and multi-objective optimization ability of protein design by integrating a variety of advanced technical means.

[0056] This method realizes the full-process automated design from obtaining the original sequence information of the target protein to obtaining the sequence of the target mutant protein, without manual intervention, greatly reducing the time and effort of relying on expert experience and manual screening in traditional protein design, significantly improving the design efficiency, being able to complete multiple rounds of optimization in a short time, and accelerating the improvement of protein function.

[0057] Through the different target attribute scoring algorithms included in the zero-shot scoring model, this method can simultaneously evaluate multiple key attributes of proteins, such as sequence rationality, structural stability, functional activity, etc., avoiding the problem of deterioration of other attributes that may be caused by single-objective optimization in traditional methods, and realizing the balanced optimization between multiple objectives. In addition, by using clustering dimensionality reduction processing and sampling screening, the Pareto optimal solution set can be searched in the multi-dimensional objective space, providing a more comprehensive and reliable optimization scheme for practical applications.

[0058] In the case of lack of prior knowledge or scarce experimental data, the zero-shot scoring model can provide reasonable mutation suggestions based on large-scale unsupervised learning, thus realizing the rapid start of protein design in the cold start stage. With the accumulation of verification experimental data, a self-learning module is introduced, and the experimental data is used to construct a protein target attribute scoring model based on a multi-modal pre-trained model, thereby guiding the protein optimization to develop towards a higher-performance target. This model is based on a multi-modal pre-trained model and an efficient parameter fine-tuning architecture, and can stably predict target attributes even in the case of only a small amount of experimental data, significantly improving the prediction effect and iteration efficiency of the model. At the same time, by integrating the sequence, structure and evolutionary information of proteins, this method shows strong generalization ability, can adapt to different protein families and mutation regions, and can also show good performance even in unseen mutation regions or new protein families, providing a wider application scope for protein design.

[0059] The target attribute scoring model can self-optimize continuously with the accumulation of experimental data based on its self-learning ability, automatically adjusting model parameters to better adapt to new data and task requirements. This self-learning mechanism not only improves the prediction accuracy of the model but also enhances its adaptability and stability under different experimental conditions. Through continuous learning, the model can continuously absorb new knowledge, thus providing more accurate and valuable suggestions in subsequent protein design, further promoting the efficiency and quality of protein design.

[0060] Through clustering and dimensionality reduction processing, the high-dimensional protein sequence space is reduced to a low-dimensional space and uniform sampling is performed. This method not only improves the computational efficiency but also ensures the diversity of mutant sequences in terms of function and structure, avoiding the local optimum problem that may occur in traditional methods. In addition, the generated candidate mutant sequences can cover different functional and structural characteristics, providing rich choices for experiments, enabling full utilization of limited experimental resources, and improving the efficiency and success rate of experiments.

[0061] Each module of this method is designed to be pluggable and can be flexibly adjusted according to different design requirements and experimental conditions, with strong flexibility and scalability. Whether in high-throughput experiment or low-throughput experiment conditions, this method can efficiently perform protein design and optimization, adapting to a variety of protein design scenarios.

[0062] Through multiple rounds of iterative optimization, this method can significantly improve the functional performance of proteins, optimize the overall performance of proteins, making them more reliable and efficient in practical applications. For example, in the fields of industrial catalysis, biomedicine, etc., the optimized proteins can exhibit higher activity, stability, and specificity, superior to traditional artificial design paths.

[0063] In summary, the multi-modal protein design method provided in this application shows significant advantages in terms of high efficiency, multi-objective optimization ability, small sample adaptability, generalization ability, optimization efficiency, flexibility, and scalability, and can be widely applied to protein design and optimization in the fields of biomedicine, industrial catalysis, materials science, etc. Brief Description of the Drawings

[0064] To more clearly illustrate the technical solutions of this application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as limiting the protection scope of this application. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.

[0065] Figure 1 It is a schematic structural diagram of the hardware operating environment related to the embodiment of the multi-modal protein design method of the present invention;

[0066] Figure 2 It is a schematic flow chart of Embodiment 1 of the multi-modal protein design method of the present invention;

[0067] Figure 3 It is a schematic flow chart of the refinement of step S200 in Embodiment 2 of the multi-modal protein design method of the present invention;

[0068] Figure 4 It is a schematic overall flow chart including the refinement of step S300 in Embodiment 3 of the multi-modal protein design method of the present invention;

[0069] Figure 5 It is a schematic flow chart of Embodiment 4 of the multi-modal protein design method of the present invention;

[0070] Figure 6 It is a schematic flow chart including the refinement of step S700 in Embodiment 4 of the multi-modal protein design method of the present invention;

[0071] Figure 7 It is a schematic flow chart including the refinement of step S400 in Embodiment 5 of the multi-modal protein design method of the present invention;

[0072] Figure 8 It is a schematic flow chart including the refinement of step S420 in Embodiment 5 of the multi-modal protein design method of the present invention;

[0073] Figure 9 It is a schematic diagram of the model performance evaluation result in Embodiment 6 of the multi-modal protein design method of the present invention;

[0074] Figure 10 It is a schematic diagram of the module connection of the multi-modal protein design device of the present invention. Detailed implementation manners

[0075] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0076] Generally, the components of the embodiments of the present application described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0077] As used hereinafter, the terms "comprising", "having" and their cognates that may be used in various embodiments of the present application are only intended to indicate specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as precluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or as precluding the possibility of adding one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0078] In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0079] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which various embodiments of the present application pertain. The terms (such as those defined in a commonly used dictionary) will be construed to have the same meaning as the contextual meaning in the relevant technical field and will not be construed to have an idealized meaning or an overly formal meaning unless clearly defined in various embodiments of the present application.

[0080] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0081] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0082] As Figure 1 shown, it is a schematic structural diagram of the hardware operating environment of the terminal related to the embodiment of the present invention.

[0083] The multimodal protein design system according to the embodiments of the present invention can be a PC, or a mobile terminal device such as a smart phone, a tablet computer, or a portable computer. The multimodal protein design system may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen, an input unit such as a keyboard, a remote control. Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a stable memory, such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the foregoing processor 1001. Optionally, the multimodal protein design system may further include an RF (Radio Frequency) circuit, an audio circuit, a WiFi module, and so on. In addition, the multimodal protein design system may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be elaborated here.

[0084] Those skilled in the art can understand that Figure 1 the multimodal protein design system shown in Figure 1 does not constitute a limitation thereto, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. As

[0085] In summary, the method provided by the present invention improves the accuracy of fall prediction by real-time monitoring of key state data, and can quickly respond to high-risk fall states and perform protection actions. This method reduces the structural damage of the robot, reduces the safety risk to surrounding personnel, improves the environmental adaptability and economy of the robot, and enhances its stability and reliability in a changing environment.

[0086] Embodiment 1:

[0087] Referring to Figure 2 , this embodiment provides a multimodal protein design method, including:

[0088] Step S100, obtaining the original sequence information of the target protein; wherein, the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence.

[0089] As described above, the target protein refers to a specific protein that researchers hope to design, modify, or optimize during the protein design or optimization process. It can be a known natural protein or a protein with specific functional requirements.

[0090] As described above, the original sequence information refers to the basic information of the target protein, including its amino acid sequence and the corresponding PDB structure information. These information are the starting point for protein design and optimization.

[0091] The amino acid sequence provides the basic composition information of the protein. The amino acid sequence refers to the linear arrangement order of amino acids in the protein. A protein is a polypeptide chain formed by amino acids connected by peptide bonds, and the amino acid sequence determines the primary structure of the protein. The amino acid sequence is the basis of protein design, determining the chemical properties and possible structures of the protein. By modifying the amino acid sequence, the function, stability, and other characteristics of the protein can be changed.

[0092] The PDB (Protein Data Bank) structure information refers to the three-dimensional spatial structure information of the protein, usually stored in the PDB file format. The PDB file contains the coordinate information of each atom in the protein, describing the spatial conformation of the protein. In the PDF structure information, each line represents an atom in the protein, including information such as atom name, residue name, chain identifier, residue number, and atom coordinates. The PDB structure information helps to understand the spatial conformation of the protein, which is crucial for predicting the function, stability, and interaction with other molecules of the protein. In protein design, the PDB structure information can be used to evaluate the impact of mutations on the protein structure, ensuring that the designed protein remains stable and functional in three-dimensional space.

[0093] In this step, the basic information of the target protein is obtained, including its amino acid sequence and the corresponding PDB structure information.

[0094] Specifically, the amino acid sequence of the target protein can be obtained by database query or experimental determination. If the target protein already has PDB structure information, it is directly obtained; if not, the predicted structure PDB can be obtained by calling a structure prediction algorithm (such as xTrimoStructure). Thus, the amino acid sequence and PDB structure information of the target protein are obtained, providing a basis for subsequent determination and design of the mutation region.

[0095] This step ensures that the starting point of the design is accurate and complete, providing a reliable data basis for subsequent multimodal analysis and optimization. Relevant information can be obtained from public protein databases (such as UniProt, PDB), or protein structures can be determined using experimental techniques (such as X-ray crystallography, nuclear magnetic resonance, etc.).

[0096] Step S200: Determine the mutatable candidate regions of the target protein according to the original sequence information.

[0097] Based on the amino acid sequence and PDB structure information of the target protein, determine which regions of amino acids can be mutated.

[0098] For example, a protein stability prediction model can be used to score single-site saturation mutations of the sequence, and the structural stability scores can be calculated for all sites based on this model. Calculate the mutation impact of each site through a formula, consider regions with less impact on stability and better overall stability, and perform a sliding window smoothing process on local regions to enhance the reliability of the scores, thereby obtaining mutatable candidate regions. At the same time, users can also customize candidate mutation regions according to their needs.

[0099] This step can accurately locate the regions suitable for mutation, avoiding the loss of protein function or structural instability caused by blind mutation. Implementation method: Use existing protein stability prediction models or train new models based on machine learning methods to evaluate the impact of mutations on protein stability.

[0100] Step S300: Based on the mutatable candidate regions and the original sequence information, obtain a recommended set of candidate protein sequences using the trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes.

[0101] As mentioned above, in protein design and optimization, the target attributes are multi-faceted. For example, they can include but are not limited to functional activity, sequence naturality, structural naturality, protein stability, affinity, etc. These attributes jointly determine the performance and applicability of the protein and are key factors that need to be considered and optimized during the protein design process.

[0102] As mentioned above, the scoring algorithm is a calculation method used to evaluate the potential performance of a protein sequence on a specific target attribute. These algorithms are based on the sequence information, structural information, or evolutionary information of the protein and predict the performance of the protein through mathematical models or machine learning models. For example, the sequence rationality scoring algorithm evaluates whether a protein sequence conforms to known biochemical laws or evolutionary conservation. For another example, multiple sequence alignment (MSA) information is used to determine whether a certain mutation is evolutionarily reasonable; the structural stability scoring algorithm is used to evaluate the stability of the protein structure; the change in protein free energy (ΔG) before and after mutation is calculated to predict the impact of the mutation on stability; the functional activity scoring algorithm is used to predict the functional activity of the protein; machine learning-based models are used to predict the catalytic efficiency of enzymes or the binding affinity of antibodies, and so on.

[0103] The above-mentioned Zero-Shot Scoring Model is a special machine learning model that can directly evaluate and score protein sequences without experimental data.

[0104] This model is based on pre-trained protein language models (such as ESM-2, PGLM) or structural models (such as RDE), and uses the knowledge learned from large-scale protein data to evaluate new protein sequences.

[0105] Zero-Shot analysis scoring and design is a highly innovative method. It can directly analyze and score the original sequence information of the target protein using a pre-trained model without any experimental data, thereby generating a recommended set of candidate protein sequences. This method breaks through the dependence of traditional design methods on a large amount of experimental data, enabling protein design to start quickly in the cold start phase without prior knowledge, greatly improving the design efficiency. In this way, researchers can evaluate and screen a large number of potential protein sequences in a very short time, quickly lock in mutant protein sequences with potential functions, and provide a solid foundation for subsequent experimental verification and optimization. This zero-shot design ability not only accelerates the process of protein design but also opens up new avenues for exploring new protein functions and applications.

[0106] The Zero-Shot Scoring Model adopted in this step does not require experimental data and can provide reasonable mutation suggestions in the cold start phase (i.e., without experimental data); it can evaluate a large number of candidate sequences in a short time, improving the design efficiency. Specifically, it can encode protein sequences based on pre-trained protein language models (such as ESM-2) or structural models (such as RDE), combine multiple sequence alignment (MSA) information, use models such as MSA Transformer to capture evolutionary information, and score protein sequences on different target attributes through fine-tuning or directly using the zero-shot ability of the pre-trained model.

[0107] In this step, a zero-shot scoring model is used to generate a candidate protein sequence recommendation set based on the mutatable candidate regions and the original sequence information. The zero-shot scoring model can evaluate the potential mutations in the mutatable regions based on a pre-trained protein language model, a structure model, and an evolutionary information model for multiple sequence alignment. The model contains scoring algorithms for different target attributes, such as sequence naturalness, structural stability, functional activity, etc. By using these algorithms to score the mutated sequences, a set of candidate protein sequence recommendation sets is obtained, and each sequence corresponds to a score, reflecting its potential performance on different target attributes. The zero-shot scoring model can directly evaluate the mutated sequences without experimental data, greatly improving the design efficiency, especially suitable for the cold start phase. Implementation method: Existing pre-trained protein language models and structure models can be adopted, and fine-tuning can be carried out on this basis or directly utilize their zero-shot capabilities.

[0108] Step S400, perform clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model, and perform sampling and screening according to the latent space representation to obtain the target mutant protein sequence.

[0109] It should be noted that the protein sequence space is usually high-dimensional because each amino acid position can have 20 different amino acid choices, which results in a very high-dimensional sequence space. Direct optimization and screening in the high-dimensional space will face: extremely high computational costs in the high-dimensional space, especially when generating and evaluating a large number of candidate sequences; data sparsity, where data points are sparse in the high-dimensional space and it is difficult to find representative regions; and, high-dimensional data is also difficult to intuitively understand and visualize.

[0110] The purpose of performing clustering and dimensionality reduction processing is to map the high-dimensional protein sequence space to a low-dimensional space while retaining key information. This can achieve more efficient calculation and optimization in the low-dimensional space; data points are denser in the low-dimensional space and it is easier to find representative regions; and, the low-dimensional space (such as two-dimensional or three-dimensional) is easier to visualize and understand.

[0111] As mentioned above, the latent space is the representation after mapping high-dimensional data to a low-dimensional space through dimensionality reduction techniques (such as PCA, t-SNE, Autoencoder, etc.). In protein design, the latent space is usually a low-dimensional embedding space generated by a pre-trained protein language model or structure model. These models compress complex high-dimensional data into a low-dimensional space by learning the internal laws of protein sequences or structures. The latent space representation refers to the representation of a protein sequence in the latent space. It is a low-dimensional vector that contains the key features and information of the protein sequence.

[0112] In this step, to adapt to different experimental throughputs and improve design efficiency, the candidate protein sequence recommendation set is further screened and optimized, including clustering and dimensionality reduction, so as to obtain the final target mutant protein sequence. Specifically, first, the latent space representation of the zero-shot scoring model can be used to cluster and reduce the dimensionality of the candidate protein sequences. Through clustering, similar sequences are grouped together to reduce duplication and redundancy. Then, uniform sampling is performed in the reduced latent space to generate a set of evenly distributed candidate mutant sequences. Finally, according to the latent space representation and scoring results, the optimal mutant protein sequence is selected.

[0113] Through the processing method in this step, optimized target mutant protein sequences can be obtained. These sequences perform well in multiple target attributes and have high diversity and innovation. This step can effectively reduce the computational amount, improve the screening efficiency, and ensure the quality and diversity of the mutant sequences through clustering and dimensionality reduction and uniform sampling.

[0114] Through the above steps, this multi-modal protein design method can efficiently and automatically generate target mutant protein sequences with excellent performance, which are applicable to a variety of application scenarios.

[0115] In summary, in this embodiment, the multi-modal protein design method starts from obtaining the original sequence information of the target protein, and uses the zero-shot scoring model and the scoring algorithms of multiple target attributes to quickly generate a candidate protein sequence recommendation set in the absence of experimental data. Through clustering and dimensionality reduction and sampling screening in the latent space representation, this method can balance and optimize key attributes such as the sequence rationality, structural stability, and functional activity of the protein. At the same time, integrating sequence, structure, and evolutionary information endows it with strong generalization ability to adapt to different protein families and mutation regions. In addition, this method efficiently utilizes experimental resources to ensure the diversity and global optimization ability of the mutant sequences. The entire process is automated and flexibly scalable, which can significantly improve protein performance, optimize the overall design effect, and make it more reliable and efficient in practical applications.

[0116] Example 2:

[0117] Referring to Figure 3 , based on the above Embodiment 1, this embodiment provides a multi-modal protein design method. In step S200, according to the original sequence information, determining the mutable candidate regions of the target protein includes:

[0118] Step S210, using a protein stability model to perform single-site saturation mutation scoring on the original sequence information to obtain a stability scoring result.

[0119] As described above, the protein stability model is a computational-based tool used to evaluate the ability of a protein to maintain the stability of its three-dimensional structure and function after changes in its amino acid sequence (such as mutations). By analyzing the sequence information, structural information, and possible physicochemical properties of the protein, the model predicts the impact of mutations on protein stability, thus helping researchers select appropriate mutation sites during protein design and optimization.

[0120] In multimodal protein design methods, the role of the protein stability model is to determine the mutatable candidate regions of the target protein. By evaluating the impact of mutations at each amino acid site on protein stability, the model can help screen out regions that can still maintain high stability after mutation, thereby improving the success rate and efficiency of protein design.

[0121] Furthermore, the calculation method of the stability score result includes:

[0122] ; where i represents the site index of the i-th amino acid in the protein sequence; S i represents the stability score result of the i-th site; μ i represents the average stability score when the i-th site mutates into 20 amino acids; σ i represents the standard deviation of the stability scores when the i-th site mutates into 20 amino acids; represents the average structural stability score of the i-th site within the sliding window range; M represents the size of the sliding window; represents the sum of the structural stability scores from the (i - M / 2)-th site to the (i + M / 2)-th site.

[0123] Step S220, screen out the mutatable candidate regions under preset conditions according to the stability score result.

[0124] As described above, the "preset conditions" refer to the conditions used to screen out amino acid sites suitable for mutation when determining mutatable candidate regions. These conditions are usually thresholds or ranges set based on the stability score result, used to judge whether a certain site is suitable for mutation.

[0125] The preset conditions are one or more screening criteria set according to the stability score result calculated by the protein stability model, used to determine which amino acid sites can be used as candidate regions for mutation. These conditions are usually to ensure that the mutated protein can achieve the expected performance optimization while maintaining its functional and structural stability.

[0126] In this step, single-site saturation mutagenesis is performed on each amino acid site of the target protein, that is, each site is mutated into the other 19 amino acids respectively. The protein stability model (such as xTrimoStability) is used to perform stability scoring on each mutated sequence, and the stability scoring results when each site is mutated into different amino acids are obtained.

[0127] According to the stability scoring results and combined with preset conditions (such as the stability score being higher than a certain threshold or within a certain range), candidate regions for mutation are screened out, so as to obtain the amino acid site regions suitable for mutation in the target protein. These regions have less impact on the overall stability of the protein after mutation and are more likely to generate mutants with optimized performance.

[0128] In this step, the candidate regions for mutation screened out through stability scoring are more targeted for mutation, reducing the risk brought by blind mutation; moreover, the mutation range is narrowed, reducing the computational cost and experimental workload of subsequent design and screening; the existing protein stability prediction model, such as the deep learning-based xTrimoStability model, is used; single-site saturation mutagenesis simulation is performed on each amino acid site in the target protein sequence, and the stability model is input for scoring; candidate regions for mutation are screened out according to the scoring results and preset conditions (such as thresholds).

[0129] Example 3:

[0130] Refer to Figure 4 , this example provides a multi-modal protein design method. Based on Example 1 above, in step S300, based on the candidate regions for mutation and the original sequence information, a candidate protein sequence recommendation set is obtained by using the trained zero-shot scoring model, including:

[0131] Step S310, mutations are performed based on the original sequence information within the candidate regions for mutation to generate different mutated sequences, and a mutated sequence library is obtained.

[0132] In this step, within the already determined candidate regions for mutation, mutation operations are performed on the amino acid sequence of the target protein to generate a series of different mutated sequences, forming a mutated sequence library.

[0133] For the regions of the amino acid sites in the target protein that have been determined to be suitable for mutation, mutation operations are performed on each site. The mutation operations can be single-site mutation, multi-site mutation, random mutation, etc. For example, the amino acid at a certain site is replaced with any one of the other 19 amino acids.

[0134] Through the above mutation operations, a large number of different mutant sequences are generated to form a mutant sequence library. As a result, a mutant sequence library containing a variety of mutant sequences is obtained, and these sequences are the basis for subsequent evaluation and optimization.

[0135] This step can achieve the generation of a large number of different mutant sequences, increasing the possibility of finding optimized sequences; mutations are only carried out within the mutable candidate regions, reducing unnecessary mutation attempts and improving efficiency.

[0136] Step S320: Score and evaluate each mutant sequence in the mutant sequence library using the zero-shot scoring model to obtain a comprehensive evaluation result.

[0137] In this step, the trained zero-shot scoring model is used to score and evaluate each mutant sequence in the mutant sequence library, evaluate the performance of these sequences on different target attributes, and obtain a comprehensive evaluation result.

[0138] First, input the mutant sequence. Each mutant sequence in the mutant sequence library is input into the zero-shot scoring model; then, the zero-shot scoring model scores each mutant sequence on different target attributes (such as sequence naturality, structural stability, functional activity, etc.) according to its internal scoring algorithm; finally, the scoring results of each mutant sequence on different target attributes are integrated to obtain the comprehensive evaluation result of each sequence. As a result, the scoring evaluation results of each mutant sequence on different target attributes and the comprehensive evaluation result are obtained.

[0139] In this step, multiple target attributes are comprehensively considered, enabling a comprehensive evaluation of the performance of mutant sequences; the zero-shot scoring model can quickly evaluate without experimental data, improving the design efficiency.

[0140] Specifically, a pre-trained model can be used: a pre-trained protein language model (such as ESM-2) or a structural model (such as RDE) can be used as the basis for the zero-shot scoring model. Multiple scoring algorithms are integrated in the model to score different target attributes respectively. Methods such as weighted average and voting mechanism can be used to integrate the scoring results of different target attributes into a comprehensive evaluation result.

[0141] Step S330: Based on the comprehensive evaluation result, use a multi-objective optimization algorithm to optimize each mutant sequence to obtain a recommended set of optimized candidate protein sequences.

[0142] In this step, according to the comprehensive evaluation result, a multi-objective optimization algorithm is used to optimize the mutant sequences, screening out the mutant sequences with the best performance to form a recommended set of candidate protein sequences.

[0143] The Multi-Objective Optimization Algorithm (MOO) is an algorithm used to find the best balance among multiple conflicting objectives. It aims to find a set of "Pareto Optimal Solutions", which achieve the best trade-off among different objectives, that is, no solution is better than other solutions in all objectives.

[0144] As mentioned above, a Pareto Optimal Solution means that if a solution improves in one objective, it must sacrifice in other objectives, then this solution is a Pareto Optimal Solution. The set of all Pareto Optimal Solutions is called the Pareto Front.

[0145] In this step, through the multi-objective optimization algorithm, it is possible to balance among multiple objective attributes and find the mutation sequence with the optimal performance; it can be flexibly adjusted according to different design requirements and constraint conditions to adapt to various application scenarios.

[0146] Specifically, multi-objective optimization algorithms such as the Non-dominated Sorting Genetic Algorithm III (NSGA-III) can be used. The scoring evaluation result in the comprehensive evaluation result is used as the objective function, and the weights and constraint conditions are designed according to actual needs.

[0147] Through the above steps, a recommended set of optimized candidate protein sequences is finally obtained. These sequences perform well in multiple objective attributes and are candidates for subsequent experimental verification and application.

[0148] In some embodiments, in step S320, each mutation sequence in the mutation sequence library is scored and evaluated using the zero-shot scoring model to obtain a comprehensive evaluation result, including:

[0149] Step S321, using the zero-shot scoring model to evaluate each mutation sequence in the mutation sequence library, obtaining the scoring evaluation result corresponding to each scoring algorithm, and integrating all the scoring evaluation results to obtain a comprehensive evaluation result.

[0150] Among them, the scoring algorithms in the zero-shot scoring model include at least one of the protein sequence naturality algorithm, the protein structure naturality algorithm, the protein stability algorithm, the protein affinity algorithm, and the functional activity algorithm.

[0151] The scoring evaluation result includes at least one of sequence naturality, sequence rationality, structure naturality, structure stability, protein affinity, functional activity, and mutation rationality.

[0152] It should be noted that in the multi-modal protein design method, the scoring algorithm and evaluation results of the zero-shot scoring model can comprehensively consider multiple aspects to comprehensively evaluate the characteristics of protein sequences and structures. The scoring algorithms can include, but are not limited to, protein sequence naturalness algorithm, structure naturalness algorithm, stability algorithm, affinity algorithm, functional activity algorithm, as well as additional sequence diversity algorithm, sequence conservation algorithm, structure prediction algorithm, protein-protein interaction algorithm, protein-ligand binding algorithm, thermodynamic stability algorithm, and kinetic stability algorithm, etc. These algorithms can evaluate the sequence rationality, structural rationality, structural stability, protein affinity, functional activity, mutation rationality, as well as sequence variability, structural variability, binding affinity, thermodynamic stability, kinetic stability, and functional prediction scores of proteins. By integrating this multi-modal information, the scoring evaluation results can not only provide an assessment of the naturalness and rationality of protein sequences and structures, but also predict the potential impact of mutations on protein functions and structures, thus providing important guidance and decision-making support in the process of protein design and optimization.

[0153] In this step, the zero-shot scoring model is used to evaluate each mutant sequence in the mutant sequence library to determine its performance on different target attributes. Each mutant sequence in the mutant sequence library is input into the zero-shot scoring model. Among them, the zero-shot scoring model evaluates each mutant sequence according to its internal multiple scoring algorithms.

[0154] Then, the evaluation results of each scoring algorithm for the mutant sequence are extracted from the zero-shot scoring model. The zero-shot scoring model may contain multiple scoring algorithms, such as protein sequence naturalness algorithm, protein structure naturalness algorithm, protein stability algorithm, protein affinity algorithm, and functional activity algorithm, etc. For each mutant sequence, the evaluation results of these scoring algorithms are extracted separately.

[0155] For example, when designing the zero-shot scoring model, ensure that the output of each scoring algorithm can be extracted separately. For each mutant sequence, record the output value of each scoring algorithm separately.

[0156] Finally, the scoring evaluation results of each mutant sequence on different target attributes are integrated to obtain a comprehensive evaluation result. For each mutant sequence, the scoring evaluation results on different target attributes are weighted averaged, voting mechanism, or other integration methods. Determine the weights of each target attribute and adjust the weight allocation according to actual needs. Result: Obtain the comprehensive evaluation results of each mutant sequence, which reflects its overall performance on multiple target attributes. This step comprehensively considers multiple target attributes and provides a more comprehensive evaluation; the weight allocation can be adjusted according to actual needs to adapt to different design goals.

[0157] Specifically, the zero-shot scoring model contains multiple scoring algorithms for evaluating the performance of mutant sequences on different target attributes. Among them, the protein sequence naturalness algorithm can evaluate whether the mutant sequence is evolutionarily reasonable. The protein structure naturalness algorithm can evaluate whether the three-dimensional structure of the mutant sequence is stable. The protein stability algorithm can evaluate the thermal stability and chemical stability of the mutant sequence. The protein affinity algorithm can evaluate the binding affinity of the mutant sequence with other molecules (such as substrates, ligands, antigens, etc.). The functional activity algorithm can evaluate the functional activity of the mutant sequence, such as the catalytic efficiency of an enzyme, the binding affinity of an antibody, etc. Correspondingly, the scoring evaluation results contain the evaluation results of multiple target attributes for comprehensively evaluating the performance of the mutant sequence.

[0158] Further, in step S330, based on the comprehensive evaluation result, using a multi-objective optimization algorithm to optimize each of the mutant sequences to obtain the optimized recommended set of candidate protein sequences, including:

[0159] Step S331, obtaining the scoring evaluation result corresponding to each mutant sequence in the comprehensive evaluation result as the objective function; and determining the constraint conditions and decision variables, and taking the objective function, the constraint conditions and the decision variables as the actual demand objectives; the constraint conditions include the candidate amino acid types that can be mutated, the mutation quantity limit, the structural stability requirement and the functional activity requirement; the decision variables include the mutation sites and the amino acid types;

[0160] As described above, the scoring evaluation result obtained by evaluating each mutant sequence in the mutant sequence library using the zero-shot scoring model is used as the objective function of the multi-objective optimization algorithm. Specifically, the scoring evaluation results of each mutant sequence on different target attributes are extracted from the comprehensive evaluation result of the zero-shot scoring model, and these results will be used as the objective function of the optimization algorithm to obtain the objective function values of each mutant sequence, and these values will be used in the subsequent optimization process. The objective function is the basis of the optimization algorithm, and a clear objective function helps the optimization algorithm to find the optimal solution more accurately. Specifically, the scoring evaluation result can be directly extracted from the output of the zero-shot scoring model and used as the value of the objective function.

[0161] In this step, the constraint conditions and decision variables in the multi-objective optimization problem are defined, and the objective function, the constraint conditions and the decision variables are integrated into the actual demand objectives.

[0162] Among them, the constraint conditions define the conditions that need to be satisfied during the optimization process, such as the candidate amino acid types that can be mutated, the mutation quantity limit, the structural stability requirement and the functional activity requirement, etc.

[0163] Among them, the decision variables define the variables that need to be optimized, such as the mutation sites and the amino acid types.

[0164] As described above, the actual demand target is to integrate the objective function, constraints, and decision variables into the actual demand target, and clarify the specific form of the optimization problem.

[0165] In this step, clarifying the specific form of the optimization problem provides clear guidance for subsequent optimization algorithms, helps the optimization algorithms search the solution space more efficiently, and find the optimal solution that meets the actual requirements.

[0166] Step S332: According to the objective function, randomly generate an initial mutation sequence as the initial population, and calculate the objective function values of the sequences in the initial population.

[0167] In this step, randomly generate a group of initial mutation sequences as the initial population of the genetic algorithm, and calculate the objective function values of these sequences. Randomly generate the initial population by specifically randomly selecting mutation sites and amino acid types within the mutatable candidate region to generate a group of initial mutation sequences. For each mutation sequence in the initial population, calculate its objective function value, thereby obtaining the initial population and its objective function values, providing a starting point for subsequent optimization processes.

[0168] The diversity of the initial population helps the optimization algorithm search the solution space more comprehensively and avoid falling into local optimal solutions. A random number generator can be used to randomly select mutation sites and amino acid types within the mutatable region to generate the initial population. Then call the objective function to calculate the value of each sequence.

[0169] Step S333: Based on the objective function values, perform non-dominated sorting on the initial mutation sequences in the initial population to divide the initial mutation sequences into multiple non-dominated levels.

[0170] In this step, perform non-dominated sorting on the mutation sequences in the initial population according to the objective function values to divide the sequences into multiple non-dominated levels.

[0171] As described above, non-dominated sorting means comparing the objective function values of each sequence in the initial population to determine which sequences are non-dominated (i.e., not worse than other sequences in one objective and at least better than other sequences in one objective). Divide the non-dominated sequences into the first level, and then continue non-dominated sorting among the remaining sequences to divide them into the second level, and so on, thereby obtaining the non-dominated levels of the initial population, providing a basis for subsequent selection operations. Non-dominated sorting helps the optimization algorithm find a balance among multiple objectives and avoid over-optimizing a certain objective while ignoring other objectives.

[0172] Specifically, use a non-dominated sorting algorithm to sort the initial population. For example, for two sequences A and B, if A is not worse than B in all objectives and at least better than B in one objective, then A dominates B.

[0173] Step S334, taking the dimension of the objective function and the set of objective function values as the target space, and determining reference points in the target space; the number of the reference points is determined according to the dimension of the objective function and the preset target number.

[0174] Step S335 , performing reference point association processing on each of the initial mutation sequences and the nearest reference point, and calculating the distance between the individual objective function value and the reference point.

[0175] In this step, a reference point is determined in the target space, and each mutant sequence is associated with the nearest reference point, and the distance between the individual objective function value and the reference point is calculated.

[0176] According to the dimension of the objective function and the preset number of targets, a set of reference points are uniformly distributed in the target space. The distance between the objective function value of each mutation sequence and each reference point is calculated, and each sequence is associated with the nearest reference point to obtain the association relationship between each mutation sequence and the reference point, as well as the distance between the individual objective function value and the reference point.

[0177] The introduction of reference points helps the optimization algorithm to distribute solutions more evenly in the target space and improve the diversity of solutions. Specifically, the distance calculation formula (such as Euclidean distance) can be used to calculate the distance between the objective function value of each sequence and the reference point.

[0178] Step S336, selecting individuals according to the non-dominated hierarchy and the distance between the individual objective function value and the reference point, performing crossover and mutation operations, generating a selected mutation sequence, and forming a progeny population.

[0179] As mentioned above, individuals with higher non-dominated levels are preferentially selected according to the non-dominated levels. For individuals at the same level, individuals with larger distance are selected according to the distance between the individual objective function value and the reference point.

[0180] Among them, a crossover operation is used, such as selecting two parent individuals and generating one or more offspring individuals by exchanging some genetic information. Also, a mutation operation is used to randomly mutate the parent individuals, introduce new genetic information, increase the diversity of the population, and thus generate offspring populations, providing new candidate solutions for the subsequent optimization process.

[0181] Selection, crossover, and mutation operations help the optimization algorithm to gradually search for better solutions while maintaining the diversity of solutions. For example, roulette selection, tournament selection, and other methods can be used for selection operations; single-point crossover, multi-point crossover, and other methods can be used for crossover operations; and random mutation and other methods can be used for mutation operations.

[0182] Step S337: Use the initial population as the parent population; and merge the parent population and the offspring population to obtain the merged population.

[0183] As described above, using the initial population as the parent population and merging the parent population and the offspring population to form a new population. The population can be merged, that is, all individuals in the parent population and the offspring population are merged into one population, so as to obtain the merged population, providing a wider range of candidate solutions for the subsequent optimization process. Merging the population helps the optimization algorithm to search the solution space more comprehensively and avoid premature convergence to local optimal solutions. Specifically, the individuals in the parent population and the offspring population can be simply merged into a list.

[0184] Step S338: Repeat the non-dominated sorting process and the reference point association process on the merged population until a preset stop condition is reached, and obtain the optimized recommended sequence set.

[0185] This step is an iterative process of repeated optimization of the recommended sequence set. The non-dominated sorting and reference point association processes are repeated on the merged population until the preset stop condition is met; the non-dominated sorting and reference point association processes are repeated on the merged population, and individuals are selected for crossover and mutation operations to generate a new offspring population; it is judged whether the preset stop condition is met, such as reaching the maximum number of iterations, the diversity of the population is lower than a certain threshold, etc., so as to obtain the optimized recommended sequence set.

[0186] Repeating the optimization process helps the optimization algorithm gradually approach the global optimal solution and improve the quality of the solution.

[0187] The specific implementation method can be to set a loop, and each loop executes the non-dominated sorting, reference point association, selection, crossover, and mutation operations once until the stop condition is met.

[0188] Step S339: Determine the Pareto optimal solutions in the recommended sequence set; and based on the actual demand target, screen sequences from the Pareto optimal solutions to obtain the candidate protein sequence recommendation set.

[0189] In this step, first, the Pareto optimal solutions are determined from the optimized recommended sequence set. The Pareto front is extracted, that is, the non-dominated individuals are extracted from the recommended sequence set to form the Pareto front, thereby obtaining the Pareto optimal solution set. These solutions achieve the best balance among multiple objectives. The Pareto optimal solution set provides multiple choices for practical applications, and the optimal solution can be selected from them according to specific requirements. Specifically, the recommended sequence set can be sorted by non-dominance, and the individuals at the first level are extracted as the Pareto optimal solutions. Then, according to the actual demand objectives, the sequences that meet specific conditions can be screened out from the Pareto optimal solutions to form a candidate protein sequence recommendation set. Specifically, according to the actual demand objectives, such as structural stability requirements, functional activity requirements, etc., the sequences that meet the conditions can be screened out from the Pareto optimal solutions.

[0190] Example 4:

[0191] Referring to Figure 5 , this embodiment provides a multimodal protein design method. Based on the above Example 3, after step S300, based on the mutable candidate region and the original sequence information, and obtaining the candidate protein sequence recommendation set by using the trained zero-shot scoring model, it further includes:

[0192] Step S500, obtaining the experimental data of the verification experiment of the mutant sequences in the mutant sequence library based on the comprehensive evaluation result.

[0193] In protein design, although the zero-shot scoring model can quickly generate a candidate sequence recommendation set without experimental data, its prediction results may be affected by the limitations of the model itself. For example, the modeling of some complex biochemical processes is not accurate enough.

[0194] In this embodiment, by introducing experimental data to iteratively optimize the multimodal objective attribute scoring model, the actual experimental results can be fed back into the model to further adjust and optimize the model parameters, thereby improving the prediction accuracy of the model for protein performance. This iterative optimization not only enhances the model's understanding of complex biochemical processes but also enables the model to more comprehensively evaluate the performance of protein sequences, considering multiple factors such as sequence naturalness, structural stability, and functional activity. In addition, through iterative optimization, the multimodal objective attribute scoring model can better adapt to new design requirements and experimental conditions, improving the generalization ability of the model. Ultimately, this process helps to reduce the number of experimental verifications, lower the experimental cost and time, and improve the overall efficiency of protein design.

[0195] In this step, the sequences in the mutant sequence library comprehensively evaluated by the zero-shot scoring model are experimentally verified to obtain experimental data. Some or all of the mutant sequences are selected from the mutant sequence library for experimental verification. The experimental verification may include, but is not limited to, functional activity tests, stability tests, affinity tests, etc., to obtain the actual performance data of the sequences, so as to obtain the performance data of each mutant sequence in the actual experiment, and these data will be used for subsequent model training and optimization.

[0196] Verifying the prediction results of the model through the experimental data of the verification experiment can improve the accuracy and reliability of the model. The experimental data provides real feedback for the iterative optimization of the model, which helps to further improve the model performance.

[0197] Specifically, an experimental scheme can be designed to test the functional activity, stability, etc. of the mutant sequences. Collect experimental data, including quantitative activity values, stability scores, etc.

[0198] Step S600, construct a multi-modal target attribute scoring model, and use the experimental data to train the multi-modal target attribute scoring model.

[0199] In this step, a multi-modal target attribute scoring model is constructed and trained based on the experimental data. This model can comprehensively consider multiple target attributes to score the mutant sequences.

[0200] First, construct a multi-modal target attribute scoring model, which can be a machine learning-based model, such as a neural network, a support vector machine, etc.; then, use the experimental data to train the model, adjust the model parameters to minimize the prediction error, so as to obtain a trained multi-modal target attribute scoring model, which can score the mutant sequences more accurately.

[0201] In this step, a multi-modal target attribute scoring model is constructed and trained using the experimental data, comprehensively considering multiple target attributes to provide a more comprehensive evaluation; training the model based on the experimental data to improve the accuracy and reliability of the model. Specifically, a machine learning framework (such as TensorFlow, PyTorch) can be used to construct the model. Use the experimental data as the training set and train the model through an optimization algorithm (such as gradient descent).

[0202] Step S700, perform iterative optimization on the multi-modal target attribute scoring model based on the verification experiment to obtain the updated multi-modal target attribute scoring model, and obtain an optimized candidate protein sequence recommendation set based on the updated multi-modal target attribute scoring model.

[0203] Iteratively optimize the multi-modal target attribute scoring model using the experimental data from the validation experiment, update the model parameters, and re-evaluate the candidate protein sequence recommendation set using the updated model.

[0204] In this step, use the experimental data to evaluate the multi-modal target attribute scoring model, calculate the prediction error of the model; adjust the model parameters according to the prediction error for iterative optimization; then, use the updated model to re-score and rank the candidate protein sequence recommendation set to obtain an optimized recommendation set, thereby obtaining an optimized multi-modal target attribute scoring model and an optimized candidate protein sequence recommendation set based on this model.

[0205] Improve the prediction accuracy of the model through iterative optimization; obtain a higher-quality candidate protein sequence recommendation set based on the optimized model.

[0206] Specifically, methods such as cross-validation can be used to evaluate the model performance and calculate the prediction error; use an optimization algorithm (such as gradient descent) to adjust the model parameters for iterative optimization; use the updated model to re-score and rank the candidate protein sequence recommendation set.

[0207] In summary, through the above steps in this embodiment, the performance of the multi-modal target attribute scoring model can be gradually improved, enabling it to play a greater role in protein design. The optimized multi-modal target attribute scoring model can more accurately predict the performance of proteins, thereby increasing the design success rate and reducing the risk of failure; the experimental data provides direct feedback for model optimization, enabling the model to continuously learn and improve, further enhancing the prediction accuracy; the multi-modal target attribute scoring model can comprehensively consider multiple target attributes, providing a more comprehensive evaluation and avoiding problems that may arise from single-target optimization; in addition, through iterative optimization, the model can flexibly adapt to different design requirements and experimental conditions, having strong scalability.

[0208] Furthermore, referring to Figure 6 , the iterative optimization includes:

[0209] Step S710, evaluate the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain a scoring evaluation result.

[0210] In this step, evaluate the candidate protein sequence recommendation set by combining the multi-modal target attribute scoring model and the zero-shot scoring model to obtain a scoring evaluation result for each sequence.

[0211] First, the multi-modal target attribute scoring model can be used to score each sequence in the candidate protein sequence recommendation set; the zero-shot scoring model can be used to score the same sequence; by integrating the scoring results of the two models, the comprehensive evaluation result of each sequence can be obtained, so as to obtain the comprehensive scoring evaluation result of each candidate protein sequence on multiple target attributes.

[0212] This step combines the advantages of the two models to provide a more comprehensive evaluation; it evaluates the sequence performance from multiple perspectives and improves the accuracy of the evaluation.

[0213] As described above, the multi-modal target attribute scoring model and the zero-shot scoring model can be used to score each sequence separately; in addition, weights can be set to perform a weighted average on the scoring results of the two models to obtain the comprehensive evaluation result.

[0214] In step S720, according to the scoring evaluation result, candidate sequences are screened to construct a candidate sequence set.

[0215] In this step, according to the comprehensive scoring evaluation result, candidate sequences with better performance are screened out to construct a candidate sequence set.

[0216] First, sequences with better performance can be selected from the candidate protein sequence recommendation set according to the set threshold or ranking; then a candidate sequence set containing these sequences is constructed for subsequent experimental verification and model optimization, so as to obtain a candidate sequence set containing sequences with better performance.

[0217] Through screening, sequences that are more likely to meet the design requirements can be obtained; the number of sequences that need to be experimentally verified is reduced, and the efficiency is improved.

[0218] Specifically, screening criteria can be set, such as sequences with a score higher than a certain threshold or ranked top N; sequences are selected from the recommendation set according to the criteria to construct a candidate sequence set.

[0219] In step S730, the experimental data of the verification experiment corresponding to each candidate sequence in the candidate sequence set is obtained, and the multi-modal target attribute scoring model is updated using the experimental data.

[0220] As described above, each sequence in the candidate sequence set is experimentally verified to obtain experimental data, and these data are used to update the multi-modal target attribute scoring model.

[0221] It should be noted that the experimental data in this step is the experimental data obtained from the current verification experiment for the sequence. If in subsequent iterative processes, it is necessary to re-perform the verification experiment for the screened sequences each time to obtain the corresponding new experimental data.

[0222] In this step, each sequence in the candidate sequence set is experimentally verified to obtain its performance data in actual experiments; and, these experimental data are used to update the multi-modal target attribute scoring model, adjusting the model parameters to improve the prediction accuracy, so as to obtain an updated multi-modal target attribute scoring model, which can more accurately predict the sequence performance.

[0223] Updating the model with experimental data improves the accuracy and reliability of the model; and, the experimental data provides real feedback for model optimization, which helps to further improve the model performance.

[0224] Functional activity, stability, etc. of the candidate sequences can be tested, experimental data can be collected, and these data are used to update the multi-modal target attribute scoring model.

[0225] Step S740, using the updated multi-modal target attribute scoring model and the zero-shot scoring model, evaluate the candidate protein sequence recommendation set to obtain a new scoring evaluation result.

[0226] In this step, the updated multi-modal target attribute scoring model and the zero-shot scoring model are used to re-evaluate the candidate protein sequence recommendation set to obtain a new scoring evaluation result.

[0227] Specifically, the updated multi-modal target attribute scoring model can be first used to score each sequence in the candidate protein sequence recommendation set. Then, the zero-shot scoring model is used to score the same sequence, and then the scoring results of the two models are combined to obtain a new comprehensive evaluation result for each sequence, so that a new comprehensive scoring evaluation result of each candidate protein sequence on multiple target attributes can be obtained.

[0228] This step improves the accuracy of the evaluation by updating the model; combines the advantages of the two models to provide a more comprehensive evaluation; scores each sequence separately using the updated multi-modal target attribute scoring model and the zero-shot scoring model; sets weights, and performs weighted averaging on the scoring results of the two models to obtain a new comprehensive evaluation result.

[0229] Step S750, using the new scoring evaluation result to screen out target sequences, construct a new candidate protein sequence recommendation set, and obtain new experimental data corresponding to the new candidate protein sequence recommendation set;

[0230] In this step, according to the new comprehensive scoring evaluation result, target sequences with better performance are screened out, a new candidate protein sequence recommendation set is constructed, and new experimental data of these sequences are obtained.

[0231] First, sequences with better performance can be selected from the candidate protein sequence recommendation set according to the set threshold or ranking; then, a new candidate protein sequence recommendation set containing these sequences is constructed; each sequence in the new candidate sequence set is experimentally verified to obtain its performance data in actual experiments, so as to obtain a new candidate protein sequence recommendation set and its corresponding experimental data.

[0232] Sequences that are more likely to meet the design requirements are obtained through screening; the number of sequences that need to be experimentally verified is reduced, improving efficiency; in addition, the model is further optimized through experimental data to improve the model performance.

[0233] Specifically, screening criteria can be set, such as sequences with a score higher than a certain threshold or ranked top N. Sequences are selected from the recommendation set according to the criteria to construct a new candidate sequence set; the new candidate sequences are tested for functional activity, stability, etc.; experimental data is collected for subsequent model optimization.

[0234] Step S760, based on the new experimental data, evaluate the iterative optimization of the multi-modal target attribute scoring model to facilitate the completion of the iterative optimization.

[0235] In this step, the performance of the updated multi-modal target attribute scoring model is evaluated using the new experimental data. This can include but is not limited to checking the prediction accuracy, generalization ability, and stability of the model for the new data, etc. By comparing the performance of the model on the new data with the previous performance metrics, it is determined whether the model has improved. After checking the convergence conditions, if the performance of the model reaches the target, or there is no significant improvement in several consecutive iterations, the iterative optimization is considered complete. At this time, the iterative process can be stopped, and the final multi-modal target attribute scoring model and the optimized candidate protein sequence recommendation set are output.

[0236] Through the above steps, it is ensured that the multi-modal target attribute scoring model can be effectively updated and optimized according to the new experimental data in each iteration, and finally the preset optimization goal is achieved, thereby improving the accuracy and efficiency of protein design.

[0237] Further, the step S760, based on the new experimental data, evaluate the iterative optimization of the multi-modal target attribute scoring model to facilitate the completion of the iterative optimization, includes:

[0238] Step S761, based on the new experimental data, determine whether the multi-modal target attribute scoring model reaches the preset optimization goal.

[0239] In this step, according to the newly obtained experimental data, evaluate whether the multi-modal target attribute scoring model reaches the preset optimization goal to decide whether to continue optimizing the model.

[0240] In this step, first, the prediction results of the multi-modal target attribute scoring model can be evaluated using new experimental data, and the prediction error or performance metrics of the model can be calculated. Then, the evaluation results are compared with the preset optimization goals to determine whether the model meets the requirements, so as to further determine whether the multi-modal target attribute scoring model has achieved the optimization goals, and thus decide whether to continue iterative optimization.

[0241] In this step, by setting optimization goals, the end point of model optimization is clarified to avoid endless optimization. According to whether the model has achieved the optimization goals, the optimization strategy is dynamically adjusted to improve the optimization efficiency.

[0242] Specifically, optimization goals can be set, such as the prediction error being lower than a certain threshold or the model performance improvement exceeding a certain percentage. The model is evaluated using new experimental data, and the prediction error or performance metrics are calculated. The evaluation results are compared with the optimization goals to determine whether the conditions are met.

[0243] It should be noted that in the process of protein design and optimization, the preset optimization goals refer to the criteria set during the iterative optimization of the model to determine whether the model has achieved the expected performance. These goals are usually based on experimental data and design requirements and are used to evaluate the prediction accuracy, stability, and generalization ability of the model. The preset optimization goals are the specific criteria used to judge whether the model meets the design requirements in the optimization process. These goals usually include the prediction accuracy of the model, performance improvement, error range, etc., and are used to decide whether to continue iterative optimization.

[0244] Step S770, if so, it is determined that the iterative optimization of the multi-modal target attribute scoring model is completed.

[0245] As described above, if the multi-modal target attribute scoring model has achieved the preset optimization goals, it is considered that the iterative optimization process of the model is completed. When the evaluation results show that the model has achieved the optimization goals, the iterative optimization process is stopped. The optimized model parameters are saved for subsequent use, and finally an optimized multi-modal target attribute scoring model is obtained, which can more accurately predict the performance of protein sequences.

[0246] Through the determination, unnecessary optimization steps are avoided, saving computing resources and time. The optimized model has better stability and reliability under the preset goals. When the evaluation results meet the optimization goals, the optimization loop is stopped. The final parameters of the model are saved for subsequent use.

[0247] Step S780, if not, it is determined that the iterative optimization of the multi-modal target attribute scoring model is not completed, and the process returns to step S720 to evaluate the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain a scoring evaluation result.

[0248] As described above, if the multi-modal target attribute scoring model fails to meet the preset optimization goal, it is considered that the iterative optimization of the model is not completed and further optimization is required.

[0249] In this step, when the evaluation result shows that the model does not meet the optimization goal, iterative optimization can be continued. Then, it is necessary to return to the step of evaluating the candidate protein sequence recommendation set using the multi-modal target attribute scoring model and the zero-shot scoring model in step S720, and re-perform the scoring evaluation, and then continue the optimization process until the model meets the preset optimization goal.

[0250] Through the above continuous optimization, the performance of the model is gradually improved until the design requirements are met; the optimization strategy is dynamically adjusted according to the performance of the model to ensure that the model can adapt to different design requirements.

[0251] Specifically, when the evaluation result does not meet the optimization goal, the optimization loop can be continued; the updated model is used to re-evaluate the candidate protein sequence recommendation set to obtain a new scoring evaluation result.

[0252] For example, the preset optimization goal is to increase the enzyme activity by at least 3 times. First, the candidate protein sequence recommendation set is evaluated using the multi-modal target attribute scoring model and the zero-shot scoring model. Currently, there is a set of candidate sequences, and each sequence is scored by these two models. The obtained scoring evaluation results include scores in multiple dimensions such as sequence naturality, structural stability, and functional activity. According to these scoring evaluation results, the candidate sequences with better performance are selected to construct a candidate sequence set, and 20 sequences with higher comprehensive scores are selected from 100 candidate sequences as the candidate sequence set.

[0253] Experimental verification is performed on these 20 candidate sequences to obtain experimental data. The experiment mainly tests the activity of the enzymes corresponding to these sequences and compares it with the activity of the seed enzyme. The experimental results show that the enzyme activities of 5 of these sequences have increased to a certain extent, but none of them have reached the goal of more than 3 times. The highest enzyme activity of one sequence has increased by 2.5 times. The experimental data (including performance indicators such as enzyme activity) of these 20 candidate sequences are used to update the multi-modal target attribute scoring model. The model parameters are adjusted through an optimization algorithm (such as gradient descent) to enable the model to better predict target attributes such as enzyme activity.

[0254] Using the updated multi-modal target attribute scoring model and zero-shot scoring model, evaluate the candidate protein sequence recommendation set (which can be a newly generated candidate sequence set here, or a re-evaluation of the sequences that were not filtered out before), and obtain a new scoring evaluation result. According to the new scoring evaluation result, screen out the target sequences again and construct a new candidate protein sequence recommendation set. This time, 15 sequences are screened out from the new candidate sequences. Conduct experimental verification on these 15 new candidate sequences to obtain new experimental data. The experimental results show that the enzyme activity of 3 of these sequences has been further improved, and the enzyme activity of one sequence has reached more than 3 times that of the seed enzyme activity, meeting the preset optimization goal. Based on these new experimental data, evaluate the iterative optimization of the multi-modal target attribute scoring model. Since one sequence has met the preset optimization goal, it can be judged that the iterative optimization of the model has achieved the expected effect and the optimization process is completed. From the sequences that meet the optimization goal, considering other factors (such as structural stability, etc.) comprehensively, determine the final target mutant protein sequence. Select the sequence with an enzyme activity increase of more than 3 times and good structural stability as the final result.

[0255] Through the above process, the optimization of the protein sequence with the goal of at least a 3-fold increase in enzyme activity is achieved, and at the same time, the application and effect of the multi-modal target attribute scoring model in the iterative optimization process are demonstrated.

[0256] Example 5:

[0257] Refer to Figure 7 , this example provides a multi-modal protein design method. Based on Example 4 above, in step S400, the candidate protein sequence recommendation set is subjected to clustering and dimensionality reduction processing according to the zero-shot scoring model, and sampling and screening are performed according to the latent space representation to obtain the target mutant protein sequence, including:

[0258] Step S410, embed the protein sequences in the candidate protein sequence recommendation set using the zero-shot scoring model or the multi-modal target attribute scoring model to obtain high-dimensional representation information in the latent space.

[0259] In this step, the optimized protein sequence is embedded using the zero-shot scoring model or the multi-modal target attribute scoring model to generate high-dimensional representation information in the latent space.

[0260] Among them, using "the zero-shot scoring model or the multi-modal target attribute scoring model" for embedding processing is mainly to provide a more flexible and comprehensive protein sequence representation method. These two models have their own characteristics and applicable scenarios. Selecting one of them or using them in combination can better meet different design requirements and optimization goals.

[0261] The zero-shot scoring model is characterized by the fact that it does not require experimental data: The zero-shot scoring model can directly evaluate and score protein sequences without experimental data. This is very useful for quickly generating candidate sequences in the initial design stage, especially in the absence of experimental verification. Quickly generate candidate sequences: The zero-shot scoring model can quickly process a large number of sequences and generate a recommended set of candidate sequences, which is suitable for large-scale preliminary screening.

[0262] For the multi-modal target attribute scoring model, it synthesizes multiple target attributes. The multi-modal target attribute scoring model can comprehensively consider multiple target attributes (such as sequence naturalness, structural stability, functional activity, etc.) and provide a more comprehensive evaluation. This is crucial for optimizing the comprehensive performance of proteins. Utilize experimental data: The multi-modal target attribute scoring model can be trained and optimized through experimental data to further improve the accuracy and reliability of the model.

[0263] Input the optimized protein sequence into the zero-shot scoring model or the multi-modal target attribute scoring model; through its internal neural network or other embedding mechanisms, the model maps each protein sequence into a high-dimensional latent space, generating high-dimensional representation information, so as to obtain the high-dimensional representation information of each protein sequence in the latent space. These information contain the key features and attributes of the sequence.

[0264] The high-dimensional representation information can capture the complex features and attributes of protein sequences; compress the complex protein sequence information into a high-dimensional space for subsequent processing; use pre-trained neural network models (such as ESM-2, ProtBERT, etc.) for embedding; through the forward propagation of the model, the protein sequence can be mapped into a high-dimensional latent space.

[0265] Step S420, perform clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space to obtain low-dimensional representation information.

[0266] In this step, clustering and dimensionality reduction processing are performed on the high-dimensional representation information in the latent space to generate low-dimensional representation information, which is convenient for subsequent sampling and screening.

[0267] Clustering algorithms (such as K-means, DBSCAN, etc.) can be used to cluster the high-dimensional representation information, grouping similar sequences into different clusters; and, dimensionality reduction methods (such as PCA, t-SNE, UMAP, etc.) are used to map the high-dimensional representation information into a low-dimensional space to obtain low-dimensional representation information, so that the representation information of each protein sequence in the low-dimensional space can be obtained. These information retain the key features of the sequence while reducing the dimension of the data.

[0268] The low-dimensional representation information simplifies the data structure and is convenient for subsequent processing; the computational cost in the low-dimensional space is lower, improving the processing efficiency.

[0269] For example, the K-means algorithm is used to cluster the high-dimensional representation information, and PCA or t-SNE is used to reduce the dimension of the high-dimensional representation information.

[0270] Step S430: In the low-dimensional space, sampling is performed according to the distribution of the low-dimensional representation information to generate a sampling and screening result; wherein, the sampling and screening result includes the selected candidate mutant sequences.

[0271] In this step, in the low-dimensional space, sampling is performed according to the distribution of the low-dimensional representation information to generate a sampling and screening result, and representative candidate mutant sequences are selected.

[0272] Specifically, analyze the distribution of the low-dimensional representation information to determine the sampling strategy; perform uniform sampling or distribution-based sampling in the low-dimensional space to generate a set of candidate mutant sequences; according to the sampling result, select the representative candidate mutant sequences, which are representative in the low-dimensional space.

[0273] As described above, the selected candidate mutant sequences are representative in the low-dimensional space and can cover different characteristic regions; through the sampling strategy, the diversity of the generated candidate sequences is ensured.

[0274] Uniform sampling or distribution-based sampling methods can be used; according to the sampling result, select the representative candidate mutant sequences.

[0275] Step S440: Determine the target mutant protein according to the sampling and screening result.

[0276] As described above, according to the sampling and screening result, determine the final target mutant protein sequence.

[0277] In this step, select the candidate mutant sequence with the best performance from the sampling and screening result; according to the actual demand target, further evaluate and verify the candidate mutant sequence; determine the final target mutant protein sequence, thereby obtaining the optimized and selected target mutant protein sequence, which performs excellently in multiple target attributes.

[0278] The finally determined mutant protein sequence performs excellently in multiple target attributes and meets the design requirements; through sampling and screening, the target mutant protein sequence is efficiently determined.

[0279] Specifically, the sampling and screening result can be evaluated to select the candidate mutant sequence with the best performance. Conduct experimental verification to ensure that the finally determined mutant protein sequence meets the design requirements.

[0280] In some embodiments, refer to Figure 8, in step S420, performing clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space, including:

[0281] Step S421, using a clustering algorithm to perform clustering on the high-dimensional representation information in the latent space, and grouping similar sequences into different clusters;

[0282] Using a clustering algorithm to perform clustering on the high-dimensional representation information in the latent space, and grouping similar protein sequences into different clusters.

[0283] In this step, select a suitable clustering algorithm (such as K-means, DBSCAN, etc.); input the high-dimensional representation information into the clustering algorithm for clustering processing. Group similar sequences into different clusters, and each cluster contains similar protein sequences.

[0284] Step S422, using a dimensionality reduction method to map the clustered high-dimensional representation information into a low-dimensional space to obtain the low-dimensional representation information; wherein, the dimensionality reduction method includes at least one of principal component analysis, t-stochastic neighborhood embedding, and UMAP non-linear dimensionality reduction technology.

[0285] In this step, use a dimensionality reduction method to map the clustered high-dimensional representation information into a low-dimensional space to generate low-dimensional representation information.

[0286] A suitable dimensionality reduction method (such as PCA, t-SNE, UMAP, etc.) can be selected. Input the clustered high-dimensional representation information into the dimensionality reduction method for dimensionality reduction processing; obtain low-dimensional representation information, thereby obtaining the representation information of each protein sequence in the low-dimensional space. These information retain the key features of the sequence while reducing the dimension of the data.

[0287] The low-dimensional representation information simplifies the data structure and facilitates subsequent processing; the computational cost in the low-dimensional space is lower, improving the processing efficiency.

[0288] Specifically, PCA or t-SNE can be used for dimensionality reduction; input the clustered high-dimensional representation information into PCA or t-SNE for dimensionality reduction processing.

[0289] Example 6:

[0290] To better illustrate the multi-modal protein design method provided in Examples 1 to 6, in this example, by constructing a multi-modal pre-training target attribute scoring model, encoding protein sequences, and then performing prediction through a fusion network.

[0291] 1. Experimental Design: (1) Dataset Preparation: Collect a set of protein sequences with known functions and their corresponding structural information (in PDB format). Ensure that the dataset has sufficient diversity to cover different protein families and functional types.

[0292] (2) Model Preparation: Prepare a multi-modal pre-trained target property scoring model (abbreviated as the multi-modal model in the figure), including a structure encoder, an MSA encoder, a sequence encoder, and a fusion network. Prepare other baseline models, including: ESM2_3B (a PLM model), RDE (a structure pre-training model), PGLM_3B (a PLM model), MSA_Transformer (a hybrid PLM model containing MSA data), etc.

[0293] (3) Experimental Settings: Divide the dataset into a training set, a validation set, and a test set. Set evaluation metrics such as Spearman rank correlation, mean squared error (MSE), etc.

[0294] 2. Experimental Process: All models are trained under the same training conditions and on the same dataset. Monitor the performance of each model on the training set and the validation set, and select the optimal model by minimizing the loss. In each iteration, record the loss values and evaluation metrics during the training process, and test the model after the training is completed.

[0295] 3. Analysis of Experimental Results: Refer to Figure 9 , which shows the performance comparison of different models in the protein sequence prediction task, using the Spearman's rank correlation coefficient as the evaluation metric. The Spearman's rank correlation coefficient is a non-parametric statistic that measures the correlation between two ranked variables. Its value ranges from -1 (perfect negative correlation) to 1 (perfect positive correlation), and the closer the value is to 1, the better the correlation between the model's prediction results and the experimental data.

[0296] The multi-modal pre-trained target property scoring model performs best among all models, with a Spearman's rank correlation coefficient reaching 0.851, significantly higher than other models. This indicates that the multi-modal model has significant advantages in integrating protein sequence, structure, and evolutionary information, and can more accurately predict properties such as protein function and stability. The MSA_Transformer model ranks second, with a Spearman's rank correlation coefficient of 0.835. Although it is lower than the multi-modal model, it is still better than the ESM2_3B and RDE models. The PGLM_3B model performs moderately, with a Spearman's rank correlation coefficient of 0.675. The ESM2_3B and RDE models perform relatively poorly, with Spearman's rank correlation coefficients of 0.589 and 0.594 respectively, indicating that these models have limited prediction accuracy when dealing with protein sequence and structure information.

[0297] It can be seen that the excellent performance of the multi-modal pre-trained target attribute scoring model verifies the effectiveness of the multi-modal integration architecture proposed in the present invention. By combining the protein sequence pre-trained model, the structure pre-trained model, and the MSA pre-trained model based on evolutionary information, and using the LORA method for efficient fine-tuning of pre-trained parameters, this architecture can efficiently optimize target attributes with the least amount of experimental data. Compared with other models, the multi-modal model not only performs well in integrating sequence information, structural information, and evolutionary information between sequences in large-scale pre-training, but also maintains high prediction accuracy in the case of limited experimental data, which makes the model more advantageous in practical applications.

[0298] The efficient performance of the multi-modal model indicates that the model can quickly and accurately predict the target attributes of proteins in the case of limited experimental data. This is particularly important for protein design and optimization, as it reduces the need for experimental verification and improves the efficiency and cost-effectiveness of research.

[0299] The Gaussian Pearson rank correlation coefficient of the multi-modal model also shows that the model has good generalization ability and can adapt to different protein families and mutation regions, and can also show good performance even in unseen mutation regions or new protein families.

[0300] In summary, Figure 9 the experimental results support the effectiveness of the multi-modal protein design method of the present invention. By integrating multiple pre-trained models and using a small amount of experimental data for fine-tuning, the multi-modal model significantly improves the efficiency and accuracy of protein design and optimization.

[0301] In addition, referring to Figure 10 , an embodiment of the present application also provides a multi-modal protein design device, including: a data acquisition module 10 for acquiring the original sequence information of the target protein; wherein, the original sequence information includes the amino acid sequence and the PDB structure information corresponding to the amino acid sequence; a candidate region module 20 for determining the mutable candidate regions of the target protein according to the original sequence information; a sequence recommendation module 30 for obtaining a candidate protein sequence recommendation set based on the mutable candidate regions and the original sequence information by using a trained zero-shot scoring model; the zero-shot scoring model includes scoring algorithms for different target attributes; a sampling generation module 40 for performing clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model, and performing sampling and screening according to the latent space representation to obtain the target mutant protein sequence.

[0302] The present application also provides a computer device. Exemplarily, the computer device includes a processor and a memory. Among them, the memory stores a computer program, and the processor executes the above multi-modal protein design method or the functions of each module in the above multi-modal protein design device by running the computer program.

[0303] Among them, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0304] The memory can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Among them, the memory is used to store the computer program, and after receiving the execution instruction, the processor can execute the computer program accordingly.

[0305] The present application also provides a computer storage medium for storing the computer program used in the above computer device. Among them, the computer storage medium can be a readable storage medium, a non-volatile storage medium, or a volatile storage medium. For example, the computer storage medium can include, but is not limited to: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which are various media that can store program codes.

[0306] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, as well as the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0307] In addition, in each embodiment of this application, each functional module or unit can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0308] If the above functions are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application.

[0309] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application.

Claims

1. A multimodal protein design method, characterized in that, Including: Obtaining the original sequence information of the target protein; wherein, the original sequence information includes the amino acid sequence and the PDB structure information corresponding to the amino acid sequence; Determining the mutatable candidate regions of the target protein according to the original sequence information; Performing mutations within the mutatable candidate regions based on the original sequence information to generate different mutant sequences, obtaining a mutant sequence library; Evaluating each mutant sequence in the mutant sequence library using a zero-shot scoring model to obtain the scoring evaluation results corresponding to each scoring algorithm, and integrating all the scoring evaluation results to obtain a comprehensive evaluation result; wherein, the zero-shot scoring model includes scoring algorithms for different target attributes, including at least one of a protein sequence naturalness algorithm, a protein structure naturalness algorithm, a protein stability algorithm, a protein affinity algorithm, and a functional activity algorithm; the scoring evaluation results include at least one of sequence naturalness, sequence rationality, structure naturalness, structure stability, protein affinity, functional activity, and mutation rationality; Optimizing each mutant sequence using a multi-objective optimization algorithm based on the comprehensive evaluation result to obtain an optimized candidate protein sequence recommendation set; Obtaining the experimental data of the verification experiments of the mutant sequences in the mutant sequence library based on the comprehensive evaluation result; Constructing a multi-modal target attribute scoring model and training the multi-modal target attribute scoring model using the experimental data; Iteratively optimizing the multi-modal target attribute scoring model based on the verification experiments to obtain the updated multi-modal target attribute scoring model, and obtaining an optimized candidate protein sequence recommendation set based on the updated multi-modal target attribute scoring model; wherein, the iterative optimization includes: Evaluating the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain the scoring evaluation results; Screening candidate sequences according to the scoring evaluation results to construct a candidate sequence set; Obtaining the experimental data of the verification experiments corresponding to each candidate sequence in the candidate sequence set and updating the multi-modal target attribute scoring model using the experimental data; Evaluating the candidate protein sequence recommendation set using the updated multi-modal target attribute scoring model and the zero-shot scoring model to obtain new scoring evaluation results; Screening out target sequences using the new scoring evaluation results to construct a new candidate protein sequence recommendation set, and obtaining new experimental data corresponding to the new candidate protein sequence recommendation set; Evaluating the iterative optimization of the multi-modal target attribute scoring model based on the new experimental data so as to complete the iterative optimization; Performing clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model, and performing sampling and screening according to the latent space representation to obtain the target mutant protein sequence.

2. The multimodal protein design method according to claim 1, wherein The determining the mutatable candidate regions of the target protein according to the original sequence information includes: Using a protein stability model, perform single-site saturation mutagenesis scoring on the original sequence information to obtain a stability scoring result; Screen out the mutatable candidate regions under preset conditions according to the stability scoring result.

3. The multimodal protein design method according to claim 2, wherein The calculation method of the stability scoring result includes: ; where i represents the site index of the i-th amino acid in the protein sequence; S i represents the stability score result of the i-th site; μ i represents the average stability score when the i-th site mutates into 20 amino acids; σ i represents the standard deviation of the stability scores when the i-th site mutates into 20 amino acids; represents the average structural stability score of the i-th site within the sliding window range; M represents the size of the sliding window; represents the sum of the structural stability scores from the (i - M / 2)-th site to the (i + M / 2)-th site.

4. The multimodal protein design method according to claim 1, wherein Evaluating the iterative optimization of the multi-modal target attribute scoring model based on the new experimental data to facilitate the completion of the iterative optimization, including: Judging whether the multi-modal target attribute scoring model reaches a preset optimization goal based on the new experimental data; If so, it is determined that the iterative optimization of the multi-modal target attribute scoring model is completed; If not, it is determined that the iterative optimization of the multi-modal target attribute scoring model is not completed, and return to evaluate the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain a scoring evaluation result.

5. The multimodal protein design method according to claim 1, wherein Performing clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model, and performing sampling and screening according to the latent space representation to obtain a target mutant protein sequence, including: Embedding the protein sequences in the candidate protein sequence recommendation set using the zero-shot scoring model or the multi-modal target attribute scoring model to obtain high-dimensional representation information in the latent space; Performing clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space to obtain low-dimensional representation information; In the low-dimensional space, sampling is performed according to the distribution of the low-dimensional representation information to generate a sampling and screening result; wherein, the sampling and screening result includes the selected candidate mutation sequences; Determine the target mutant protein according to the sampling and screening result.

6. The multimodal protein design method according to claim 1, wherein Performing clustering and dimensionality reduction processing on the high-dimensional representation information in the latent space, including: Using a clustering algorithm to perform clustering processing on the high-dimensional representation information in the latent space, and grouping similar sequences into different clusters; Using a dimensionality reduction method, mapping the clustered high-dimensional representation information to a low-dimensional space to obtain the low-dimensional representation information; wherein, the dimensionality reduction method includes at least one of principal component analysis, t-stochastic neighborhood embedding, and UMAP non-linear dimensionality reduction technology.

7. A multimodal protein design device, characterized in that, Including: A data acquisition module for acquiring the original sequence information of the target protein; wherein, the original sequence information includes an amino acid sequence and PDB structure information corresponding to the amino acid sequence; A candidate region module for determining the mutatable candidate regions of the target protein according to the original sequence information; A sequence recommendation module is used to perform mutations within the mutable candidate region based on the original sequence information to generate different mutant sequences, obtaining a mutant sequence library; use a zero-shot scoring model to evaluate each mutant sequence in the mutant sequence library to obtain the scoring evaluation results corresponding to each scoring algorithm, and integrate all the scoring evaluation results to obtain a comprehensive evaluation result; wherein, the zero-shot scoring model includes scoring algorithms for different target attributes, including at least one of a protein sequence naturalness algorithm, a protein structure naturalness algorithm, a protein stability algorithm, a protein affinity algorithm, and a functional activity algorithm; the scoring evaluation results include at least one of sequence naturalness, sequence rationality, structure naturalness, structure stability, protein affinity, functional activity, and mutation rationality; based on the comprehensive evaluation result, use a multi-objective optimization algorithm to optimize each mutant sequence to obtain an optimized candidate protein sequence recommendation set; obtain experimental data of the verification experiment of the mutant sequences in the mutant sequence library based on the comprehensive evaluation result; construct a multi-modal target attribute scoring model and use the experimental data to train the multi-modal target attribute scoring model; perform iterative optimization on the multi-modal target attribute scoring model based on the verification experiment to obtain the updated multi-modal target attribute scoring model, and obtain an optimized candidate protein sequence recommendation set based on the updated multi-modal target attribute scoring model; wherein, the iterative optimization includes: evaluating the candidate protein sequence recommendation set according to the multi-modal target attribute scoring model and the zero-shot scoring model to obtain scoring evaluation results; screening candidate sequences according to the scoring evaluation results to construct a candidate sequence set; obtaining the experimental data of the verification experiment corresponding to each candidate sequence in the candidate sequence set and using the experimental data to update the multi-modal target attribute scoring model; using the updated multi-modal target attribute scoring model and the zero-shot scoring model to evaluate the candidate protein sequence recommendation set to obtain new scoring evaluation results; using the new scoring evaluation results to screen out target sequences to construct a new candidate protein sequence recommendation set and obtaining new experimental data corresponding to the new candidate protein sequence recommendation set; evaluating the iterative optimization of the multi-modal target attribute scoring model based on the new experimental data to facilitate the completion of the iterative optimization; A sampling generation module performs clustering and dimensionality reduction processing on the candidate protein sequence recommendation set according to the zero-shot scoring model, and performs sampling and screening according to the latent space representation to obtain a target mutant protein sequence.

8. A computer device, characterized in that, The computer device includes a processor and a memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the multi-modal protein design method according to any one of claims 1-6.

9. A computer storage medium, characterized in that, It stores a computer program, and when the computer program is executed on a processor, it implements the multi-modal protein design method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and apparatus for evolutionary data driven design of protein and other sequence defined biomolecules using machine learning

    CN114651064A

  • Generation method of target antibacterial peptide

    CN118430654A

  • Antibody modification method, device, equipment and storage medium

    CN119007800A

  • Mutant protein screening method and device and related equipment

    CN119296640A