Multifunctional nuclease and oxidoreductase reverse design method and system
By constructing a Hamiltonian matrix and a two-dimensional spectral model, and combining conditional diffusion and spectral analysis, a precise mapping from properties to structure was achieved, solving the difficulties in designing multifunctional enzymes in existing technologies, generating multifunctional enzymes that meet industrial needs, and improving the success rate of design, enzyme stability, and catalytic activity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2026-03-31
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies lack physical interpretability, and the black-box problem of property-structure mapping leads to imprecise property control and difficulties in multifunctional synergistic design, making it difficult to meet the structural requirements of nucleases and oxidoreductases in a single enzyme backbone.
A reverse design approach using multifunctional nucleases and oxidoreductases was employed. By obtaining the protein sequence and three-dimensional structure of the enzymes, a Hamiltonian matrix was constructed and a two-dimensional spectrum was calculated. The model was trained using conditional diffusion and spectral analysis models, and combined with geometric optimization and sequence design, to achieve a precise mapping from properties to structure.
This technology enables the design of highly efficient and interpretable multifunctional enzymes, which can rapidly generate enzymes that meet the needs of specific industrial environments, reduce R&D costs, and improve the success rate of design, as well as the stability and catalytic activity of the enzymes.
Smart Images

Figure CN122435997A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary fields of bioinformatics, computational biology, synthetic biology, and artificial intelligence, and specifically to a method and system for reverse design of multifunctional nucleases and oxidoreductases. Background Technology
[0002] Enzymes, as biocatalysts, play an irreplaceable role in pharmaceutical manufacturing, environmental remediation, bioenergy, and food processing. Nucleases, responsible for the hydrolysis and recombination of nucleic acids, are crucial in gene editing and biosensoring; oxidoreductases, widely involved in energy metabolism and redox reactions within organisms, are core components of industrial biocatalysis. However, naturally occurring enzymes often struggle to withstand the harsh physicochemical environments of industrial applications (such as high temperatures, extreme pH levels, and organic solvent environments) and often lack highly efficient catalytic capabilities for specific non-natural substrates. Therefore, obtaining multifunctional enzymes that combine high stability with specific catalytic activities is a core problem urgently needing to be solved in the field of enzyme engineering.
[0003] Currently, some technologies have attempted to utilize machine learning methods for protein sequence design. For example, prior art document CN114651064A discloses a biomolecular design method driven by evolutionary data, which generates new sequences by analyzing evolutionary conservation information in multiple sequence alignment (MSA) and constructing statistical energy landscape or variational autoencoder (VAE) models. However, these evolutionary information-based methods have the following limitations: Functional correlation is not intuitive: Evolutionary data mainly reflects the survival pressure of sequences and is difficult to accurately correspond to specific physicochemical catalytic processes (such as the hydrolysis kinetics of nucleases or the electron transfer process of oxidoreductases); Sample dependence is strong: For multifunctional chimeric enzymes or non-natural enzymes lacking a large number of homologous sequences, evolution-driven models cannot effectively capture their complex synergistic mechanisms; Spatial geometric constraints are vague: Sequence distribution-based generation models often lack direct constraints on fine three-dimensional spatial geometric features (such as inter-residue distances and dihedral angles), resulting in low folding success rates of the designed sequences after laboratory synthesis. Therefore, there is an urgent need for a reverse design scheme that can directly correlate functional characterization signals (such as spectral features) with three-dimensional spatial geometric constraints and has higher design accuracy.
[0004] Traditional enzyme modification techniques mainly include directed evolution and rational design. Directed evolution simulates the natural selection process, constructing mutant libraries through error-prone PCR or DNA shuffling techniques, followed by high-throughput screening. While this method has achieved great success, it relies on a massive screening workload, is time-consuming and labor-intensive, and is often limited to local optima, making it difficult to achieve disruptive reconstruction of the enzyme backbone. Rational design, on the other hand, relies on researchers' deep understanding of the enzyme structure-function relationship, modifying it through site-directed mutagenesis or homology modeling. However, when dealing with complex bifunctional enzymes (such as chimeric enzymes possessing both nuclease and oxidoreductase activities) or when global properties such as the need for significant changes in optimal temperature are required, the success rate of rational design is often low.
[0005] In recent years, artificial intelligence (AI) technology, especially deep learning, has made groundbreaking progress in the field of protein science. Structure prediction models, exemplified by AlphaFold, have solved the forward folding problem "from sequence to structure"; sequence design models, such as ProteinMPNN, have solved the reverse folding problem "from structure to sequence". However, current AI protein design methods still have significant limitations:
[0006] 1. Lack of physical interpretability: Most generative models directly sample structures or sequences in the latent space, ignoring the physical kinetics and thermodynamic constraints of the protein folding process. As a result, the generated designs often perform well on computers, but fail to fold or lack activity in wet experiments.
[0007] 2. The "black box" problem of property-structure mapping: Existing conditional generation models struggle to accurately map continuous physical quantities (such as optimal temperature values and catalytic rate constant kcat) directly to discrete atomic coordinates. Direct end-to-end generation often leads to imprecise property control.
[0008] 3. Challenges of multifunctional synergistic design: Nucleases and oxidoreductases have drastically different catalytic mechanisms and active site geometry, making it difficult for traditional single models to simultaneously meet the structural requirements of both functions within a single framework.
[0009] In summary, existing technologies suffer from technical problems such as a lack of physical interpretability, inaccurate property control due to the black-box nature of property-structure mapping, and difficulties in multifunctional collaborative design. Summary of the Invention
[0010] The technical problem to be solved by this invention is: how to solve the technical problems of imprecise property control and difficulty in multifunctional collaborative design caused by the lack of physical interpretability, the black box problem of property-structure mapping in the prior art.
[0011] This invention solves the above-mentioned technical problems by employing the following technical solution: A method for reverse design of multifunctional nucleases and oxidoreductases, comprising: S1. Obtain the protein sequences and three-dimensional structure files of nucleases and oxidoreductases, predict biological activity parameters and optimal temperatures based on the protein sequences, construct conditional prompts, and obtain the two-dimensional spectra of proteins by constructing Hamiltonian matrices based on the three-dimensional structure files. S2. Based on the conditional prompt word Prompt and the two-dimensional spectrum of the protein, the preset conditional diffusion model DiffusionModel is trained to obtain the trained conditional diffusion model. Based on the two-dimensional spectrum of the protein and the three-dimensional structure of the protein, the preset spectral analysis model is trained to obtain the trained spectral analysis model. The trained conditional diffusion model and the trained spectral analysis model are integrated to obtain the protein generation model. S3. Input the conditional prompts for the target properties of the enzyme to be designed into the protein generation model. After the process of spectral generation and structural reconstruction, the three-dimensional backbone of the protein to be optimized is obtained. S4. Perform geometric and energy optimization on the three-dimensional backbone of the protein to be optimized, obtain and, based on the optimized backbone, reverse design the amino acid sequence, and perform multi-dimensional functional and structural verification screening to obtain the protein molecule to be designed.
[0012] This invention provides a de novo enzyme design method based on spectral diffusion and physical mapping, developed using Python and Shell languages. By establishing quantitative relationships between enzymatic properties and two-dimensional spectra, and between two-dimensional spectra and protein structures, it provides physically driven guidance for the reverse design of multifunctional enzymes. It automatically generates spectral data based on temperature and activity requirements, effectively utilizing the sensitivity of spectra to structure, and provides a stable spectral generation and structure reconstruction model that can be trained multiple times. This invention guides the design of nucleases and oxidoreductases, enabling the rapid generation of enzymes that meet the requirements of specific industrial environments (such as high temperatures), accelerating the development cycle of biocatalysts and reducing R&D costs. Compared with traditional directed evolution methods, this invention is characterized by high efficiency, interpretability, and strong physical rationality.
[0013] In a more specific technical solution, in S1, the spatial coordinates of key atoms in the protein backbone are extracted from the PDB file, and the Hamiltonian matrix of the system is constructed based on exciton coupling theory. By numerically solving the Schrödinger equation, the system response of molecules at a specific temperature is simulated and calculated, and discrete two-dimensional infrared spectral data are obtained and saved as binary files; Spectral frequency and intensity information is extracted from the output file, and the data is normalized into a fixed-dimensional matrix using interpolation. The spectral intensity is normalized to create a spectral dataset corresponding to the conditional prompt word "Prompt".
[0014] This invention employs a physics-driven rational design, introducing "Hamiltonian matrix calculation" and "two-dimensional spectroscopy" as intermediate modes. Compared to generating structures directly from text, spectroscopy directly corresponds to the vibrational energy levels and conformational distribution of proteins, possessing clear physical meaning. This "physical constraint" significantly improves the thermodynamic rationality of the generated framework.
[0015] In a more specific technical solution, the local mode frequency of the peptide bond amide-I band is calculated, the influence of electrostatic potential fluctuations on the frequency is considered, and the transition dipole coupling between different peptide bonds is calculated.
[0016] In a more specific technical solution, in S1, during the operation of constructing the conditional prompt word Prompt, protein data with pre-set enzymology committee numbers are downloaded in batches from the BRENDA database and RCSB PDB database through an automated script to clean low-quality data. Using a pre-trained deep learning model, the activity level and optimal temperature values of sequences are predicted in batches. The activity level, optimal temperature, and enzyme function classification are combined into a label for the conditional prompt word "Prompt".
[0017] This invention employs multi-objective synergistic optimization. By embedding a multi-dimensional Prompt of temperature and activity into the diffusion model, this system can simultaneously optimize enzyme stability (controlled by temperature parameters) and catalytic efficiency (controlled by activity parameters), effectively solving the "trade-off" problem of traditional methods that struggle to balance activity and stability.
[0018] In more specific technical solutions, the forms of tags include: natural language form and structured vector form.
[0019] In a more specific technical solution, in S2, a conditional latent diffusion model based on the U-Net architecture is constructed, and a cross-attention mechanism layer is embedded in the encoder and decoder parts of the U-Net to receive and process the conditional cue word vectors transformed by the text encoder. Using the conditional prompt "Prompt" as the guiding condition, a two-dimensional spectral matrix with added Gaussian noise as the input, and a pre-defined conditional diffusion model, the model is trained for epochs. The difference between the model's predicted noise and the actual added noise is calculated, and the model's parameters are updated through the backpropagation algorithm. The model learns the inverse mapping from noise to a specific spectrum, thus obtaining the trained conditional diffusion model.
[0020] This invention achieves truly physical-driven de novo protein design by constructing a reverse design chain from "property prompt to two-dimensional spectrum to three-dimensional structure to sequence" and combining the physical calculation of Hamiltonian matrix with the generation capability of diffusion model.
[0021] In a more specific technical solution, S2 constructs a spectral analysis model with a two-dimensional convolutional feature extraction layer and a visual Transformer architecture to extract cross-peak patterns and diagonal feature information from the input two-dimensional spectral image. The geometric mapping objective is set, the two-dimensional spectrum of the protein is used as input, and the inter-residue distance matrix and dihedral angle data calculated from the three-dimensional structure of the protein are used as labels. The training is carried out for epochs, the predicted differences are calculated by regression loss function and the model parameters are updated, and the mapping relationship from spectral features to three-dimensional geometric constraints is established to obtain the trained spectral analysis model.
[0022] This invention introduces "two-dimensional spectroscopy (2D Spectroscopy)" as a physical intermediate layer. The two-dimensional infrared spectrum (2DIR) of proteins contains extremely rich structural dynamics information, is extremely sensitive to the coupling of secondary structures (such as alpha-helices and beta-sheets), and is highly correlated with physical parameters such as temperature.
[0023] In a more specific technical solution, in S4, the three-dimensional backbone of the protein to be optimized is input into the SCUBA-D model, the main chain geometry is smoothed by the denoising diffusion process, non-physical bond lengths and bond angles are repaired, and the side chain rotator isomer conformation is initially optimized to minimize energy, thus obtaining the optimized protein backbone. The optimized protein backbone is input into the ProteinMPNN model. Based on the graph neural network message passing mechanism, the sequence reverse design with a fixed backbone is performed, and candidate amino acid sequences with appropriate confidence levels are output.
[0024] This invention combines the geometric optimization of SCUBA-D, the efficient sequence design of ProteinMPNN, and the high-precision verification of AlphaFold3 to form a rigorous "generation-optimization-verification" closed loop. In particular, the introduction of CLEAN and EasIFA for dual functional-level verification significantly reduces the failure rate of wet experimental screening, resulting in an extremely high design success rate.
[0025] In a more specific technical solution, in S4, the CLEAN model is used to predict the enzymatic function of candidate amino acid sequences and screen out sequences whose EC numbers belong to the nuclease or oxidoreductase category. AlphaFold3 was used to predict the atomic structure of the screened sequences, and the root mean square deviation (RMSD) between the predicted structure and the optimized protein backbone was calculated. Sequences with RMSD below a preset threshold were screened out. For proteins that pass the structural consistency test, the PrankWeb tool is used to predict surface binding pockets, and the EasIFA model is used to predict key active site residues to verify whether the key active site residues possess the geometric features and key amino acids required for catalysis.
[0026] This invention efficiently designs multifunctional enzymes that possess both nuclease and oxidoreductase properties, and constructs a high-confidence, closed-loop system from design to verification. This invention is not only applicable to nucleases and oxidoreductases, but its core framework (spectrally-based inverse design) can also be applied to the design of other enzymes, possessing extremely high platform technology value.
[0027] In a more specific technical solution, a multifunctional nuclease and oxidoreductase reverse design system includes: The data spectral calculation module is used to obtain the protein sequences and three-dimensional structure files of nucleases and oxidoreductases, predict biological activity parameters and optimal temperatures based on the protein sequences, construct conditional prompts, and obtain the two-dimensional spectrum of proteins by constructing a Hamiltonian matrix based on the three-dimensional structure files. The model building and generation module is used to train a preset conditional diffusion model based on the conditional prompt word Prompt and the two-dimensional spectrum of proteins to obtain a trained conditional diffusion model. It also trains a preset spectral analysis model based on the two-dimensional spectrum of proteins and the three-dimensional structure of proteins to obtain a trained spectral analysis model. The trained conditional diffusion model and the trained spectral analysis model are integrated to obtain a protein generation model. The model building and generation module is connected to the data spectral calculation module. The optimization and sequence design module is used to input the conditional prompts (Prompts) indicating the target properties of the enzyme to be designed into the protein generation model. After spectral generation and structural reconstruction, the three-dimensional backbone of the protein to be optimized is obtained. The optimization and sequence design module is connected to the model construction and generation module. The multidimensional validation and screening module is used to perform geometric and energy optimization on the three-dimensional backbone of the protein to be optimized, obtain and reverse design amino acid sequences based on the optimized backbone, perform multidimensional functional and structural validation and screening, and obtain the protein molecule to be designed. The multidimensional validation and screening module is connected to the optimization and sequence design module.
[0028] The present invention has the following advantages over the prior art: Compared to prior art such as CN114651064A, this invention has the following significant technical advancements and inventiveness: Fundamental difference in the driving source: Prior art 1 relies on static sequence evolution data, while this application uses multidimensional spectral images (2D-IR / Raman) as the driving source. Spectral features can directly reflect the chemical environment of enzyme molecules, the kinetics of active centers, and the interactions between functional groups, which has stronger physical guidance significance for designing multifunctional enzymes involving complex electron transfer (oxidoreductases) and phosphodiester bond cleavage (nucleases). Cross-modal upgrade of the generation mechanism: This application is not a simple sequence distribution fitting, but rather constructs a multi-level mapping of "functional keywords - two-dimensional spectrum - three-dimensional geometric constraints - amino acid sequence". By capturing subtle spectral changes through a conditional diffusion model and combining it with a visual Transformer architecture to analyze cross-peak patterns in the spectrum, precise decoupling from functional characterization to spatial topology is achieved, solving the problem of unclear structure-function mapping in multifunctional enzyme design in existing technologies. Improved precision in structural optimization: This application introduces SCUBA-D main chain smoothing and ProteinMPNN fixed backbone design during the sequence generation stage, combined with AlphaFold3 structure verification. Compared with the energy density-based generation method in Reference 1, this application performs non-physical bond length and bond angle repair on the main chain geometry at the physical level, which greatly improves the stability and catalytic activity of the final designed sequence in real biochemical environments.
[0029] This invention provides a de novo enzyme design method based on spectral diffusion and physical mapping, developed using Python and Shell languages. By establishing quantitative relationships between enzymatic properties and two-dimensional spectra, and between two-dimensional spectra and protein structures, it provides physically driven guidance for the reverse design of multifunctional enzymes. It automatically generates spectral data based on temperature and activity requirements, effectively utilizing the sensitivity of spectra to structure, and provides a stable spectral generation and structure reconstruction model that can be trained multiple times. This invention guides the design of nucleases and oxidoreductases, enabling the rapid generation of enzymes that meet the requirements of specific industrial environments (such as high temperatures), accelerating the development cycle of biocatalysts and reducing R&D costs. Compared with traditional directed evolution methods, this invention is characterized by high efficiency, interpretability, and strong physical rationality.
[0030] This invention employs a physics-driven rational design, introducing "Hamiltonian matrix calculation" and "two-dimensional spectroscopy" as intermediate modes. Compared to generating structures directly from text, spectroscopy directly corresponds to the vibrational energy levels and conformational distribution of proteins, possessing clear physical meaning. This "physical constraint" significantly improves the thermodynamic rationality of the generated framework.
[0031] This invention employs multi-objective synergistic optimization. By embedding a multi-dimensional Prompt of temperature and activity into the diffusion model, this system can simultaneously optimize enzyme stability (controlled by temperature parameters) and catalytic efficiency (controlled by activity parameters), effectively solving the "trade-off" problem of traditional methods that struggle to balance activity and stability.
[0032] This invention achieves truly physical-driven de novo protein design by constructing a reverse design chain from "property prompt to two-dimensional spectrum to three-dimensional structure to sequence" and combining the physical calculation of Hamiltonian matrix with the generation capability of diffusion model.
[0033] This invention introduces "two-dimensional spectroscopy (2D Spectroscopy)" as a physical intermediate layer. The two-dimensional infrared spectrum (2DIR) of proteins contains extremely rich structural dynamics information, is extremely sensitive to the coupling of secondary structures (such as alpha-helices and beta-sheets), and is highly correlated with physical parameters such as temperature.
[0034] This invention combines the geometric optimization of SCUBA-D, the efficient sequence design of ProteinMPNN, and the high-precision verification of AlphaFold3 to form a rigorous "generation-optimization-verification" closed loop. In particular, the introduction of CLEAN and EasIFA for dual functional-level verification significantly reduces the failure rate of wet experimental screening, resulting in an extremely high design success rate.
[0035] This invention solves the technical problems in the prior art, such as the lack of physical interpretability, the black box problem of property-structure mapping leading to inaccurate property control, and the difficulty of multifunctional collaborative design. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the basic steps of a reverse design method for multifunctional nucleases and oxidoreductases according to Embodiment 1 of the present invention; Figure 2 This is a detailed data flow diagram of the first stage (spectral calculation and data construction) of the present invention, showing the transformation process from the database to the Hamiltonian matrix and then to the two-dimensional spectrum.
[0037] Figure 3 This is a detailed data flow diagram for the second stage (user requirement input) of the present invention, showing what conditions the user should input.
[0038] Figure 4 This is a schematic diagram of the model principle of the third stage of the present invention (spectral generation based on diffusion model), which shows how the conditional Prompt guides noise denoising to generate the target spectrum.
[0039] Figure 5 This is a cascaded flowchart for the fourth stage of optimization and verification of this invention, which includes the interaction logic of models such as SCUBA-D, ProteinMPNN, CLEAN, and AlphaFold 3.
[0040] Figure 6 This is a cascading flowchart for the fifth stage (optimization and verification) of the present invention, which includes the interaction logic of models such as PrankWeb and EasIFA.
[0041] Figure 7 This is a three-dimensional structural diagram of a thermostable nuclease generated in an embodiment of the present invention.
[0042] Figure 8 This is a schematic diagram showing the predicted active site of a certain oxidoreductase generated in an embodiment of the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Example 1 like Figure 1 As shown, the present invention provides a method for reverse design of multifunctional nucleases and oxidoreductases, comprising the following basic steps: S1. Obtain the protein sequences and PDB files of nucleases and oxidoreductases, predict their activity and temperature based on the protein sequences and put them into the Prompt for Prompt construction, and complete the spectral calculation through the Hamiltonian matrix.
[0045] In this embodiment, the properties predicted based on the protein sequence include biological activity (kcat / Km), optimal temperature, and enzyme functional classification (EC number).
[0046] like Figure 2As shown in this embodiment, the transformation process from database to Hamiltonian matrix to two-dimensional spectrum is demonstrated in the first stage of spectral calculation and data construction, including data construction and physical spectrum mapping. Sequence and structural data of nucleases and oxidoreductases are obtained from the BRENDA and RCSBPDB databases, and their activity and temperature properties are predicted, and conditional prompts are constructed. The core lies in using the Hamiltonian matrix to physically model the known protein structure, simulating and calculating its two-dimensional spectral (2D Spectroscopy) data, thereby establishing a multimodal basic dataset of "property prompt - three-dimensional structure - two-dimensional spectrum".
[0047] In this embodiment, the specific operations for performing spectral calculations and constructing the Prompt based on the protein file include, but are not limited to: S11. By accessing the BRENDA and RCSB PDB databases, batch download the protein sequences (.fasta format) and 3D structure files (.pdb format) of nucleases (EC 3.1.xx) and oxidoreductases (EC 1.xxx). Use a Python script to clean the downloaded data, removing low-quality data with sequences shorter than 50 amino acids or structural resolutions below 2.5 Å, and construct standardized metadata tags for each protein.
[0048] S12. Based on a pre-trained deep learning prediction model (such as a regression model fine-tuned based on ESM-2), perform batch predictions of the activity and optimal temperature of the downloaded protein sequences, and obtain the predicted activity value and predicted temperature value for each protein.
[0049] The specific method for constructing a Prompt is as follows: Combine the predicted activity level (e.g., "High Activity"), the optimal temperature value (e.g., "65°C"), and the functional description of the enzyme (e.g., "Nuclease") into a natural language Prompt, for example: "Athermostable nuclease with high activity at 65 degrees Celsius".
[0050] S13. Perform physical modeling on the downloaded PDB file, calculating the two-dimensional spectrum of the protein using the Hamiltonian matrix. Step S13 specifically includes: S131. A one-click batch processing script based on the Slurm scheduling system, allowing for one-click submission to the task queue. First, the coordinate information of the protein backbone is extracted from the PDB file, and the Hamiltonian matrix of the system is constructed based on exciton coupling theory. (See also...) Figure 3This displays the conditions that the user should input during the second phase of the user needs input process; S132. In the construction of the Hamiltonian matrix, calculate the local mode frequency of the amide-I band of the peptide bond and consider the influence of electrostatic potential fluctuations on the frequency; at the same time, calculate the transition dipole coupling between different peptide bonds.
[0051] S133. By numerically solving the Schrödinger equation, the system response of molecules at a specific temperature is simulated and calculated, thereby obtaining discrete two-dimensional infrared (2DIR) spectral data, and the spectral data is saved in the output binary file.
[0052] S14. Extract the corresponding frequency and intensity information from the output file, and normalize the discrete two-dimensional spectral data to 64 using linear interpolation. 64 or 128 A 128 matrix was used to normalize the spectral intensity. Spectral feature data was extracted using a Python script and mapped one-to-one with the corresponding Prompt tags, then saved into a dataset in HDF5 format.
[0053] S2. Based on spectral diffusion and protein backbone generation; wherein, a preset diffusion model is trained based on Prompt and two-dimensional spectroscopy to obtain a trained diffusion model; a preset 2DIR analytical model is trained based on two-dimensional spectroscopy and protein three-dimensional structure to obtain a trained 2DIR analytical model. Generative spectral design and backbone reconstruction: Using a pre-trained conditional diffusion model, novel two-dimensional spectra with specific physical characteristics are generated based on the input target activity and temperature prompt. Subsequently, the generated spectra are input into a 2DIR analytical model, which maps them into a protein residue distance matrix heatmap, and based on this, an initial rough three-dimensional backbone of the unoptimized enzyme is constructed.
[0054] In this embodiment, the prompt, two-dimensional spectrum, and protein three-dimensional structure are all stored in the dataset. The specific operations for training the preset Diffusion model based on the prompt and two-dimensional spectrum include, but are not limited to: S211. Construct a Conditional Latent Diffusion Model. Improve the existing U-Net structure of the Diffusion model by embedding a cross-attention mechanism in the Encoder and Decoder parts of the U-Net. Set up a text encoder (such as CLIP Text Encoder) to convert the Prompt constructed in step 1 into word vectors, and inject them into the U-Net through the cross-attention layer to guide the generation of spectra.
[0055] Parameter settings: Parameters are set to control the number of model layers, the initial dimensionality of U-Net, and the dimensionality multiplier based on the number of layers. The skip connection part uses fully connected layers and a self-attention mechanism for feature extraction, capturing the global correlation of the spectrum.
[0056] S212. Using the Prompt as a condition and the two-dimensional spectrum with added Gaussian noise as input, input the preset Diffusion model; perform epoch-times of training, calculate the mean squared error (MSE) between the model's predicted noise and the actual added noise, and update the parameters of the preset Diffusion model through the backpropagation algorithm to obtain the trained Diffusion model. (See also...) Figure 4 This demonstrates the principle of generating the target spectrum by using conditional Prompt to guide noise denoising during the third stage of the diffusion model-based spectrum generation process.
[0057] In the training process of this embodiment, the hyperparameters of the Diffusion model are adjusted based on the performance of the training process to optimize the model. Specifically, this includes: defining the AdamW optimizer and the learning rate decay strategy; using the structural similarity index (SSIM) between the generated spectrum and the real spectrum as the evaluation metric; setting an inference generation every 10 epochs to generate a two-dimensional spectrum to check whether the model can generate a spectrum with specific characteristics (such as narrowing of the spectral linewidth at high temperatures) based on the prompt.
[0058] The process of training a pre-defined 2DIR analytical model based on two-dimensional spectra and three-dimensional protein structures includes: S221. Construct a 2DIR analysis model, which adopts an architecture similar to VisionTransformer (ViT). Add a two-dimensional convolutional feature extraction layer before the model to extract cross-peak patterns and diagonal feature information from the input two-dimensional spectrum.
[0059] S222. Set up the geometric mapping by inputting the two-dimensional spectrum into the preset 2DIR analytical model, using the protein distance matrix and dihedral angles as labels. Perform epochs of training, updating the model parameters through the backpropagation algorithm to obtain the trained 2DIR analytical model.
[0060] S223. In the model data input, for the protein structure in the Label, it is converted into an image form by calculating the $C\alpha-C\alpha$ distance matrix. In the model prediction output, the difference between the model output matrix and the true distance matrix is calculated using a regression loss function (such as L2 Loss).
[0061] S3. Input the properties of the molecule to be designed into the molecular prediction model to obtain the structure of the unoptimized enzyme; Structural geometry optimization and sequence reverse design: The unoptimized, coarse backbone is input into the SCUBA-D model for fine-tuning of geometry and energy, resulting in a thermodynamically stable protein backbone. Based on the optimized backbone, the ProteinMPNN model is used for reverse folding to design the amino acid sequence that best matches the backbone.
[0062] S31. Write the properties of the enzyme to be designed (e.g., "possesses nuclease activity and tolerates 80°C high temperature") as a Prompt. See also Figure 6 .
[0063] S32. Input the Prompt into the trained Diffusion model. The model starts with random Gaussian noise and performs multi-step denoising sampling under the guidance of the Prompt, finally mapping to obtain a new two-dimensional spectrum that conforms to the property description.
[0064] S33. Input the generated two-dimensional spectrum into the trained 2DIR analytical model. The model analyzes the spectral features and maps them to obtain a heatmap of the distance matrix between protein residues.
[0065] S34. Using multidimensional scaling analysis (MDS) or geometric reconstruction algorithms, the distance matrix heatmap is transformed into an unoptimized three-dimensional coordinate structure of the enzyme.
[0066] This invention first inputs the enzyme's functional properties (activity, temperature) required by the user into the Diffusion model via Prompt to obtain spectral data. This step maps from a limited property dimension to a physical spectral dimension, where the dimension is relatively controllable, and the spectrum contains rich kinetic information. Then, the spectral data is input into the 2DIR model to obtain the protein structure, and a suitable backbone is generated through physical spectroscopy. This mapping is performed in a space with clear physicochemical meaning, avoiding the "black box" mapping from text to three-dimensional structure. It decomposes the large dimensionality transition problem into two relatively simple physical feature conversion problems, thereby reducing the risk that the designed backbone will not conform to thermodynamic principles.
[0067] S4. Protein Structure Design and Functional Validation. The structure of the unoptimized enzyme was optimized using SCUBA-D, the sequence was predicted using ProteinMPNN, the EC number was predicted using CLEAN, and the results were compared with AlphaFold 3 results to predict pockets and active sites. (See also...) Figure 5 It demonstrates the interaction logic of models such as SCUBA-D, ProteinMPNN, CLEAN, and AlphaFold 3 in the fourth stage of optimization and verification.
[0068] Multidimensional functional verification and closed-loop screening were performed. The designed sequences underwent rigorous validation: the CLEAN model was used to predict the EC number to confirm the functional classification (nuclease / oxidoreductase); AlphaFold 3 was used to predict the full-atomic structure and compare it with the designed backbone using RMSD to verify structural stability; finally, PrankWeb and EasIFA were used to predict whether the pocket and active site met catalytic requirements. Based on the above results, the final target protein molecule was selected. (See also...) Figure 6 This demonstrates the interaction logic of models such as PrankWeb and EasIFA during the fifth phase of optimization and verification.
[0069] In this embodiment, the specific operations for fine-tuning and validating based on the unoptimized enzyme structure include, but are not limited to: S41. Structural Geometry Optimization: The unoptimized enzyme structure obtained in S34 was processed using the SCUBA-D model. SCUBA-D uses a denoised diffusion probability model to smooth the main chain, repair non-physical bond lengths and angles, and preliminarily optimize the side chain conformation.
[0070] Specifically, if the energy of the structure after SCUBA-D optimization is still high, Rosetta's FastRelax protocol is called to further minimize the energy and obtain the screened stable molecular conformation.
[0071] S42. Sequence Reverse Design: Based on the screened molecular conformation, the ProteinMPNN model is used for sequence design. The optimized PDB backbone is input, and ProteinMPNN outputs multiple possible amino acid sequences.
[0072] S43. Functional classification verification: Extract the protein sequence generated by ProteinMPNN and input it into the CLEAN (Contrastive Learning-enabled Enzyme Annotation) model.
[0073] The S431 CLEAN model is based on the EC number (Enzyme Committee number) of the predicted sequence through contrastive learning.
[0074] S432. Select sequences whose EC number prediction results belong to 3.1.xx (nuclease) or 1.xxx (oxidoreductase) and have a confidence level greater than 0.9. See [link to relevant documentation]. Figure 8 .
[0075] S44. Structural Consistency Verification: Input the sequences selected through S43 into AlphaFold 3 for full-atom structure prediction.
[0076] S441. Calculate the root mean square deviation (RMSD) between the AlphaFold 3 predicted structure and the SCUBA-D optimized structure.
[0077] S442. Proteins with an RMSD of less than 2.0 Å were screened, which means that the designed sequence can fold into the expected physical backbone with high fidelity.
[0078] S45. Active site depth validation: Proteins with good results and meeting the Prompt criteria are subjected to final screening.
[0079] S451. Use the PrankWeb tool to predict the location of binding pockets on the protein surface to confirm whether there are deep pockets suitable for substrate binding.
[0080] S452. Predict active site residues using the EasIFA model. Check whether the predicted residues contain key amino acids required for catalysis (such as the His-Asp-Ser triplet of nucleases or the Cys-His pair of oxidoreductases).
[0081] S453. After verification, save the final protein sequence, structure file and related verification report as the output of the molecule to be designed.
[0082] In summary, this invention overcomes the nonlinearity and complexity of the "sequence-structure-function" mapping in traditional protein design methods, enabling precise pre-setting and control of enzyme physicochemical properties (such as heat resistance and high activity). It addresses the lack of physical plausibility and thermodynamic stability in protein backbones generated by generative AI models. This invention introduces physical spectral features as an intermediate bridge connecting macroscopic properties and microscopic structures, utilizing a deep generative model to achieve reverse de novo design from specific physicochemical properties (such as temperature and activity) to the three-dimensional structure and sequence of proteins.
[0083] This invention solves the technical problems in the prior art, such as the lack of physical interpretability, the black box problem of property-structure mapping leading to inaccurate property control, and the difficulty of multifunctional collaborative design.
[0084] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for reverse design of multifunctional nucleases and oxidoreductases, characterized in that, The method includes: S1. Obtain the protein sequences and three-dimensional structure files of nucleases and oxidoreductases, predict the biological activity parameters and optimal temperature based on the protein sequences, construct the conditional prompt word Prompt, and obtain the two-dimensional spectrum of the protein by constructing the Hamiltonian matrix based on the three-dimensional structure file. S2. Based on the conditional prompt word Prompt and the protein two-dimensional spectrum, a preset conditional diffusion model is trained to obtain a trained conditional diffusion model. Based on the protein two-dimensional spectrum and the protein three-dimensional structure, a preset spectral analysis model is trained to obtain a trained spectral analysis model. The trained conditional diffusion model and the trained spectral analysis model are integrated to obtain a protein generation model. S3. Input the conditional prompt word Prompt of the target property of the enzyme to be designed into the protein generation model, and obtain the three-dimensional backbone of the protein to be optimized through the process of spectrum generation and structure reconstruction. S4. Perform geometry and energy optimization on the three-dimensional backbone of the protein to be optimized, obtain and, based on the optimized backbone, reverse design the amino acid sequence, perform multi-dimensional functional and structural verification and screening, and obtain the protein molecule to be designed.
2. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 1, characterized in that, In S1, the spatial coordinates of key atoms in the protein backbone are extracted from the PDB file, and the Hamiltonian matrix of the system is constructed based on exciton coupling theory. By numerically solving the Schrödinger equation, the system response of molecules at a specific temperature is simulated and calculated, and discrete two-dimensional infrared spectral data are obtained and saved as binary files; Spectral frequency and intensity information are extracted from the output file, and the data is normalized into a fixed-dimensional matrix using interpolation. The spectral intensity is normalized to establish a spectral dataset corresponding to the conditional prompt word "Prompt".
3. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 2, characterized in that, Calculate the local mode frequencies of the amide-I band of the peptide bond, considering the influence of electrostatic potential fluctuations on the frequencies, and calculate the transition dipole coupling between different peptide bonds.
4. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 1, characterized in that, In S1, during the operation of constructing the conditional prompt word Prompt, protein data with pre-set enzymology committee numbers are downloaded in batches from the BRENDA database and RCSB PDB database using an automated script, and low-quality data is cleaned. Using a pre-trained deep learning model, the activity level and optimal temperature values of sequences are predicted in batches. The activity level value, the optimal temperature value, and the enzyme function classification description are combined into a label for the conditional prompt word Prompt.
5. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 4, characterized in that, The labels can take the form of natural language or structured vector.
6. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 1, characterized in that, In S2, a conditional latent diffusion model based on the U-Net architecture is constructed, and a cross-attention mechanism layer is embedded in the encoder and decoder parts of the U-Net to receive and process the conditional cue word vectors transformed by the text encoder. Using the conditional prompt word "Prompt" as the guiding condition, a two-dimensional spectral matrix with added Gaussian noise as the input, and the preset conditional diffusion model as the input, the model is trained for epochs. The difference between the model's predicted noise and the actual added noise is calculated, and the model's parameters are updated through the backpropagation algorithm. The model learns the inverse mapping from noise to a specific spectrum to obtain the trained conditional diffusion model.
7. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 1, characterized in that, In S2, a spectral analysis model with a two-dimensional convolutional feature extraction layer and a visual Transformer architecture is constructed to extract cross-peak patterns and diagonal feature information from the input two-dimensional spectral image. The geometric mapping target is set, the two-dimensional spectrum of the protein is used as input, and the inter-residue distance matrix and dihedral angle data calculated from the three-dimensional structure of the protein are used as labels. The training is carried out for epochs, the predicted difference is calculated by regression loss function and the model parameters are updated, and the mapping relationship from spectral features to three-dimensional geometric constraints is established to obtain the trained spectral analysis model.
8. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 1, characterized in that, In step S4, the three-dimensional backbone of the protein to be optimized is input into the SCUBA-D model. The main chain geometry is smoothed by the denoising diffusion process, non-physical bond lengths and bond angles are repaired, and the side chain rotator isomer conformation is initially optimized to minimize energy, thus obtaining the optimized protein backbone. The optimized protein backbone is input into the ProteinMPNN model. Based on the graph neural network message passing mechanism, the sequence reverse design with a fixed backbone is performed, and candidate amino acid sequences with appropriate confidence levels are output.
9. The method for reverse design of multifunctional nucleases and oxidoreductases according to claim 1, characterized in that, In step S4, the CLEAN model is used to predict the enzymatic function of candidate amino acid sequences and screen out sequences whose EC numbers belong to the nuclease or oxidoreductase category. AlphaFold3 was used to predict the atomic structure of the screened sequences, and the root mean square deviation (RMSD) between the predicted structure and the optimized protein backbone was calculated. Sequences with RMSD below a preset threshold were screened out. For proteins that pass the structural consistency test, the surface binding pocket is predicted using the PrankWeb tool, and the key active site residues are predicted using the EasIFA model to verify whether the key active site residues possess the geometric features and key amino acids required for catalysis.
10. A multifunctional nuclease and oxidoreductase reverse design system, characterized in that, The system includes: The data spectral calculation module is used to obtain the protein sequences and three-dimensional structure files of nucleases and oxidoreductases, predict biological activity parameters and optimal temperatures based on the protein sequences, construct conditional prompts, and obtain the two-dimensional spectrum of proteins by constructing a Hamiltonian matrix based on the three-dimensional structure files. The model building and generation module is used to train a preset conditional diffusion model based on the conditional prompt word Prompt and the two-dimensional spectrum of the protein to obtain a trained conditional diffusion model; to train a preset spectral analysis model based on the two-dimensional spectrum of the protein and the three-dimensional structure of the protein to obtain a trained spectral analysis model; and to integrate the trained conditional diffusion model and the trained spectral analysis model to obtain a protein generation model. The model building and generation module is connected to the data spectral calculation module. The optimization and sequence design module is used to input the conditional prompt word Prompt of the target property of the enzyme to be designed into the protein generation model, and obtain the three-dimensional backbone of the protein to be optimized through the process of spectrum generation and structure reconstruction. The optimization and sequence design module is connected to the model construction and generation module. The multidimensional verification and screening module is used to perform geometric and energy optimization on the three-dimensional backbone of the protein to be optimized, obtain and reverse design amino acid sequences based on the optimized backbone, perform multidimensional functional and structural verification and screening, and obtain the protein molecule to be designed. The multidimensional verification and screening module is connected to the optimization and sequence design module.
Citation Information
Patent Citations
Method and apparatus for evolutionary data driven design of protein and other sequence defined biomolecules using machine learning
CN114651064A