A method and system for de novo protein reverse design based on spectrum structure-activity relationship
By combining a multi-task neural network model and molecular dynamics simulation based on spectrum structure-activity relationship, and SCUBA-D and ProteinMPNN algorithms, the challenges of multi-property fusion and dynamic behavior in protein design are solved, achieving efficient and flexible protein design and drug development.
Patent Information
- Application Number
- CN202411316576.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Existing protein design methods have difficulty in effectively integrating multiple properties and ignore dynamic behaviors, resulting in limited design flexibility and functional diversity.
A multi-task neural network model based on spectrum structure-activity relationship is used, combined with molecular dynamics simulation and spectral data, to automatically design protein structure and amino acid sequence through SCUBA-D protein structure optimization and ProteinMPNN algorithm.
It improves the efficiency and accuracy of protein design, can quickly screen and verify proteins of specified needs, reduce drug development costs, and achieve flexible protein design.
Smart Images

Figure CN119296655B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spectral inversion quantum chemistry theory, and specifically to a method and system for de novo reverse design of proteins based on spectral structure-activity relationship. Background Art
[0002] In recent years, the field of protein design, especially de novo design, has made significant progress. However, the field still faces many challenges. Traditional protein design methods mainly rely on two types of models: sequence-based models and structure-based models. Sequence-based models design by learning the statistical features and patterns of protein sequences, while structure-based models focus on the three-dimensional structural information of proteins. Although these methods have been successful in certain scenarios, they still have many limitations when dealing with the design of complex protein functions, such as the fusion of multiple properties. The design of protein function usually requires the optimization of multiple objectives, such as stability, catalytic efficiency and specificity.
[0003] How to effectively integrate these properties in a high-dimensional structural space has become a difficult problem, and existing methods have found it difficult to make breakthroughs in multi-objective optimization. End-to-end design: The vastness of the sequence space, the complexity of structure prediction, and the challenges of multi-objective optimization make end-to-end protein design very difficult. Existing methods are unable to simultaneously meet the multiple requirements of sequence, structure, and function, and the designed proteins often cannot effectively balance these factors. Dynamic protein design: Current design methods mainly focus on the design of fixed skeletons, ignoring the dynamic behavior of proteins in different states. However, the dynamic changes of proteins are crucial to the realization of their functions, and existing technologies have insufficient consideration of this dynamic behavior, limiting the flexibility and functional diversity of the design. Therefore, how to overcome the challenges of multi-property fusion, dynamic behavior simulation, and end-to-end design has become an urgent problem to be solved in the current field of protein design. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to solve the problem that the existing technology does not take the dynamic behavior of proteins into consideration sufficiently, which limits the flexibility of design and functional diversity.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] A protein de novo reverse design method based on spectrum structure-activity relationship, comprising:
[0007] S10, acquiring a data set, wherein the data set includes protein structural feature data, protein simulated spectrum data, and protein-ligand interaction data;
[0008] S20, establishing a multi-task neural network model, training it using the data set, inputting protein simulated spectral data into the trained multi-task neural network model, and simultaneously obtaining the distance matrix between protein backbone atoms and the number of protein amino acids;
[0009] S30, integrating the distance matrix between protein backbone atoms and the number of protein amino acids to obtain the final protein distance matrix;
[0010] S40, obtaining the initial protein PDB structure based on the final protein distance matrix;
[0011] S50, using the SCUBA-D protein structure optimization model to optimize the initial protein PDB structure and obtain the all-atom protein structure;
[0012] S60, applying the ProteinMPNN algorithm, takes the all-atom protein structure as input and obtains multiple amino acid sequences.
[0013] In one embodiment of the present invention, before obtaining protein simulation spectrum data, initial protein simulation spectrum data is first obtained, including the following steps:
[0014] Obtain the initial PDB structure file from the RCSB protein PDB database; wherein the initial PDB structure file includes protein three-dimensional structure data;
[0015] Perform molecular dynamics simulation of the protein on the initial PDB structure file to obtain molecular dynamics simulation data;
[0016] AIM was used to calculate the molecular dynamics simulation data to obtain the intermolecular interaction data; NISE was used to convert the intermolecular interaction data into converted spectral data to obtain the initial protein simulation spectral data.
[0017] In one embodiment of the present invention, protein structural feature data, protein simulated spectral data, and protein-ligand interaction data are obtained by:
[0018] The protein three-dimensional structure data is converted into a protein two-dimensional distance matrix; at the same time, the initial protein simulated spectrum data is formatted to obtain a protein simulated spectrum two-dimensional matrix; and the protein two-dimensional distance matrix and the protein simulated spectrum two-dimensional matrix are normalized;
[0019] According to the protein topology, the normalized protein two-dimensional distance matrix and protein simulated spectrum two-dimensional matrix are simultaneously classified to obtain protein structural feature data and protein simulated spectrum data;
[0020] A mapping relationship is established between protein structural feature data and protein simulated spectral data with the same protein structural features and properties to obtain protein-ligand interaction data.
[0021] In one embodiment of the present invention, the network architecture of the multi-task neural network model includes: an input layer, a convolutional layer, a fully connected layer, a multi-layer feature extraction layer, a feature fusion layer and an output layer; wherein,
[0022] Input layer, used for normalization preprocessing of protein simulation spectral data;
[0023] Convolutional layer, used to extract spatial features from protein simulation spectral data;
[0024] The fully connected layer is used to map the spatial features of the protein simulated spectral data extracted by the convolutional layer to the structural feature space of the protein and obtain the distance matrix between the initial protein skeleton atoms;
[0025] Multiple feature extraction layers, including first-layer feature extraction, second-layer feature extraction, and third-layer feature extraction; and the extraction scales from the first layer to the third layer increase in sequence; the three-layer feature extraction is used to extract features of different scales from the initial protein backbone atom distance matrix;
[0026] The feature fusion layer is used to use the splicing function to splice and fuse the features of different scales extracted from multiple layers of feature extraction layers to obtain the final distance matrix between protein backbone atoms;
[0027] The output layer includes a protein structure output layer and an amino acid number output layer; the protein structure output layer directly outputs the final distance matrix between protein backbone atoms; the protein structure is obtained based on the distance matrix between protein backbone atoms, and the characteristic lengths of all output protein structures are unified. The amino acid number output layer uses a filling method to fill in the missing amino acids with invalid information, so that the characteristic lengths of all output protein structures are unified, and then the number of amino acids is output.
[0028] In one embodiment of the present invention, obtaining the final protein distance matrix includes: adjusting the dimension of the distance matrix between protein backbone atoms according to the number of protein amino acids, and using the PyRosetta tool to generate the initial protein PDB structure from the distance matrix between protein backbone atoms after the dimension adjustment.
[0029] In one embodiment of the present invention, obtaining a full-atom protein structure includes: establishing a protein index mapping and protein structure optimization configuration to ensure that the SCUBA-D protein structure optimization model can process every part of the protein; the SCUBA-D protein structure optimization model is based on molecular dynamics and force field energy calculations, and continuously adjusts the main chain and side chain of the protein to achieve energy minimization to optimize the full-atom structure of the protein and obtain the full-atom protein structure.
[0030] In one embodiment of the present invention, obtaining multiple amino acid sequences includes: performing global energy calculation on the atomic protein structure using a ProteinMPNN algorithm, and generating multiple amino acid sequences according to the energy minimization principle.
[0031] In one embodiment of the present invention, the method for de novo reverse design of proteins based on spectrum-structure-activity relationship also includes step S100, which is located between step S10 and step S20; wherein, step S100 is: establishing a spectral diffusion generation model constructed based on a neural network; using protein simulated spectral data and input custom protein description as inputs of the spectral diffusion generation model for training, and outputting custom diffuse protein spectral data; wherein, the input custom protein description is obtained in the following manner: using a script to generate a mapping relationship file between the spectrum and the custom description, and the mapping relationship file records the matching relationship between each spectral data and its corresponding protein attribute, function and topological description.
[0032] In one embodiment of the present invention, according to user needs, a custom protein description is input into the trained spectral diffusion generation model to output custom diffuse protein spectral data, and then the custom diffuse protein spectral data is processed through steps S20 to S60.
[0033] The present invention also provides a protein de novo reverse design system based on spectrum structure-activity relationship, which uses the above-mentioned protein de novo reverse design method based on spectrum structure-activity relationship, including:
[0034] Dataset module: used to obtain data sets, where the data sets include protein structural feature data, protein simulation spectrum data, and protein-ligand interaction data;
[0035] The multi-task prediction module is used to build a multi-task neural network model and train it using a data set. The protein simulation spectrum data is input into the trained multi-task neural network model, and the distance matrix between protein backbone atoms and the number of protein amino acids are obtained at the same time.
[0036] Protein distance matrix module, which is used to integrate the distance matrix between protein backbone atoms and the number of protein amino acids to obtain the final protein distance matrix;
[0037] Initial PDB module, used to obtain the initial protein PDB structure based on the final protein distance matrix;
[0038] All-atom protein structure module, used to optimize the initial protein PDB structure using the SCUBA-D protein structure optimization model to obtain the all-atom protein structure;
[0039] The amino acid sequence module is used to apply the ProteinMPNN algorithm to obtain multiple amino acid sequences using all-atom protein structures as input.
[0040] Compared with existing technologies, the present invention offers the following advantages: by establishing the underlying relationship between protein spectral data and protein properties and topological structure, it provides data-driven guidance for efficient and rapid protein design. It also automates the processing and analysis of spectral data, effectively improving the efficiency and accuracy of spectral data analysis. Furthermore, it provides a reproducibly trained and robust spectral diffusion generation model capable of designing proteins that meet specific requirements, simply by providing user-provided design goals. This system, used to guide protein design and optimization, can rapidly screen and validate proteins of specific interest, accelerating drug development cycles and reducing R&D costs.
[0041] The system is easy to operate and can automatically process user needs and design specified protein structures. This automated feature reduces the user's workload during data processing and improves processing efficiency. Combined with the latest diffusion models, it enables efficient de novo protein design. Results are validated using the most advanced AF2 and AF3 methods, significantly accelerating drug development and design, as well as the discovery and exploration of novel functional proteins, and holds great potential.
[0042] The initial protein simulation spectral data are based on molecular dynamics simulation to study the dynamic behavior of proteins, and the spectral diffusion generation model is user-defined, enabling flexible protein design. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flow chart of a method for de novo protein reverse design based on spectrum-structure-activity relationship according to an embodiment of the present invention.
[0044] Figure 2 2 is a network architecture diagram of a multi-task neural network model according to an embodiment of the present invention.
[0045] Figure 3 This is a graph showing the accuracy of the protein design results according to an embodiment of the present invention.
[0046] Figure 4 A block diagram of a protein de novo reverse design system based on spectrum-structure-activity relationship according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] To facilitate those skilled in the art to understand the technical solution of the present invention, the technical solution of the present invention is further described with reference to the accompanying drawings.
[0048] The terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0049] See also Figure 1 As shown, the present invention discloses a method for de novo reverse design of proteins based on spectrum structure-activity relationship, comprising the following steps:
[0050] S10, acquiring a data set, wherein the data set includes protein structural feature data, protein simulated spectral data, and protein-ligand interaction data.
[0051] In one embodiment of the present invention, before obtaining protein simulation spectrum data, initial protein simulation spectrum data is first obtained, including the following steps:
[0052] S11, obtaining an initial PDB structure file from the RCSB protein PDB database; wherein the initial PDB structure file includes protein three-dimensional structure data.
[0053] In this embodiment, the RCSB protein PDB database is an open source database. As needed, the three-dimensional structure of the required protein can be downloaded from the RCSB protein PDB database and saved in PDB format to provide an initial structure for molecular dynamics simulation.
[0054] In the present embodiment, the pdbfixer package can be used to repair the initial PDB structure file, to handle problems such as missing atoms, residues or chains of proteins and residue naming problems, to ensure that the PDB structure file is complete and correct, and to prepare for subsequent simulation and analysis. This process involves the arrangement and standardization of the PDB structure file to ensure that it meets the requirements of subsequent calculations and simulations. The initial PDB structure file contains the atomic coordinate information of the protein and is the basis for all subsequent calculations. Therefore, in this step, the present invention ensures that the file format is correct and pre-processes any potential problems to lay a solid foundation for molecular dynamics simulations and spectral calculations.
[0055] S12, performing molecular dynamics simulation of the protein on the initial PDB structure file to obtain molecular dynamics simulation data.
[0056] In this embodiment, molecular dynamics simulations were performed using GROMACS software to study the dynamic behavior of proteins. Molecular dynamics simulations are computer simulation techniques used to study the physical motion of molecular systems over time. Specifically, molecular dynamics simulations predict the trajectory and behavior of molecules by calculating the interaction forces between molecules based on the equations of motion of classical mechanics. The process includes constructing a molecular system, such as a protein-ligand complex, setting initial conditions such as temperature, pressure, and a simulated environment, minimizing energy to ensure system stability, and then applying a corresponding integral algorithm, such as the Verlet integral method, to iteratively calculate the displacement and velocity of molecules at different time points. Through long-term simulations, the dynamic changes of the molecular system can be observed, and information such as the conformational evolution of the protein and the binding mechanism with the ligand can be obtained. This step provides a verification basis for the dynamic behavior of the protein structure after design optimization, ensuring that the designed protein not only has the expected function under static conditions, but also remains stable in a dynamic environment.
[0057] S13, using AIM to calculate the molecular dynamics simulation data to obtain intermolecular interaction data; using NISE to convert the intermolecular interaction data into converted spectral data to obtain initial protein simulation spectral data.
[0058] In one embodiment of the present invention, AIM (Atoms in Molecules) is primarily used to calculate and analyze interatomic interactions in molecular structures. Specifically, AIM analyzes bonding and non-bonding interactions in molecular systems by calculating the electron density distribution and using the Hamiltonian matrix in quantum chemistry. First, the electron density distribution of the molecule is calculated through quantum chemical calculations. Then, AIM uses Bader theory to divide the electron density into atomic regions. By analyzing the boundaries of these regions, namely the zero-pass points of the electron density and its gradient, the interactions between atoms, including covalent bonds, hydrogen bonds, and van der Waals forces, can be accurately identified. Finally, AIM uses the Hamiltonian matrix to calculate the bond strength, potential energy, and interaction energy of the molecular system, thereby evaluating the stability and structural properties of the molecular system.
[0059] In this real-time example, NISE (Nonlinear Infrared Spectroscopy Estimator) is used to process and analyze the spectral data in molecular dynamics simulation, particularly for calculating two-dimensional infrared spectra (2DIR). The NISE method, by combining molecular dynamics trajectory, first calculates the dynamic behavior of molecules evolving over time. Then, the molecular configuration in each frame trajectory is utilized to simulate spectral characteristics by calculating the second-order response function of the molecule. NISE adopts efficient computing algorithms to process a large amount of trajectory data, and by calculating the transition dipole moments of electronic and vibrational states, accurately simulates the influence of intermolecular interactions on the spectrum, particularly energy transfer, coupling effects, etc. The method can also simulate the dynamic behavior of molecules on different time scales, thereby obtaining spectral simulation results with physical reality. In the present embodiment, AIM and NISE are prior art, and the present application is not improved.
[0060] In one embodiment of the present invention, we can visualize the NISE-generated results using Python scripts and specialized visualization tools. This process not only provides insight into the displayed data but also reveals underlying scientific discoveries and trends. Detailed analysis of the results allows us to draw important conclusions about protein behavior, providing solid support for further research.
[0061] In one embodiment of the present invention, protein structural feature data, protein simulated spectral data, and protein-ligand interaction data are obtained by:
[0062] S14, converting the protein three-dimensional structure data into a protein two-dimensional distance matrix; at the same time, converting the format of the initial protein simulated spectrum data to obtain a protein simulated spectrum two-dimensional matrix; and normalizing the protein two-dimensional distance matrix and the protein simulated spectrum two-dimensional matrix.
[0063] In this example, a script was used to calculate the corresponding residue distance matrix from the data in the original PDB structure file. Specifically, the residue distance matrix was calculated based on the Euclidean distances between the CA atoms (i.e., α-carbon atoms) in the original PDB structure file. The calculation process was as follows: First, the three-dimensional coordinates of the CA atoms of all residues in the original PDB structure file were extracted. Then, the distance between any two CA atoms was calculated, i.e., the protein three-dimensional structure data was converted into a protein two-dimensional distance matrix using the following formula:
[0064]
[0065] Where d(i, j) is the distance between two CA atoms i and j, (x i ,y i, z i ) and (x j ,y j , z j ), which are the three-dimensional coordinates of the two CA atoms i and CA atom j respectively. By calculating the distances between all residues in the protein, a symmetric two-dimensional protein distance matrix can be generated. This two-dimensional protein distance matrix represents the geometric structure of the protein skeleton and can be used to reconstruct the protein skeleton model. At the same time, the corresponding initial protein simulation spectral data is converted into a format acceptable to the model, such as a two-dimensional image or matrix, and normalized to ensure that the spectral data matches the protein structural features. The purpose of this step is to enable the multi-task neural network model to not only receive the structural information of the protein, but also process its corresponding spectral data.
[0066] S15, according to the topological structure of the protein, the normalized protein two-dimensional distance matrix and the protein simulated spectrum two-dimensional matrix are simultaneously classified to obtain protein structural feature data and protein simulated spectrum data.
[0067] In one embodiment of the present invention, the normalized two-dimensional protein distance matrix is categorized according to protein topologies, such as ring-shaped and barrel-shaped proteins. Simultaneously, the corresponding two-dimensional matrices of protein simulated spectra are processed and mapped. By combining spectral data with the structural features and properties of proteins, a multi-task neural network model can better learn the relationship between the two.
[0068] S16, establishing a mapping relationship between protein structural feature data and protein simulated spectral data with the same protein structural features and properties to obtain protein-ligand interaction data.
[0069] In one embodiment of the present invention, spectral data is acquired based on the initial PDB structure file, and when classified according to the topological structure of the protein, the spectral data is classified synchronously, so that a mapping relationship can be established between protein structural feature data and protein simulation spectral data with the same protein structural features and properties. The ligands of the protein are such as atoms, molecules and ions, and the structure of the protein can be understood as including the ligand. The protein-ligand interaction data also contains spectral information of the protein and the ligand. For example, by calculating and analyzing the spectral changes of the protein after ligand binding, it is ensured that the protein-ligand interaction relationship can be reflected in the spectral data. This step enables the spectral data to not only describe a single protein, but also reflect its interaction with the ligand.
[0070] In this example, a detailed analysis of the classification results and the matching of protein and spectral data was performed to construct a dataset for protein design. This dataset includes protein structural feature data, protein simulated spectral data, and protein-ligand interaction data. This ensures accurate correlation between spectral and protein features and properties within the dataset, facilitating subsequent model learning and training.
[0071] In one embodiment of the present invention, protein simulation spectral data is generated through theoretical calculations, avoiding the high cost and difficulty of experimental acquisition. The core of this process is to accurately match each spectrum with its corresponding protein structure, ensuring a one-to-one correspondence between the spectra in the dataset and their associated molecular properties.
[0072] S20, establish a multi-task neural network model, train it using the data set, input protein simulation spectral data into the trained multi-task neural network model, and simultaneously obtain the distance matrix between protein skeleton atoms and the number of protein amino acids.
[0073] See also Figure 1 and Figure 2 As shown, in one embodiment of the present invention, the network architecture of the multi-task neural network model includes: an input layer, a convolutional layer, a fully connected layer, a multi-layer feature extraction layer, a feature fusion layer and an output layer.
[0074] The input layer is used to perform normalization preprocessing on the protein simulation spectrum data. During training, the data in the dataset is input into the input layer.
[0075] Convolutional layers are used to extract spatial features from protein simulation spectral data. These layers identify local patterns and features in the spectra, helping the model understand the potential relationship between spectral information and protein structure.
[0076] The fully connected layer maps the spatial features of the simulated protein spectral data extracted by the convolutional layer to the protein's structural feature space, obtaining the initial protein backbone atom distance matrix. This involves mapping the spatial features extracted by the convolutional layer to the protein's structural feature space, such as the backbone atom distance matrix or three-dimensional structural information. The fully connected layer is responsible for converting high-dimensional spectral features into structurally relevant outputs.
[0077] In this embodiment, before output, the present invention uses a special fusion method to achieve the best performance. In this model, feature fusion is completed through the torch.cat (concat) operation. Specifically, the model first extracts features from different layers (layer1, layer2, layer3), and then splices the features of these layers. The innovation of feature fusion lies in combining low-level fine-grained features such as local patterns extracted by convolutional layers with high-level abstract features, so that the final output not only contains local information, but also captures global structural information. The details are as follows:
[0078] The multi-layer feature extraction layer includes a first-layer feature extraction, a second-layer feature extraction and a third-layer feature extraction; and the extraction scale increases successively from the first-layer feature extraction to the third-layer feature extraction; the three-layer feature extraction is used to extract features of different levels of scale from the distance matrix between the initial protein skeleton atoms.
[0079] In this embodiment, three feature extraction layers propose features at different levels, which represent gradual abstraction from the bottom layer to the top layer.
[0080] The feature fusion layer is used to use the splicing function to splice and fuse the features of different scales extracted from multiple layers of feature extraction layers to obtain the final distance matrix between protein skeleton atoms.
[0081] In this embodiment, torch.cat(concat) is used to concatenate and fuse features of different scales extracted from multiple feature extraction layers. Through this combination of features, information from different scales can be integrated, thereby improving the perception of complex protein features.
[0082] The output layer includes a protein structure output layer and an amino acid number output layer; the protein structure output layer directly outputs the final distance matrix between protein backbone atoms; the protein structure is obtained based on the distance matrix between protein backbone atoms, and the characteristic lengths of all output protein structures are unified. The amino acid number output layer uses a filling method to fill in the missing amino acids with invalid information, so that the characteristic lengths of all output protein structures are unified, and then the number of amino acids is output.
[0083] In this embodiment, for the prediction of known proteins, we only need the distance matrix between the protein backbone atoms to obtain the protein structure. The prediction of unknown proteins lacks protein length (number of amino acids) information, so the output layer needs to additionally predict a single value, which represents the number of amino acids.
[0084] Since the input protein simulation spectral data is fixed as a 3x224x224 image, and the output protein lengths vary, we use a padding method to unify the lengths of all output protein structure features to a fixed size. During the training process, we use a custom mask loss function (MaskLoss) to ignore the invalid information in the padding part. The working principle of MaskLoss is: first generate a mask to mark the valid parts of the input and target that are not zero. Then, the loss is calculated only for these valid parts, ignoring the padding part. In this way, the model can focus on learning valid information and improve the accuracy of the prediction. In addition, during the training process, we used a dynamic learning rate strategy, which can automatically change the learning rate size according to the real-time performance of the model to help the model fit and converge faster.
[0085] S30, integrating the distance matrix between protein backbone atoms and the number of protein amino acids to obtain the final protein distance matrix.
[0086] In this embodiment, the integration process includes adjusting the dimension of the distance matrix between protein backbone atoms according to the number of amino acids, ensuring that its size matches the predicted amino acid sequence.
[0087] S40, obtaining the initial protein PDB structure based on the final protein distance matrix.
[0088] In this embodiment, the dimension of the distance matrix between the protein backbone atoms is adjusted according to the number of protein amino acids, and the PyRosetta tool is used to generate the initial protein PDB structure from the distance matrix between the protein backbone atoms after the dimension adjustment, ensuring that the dimension size of the distance matrix between the protein backbone atoms matches the predicted amino acid sequence.
[0089] S50, the initial protein PDB structure was optimized using the SCUBA-D protein structure optimization model to obtain the all-atom protein structure.
[0090] In one embodiment of the present invention, the SCUBA-D protein structure optimization model is used to optimize an initial low-resolution protein structure. The low-resolution initial protein PDB structure is input, and using known protein energy functions and force fields, the protein's main chain and side chain conformations are gradually adjusted to output a higher-resolution all-atom protein structure. This process, based on molecular dynamics simulations and energy minimization algorithms, converts the low-resolution rough structure into a full-atom model with more precise atomic coordinates. Specifically, it includes:
[0091] S51: Establishing the protein index mapping and protein structure optimization configuration ensures that the SCUBA-D protein structure optimization model can process every part of the protein. The core of this step is to ensure that the protein's individual residues, atoms, and other information are consistent with the optimization results.
[0092] The S52, SCUBA-D protein structure optimization model, uses molecular dynamics and force field energy calculations to minimize energy by continuously adjusting the protein's backbone and side chains to optimize the protein's all-atom structure. This ultimately outputs a high-resolution, all-atom protein structure whose conformation is more consistent with actual biological function.
[0093] In this example, the SCUBA-D protein structure optimization model further optimizes protein folding by analyzing the protein's global topology and local interactions. Given a protein PDB structure file as input, the model predicts high-resolution, foldable protein structures, ensuring that the resulting protein structures are physically stable and biologically feasible. The SCUBA-D protein structure optimization model is an existing open-source model.
[0094] S60, applied ProteinMPNN, taking the all-atom protein structure as input to obtain multiple amino acid sequences.
[0095] In one embodiment of the present invention, ProteinMPNN performs global energy calculations on atomic protein structures and generates multiple amino acid sequences based on the energy minimization principle. A desired amino acid sequence can be selected from the multiple amino acid sequences as needed. Specifically, the ProteinMPNN algorithm is an open source model.
[0096] In one embodiment of the present invention, the method for de novo protein reverse design based on spectrum-structure-activity relationship further includes step S100, which is located between step S10 and step S20. Step S100 comprises: establishing a spectral diffusion generative model based on a neural network; training the spectral diffusion generative model using simulated protein spectral data and a custom protein description as inputs, and outputting custom diffuse protein spectral data. Based on user needs, the custom protein description is input into the trained spectral diffusion generative model to output the custom diffuse protein spectral data, and then executing steps S20 to S60 on the custom diffuse protein spectral data.
[0097] In one embodiment of the present invention, a custom protein description is input and obtained by using a script to generate a mapping file between spectra and custom descriptions. This mapping file records the matching relationship between each spectral data and its corresponding protein property, function, and topological description. For example, the custom description corresponding to a spectrum may include information such as the protein's hydrophobicity, amino acid sequence, and structural morphology. Through this mapping relationship, the model can associate spectra with descriptions one by one, helping to learn the intrinsic connection between spectra and protein functional and structural characteristics.
[0098] In this example, the spectral diffusion generative model generates a log file during training, which records detailed training parameters. This log file is loaded and visualized in TensorBoard, allowing real-time monitoring of the loss function and selecting the checkpoint with the lowest loss during training. This visualization helps monitor the performance of the spectral diffusion generative model and ensure its stability during training.
[0099] In this example, users can input specific requirements, such as the properties, shape, and function of the target protein. The spectral diffusion generation model then generates protein spectral data that meets these requirements. This process enables users to generate diverse protein spectra with specific functional characteristics based on specific requirements, providing a foundation for further protein design.
[0100] See also Figure 1 As shown, in one embodiment of the present invention, the verification step is to perform structural verification on the amino acid sequence of the designed protein using AF2 (AlphaFold2) or AF3 (AlphaFold3). AF2 and AF3 are advanced protein structure prediction algorithms developed by DeepMind. They are based on deep learning technology and can predict the three-dimensional structure of a protein by inputting its amino acid sequence.
[0101] In current protein design and research, AF2 and AF3 are widely used for structure verification because of their high accuracy. Compared with traditional experimental methods such as X-ray crystallography and nuclear magnetic resonance (NMR), AF2 and AF3 can not only significantly speed up the prediction of protein structure, but also their prediction accuracy is close to the experimental level in some cases. Although their prediction results cannot completely replace experimental verification, the root mean square deviation (RMSD) analysis shows that the structure predictions of AF2 and AF3 are very close to the actual structure determined by experiment in most proteins, with the error usually within a few angstroms. By inputting the amino acid sequence of a designed protein, the predicted structure generated by AF2 / AF3 can be compared with the designed protein, and the RMSD value of the two can be calculated to assess the closeness of the designed protein to the actual physical structure. This process can verify whether the designed protein meets the expected conformation and function.
[0102] In this example, the model used in the protein de novo reverse design method based on spectrum structure-activity relationship is integrated into an AI system to establish a fully automated process from user requirements directly to the final design of a high-confidence protein structure. Through this system, users can input the target function or structural requirements, and the system will automatically generate and optimize the protein design, such as Figure 3 The structure of the protein was verified by AF2 / AF3 to ensure that the generated protein met the expected requirements.
[0103] See also Figures 1 to 4 As shown, the present invention also provides a protein de novo reverse design system based on spectrum structure-activity relationship, which applies the above-mentioned protein de novo reverse design method based on spectrum structure-activity relationship, including:
[0104] Dataset module 10: used to obtain a data set, wherein the data set includes protein structural feature data, protein simulated spectral data, and protein-ligand interaction data;
[0105] The multi-task prediction module 20 is used to establish a multi-task neural network model, train it using a data set, input protein simulated spectral data into the trained multi-task neural network model, and simultaneously obtain the distance matrix between protein backbone atoms and the number of protein amino acids;
[0106] The protein distance matrix module 30 is used to integrate the distance matrix between protein backbone atoms and the number of protein amino acids to obtain the final protein distance matrix;
[0107] An initial PDB module 40 is used to obtain an initial protein PDB structure according to the final protein distance matrix;
[0108] The all-atom protein structure module 50 is used to optimize the initial protein PDB structure using the SCUBA-D protein structure optimization model to obtain the all-atom protein structure;
[0109] The amino acid sequence module 60 is used to apply ProteinMPNN to obtain multiple amino acid sequences by taking the all-atom protein structure as input.
[0110] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description. It is intended that all variations within the meaning and range of equivalents of the claims be embraced herein, and any reference signs in the claims should not be construed as limiting the claims to which they relate.
[0111] The above-mentioned embodiments merely represent the implementation methods of the invention. The protection scope of the present invention is not limited to the above-mentioned embodiments. For those skilled in the art, several variations and improvements can be made without departing from the concept of the present invention, which all fall within the protection scope of the present invention.
Claims
1. A method for de novo protein reverse design based on spectrum structure-activity relationship, characterized in that: include: S10, acquiring a data set, wherein the data set includes protein structural feature data, protein simulated spectrum data, and protein-ligand interaction data; S20, establishing a multi-task neural network model, training it using the data set, inputting protein simulated spectral data into the trained multi-task neural network model, and simultaneously obtaining the distance matrix between protein backbone atoms and the number of protein amino acids; S30, integrating the distance matrix between protein backbone atoms and the number of protein amino acids to obtain the final protein distance matrix; S40, obtaining the initial protein PDB structure based on the final protein distance matrix; S50, using the SCUBA-D protein structure optimization model to optimize the initial protein PDB structure and obtain the all-atom protein structure; S60, applying ProteinMPNN, takes the all-atom protein structure as input and obtains multiple amino acid sequences; The network architecture of the multi-task neural network model includes: input layer, convolution layer, fully connected layer, multi-layer feature extraction layer, feature fusion layer and output layer; among them, Input layer, used for normalization preprocessing of protein simulation spectral data; Convolutional layer, used to extract spatial features from protein simulation spectral data; The fully connected layer is used to map the spatial features of the protein simulated spectral data extracted by the convolutional layer to the structural feature space of the protein and obtain the distance matrix between the initial protein skeleton atoms; Multiple feature extraction layers, including first-layer feature extraction, second-layer feature extraction, and third-layer feature extraction; and the extraction scales from the first layer to the third layer increase in sequence; the three-layer feature extraction is used to extract features of different scales from the initial protein backbone atom distance matrix; The feature fusion layer is used to use the splicing function to splice and fuse the features of different scales extracted from multiple layers of feature extraction layers to obtain the final distance matrix between protein backbone atoms; The output layer includes a protein structure output layer and an amino acid number output layer; the protein structure output layer directly outputs the final distance matrix between protein backbone atoms; the protein structure is obtained based on the distance matrix between protein backbone atoms, and the characteristic lengths of all output protein structures are unified. The amino acid number output layer uses a filling method to fill in the missing amino acids with invalid information, so that the characteristic lengths of all output protein structures are unified, and then the number of amino acids is output.
2. The method for de novo protein reverse design based on spectrum structure-activity relationship according to claim 1, characterized in that: Before obtaining protein simulation spectral data, initial protein simulation spectral data is first obtained, including the following steps: Obtain the initial PDB structure file from the RCSB protein PDB database; wherein the initial PDB structure file includes protein three-dimensional structure data; Perform molecular dynamics simulation of the protein on the initial PDB structure file to obtain molecular dynamics simulation data; AIM was used to calculate the molecular dynamics simulation data to obtain the intermolecular interaction data; NISE was used to convert the intermolecular interaction data into converted spectral data to obtain the initial protein simulation spectral data.
3. The method for protein de novo reverse design based on spectrum structure-activity relationship according to claim 2, characterized in that: Obtain protein structural feature data, protein simulation spectral data, and protein-ligand interaction data through the following methods: The protein three-dimensional structure data is converted into a protein two-dimensional distance matrix; at the same time, the initial protein simulated spectrum data is formatted to obtain a protein simulated spectrum two-dimensional matrix; and the protein two-dimensional distance matrix and the protein simulated spectrum two-dimensional matrix are normalized; According to the topological structure of the protein, the normalized two-dimensional protein distance matrix and the two-dimensional matrix of protein simulated spectra are simultaneously classified to obtain protein structural feature data and protein simulated spectrum data; A mapping relationship is established between protein structural feature data and protein simulated spectral data with the same protein structural features and properties to obtain protein-ligand interaction data.
4. The method for de novo protein reverse design based on spectrum structure-activity relationship according to claim 1, characterized in that: Obtaining the final protein distance matrix includes: adjusting the dimension of the distance matrix between protein backbone atoms according to the number of protein amino acids, and using the PyRosetta tool to generate the initial protein PDB structure from the adjusted dimension of the distance matrix between protein backbone atoms.
5. The method for de novo protein reverse design based on spectrum structure-activity relationship according to claim 1, characterized in that: Obtaining the all-atom protein structure includes: establishing the protein index mapping and protein structure optimization configuration to ensure that the SCUBA-D protein structure optimization model can process every part of the protein; the SCUBA-D protein structure optimization model is based on molecular dynamics and force field energy calculations, and continuously adjusts the main chain and side chain of the protein to achieve energy minimization to optimize the protein's all-atom structure and obtain the all-atom protein structure.
6. The method for de novo protein reverse design based on spectrum structure-activity relationship according to claim 1, characterized in that: Obtain multiple amino acid sequences, including: ProteinMPNN performs global energy calculations on atomic protein structures and generates multiple amino acid sequences based on the energy minimization principle.
7. The method for de novo protein reverse design based on spectrum structure-activity relationship according to claim 1, characterized in that: The protein de novo reverse design method based on spectrum-structure-activity relationship also includes step S100, which is located between step S10 and step S20; wherein, step S100 is: establishing a spectral diffusion generation model constructed based on a neural network; using protein simulated spectral data and input custom protein description as inputs of the spectral diffusion generation model for training, and outputting custom diffused protein spectral data; wherein, the input custom protein description is obtained in the following manner: using a script to generate a mapping relationship file between the spectrum and the custom description, and the mapping relationship file records the matching relationship between each spectral data and its corresponding protein attribute, function and topological description.
8. The method for de novo protein reverse design based on spectrum structure-activity relationship according to claim 7, characterized in that: According to user needs, a custom protein description is input into the trained spectral diffusion generation model to output custom diffuse protein spectral data, and then the custom diffuse protein spectral data is processed through steps S20 to S60.
9. A protein de novo reverse design system based on spectrum structure-activity relationship, characterized by: The method for de novo protein reverse design based on spectrum structure-activity relationship according to any one of claims 1 to 8 comprises: Dataset module: used to obtain data sets, where the data sets include protein structural feature data, protein simulation spectrum data, and protein-ligand interaction data; The multi-task prediction module is used to build a multi-task neural network model and train it using a data set. The protein simulation spectrum data is input into the trained multi-task neural network model, and the distance matrix between protein backbone atoms and the number of protein amino acids are obtained at the same time. Protein distance matrix module, which is used to integrate the distance matrix between protein backbone atoms and the number of protein amino acids to obtain the final protein distance matrix; Initial PDB module, used to obtain the initial protein PDB structure based on the final protein distance matrix; All-atom protein structure module, used to optimize the initial protein PDB structure using the SCUBA-D protein structure optimization model to obtain the all-atom protein structure; The amino acid sequence module is used to apply ProteinMPNN to obtain multiple amino acid sequences using all-atom protein structures as input.