Natural product molecular activity prediction method, device and equipment based on molecular element structure, medium and product
By dismantling natural product molecules into metastructures and using deep learning models for activity prediction, the problem of failure to effectively consider the local characteristics and interactions of molecules in the prior art is solved, and higher prediction accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510058404.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The existing deep learning-based prediction methods for the molecular activity of natural product fail to effectively consider the local characteristics and interactions of molecules, resulting in low accuracy of prediction results and cannot reflect the true activity of natural product molecules.
By disassembling natural product molecules into multiple fragments and using the structure of each fragment as a metastructure, each metastructure is represented by a SMILES string, and the molecular attribute prediction model is input for activity prediction. The model includes a sequence information characterization module, a structural feature extraction module and an MLP module, which can capture the local features and interactions of molecules more carefully.
It improves the accuracy and efficiency of molecular activity prediction of natural products, reduces R&D costs, and accelerates the discovery process of active ingredients.
Smart Images

Figure CN119993315A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural product molecular activity prediction, and in particular to a natural product molecular activity prediction method, device, equipment, medium and product based on molecular metastructure. Background Art
[0002] In the field of natural product molecular activity prediction, traditional methods mainly rely on experimental tests and computational models based on the whole molecular structure. Although the experimental test method is accurate, it has significant disadvantages such as high cost, long time consumption, and high resource consumption, which makes it difficult to meet the needs of large-scale screening and rapid discovery of active ingredients. In recent years, with the rapid development of artificial intelligence and machine learning technology, the use of deep learning models to predict molecular activity has gradually become a research hotspot. These methods learn the relationship between molecular structure and activity by training models, thereby realizing the activity prediction of new molecules. However, most of the existing deep learning-based prediction methods still use the whole molecular structure as input, without considering the local characteristics and interactions of the molecule, and the prediction results have low accuracy and cannot reflect the true activity of natural product molecules. Summary of the invention
[0003] The purpose of this application is to provide a method, device, equipment, medium and product for predicting the molecular activity of natural products based on molecular metastructure, so as to improve the accuracy of predicting the molecular activity of natural products.
[0004] To achieve the above objectives, this application provides the following solutions.
[0005] In a first aspect, the present application provides a method for predicting the molecular activity of natural products based on molecular metastructures, comprising:
[0006] The natural product molecule to be predicted is disassembled into multiple fragments, and the structure of each fragment is used as the meta-structure;
[0007] Determine the SMILES string representation of each metastructure separately;
[0008] The SMILES string representation of each meta-structure is input into a molecular property prediction model to obtain an activity prediction result of the natural product molecule to be predicted; the molecular property prediction model is obtained by training a large prediction model, and the large prediction model includes a sequence information representation module, a structural feature extraction module and an MLP module; the sequence information representation module and the structural feature extraction module are both connected to the MLP module, the sequence information representation module is used to process the sequence information of each meta-structure based on the SMILES string of each meta-structure, the structural feature extraction module is used to process the structural features of each meta-structure based on the SMILES string of each meta-structure, and the MLP module is used to connect the sequence information processing results and the structural feature processing results, and perform activity prediction to obtain an activity prediction result.
[0009] In a second aspect, the present application provides a natural product molecular activity prediction device based on a molecular meta-structure, wherein the natural product molecular activity prediction device based on a molecular meta-structure applies the above-mentioned natural product molecular activity prediction method based on a molecular meta-structure, and the natural product molecular activity prediction device based on a molecular meta-structure comprises:
[0010] The molecular disassembly module is used to disassemble the natural product molecule to be predicted into multiple fragments, and use the structure of each fragment as the meta-structure;
[0011] A metastructure representation module, used to determine the SMILES string representation of each metastructure respectively;
[0012] The activity prediction module is used to input the SMILES string representation of each meta-structure into a molecular property prediction model to obtain the activity prediction result of the natural product molecule to be predicted; the molecular property prediction model is obtained by training a large prediction model, and the large prediction model includes a sequence information representation module, a structural feature extraction module and an MLP module; the sequence information representation module and the structural feature extraction module are both connected to the MLP module, the sequence information representation module is used to perform sequence information processing of each meta-structure based on the SMILES string of each meta-structure, the structural feature extraction module is used to perform structural feature processing of each meta-structure based on the SMILES string of each meta-structure, and the MLP module is used to connect the sequence information processing results and the structural feature processing results, and perform activity prediction to obtain the activity prediction result.
[0013] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for predicting the molecular activity of natural products based on molecular meta-structure.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for predicting the molecular activity of natural products based on molecular meta-structures.
[0015] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for predicting the molecular activity of natural products based on molecular meta-structure.
[0016] According to the specific embodiments provided in this application, this application has the following technical effects.
[0017] The present application provides a method, device, equipment, medium and product for predicting the molecular activity of natural products based on molecular meta-structure. The present application first splits the natural product molecules into fragments and uses the structure of each fragment as the meta-structure; then the natural product molecular activity is predicted based on the meta-structure. Through the local characteristics and interactions reflected by the meta-structure, the model can capture the key characteristics of the natural product molecules in more detail, thereby improving the accuracy of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 A schematic flow chart of a method for predicting the molecular activity of natural products based on molecular meta-structures provided in one embodiment of the present application.
[0020] Figure 2 A schematic diagram of a method for predicting the molecular activity of natural products based on molecular meta-structures provided in one embodiment of the present application.
[0021] Figure 3 A schematic diagram of the architecture of a sequence information representation module provided in one embodiment of the present application.
[0022] Figure 4 A schematic diagram of the architecture of a structural feature extraction module provided in one embodiment of the present application.
[0023] Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0025] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0026] As a component of a molecule, molecular metastructure can reflect the local characteristics and interactions of the molecule in more detail, which is of great significance for improving prediction accuracy.
[0027] In view of the many shortcomings of the prior art in predicting the molecular activity of natural products, the embodiments of the present application provide a method for predicting the molecular activity of natural products based on molecular meta-structures, so as to improve the accuracy and efficiency of prediction, reduce R&D costs, and accelerate the discovery process of active ingredients, which has important research value and practical significance.
[0028] In an exemplary embodiment, Figure 1 As shown, a method for predicting the molecular activity of natural products based on molecular meta-structure is provided, including the following steps 101 to 103.
[0029] Step 101, decomposing the natural product molecule to be predicted into multiple fragments, and taking the structure of each fragment as a meta-structure.
[0030] Step 102, determining the SMILES string representation of each metastructure respectively.
[0031] Step 103, input the SMILES string representation of each meta-structure into a molecular property prediction model to obtain an activity prediction result of the natural product molecule to be predicted; the molecular property prediction model is obtained by training a large prediction model, and the large prediction model includes a sequence information representation module, a structural feature extraction module and an MLP module; the sequence information representation module and the structural feature extraction module are both connected to the MLP module, the sequence information representation module is used to perform sequence information processing of each meta-structure based on the SMILES string of each meta-structure, the structural feature extraction module is used to perform feature processing of each meta-structure based on the SMILES string of each meta-structure, and the MLP module is used to connect the sequence information processing result and the structural feature processing result, and perform activity prediction to obtain an activity prediction result.
[0032] Implementing the above steps 101 to 103 can improve the accuracy of the prediction of the molecular activity of natural products.
[0033] In another exemplary embodiment, the above step 101 specifically includes the following steps 201 and 202 .
[0034] Step 201, by simulating the cleavage process of the natural product to be predicted under mass spectrometry conditions, identifying the chemical bond breakage combination of the natural product to be predicted during the cleavage process.
[0035] Step 202: Disassemble the natural product molecule to be predicted into multiple fragments based on the chemical bond breakage combination.
[0036] In another exemplary embodiment, the above step 103 includes the following steps 301 to 305 .
[0037] Step 301: Data collection and preprocessing.
[0038] The present application embodiment needs to collect molecular structure data of natural products from public databases and literature. These data should cover various types of natural products to ensure the generalization ability of subsequent models. The collected data needs to be preprocessed, including removing duplicates, cleaning noise data, etc., to ensure data quality.
[0039] The data of this application mainly comes from public databases and literature, such as PubChem, ChEMBL, etc. These databases provide a large amount of natural product molecular structure data, including molecular formula, SMILES string, InChI key and other information. In order to ensure the diversity and representativeness of the data, this application has collected more than 400,000 natural product molecular structure data from multiple databases.
[0040] The BBBP (Blood-Brain Barrier Penetration), Tox21, ClinTox, BACE (Beta-secretase Cleavage Site Prediction) and SIDER (Side Effect Resource) datasets in the MoleculeNet database were used as high-throughput screening datasets. The BBBP task aims to predict whether a small molecule drug can penetrate the blood-brain barrier (BBB); Tox21 contains 12 independent target prediction subtasks to evaluate the potential effects of molecules on a series of toxic pathways, including nuclear receptor activity, stress response, ion channel blockade, etc.; the ClinTox dataset contains two prediction tasks, the goal of which is to predict whether a molecule is clinically toxic (whether an FDA-approved molecule has been withdrawn due to toxicity) and whether it is mutagenic (Ames test); BACE1 is an enzyme associated with Alzheimer's disease, and the task goal is to predict whether the molecule has BACE1 inhibitory ability; the SIDER dataset contains 27 different side effect prediction subtasks, the goal of which is to predict whether a given drug may cause specific side effects, which is crucial for evaluating drug safety.
[0041] The collected data may have problems such as duplication, missing or inconsistent format, so data cleaning is required. First, use the hash algorithm to remove duplicate data items. Second, for missing data, fill it in based on context information or use interpolation. Finally, unify the format of the data, such as converting the SMILES string into a unified representation for subsequent processing.
[0042] Step 302: Generate molecular metastructure.
[0043] After data preprocessing, this application uses a specific algorithm to decompose natural product molecules into multiple molecular metastructures. These molecular metastructures should contain specific chemical information and connection relationships to reflect the local characteristics and interactions of the molecules, so that the model can capture the key features of the molecules in more detail, thereby improving the accuracy of the prediction.
[0044] The present application embodiment adopts a rule-based fragmentation algorithm. The algorithm simulates the cleavage process of natural products under mass spectrometry conditions, identifies specific chemical bond breakage combinations, and then disassembles the molecule into multiple fragments based on these structural units. Each fragment retains part of the structure and chemical information of the original molecule, which is the metastructure.
[0045] The generated metastructures may contain redundant or irrelevant information, so they need to be screened and optimized. This application screens the metastructures based on factors such as size, complexity, and relevance to molecular activity to ensure that each metastructure effectively reflects the local features and interactions of the molecule.
[0046] In order to input the metastructures into the model, they need to be converted into appropriate representations. In the embodiment of the present application, SMILES strings are used to represent the metastructures. The SMILES string of each metastructure contains information such as the atomic connection order and chemical bond type.
[0047] Step 303: Construct a molecular property prediction model.
[0048] Based on the generated meta-structure data, the embodiment of the present application designs a large prediction model architecture, including a sequence information representation module, a structural feature extraction module and an MLP module. The sequence information representation module is used to process the sequence information of the molecular meta-structure, and the structural feature extraction module focuses on processing the structural features of the molecular meta-structure, and further mines the structural information of the molecule by capturing the spatial relationship within and between the meta-structures. The two modules work together to achieve efficient encoding and decoding of the molecular meta-structure.
[0049] The sequence information encoding module consists of 2-12 Transformer encoder units (i.e., the first Transformer encoder), which is responsible for processing the sequence information of molecular metastructures. It uses a self-attention mechanism to capture the sequential associations and interactions between metastructures. Specifically, the module converts the input SMILES string into an embedding vector sequence, which is then processed through multiple layers of self-attention layers and feedforward neural networks.
[0050] The first Transformer encoder focuses on processing the features of the molecular meta-structure. The meta-structure vector representation obtained in the previous step is input and output as a feature tensor of equal dimensions. Then, the meta-structure number dimension is compressed to 1 through global feature average pooling, and the final output is the feature representation of each meta-structure.
[0051] The first Transformer encoder includes a plurality of transformer encoder blocks, which, combined with an atomic level sequence tokenizer, can encode any SMILES (Simplified molecular input line entry system) into a fixed-length feature vector.
[0052] like Figure 3 As shown, the sequence information representation module specifically includes:
[0053] 1. Tokenizer.
[0054] Input: SMILES string.
[0055] Output: tokenized SMILES sequence.
[0056] Computation process: Use a preset vocabulary to convert each character or special symbol in the SMILES string into a corresponding token (integer index).
[0057] 2. Embedding layer.
[0058] Input: tokenized SMILES sequence (length L, where L is the number of tokens in the sequence).
[0059] Output: A sequence of embedding vectors (one d-dimensional embedding vector for each token).
[0060] Calculation process: For each token ti (i = 1, 2, ..., L) in the sequence, find its corresponding embedding vector through the embedding matrix E∈R|V|×d, where |V| is the size of the vocabulary and d is the dimension of the embedding vector.
[0061] Calculation formula: ei=E[ti], where ei∈Rd is the embedding vector of label ti.
[0062] 3. Transformer encoder layer.
[0063] Input: Sequence of embedding vectors.
[0064] Output: Sequence of encoded vectors after self-attention processing.
[0065] The calculation process is as follows:
[0066] Position encoding: In order to introduce the position information in the sequence, rotary positional embeddings are used. For each embedded vector ei, its position encoding pi is calculated, and then the two are added to obtain a vector xi=ei+pi with position information.
[0067] Self-attention mechanism: Calculate the query, key, and value matrices: Q = XWQ, K = XWK, V = XWV, where X = [x1, x2, ..., xL] is the stack of input vector sequences, WQ, WK, WV∈Rd×dk are learnable parameter matrices, and dk is the dimension of the query, key, and value vectors.
[0068] Use the linear attention mechanism to calculate the attention score and output:
[0069]
[0070] Feedforward Neural Network: Apply a two-layer feedforward neural network to the output of the self-attention mechanism. The first layer is a linear transformation plus an activation function (such as ReLU), and the second layer is another linear transformation.
[0071] 4. Multiple first Transformer encoder stacking
[0072] Input: Sequence of output encoded vectors from the previous layer.
[0073] Output: A sequence of deeper encoding vectors.
[0074] Calculation process:
[0075] The computation process is the same as that of a single first Transformer encoder, but it is repeated multiple times (multiple layers are stacked), with each layer receiving the output of the previous layer as input and producing a new output. By stacking multiple layers, the model is able to capture more complex features and dependencies.
[0076] The above sequence information representation module uses the Masked Language Model (MLM) training method. During the training process, a portion of the tokens in the input sequence are randomly masked, and then the model predicts the loss of these masked tokens (such as cross entropy loss) and updates the model parameters through the back propagation algorithm.
[0077] like Figure 4 As shown in the figure, the second transformer encoder part in the structural feature extraction module is composed of several transformer encoder modules, which takes the entire meta-structure encoding vector of the molecule obtained in the previous step as input and outputs a specific task (such as activity prediction, toxicity prediction, etc.) to extract the hidden structure and function relationship in the meta-structure feature vector. Through the processing of multiple encoding layers, the multi-head self-attention mechanism is used to capture the dependencies in the sequence, and the feature representation is further processed and transformed through a feedforward neural network. The structure and working principle of the second transformer encoder are consistent with those of the first transformer encoder, which will not be repeated here.
[0078] The MLP module uses a multivariate perceptron with a fully connected layer to connect the two feature representations. Then, a fully connected layer and a sigmoid activation function are used to output the activity prediction results of the natural product molecules.
[0079] Step 304: Model training and verification.
[0080] After the model architecture is built, the embodiment of the present application uses the collected molecular meta-structure data to train the model. During the training process, by adjusting the model parameters and optimizing the algorithm, the model can accurately predict the active labels corresponding to the molecular meta-structure. In order to verify the predictive ability of the model, the present application is also tested on multiple independent data sets and compared with existing methods. The test results show that the method proposed in this application performs well in terms of prediction accuracy and generalization ability.
[0081] 1. Data set division.
[0082] In order to train and validate the model, the sample data set was divided into training set, validation set and test set in a ratio of 8:1:1. The training set was used to train the model parameters, the validation set was used to adjust the hyperparameters and prevent overfitting, and the test set was used to evaluate the final performance of the model.
[0083] 2. Loss function and optimization algorithm.
[0084] The loss function of the embodiment of the present application adopts the cross entropy loss function to measure the difference between the model prediction result and the true label. The optimization algorithm adopts the Adam optimizer, which combines the advantages of the momentum method and the RMSprop optimizer, and has a faster convergence speed and better robustness.
[0085] 3. Training process.
[0086] During the training process, the mini-batch gradient descent method was used to update the model parameters. Each batch contains a certain number of molecular metastructure samples. An early stopping mechanism was set to prevent overfitting, that is, training was stopped when the performance on the validation set no longer improved. In addition, a learning rate decay strategy was used to gradually reduce the learning rate to speed up the convergence speed and improve the generalization ability of the model, such as Figure 2 shown.
[0087] 4. Model verification.
[0088] In order to verify the performance of the model, an evaluation was performed on the test set. The evaluation indicator was the area under the receiver operating characteristic curve (AUROC). In addition, a comparative experiment was conducted with other existing molecular activity prediction methods to demonstrate the superiority of the method of this application. The results are shown in Table 1.
[0089] Table 1 Performance comparison of the embodiment model with other models (AUROC)
[0090] BBBP Tox21 ClinTox2 HIV BACE SIDER RF 71.4 76.9 71.3 78.1 86.70 68.4 SVM 72.9 81.8 66.9 79.2 86.20 68.2 MGCN 85.0 70.7 63.4 73.8 73.4 55.2 D-MPNN 71.2 68.9 90.5 75.0 85.3 63.2 DimeNet - 78.0 76.0 - - 61.6 Hu 70.8 78.7 78.9 80.2 85.9 65.2 N-gram 91.2 76.9 85.5 83 87.6 63.2 GraphMVP-C 72.4 74.4 77.5 77.0 81.2 63.9 GEM 72.4 78.1 90.1 80.6 85.6 67.2 ChemBERTa 64.3 - 90.6 62.2 - - MOLFormer 92.6 82.4 89.2 80.9 87.9 61.1 This embodiment 93.34 82.41 91.57 79.62 88.28 66.81
[0091] Step 305: Post-processing.
[0092] The trained model can be applied to the actual prediction of the activity of natural product molecules. By inputting the new natural product molecular structure, the model can quickly output its activity prediction results. In order to further improve the practicality and reliability of the prediction results, this application also includes the step of post-processing the prediction results. For example, removing noise data, screening highly active molecules, etc., to ensure the accuracy and application value of the prediction results, specifically including:
[0093] First, the noise data in the prediction results are removed, that is, those samples with extremely low or extremely high prediction probabilities.
[0094] Secondly, the prediction results are screened and ranked according to business needs to find molecules with high activity.
[0095] Then, the screened molecules were further experimentally verified and analyzed.
[0096] Compared with the prior art, the molecular activity prediction method of natural products based on molecular metastructure in the embodiments of the present application has the following significant advantages:
[0097] Improve prediction accuracy: By introducing the concept of molecular metastructure and combining it with deep learning technology, this application can capture the local characteristics and interactions of molecules in more detail, thereby improving the accuracy of predictions.
[0098] Reduce R&D costs: Compared with traditional experimental testing methods, this application does not require a large amount of experimental operations and material consumption, thereby significantly reducing R&D costs.
[0099] Improve prediction efficiency: By utilizing the powerful computing power of deep learning models, this application can predict the activity of a large number of natural product molecules in a short period of time, significantly improving prediction efficiency.
[0100] Enhanced generalization capability: By constructing a diverse molecular meta-structure dataset and fully training and validating the model, the method proposed in this application has strong generalization capability and can be applied to different types of natural product molecular activity prediction tasks.
[0101] Promote new drug research and development: This application provides an efficient and accurate method for predicting molecular activity in the field of new drug research and development, which helps to accelerate the process of new drug research and development.
[0102] In summary, the molecular meta-structure-based natural product molecular activity prediction method proposed in this application has significant technical advantages and broad application prospects, and is expected to bring revolutionary changes to new drug research and development.
[0103] Based on the same inventive concept, the embodiment of the present application also provides a natural product molecular activity prediction device based on molecular metastructure for realizing the natural product molecular activity prediction method based on molecular metastructure involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more embodiments of the natural product molecular activity prediction device based on molecular metastructure provided below can refer to the limitations of the natural product molecular activity prediction method based on molecular metastructure above, and will not be repeated here.
[0104] In an exemplary embodiment, a natural product molecular activity prediction device based on molecular metastructure is provided, comprising:
[0105] The molecular disassembly module is used to disassemble the natural product molecule to be predicted into multiple fragments, and use the structure of each fragment as the meta-structure.
[0106] The metastructure representation module is used to determine the SMILES string representation of each metastructure respectively.
[0107] The activity prediction module is used to input the SMILES string representation of each meta-structure into a molecular property prediction model to obtain the activity prediction result of the natural product molecule to be predicted; the molecular property prediction model is obtained by training a large prediction model, and the large prediction model includes a sequence information representation module, a structural feature extraction module and an MLP module; the sequence information representation module and the structural feature extraction module are both connected to the MLP module, the sequence information representation module is used to perform sequence information processing of each meta-structure based on the SMILES string of each meta-structure, the structural feature extraction module is used to perform feature processing of each meta-structure based on the SMILES string of each meta-structure, and the MLP module is used to connect the sequence information processing results and the structural feature processing results, and perform activity prediction to obtain the activity prediction result.
[0108] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Wherein, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Wherein, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for predicting the molecular activity of a natural product based on a molecular metastructure is implemented.
[0109] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0110] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0111] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0113] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0114] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0115] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0116] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for predicting the molecular activity of natural products based on molecular metastructure, characterized in that: include: The natural product molecule to be predicted is disassembled into multiple fragments, and the structure of each fragment is used as the meta-structure; Determine the SMILES string representation of each metastructure separately; The SMILES string representation of each meta-structure is input into a molecular property prediction model to obtain an activity prediction result of the natural product molecule to be predicted; the molecular property prediction model is obtained by training a large prediction model, and the large prediction model includes a sequence information representation module, a structural feature extraction module and an MLP module; the sequence information representation module and the structural feature extraction module are both connected to the MLP module, the sequence information representation module is used to perform sequence information processing of each meta-structure based on the SMILES string of each meta-structure, the structural feature extraction module is used to perform structural feature processing of each fragment based on the SMILES string of each meta-structure, and the MLP module is used to connect the sequence information processing results and the structural feature processing results, and perform activity prediction to obtain an activity prediction result.
2. The method for predicting the molecular activity of natural products based on molecular metastructure according to claim 1, characterized in that: Decompose the natural product molecules to be predicted into multiple fragments, including: By simulating the cleavage process of the natural product to be predicted under mass spectrometry conditions, the chemical bond breakage combination of the natural product to be predicted in the cleavage process is identified; The natural product molecule to be predicted is disassembled into multiple fragments based on the chemical bond breakage combination.
3. The method for predicting the molecular activity of natural products based on molecular metastructure according to claim 1, characterized in that: The sequence information representation module includes: a plurality of first Transformer encoders connected in sequence; The first Transformer encoder includes multiple layers of self-attention layers and feed-forward neural network layers.
4. The method for predicting the molecular activity of natural products based on molecular metastructure according to claim 1, characterized in that: The structural feature extraction module includes a global feature average pooling layer, an output layer, and a plurality of second Transformer encoders connected in sequence; The last second Transformer encoder among the plurality of second Transformer encoders connected in sequence is connected to the global feature average pooling layer, and the global feature average pooling layer is connected to the output layer.
5. The method for predicting the molecular activity of natural products based on molecular metastructure according to claim 1, characterized in that: The process of training the large prediction model includes: Obtain the structural data and activity data of natural product molecules with known activity from the data source and construct a data set; A hash algorithm is used to remove duplicate data items in a data set to obtain a deduplicated data set; the data items include structural data and activity data of natural product molecules with known activity; Using context information or interpolation method to fill in the missing data of each data item in the deduplicated data set, to obtain a filled data set; Each natural product molecule with known activity in the filled data set is split into multiple fragments, and the SMILES string representation of each fragment of each natural product molecule with known activity is used as the sample input. The activity data of each natural product molecule with known activity is used as the sample label to construct the sample data set; Dividing the sample data set into a training set, a test set and a validation set; The large prediction model is trained, tested and verified based on the training set, the test set and the verification set to obtain a trained large prediction model as a molecular property prediction model.
6. The method for predicting the molecular activity of natural products based on molecular metastructure according to claim 5, characterized in that: In the process of training, testing and verifying the large prediction model, the cross entropy loss function is used to calculate the loss of the large prediction model obtained in each iterative training; In the process of training, testing and validating the large prediction model, the small batch gradient descent method is used to update the parameters of the large prediction model; During the training, testing and validation of the large oracle model, the learning rate decay strategy is used to gradually reduce the learning rate.
7. A natural product molecular activity prediction device based on molecular metastructure, characterized in that: The molecular activity prediction device for natural products based on molecular meta-structures applies the molecular activity prediction method for natural products based on molecular meta-structures according to any one of claims 1 to 6, and the molecular activity prediction device for natural products based on molecular meta-structures comprises: The molecular disassembly module is used to disassemble the natural product molecule to be predicted into multiple fragments, and use the structure of each fragment as the meta-structure; A metastructure representation module, used to determine the SMILES string representation of each metastructure respectively; The activity prediction module is used to input the SMILES string representation of each meta-structure into a molecular property prediction model to obtain the activity prediction result of the natural product molecule to be predicted; the molecular property prediction model is obtained by training a large prediction model, and the large prediction model includes a sequence information representation module, a structural feature extraction module and an MLP module; the sequence information representation module and the structural feature extraction module are both connected to the MLP module, the sequence information representation module is used to perform sequence information processing of each meta-structure based on the SMILES string of each meta-structure, the structural feature extraction module is used to perform structural feature processing of each meta-structure based on the SMILES string of each meta-structure, and the MLP module is used to connect the sequence information processing results and the structural feature processing results, and perform activity prediction to obtain the activity prediction result.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for predicting the molecular activity of natural products based on molecular meta-structures as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting the molecular activity of natural products based on molecular meta-structures as described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for predicting the molecular activity of natural products based on molecular meta-structures as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Database construction method for mass spectrum analysis of natural product
CN105095448A
Drug small molecule property prediction method and device based on deep learning
CN112164428A
Intelligent drug molecule generation method based on reinforcement learning and docking
CN113488116A
Zero sample learning-based drug virtual screening system for newly discovered target spot
CN114974409A
Method and system for intelligently identifying chemically unstable natural products
CN117854635A