Methods for predicting the TCM functions of single Chinese medicinal substances and for drug screening
Patent Information
- Application Number
- CN202611222231.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-12
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]本发明提供一种中药单体的中医功能预测及药物筛选方法,用以解决现有技术中很难准确预测中药单体的中医功能的缺陷,实现提高中药单体的中医功能预测准确性
[0036] 1. Optimization of feature extraction for complex molecular structures of traditional Chinese medicine significantly improves prediction accuracy.
Smart Images

Figure CN122738679A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence drug technology, and in particular to a method for predicting the TCM functions of a single Chinese herbal medicine and for drug screening. Background Technology
[0002] Traditional Chinese medicine (TCM) has enormous potential in the field of new drug development, but traditionally, screening for active ingredients with specific functions from a vast number of TCM monomers relies mainly on experience, literature review, and extensive in vitro / in vivo experiments. This method is time-consuming, costly, and inefficient, making it difficult to meet the ever-increasing demand for drug screening.
[0003] With the development of artificial intelligence technology, predicting molecular properties using computational methods has become a research hotspot. However, most existing molecular property prediction methods are based on linear representations of molecules (such as SMILES strings) or hand-designed molecular fingerprints. These methods struggle to fully capture the complex topological information and spatial relationships between atoms in molecular structures, resulting in limited prediction accuracy.
[0004] Traditional Chinese medicine (TCM) encompasses a vast array of species, with its active ingredients being numerous and diverse monomeric compounds. Unlike conventional Western medicines or environmental toxins, TCM monomers exhibit an extremely wide range of structural characteristics, including: skeletal complexity: ranging from rigid three-dimensional skeletons composed of complex fused rings, bridged rings, and polycyclic systems (such as steroids) to the flexible long-chain structures unique to terpenes; bond diversity: a greater variety of atomic bonding types, frequently including special structures such as macrocycles and lactones; and functional group richness: including both polar acid-base groups and groups such as hydroxyl, carboxyl, carbonyl, and amino groups primarily based on hydrogen bond networks. The complex and varied composition of TCM monomers at all levels—structure, atomic bonding, and functional groups—makes it difficult to accurately predict their TCM functions. Summary of the Invention
[0005] This invention provides a method for predicting the TCM functions of TCM monomers and screening drugs, which solves the problem that it is difficult to accurately predict the TCM functions of TCM monomers in the existing technology, and improves the accuracy of predicting the TCM functions of TCM monomers.
[0006] This invention provides a method for predicting the TCM functions of monomers in traditional Chinese medicine, comprising:
[0007] Each atom in the molecule of the Chinese herbal monomer to be predicted is regarded as a node in a graph. Based on the physical and chemical properties of each atom, a feature vector is constructed for the node corresponding to each atom. Based on the feature vectors of all atoms, a feature vector matrix is constructed.
[0008] The chemical bonds between the atoms are considered as edges of the graph. An adjacency matrix is constructed for the edges based on the bonding types between the atoms. A graph structure matrix is constructed based on the eigenvector matrix and the adjacency matrix.
[0009] The graph structure matrix is input into the graph convolutional neural network model to obtain the output function score vector of the graph structure matrix. The length of the function score vector is equal to the total number of preset TCM functions. Each score in the function score vector represents the probability that the TCM monomer to be predicted belongs to the corresponding preset TCM function.
[0010] Each score in the functional score vector is compared with a preset threshold, and the predicted TCM function value of the TCM monomer to be predicted is determined based on the comparison result.
[0011] The graph convolutional neural network model is obtained by training a training set, which includes sample Chinese herbal medicine monomers and the true TCM functional label vectors of the sample Chinese herbal medicine monomers. Each label in the true TCM functional label vector corresponds one-to-one with each score in the functional score vector.
[0012] According to the method for predicting the TCM function of a single herb provided by the present invention, the feature vector of the node includes atom type, level, charge number, number of hydrogen bonds and atomic radius.
[0013] According to the present invention, a method for predicting the TCM function of a monomer of a traditional Chinese medicine, an adjacency matrix is constructed for the edges based on the bonding type between the atoms, including:
[0014] The bonding type between atoms corresponding to rows and columns at each position in the adjacency matrix is encoded as an element at each position, and the rows and columns in the adjacency matrix correspond one-to-one with the atoms.
[0015] According to the present invention, a method for predicting the TCM function of a monomer of a traditional Chinese medicine is provided, wherein the bonding types include single bonds, double bonds, triple bonds, aromatic bonds, self-bonds, and unconnected bonds.
[0016] According to the present invention, a method for predicting the TCM function of a single herb is provided, which constructs a graph structure matrix based on the feature vector matrix and the adjacency matrix, including:
[0017] Calculate the connectivity of each atom based on the adjacency matrix;
[0018] Based on the degree of connectivity of all atoms in the molecule of the Chinese herbal monomer to be predicted, a degree matrix is constructed;
[0019] Construct a graph structure matrix based on the eigenvector matrix, adjacency matrix, and degree matrix.
[0020] According to the method for predicting the TCM function of a single Chinese herbal medicine provided by the present invention, the degree matrix is a diagonal matrix, the diagonal matrix has the same dimension as the adjacency matrix, and each diagonal element in the diagonal matrix is obtained according to the element at the corresponding position in the adjacency matrix of each diagonal element's row or column.
[0021] According to the method for predicting the TCM function of a single herb provided by the present invention, a graph structure matrix is constructed based on the feature vector matrix, adjacency matrix, and degree matrix, including:
[0022] Calculate the outer product of the adjacency matrix and the eigenvector matrix;
[0023] The inner product of the degree matrix and the outer product is calculated to obtain the graph structure matrix.
[0024] According to the method for predicting the TCM function of a single herb provided by the present invention, before inputting the graph structure matrix into a graph convolutional neural network model, the method further includes:
[0025] The functional score vector of each sample Chinese herbal medicine monomer and the real Chinese medicine functional label vector are input into the loss function to obtain the error between the functional score vector and the real Chinese medicine functional label vector.
[0026] The gradient is calculated using the backpropagation algorithm based on the error.
[0027] The optimizer updates the weight parameters in the graph convolutional neural network model based on the gradient.
[0028] According to the method for predicting the TCM function of a single herb provided by the present invention, the loss function for training the graph convolutional neural network model is:
[0029] Loss= ;
[0030] Where Loss is the loss function, and m is the number of drug monomers in the sample. This represents the true TCM functional label of the monomer of the i-th sample. Let i be the functional score predicted for the drug monomer in the i-th sample. This is the Sigmoid activation function.
[0031] This invention also provides a drug screening method, comprising:
[0032] The Chinese herbal monomers that were not marked as having the target TCM function were selected from the database and used as the Chinese herbal monomers to be screened.
[0033] Based on any of the above-described methods for predicting the TCM function of individual TCM monomers, the predicted TCM function value of the individual TCM monomer to be screened is obtained.
[0034] If the predicted TCM function of the target TCM monomer includes the target TCM function, the target TCM monomer will be selected as a candidate drug.
[0035] The method for predicting the TCM functions of single Chinese medicinal substances and screening drugs provided by this invention has the following beneficial effects:
[0036] 1. Optimization of feature extraction for complex molecular structures of traditional Chinese medicine significantly improves prediction accuracy.
[0037] Existing technologies for processing molecular data often employ general chemical fingerprints or simple one-hot encoding of atomic types. These methods struggle to capture the intricate physicochemical properties of traditional Chinese medicine monomers. This invention delves into the structural characteristics of these monomers and designs a targeted feature extraction mechanism:
[0038] First, this invention innovatively introduces continuous physicochemical parameters such as atomic radius and partial charge number. The active ingredients of traditional Chinese medicine (TCM) (such as flavonoids and alkaloids) typically possess complex fused-ring and bridged-ring rigid skeletons and are rich in polar functional groups. Discrete atomic type encoding alone cannot distinguish the differences of the same element under different chemical environments. This invention, by introducing atomic radius, enables the model to perceive the steric hindrance effect prevalent in TCM molecules, determining whether drug molecules can enter the target binding pocket; by introducing partial charge number, it accurately describes the electron cloud distribution and electrostatic potential of the molecular surface, effectively capturing the interaction forces between hydrogen bond donors and acceptors. This hybrid "discrete + continuous" encoding strategy greatly enhances the model's ability to characterize the microscopic physicochemical properties of TCM monomers.
[0039] Secondly, this invention reconstructs adjacency relationships based on the degree matrix, enhancing local topological features. Considering the diverse bonding modes of traditional Chinese medicine molecules, this invention introduces a degree matrix for normalization during graph structure construction, strengthening the differentiated expression between central atoms and peripheral functional groups. This allows the algorithm to adaptively focus on key active substructures that determine efficacy, rather than being overwhelmed by massive inert framework noise.
[0040] 2. Unique “TCM function”-oriented prediction, with excellent interpretability and potential for new drug discovery.
[0041] Unlike existing technologies that primarily focus on the single safety indicator of "toxicity," this invention focuses on multi-label prediction of "traditional Chinese medicine functions," a goal that endows the model with unique application value.
[0042] On the one hand, it represents a leap from "black-box prediction" to "mechanism discovery." Traditional toxicity predictions often only provide a risk probability, making it difficult to explain the underlying causes. In contrast, the multi-label classification model of this invention can output the specific distribution of monomers across 38 TCM functional dimensions. Since TCM functional labels inherently possess clear pharmacological orientations (for example, "clearing heat" often corresponds to anti-inflammatory and antibacterial mechanisms, and "activating blood circulation" often corresponds to anticoagulant mechanisms), the model's prediction results can be directly mapped to specific biological pathways, exhibiting strong interpretability.
[0043] On the other hand, the successful discovery of previously unrecorded monomers validated the algorithm's discovery capabilities. In practical embodiments of this invention, the model successfully screened and predicted 12 potential active monomers not included in existing databases. These monomers not only structurally conform to the structure-activity relationship rules of specific TCM functions, but subsequent analysis also confirmed their potential pharmacodynamic material basis. This demonstrates that this invention is not merely a classifier, but a powerful tool capable of guiding the development of new TCM drugs and discovering new pharmacodynamic mechanisms.
[0044] 3. The graph-based convolutional neural network backbone architecture and binary cross-entropy loss function are well-suited to the complexity and multi-target characteristics of traditional Chinese medicine theory.
[0045] To address the practical application characteristics of traditional Chinese medicine (TCM) monomer molecules, which often possess multiple TCM functions (e.g., a monomer can both "detoxify" and "invigorate blood"), this invention features a specially designed model architecture. It abandons the shallow networks of traditional methods that are only suitable for single-task prediction, employing a Graph Convolutional Neural Network (GCN) as the backbone network. Multiple layers of graph convolutional layers are used to perform deep feature mining on molecular graph data. This architecture effectively aggregates node and neighborhood information, thereby simultaneously extracting the topological features (e.g., ring structure, connectivity) and attribute features (e.g., atom type, functional group distribution) of molecules. At the network output, a fully connected layer is added to output multi-label prediction probabilities for 38 TCM functions. This architectural design perfectly suits the "one medicine, multiple effects" characteristic of TCM, enabling the model to simultaneously evaluate the potential of monomers across multiple therapeutic dimensions in a single inference, avoiding the redundancy of repeatedly constructing multiple networks required by traditional single-task models, and significantly improving prediction efficiency and system consistency.
[0046] 4. In the model training and optimization phase, this invention addresses the unique characteristics of multi-label classification of TCM functions by making key improvements to the loss function, thus solving the failure problem of traditional methods in the face of the complexity of TCM monomer features: Because the pharmacological effects of TCM monomers have broad overlap, the functional labels are not mutually exclusive. For example, monomers with "heat-clearing" function often also have "detoxifying" function. If the traditional Softmax loss function combined with the cross-entropy loss function is used, the model is forced to treat each label as a mutually exclusive category, leading to information competition during probability normalization and severely distorting the prediction results. Therefore, this invention abandons Softmax and instead adopts the binary cross-entropy (BCE) loss function. The BCE loss function decomposes the multi-label classification problem into multiple independent binary classification problems, enabling independent calculation of the loss value for each TCM functional label and backpropagation. This design ensures that when dealing with the complex features of Chinese medicine monomers, the model can objectively learn the coexistence relationship between different functional labels, rather than forcibly distinguishing them, thereby significantly improving the convergence speed and prediction accuracy of the model in complex multi-label scenarios. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0048] Figure 1 This is a flowchart illustrating the method for predicting the TCM functions of single Chinese medicinal substances provided by the present invention.
[0049] Figure 2 This is a complete flowchart of the TCM function prediction method provided by the present invention;
[0050] Figure 3 This is a schematic flowchart of the drug screening method provided by the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0052] The following is combined with Figure 1 The present invention describes a method for predicting the TCM function of a single herb, comprising:
[0053] Step 101: Treat each atom in the molecule of the Chinese herbal monomer to be predicted as a node in a graph, construct a feature vector for each node corresponding to each atom based on the physical and chemical properties of each atom, and construct a feature vector matrix V based on the feature vectors of all atoms.
[0054] Step 102: Treat the chemical bonds between the atoms as edges of the graph, construct an adjacency matrix A for the edges based on the bonding types between the atoms, and construct a graph structure matrix X (e.g., X = A⨂V) based on the eigenvector matrix V and the adjacency matrix A.
[0055] Step 103: Input the graph structure matrix into the graph convolutional neural network (GCN) model to obtain the output function score vector of the graph structure matrix. The length of the function score vector is equal to the total number of preset TCM functions. Each score in the function score vector represents the probability that the predicted TCM monomer belongs to the corresponding preset TCM function. The score range is between 0 and 1.
[0056] Step 104: Compare each score in the functional score vector with a preset threshold (e.g., 0.5), and determine the TCM functional prediction value of the TCM monomer to be predicted based on the comparison result;
[0057] The graph convolutional neural network model is obtained by training a training set, which includes sample Chinese herbal medicine monomers and the true TCM functional label vectors of the sample Chinese herbal medicine monomers. Each label in the true TCM functional label vector corresponds one-to-one with each score in the functional score vector.
[0058] Molecular graph data encoding of the monomers of traditional Chinese medicine to be predicted refers to converting their molecular structure into a numerical representation that can be processed by a graph neural network, including node feature encoding, edge feature encoding, and the construction of a graph structure matrix X. The atomic feature vectors obtained from node feature encoding are numerical vectors used to characterize the properties of individual atoms, while the adjacency matrix obtained from edge feature encoding is a matrix used to characterize the inter-atomic connections in the molecular structure.
[0059] Graph Convolutional Neural Networks (GNNs) are deep learning models specifically designed for processing graph-structured data. In this embodiment, they are used to extract features from compound molecular graphs and perform functional prediction. A GNN model is constructed. This model receives an input matrix X, extracts the topological and attribute features of the molecule through multiple graph convolutional layers, and finally outputs a functional score vector through a fully connected layer. This vector represents the multi-label prediction probability of multiple (e.g., 38) TCM functions, adapting to the practical application characteristics of TCM monomer molecules often possessing multiple TCM functions.
[0060] If a certain score in the function score vector is greater than a preset threshold, then the entity is predicted to have the function corresponding to that score.
[0061] TCMID (Traditional Chinese Medicine Integrated Database) can be used as a data source for compound structural formulas and their TCM functional labels. The database samples can be two-dimensional structural formulas (e.g., Mol files or SDF formats) of all TCM compound monomers collected from TCMID, covering various molecular structures, atomic bonds, and functional group characteristics. Simultaneously, TCM functional labels corresponding to each compound monomer are obtained from the TCMID database, such as "promoting blood circulation" and "clearing heat and detoxifying," totaling 38 categories. All appearing functional labels are combined into a functional category set, serving as the multi-label classification target for the prediction task. The structural formula and functional label of each compound monomer are combined into a data pair, serving as the training set for the multi-label classification task.
[0062] The actual function label vector is constructed based on the collected data, and its length is the same as the function score vector. If a function is recorded in the TCMID database for a given entity, the value at the corresponding position is 1; otherwise, it is 0.
[0063] This embodiment employs a graph convolutional neural network, which can directly perform deep learning on the graph structure of compounds, effectively extracting atomic features and topological structure information from molecules. Compared with traditional methods, it can more accurately predict the function of traditional Chinese medicine monomers. Through the learning process of the graph neural network, atoms or functional groups that contribute significantly to the prediction results can be traced, providing clues for subsequent drug structure optimization and mechanism of action research.
[0064] Based on the above embodiments, the feature vector of the node in this embodiment includes atom type, level, charge number, number of hydrogen bonds, and atomic radius.
[0065] To overcome the challenges posed by the complex structures and wide distribution of characteristics of traditional Chinese medicine monomers, as well as the need to consider both conformation and the electron cloud of group atomic interactions in predicting TCM therapeutic functions, this embodiment makes targeted improvements to the molecular diagram encoding method:
[0066] Unlike existing technologies that only use one-hot discrete encoding such as atom type and number of connections, this embodiment introduces physicochemical continuous parameters such as atomic radius and charge number into the feature vector of the node.
[0067] Reason for introduction: The bioactivity of Chinese herbal monomers (such as flavonoids and alkaloids) is highly dependent on their three-dimensional spatial conformation (steric hindrance) and electron cloud effect (electron transfer caused by differences in the number of electrons and distance between interacting groups).
[0068] Technical benefits: Atomic radius directly reflects the magnitude of steric hindrance, and charge number accurately describes the electron cloud distribution. This hybrid encoding method, which integrates discrete and continuous features, significantly enhances the characterization ability of complex traditional Chinese medicine molecules.
[0069] Based on the above embodiments, this embodiment constructs an adjacency matrix for the edges according to the bonding type between the atoms, including:
[0070] The bonding type between atoms corresponding to rows and columns at each position in the adjacency matrix is encoded as an element at each position, and the rows and columns in the adjacency matrix correspond one-to-one with the atoms.
[0071] like Figure 2 As shown, the molecule of the herbal monomer to be predicted contains five atoms: A, B, C, D, and E. The rows and columns of the adjacency matrix correspond one-to-one with these five atoms, resulting in a 5x5 adjacency matrix. The bonding type between any two atoms among these five atoms is encoded to obtain the elements in the adjacency matrix.
[0072] Based on the above embodiments, the bonding types described in this embodiment include single bonds, double bonds, triple bonds, aromatic bonds, self-bonding, and no bonding.
[0073] like Figure 2 As shown, single bonds, double bonds, triple bonds, aromatic bonds, themselves (diagonal elements), and unconnected bonds are encoded as 1, 2, 3, 0.5, 1, and 0, respectively.
[0074] Based on the above embodiments, this embodiment constructs a graph structure matrix according to the feature vector matrix and the adjacency matrix, including:
[0075] Calculate the connectivity of each atom based on the adjacency matrix A;
[0076] Based on the degree of connectivity of all atoms in the molecule of the Chinese herbal monomer to be predicted, construct the degree matrix D;
[0077] Based on the eigenvector matrix V, the adjacency matrix A, and the degree matrix D, construct the graph structure matrix X.
[0078] The degree of atomic connectivity refers to the number of chemical bonds that an atom directly forms with other atoms in a molecule.
[0079] From the perspective of graph structure processing, the atomic degree of Chinese medicine molecules varies greatly, requiring special matrix processing methods to balance local and global features. A degree matrix is introduced to specially process "self-connection (diagonal elements)" and atomic degree weights. This method strengthens the differentiated expression of inter-atomic connection relationships with a large or small number of connections, thereby enabling the algorithm to have a stronger feature extraction capability to deal with the more complex molecular features of Chinese medicine monomers.
[0080] Based on the above embodiments, the degree matrix in this embodiment is a diagonal matrix, and the diagonal matrix has the same dimension as the adjacency matrix. Each diagonal element in the diagonal matrix is obtained according to the element at the corresponding position in the adjacency matrix of each diagonal element's row or column.
[0081] like Figure 2 As shown, if a molecule contains five atoms, then both the degree matrix D and the adjacency matrix A are 5x5 matrices. The degree matrix D is a diagonal matrix, and its diagonal element D in the i-th row and i-th column is... ii The results are obtained by statistical analysis of the elements in the i-th row (or i-th column) of the adjacency matrix A.
[0082] Based on the above embodiments, this embodiment constructs a graph structure matrix according to the feature vector matrix, adjacency matrix, and degree matrix, including:
[0083] Calculate the outer product of the adjacency matrix A and the eigenvector matrix V;
[0084] The inner product of the degree matrix D and the outer product is calculated to obtain the graph structure matrix X.
[0085] X = D -1 *(A⨂V), where ⨂ represents the outer product operation.
[0086] This embodiment introduces a degree matrix D to weight the connection relationships.
[0087] Processing method: Calculate the connectivity (Degree) of each atom and construct a diagonal matrix D. During feature propagation, use D... -1 / 2 Perform symmetric normalization on the adjacency matrix.
[0088] Technical effect: This processing method strengthens the differentiated expression of the connection relationship between the number of connections (central atom) and the number of connections (edge atom), avoids numerical instability in the feature aggregation process, and enables the algorithm to better adapt to the complex fused ring and macrocyclic structures commonly found in Chinese medicine monomers.
[0089] Based on the above embodiments, this embodiment further includes the following step before inputting the graph structure matrix into the graph convolutional neural network model:
[0090] The functional score vector of each sample Chinese herbal medicine monomer and the real Chinese medicine functional label vector are input into the loss function to obtain the error between the functional score vector and the real Chinese medicine functional label vector.
[0091] The gradient is calculated using the backpropagation algorithm based on the error.
[0092] The weight parameters in the graph convolutional neural network model are updated using an optimizer (such as Adam) based on the gradient.
[0093] Repeat the above process for multiple rounds of training until the model's performance on the validation set no longer improves or reaches the preset number of training rounds.
[0094] After training, a batch of test data that did not appear during training is loaded and input into the trained model for prediction. The prediction results are compared with the true labels, and evaluation metrics such as accuracy and precision are calculated.
[0095] Based on the above embodiments, the loss function for training the graph convolutional neural network model in this embodiment is:
[0096] Loss= ;
[0097] Where Loss is the loss function, and m is the number of drug monomers in the sample. This is the true TCM function label for the i-th sample herb monomer (0 or 1, indicating whether it has the corresponding function). Let i be the functional score predicted for the drug monomer in the i-th sample. The Sigmoid activation function maps the output to the (0,1) interval. For the raw output of the model After negating the value, calculate the power of e.
[0098] The loss function in this embodiment is optimized for multi-label classification of TCM functions and the complexity of individual herbal characteristics. Since individual herbal substances often possess multiple functions simultaneously (e.g., a substance can both "detoxify" and "invigorate blood"), and the labels are not mutually exclusive, this embodiment abandons the Softmax and Cross-Entropy loss functions suitable for single-label classification and instead adopts the Binary Cross-Entropy (BCE) loss function.
[0099] Advantages of loss function techniques:
[0100] The loss for each label is calculated independently, allowing the model to predict multiple positive labels simultaneously. This perfectly aligns with the "multi-functional synergy" characteristic of traditional Chinese medicine and solves the limitations of traditional multi-classification loss functions in handling the prediction of complex efficacy of Chinese medicine.
[0101] Taking the negative and calculating the power of e makes the loss function differentiable and smooth, which facilitates gradient descent optimization during algorithm training. This addresses the pain point of difficult learning and training caused by the high-dimensional and widely distributed complexity of the molecular features of traditional Chinese medicine monomers.
[0102] The specific steps of model training and performance evaluation include:
[0103] Data preparation: 49,677 compound monomers and their corresponding 38 TCM functional labels were obtained from the TCMID database.
[0104] Data encoding: Each molecule is encoded. The atomic property vector has a dimension of 5 (atomic type, atomic degree, charge number, number of hydrogen bonds, atomic radius), and the bonding type is encoded by the adjacency matrix plus the degree matrix.
[0105] Model Training: Construct a neural network with 3 graph convolutional layers. Use 60% of the data as the training set, 20% as the validation set, and 20% as the test set. Set the batch size to 32, the learning rate to 0.001, and train for 100 epochs.
[0106] Performance Results: After training, the model was evaluated on the test set. The final model achieved 88% accuracy and 82% precision on the multi-label prediction task of predicting the TCM functions of herbal monomers.
[0107] The steps for evaluating the effectiveness of feature engineering include:
[0108] Data preparation: 49,677 compound monomers and their corresponding 38 TCM functional labels were obtained from the TCMID database.
[0109] Data Encoding: Each molecule is encoded. The atomic property vector has a dimension of 5 (atomic type, atomic degree, charge number, number of hydrogen bonds, atomic radius), and the adjacency matrix of bonding type is combined with the degree matrix encoding. At the same time, the adjacency matrix encoding of atomic properties with a dimension of 4 (atomic type, atomic degree, number of connections, neighbor heteroatom type) and bonding type is performed using the one-heat encoding scheme in the toxicity file.
[0110] Model Training: Construct a neural network with 3 graph convolutional layers. Use 60% of the data as the training set, 20% as the validation set, and 20% as the test set. Set the batch size to 32, the learning rate to 0.001, and train for 100 epochs.
[0111] Performance Results: After training, the results were evaluated on the test set. The final models of both methods achieved the accuracy and precision shown in Table 1 for the multi-label prediction task of predicting the TCM functions of herbal monomers.
[0112] Table 1
[0113] method accuracy accuracy Comparison with existing methods 64% 52% Data encoding method of the present invention 88% 82%
[0114] like Figure 3As shown, the present invention also provides a drug screening method, comprising:
[0115] Step 301: Select Chinese herbal medicine monomers that are not marked as having the target TCM function from the database as Chinese herbal medicine monomers to be screened;
[0116] Step 302: Based on the method for predicting the TCM function of TCM monomers in any of the above embodiments, obtain the predicted TCM function value of the TCM monomer to be screened.
[0117] Step 303: If the predicted TCM function of the TCM monomer to be screened includes the target TCM function, the TCM monomer to be screened is selected as a candidate drug.
[0118] The specific steps involved in new drug screening applications include:
[0119] 1. Determine the target function, such as "sedating the nerves".
[0120] 2. Screen out all compound monomers not labeled as having "nerve-sedating" function from the TCMID database.
[0121] 3. Encode the molecular structures of these monomeric compounds and input them into the model trained in step 3 for prediction.
[0122] 4. Collect all monomeric compounds predicted to have "nerve-sedating" function (i.e., corresponding functional score > 0.5) as potential drug candidates.
[0123] This embodiment can automatically and in batches perform functional prediction on massive compound databases, quickly screening potential active monomers with target functions from tens of thousands of candidates, greatly shortening the initial screening cycle of new drug discovery.
[0124] Literature review and experimental verification were conducted on the selected candidate drugs to confirm their actual functions.
[0125] The steps for screening sedative neurofunctional monomers include:
[0126] Screening objective: To find potential Chinese herbal monomers with "nerve-calming" function.
[0127] Screening process: Approximately 8,000 individuals not labeled as “sedating nerves” in the TCMID database were input into the model trained in Example 1 for prediction.
[0128] Screening results: The model predicted that 101 monomers had a greater than 0.5 probability of having a "nerve-sedating" function.
[0129] Results Verification: Literature review revealed that 12 of the 101 monomers have been experimentally confirmed to have sedative effects in recent studies, but this information is not yet included in the TCMID database. Another 25 monomers have been speculated to have sedative effects in some studies and can be considered as key targets for further experimental verification.
[0130] Conclusion: This embodiment demonstrates that the method of the present invention can effectively mine unlabeled potential active compounds in the database, providing an efficient and reliable predictive tool for new drug screening.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting the TCM functions of a single herb, characterized in that, include: Each atom in the molecule of the Chinese herbal monomer to be predicted is regarded as a node in a graph. Based on the physical and chemical properties of each atom, a feature vector is constructed for the node corresponding to each atom. A feature vector matrix is constructed based on the feature vectors of all atoms. The chemical bonds between the atoms are considered as edges of the graph. An adjacency matrix is constructed for the edges based on the bonding types between the atoms. A graph structure matrix is constructed based on the eigenvector matrix and the adjacency matrix. The graph structure matrix is input into the graph convolutional neural network model to obtain the output function score vector of the graph structure matrix. The length of the function score vector is equal to the total number of preset TCM functions. Each score in the function score vector represents the probability that the TCM monomer to be predicted belongs to the corresponding preset TCM function. Each score in the functional score vector is compared with a preset threshold, and the predicted TCM function value of the TCM monomer to be predicted is determined based on the comparison result. The graph convolutional neural network model is trained using a training set, which includes sample Chinese herbal medicine monomers and the true TCM functional label vectors of the sample Chinese herbal medicine monomers. Each label in the true TCM functional label vector corresponds one-to-one with each score in the functional score vector.
2. The method for predicting the TCM function of monomers of traditional Chinese medicine according to claim 1, characterized in that, The feature vector of the node includes atom type, level, charge number, number of hydrogen bonds, and atomic radius.
3. The method for predicting the TCM function of monomers of traditional Chinese medicine according to claim 1, characterized in that, Constructing an adjacency matrix for the edges based on the bonding type between the atoms includes: The bonding type between atoms corresponding to rows and columns at each position in the adjacency matrix is encoded as an element at each position, and the rows and columns in the adjacency matrix correspond one-to-one with the atoms.
4. The method for predicting the TCM function of monomers of traditional Chinese medicine according to claim 3, characterized in that, The bonding types include single bonds, double bonds, triple bonds, aromatic bonds, self bonds, and bonds without linkage.
5. The method for predicting the TCM function of monomers of traditional Chinese medicine according to claim 1, characterized in that, Constructing a graph structure matrix based on the eigenvector matrix and the adjacency matrix includes: Calculate the connectivity of each atom based on the adjacency matrix; Based on the degree of connectivity of all atoms in the molecule of the Chinese herbal monomer to be predicted, a degree matrix is constructed; Construct a graph structure matrix based on the eigenvector matrix, adjacency matrix, and degree matrix.
6. The method for predicting the TCM function of monomers according to claim 5, characterized in that, The degree matrix is a diagonal matrix with the same dimension as the adjacency matrix. Each diagonal element in the diagonal matrix is obtained based on the element at the corresponding position in the adjacency matrix of the row or column of each diagonal element.
7. The method for predicting the TCM function of monomers of traditional Chinese medicine according to claim 5, characterized in that, Based on the eigenvector matrix, adjacency matrix, and degree matrix, a graph structure matrix is constructed, including: Calculate the outer product of the adjacency matrix and the eigenvector matrix; The inner product of the degree matrix and the outer product is calculated to obtain the graph structure matrix.
8. The method for predicting the TCM function of monomers of traditional Chinese medicine according to claim 1, characterized in that, Before inputting the graph structure matrix into the graph convolutional neural network model, the following steps are also included: The functional score vector of each sample Chinese herbal medicine monomer and the real Chinese medicine functional label vector are input into the loss function to obtain the error between the functional score vector and the real Chinese medicine functional label vector. The gradient is calculated using the backpropagation algorithm based on the error. The optimizer updates the weight parameters in the graph convolutional neural network model based on the gradient.
9. The method for predicting the TCM function of monomers of traditional Chinese medicine according to claim 8, characterized in that, The loss function for training the graph convolutional neural network model is: Loss= ; Where Loss is the loss function, and m is the number of drug monomers in the sample. This represents the true TCM functional label of the monomer of the i-th sample. Let i be the functional score predicted for the drug monomer in the i-th sample. This is the Sigmoid activation function.
10. A drug screening method, characterized in that, include: The Chinese herbal monomers that were not marked as having the target TCM function were selected from the database and used as the Chinese herbal monomers to be screened. Based on the method for predicting the TCM function of TCM monomers according to any one of claims 1-9, the predicted TCM function value of the TCM monomer to be screened is obtained. If the predicted TCM function of the target TCM monomer includes the target TCM function, the target TCM monomer will be selected as a candidate drug.