Traditional Chinese medicine compound adaptation disease prediction method based on high-order network
By constructing a network of Chinese medicine compound prescriptions and indications and using the high-order network similarity algorithm HODDA, the accuracy and precision problems of Chinese medicine compound indication prediction were solved, and accurate prediction of Chinese medicine compound indications and cost reduction were achieved.
Patent Information
- Application Number
- CN202510709864.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies make it difficult to effectively predict the multi-component-multi-target-multi-pathway associations of traditional Chinese medicine compound prescriptions, resulting in a wide range of indications and unclear positioning of applicable groups, which can easily lead to drug mixing, misuse and adverse reactions.
The Chinese medicine compound network and indication network were constructed, the first-order network and N-order subnetwork were defined, and the association between Chinese medicine compound and indication was predicted using the high-order network similarity algorithm HODDA by calculating the first-order network similarity and N-order subnetwork similarity.
It improves the accuracy and precision of the prediction of indications for traditional Chinese medicine compounds, solves the problems of a wide range of indications and unclear positioning of applicable groups, and reduces the cost of drug research and development.
Smart Images

Figure CN120613154A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of Chinese medicine compound indication prediction, and in particular relates to a Chinese medicine compound indication prediction method based on a high-order network. Background Art
[0002] Clarifying the indications of traditional Chinese medicine (TCM) compounds can provide a basis for doctors to prescribe rationally and clarify patient medication use, helping them maximize their effectiveness in clinical practice. Existing TCM compounds often suffer from broad indications in their instructions and unclear target groups. This not only affects their efficacy but also easily leads to drug mix-ups and misuse, and even adverse reactions. Unlike existing chemical drugs, which typically have a single component, TCM compounds possess a multi-component, multi-target, and multi-pathway nature, making the prediction of TCM compound indication associations difficult and complex.
[0003] Computational biology methods have attracted widespread attention due to their ability to predict drug-indication associations more quickly and at a lower cost. However, existing research methods primarily focus on predicting associations between the active ingredients of traditional Chinese medicine compounds and target proteins, or between single herbs and target proteins, without considering the multi-ingredient, multi-target nature of traditional Chinese medicine compounds. Alternatively, they focus on extracting similarity features from local and single networks, ignoring the global structure of multiple networks and the high-order characteristics of their nodes. Summary of the Invention
[0004] To solve the above technical problems, the present invention combines complex networks, artificial intelligence, optimization calculation and other methods to propose a high-order network structure-based Chinese medicine compound indication prediction research algorithm HODDA (High-Order Drug-Disease Associations) to predict the association between Chinese medicine compounds and indications.
[0005] To achieve the above objectives, the present invention provides a method for predicting the indications of traditional Chinese medicine compound prescriptions based on a high-order network, comprising:
[0006] Constructing a traditional Chinese medicine compound network and an indication network, and defining a first-order network and an N-order subnetwork in the traditional Chinese medicine compound network and the indication network;
[0007] Calculating the first-order network similarity and the N-order subnetwork similarity in the traditional Chinese medicine compound network and the indication network;
[0008] Based on the first-order network similarity and the N-order subnetwork similarity, the association between the traditional Chinese medicine compound and the indication is predicted.
[0009] Preferably, the process of constructing the Chinese medicine compound network and indication network includes:
[0010] Obtain targets corresponding to ingredients in traditional Chinese medicine compound prescriptions and targets related to indications;
[0011] Using protein interaction network information, establish a traditional Chinese medicine compound network and indication network;
[0012] The nodes in the TCM compound network represent the targets corresponding to the medicinal ingredients in the TCM compound, and the edges represent the associations between the targets.
[0013] The nodes in the indication network represent indication-related targets, and the edges represent the associations between targets.
[0014] Preferably, the process of obtaining targets corresponding to the components of a traditional Chinese medicine compound and targets related to indications includes:
[0015] Search for drug ingredient targets and indication-related targets in traditional Chinese medicine compounds based on the database;
[0016] The acquired targets were integrated and deduplicated, and the TCM compound network and indication network were established using the STRING database.
[0017] Preferably, the process of defining the first-order network and the N-order subnetwork in the Chinese medicine compound network and the indication network includes:
[0018] The target set of TCM compound perturbation and the target set of indication perturbation are respectively regarded as first-order networks;
[0019] In the protein interaction network, for any traditional Chinese medicine compound or indication, there are N nodes. If any two nodes in the node are connected, then the N nodes constitute an N-order subnetwork of the traditional Chinese medicine compound or indication;
[0020] Wherein, N is greater than 1, and the N-order subnetwork does not contain similar nodes of 1 to N-1 subnetworks.
[0021] Preferably, the process of defining the N-order sub-network further includes:
[0022] In the process of determining an N-order subnetwork, if the current-order subnetwork contains nodes similar to those of a lower-order subnetwork, the similar nodes are removed and a higher-order subnetwork is determined until no higher-order subnetwork can be constructed.
[0023] Preferably, the process of calculating the first-order network similarity includes:
[0024] For any traditional Chinese medicine compound and indication, determine the set of perturbed targets respectively;
[0025] The Jaccard similarity formula was used to calculate the first-order network similarity between the target sets.
[0026] Preferably, the process of calculating the similarity of the N-order subnetwork includes:
[0027] Determine the N-order subnetwork of traditional Chinese medicine compound prescriptions and indications;
[0028] For any two nodes, if the distance between them in the protein interaction network is less than or equal to N, the corresponding subnetwork is regarded as the N-order similarity subnetwork;
[0029] The N-order similarity between the Chinese herbal compound and its indications is calculated based on preset parameters and formulas.
[0030] Preferably, the process of calculating the N-order similarity between the Chinese herbal compound and the indication according to the preset parameters and the HODDA algorithm includes:
[0031] The parameters of the HODDA algorithm are optimized and selected using an n-fold cross-validation method.
[0032] Preferably, the value range of the preset parameter is between 0 and 1.
[0033] Preferably, the process of predicting the association between the Chinese herbal compound and the indication based on the first-order network similarity and the N-order subnetwork similarity includes:
[0034] Performing weighted summation on the first-order network similarity and the N-order sub-network similarity to obtain a comprehensive similarity between the Chinese medicine compound and the indication;
[0035] The association between the Chinese herbal compound and the indication is predicted based on the comprehensive similarity, and the comprehensive similarity scores are ranked. The higher the score, the greater the association between the Chinese herbal compound and the indication.
[0036] Compared with the prior art, the present invention has the following advantages and technical effects:
[0037] The prediction method of the present invention takes into account the "multi-component-multi-target-multi-pathway" of traditional Chinese medicine compounds. By defining first-order networks and N-order subnetworks, as well as a high-order network similarity calculation method, it can capture the global structure of the network and the high-order characteristics of the nodes in the network, thereby improving the accuracy and precision of the prediction of the indications of traditional Chinese medicine compounds, and providing better technical support for the prediction research of the indications of traditional Chinese medicine compounds.
[0038] The present invention has significantly improved the accuracy and precision of predicting the indications of traditional Chinese medicine compounds. It can solve the problems of a wide range of indications for traditional Chinese medicine compounds and unclear positioning of applicable groups, realize potential prediction of traditional Chinese medicine compound-indication associations, and provide technical support for the discovery of new indications for existing traditional Chinese medicine compounds, expand the scope of application, and reduce the cost of drug research and development. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0040] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention. DETAILED DESCRIPTION
[0041] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0042] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0043] like Figure 1 As shown, this embodiment provides a method for predicting the indications of a traditional Chinese medicine compound based on a high-order network, comprising:
[0044] Construct a Chinese medicine compound network and an indication network, and define the first-order network and N-order subnetwork in the Chinese medicine compound network and the indication network;
[0045] Calculate the first-order network similarity and N-order sub-network similarity in the traditional Chinese medicine compound network and indication network;
[0046] The association between Chinese herbal compound prescriptions and indications was predicted based on the first-order network similarity and N-order subnetwork similarity.
[0047] Specifically, by selecting targets corresponding to the ingredients of the Chinese herbal compound and targets related to indications, a Chinese herbal compound network and an indication network are constructed; wherein, there can be multiple indications and corresponding multiple indication-related targets;
[0048] According to the connection relationship between the nodes in the TCM compound network or indication network in the protein interaction network PPI, the first-order network and N-order sub-network are defined;
[0049] According to the distance between the nodes in the two heterogeneous networks of TCM compound network and indication network, the first-order network similarity and N-order sub-network similarity (N is greater than or equal to 2) are calculated;
[0050] Based on the N-order sub-network, a high-order network similarity calculation method is constructed, that is, the association prediction of traditional Chinese medicine compound-indications is performed based on the calculation results.
[0051] Furthermore, a high-order network similarity calculation method is constructed as a similarity algorithm HODDA based on a high-order network structure.
[0052] Furthermore, the process of constructing the TCM compound network and indication network includes:
[0053] Obtain targets corresponding to ingredients in traditional Chinese medicine compound prescriptions and targets related to indications;
[0054] Using protein interaction network information, establish a traditional Chinese medicine compound network and indication network;
[0055] The nodes in the TCM compound network represent the targets corresponding to the drug components in the TCM compound, and the edges represent the associations between the targets.
[0056] The nodes in the indication network represent indication-related targets, and the edges represent the associations between targets.
[0057] Furthermore, the process of obtaining targets corresponding to the components of Chinese herbal compound medicines and targets related to indications includes:
[0058] Search for drug component targets and indication-related targets in traditional Chinese medicine compounds from relevant databases;
[0059] The acquired targets were integrated and deduplicated, and the TCM compound network and indication network were established using the STRING database.
[0060] In this step, as an additional implementation method, the process of constructing the target network of the traditional Chinese medicine compound and the target network of the indication includes:
[0061] We searched for drug ingredient targets and indication-related targets in TCM formulas using the corresponding databases, and constructed TCM formula networks and indication networks using the STRING database. Nodes in the TCM formula network represent targets corresponding to drug ingredients in the formula, and edges represent relationships between targets. Nodes in the indication network represent indication-related targets, and edges represent relationships between indication-related targets.
[0062] Furthermore, the process of defining the first-order network and N-order subnetworks in the TCM compound network and indication network includes:
[0063] The target set of TCM compound perturbation and the target set of indication perturbation are respectively regarded as first-order networks;
[0064] In the protein interaction network, for any Chinese herbal compound or indication, there are N nodes. If any two nodes in the node are connected, then the N nodes constitute an N-order subnetwork of the Chinese herbal compound or indication.
[0065] Among them, N is greater than 1, and the N-order subnetwork does not contain similar nodes in 1 to N-1 subnetworks.
[0066] Furthermore, the process of defining an N-order subnetwork also includes:
[0067] In the process of determining the N-order subnetwork, if the current-order subnetwork contains nodes similar to those of the lower-order subnetwork, the similar nodes are removed and the higher-order subnetworks are determined until no higher-order subnetworks can be formed.
[0068] In this step, as an additional implementation method, the first-order network is: For any Chinese herbal formula F and indication D, assuming that the target set perturbed by the Chinese herbal formula F is Target f Target of indication D perturbation is Target d , then the first-order network is the original Target f and Target d .
[0069] In this step, as an additional implementation method, the N-order subnetwork means that for any Chinese herbal compound F or indication D, there are N nodes, so that any two points among these nodes are connected in the protein-protein interaction network (PPI), then these N nodes are called an N-order subnetwork of the Chinese herbal compound F or indication D, and satisfy N greater than 1, and the N-order subnetwork does not contain similar nodes of 1 to N-1 subnetworks.
[0070] Furthermore, the process of calculating the first-order network similarity includes:
[0071] For any traditional Chinese medicine compound and indication, determine the set of perturbed targets respectively;
[0072] The Jaccard similarity formula was used to calculate the first-order network similarity between target sets.
[0073] Furthermore, the process of calculating the similarity of the N-order subnetwork includes:
[0074] Determine the N-order subnetwork of traditional Chinese medicine compound prescriptions and indications;
[0075] For any two nodes, if the distance between them in the protein interaction network is less than or equal to N, the corresponding subnetwork is regarded as the N-order similarity subnetwork;
[0076] The N-order similarity between the Chinese herbal compound and its indications is calculated based on the preset parameters and the HODDA algorithm.
[0077] Furthermore, the process of calculating the N-order similarity between the Chinese herbal compound and the indication according to the preset parameters and formula includes:
[0078] The n-fold cross validation method is used to optimize the parameters of the HODDA algorithm.
[0079] Furthermore, the value range of the preset parameter is between 0 and 1.
[0080] Further optimization scheme, in order to further optimize the effectiveness of the algorithm, this embodiment optimizes the parameters of the HODDA algorithm.
[0081] When the parameter θ is 1,
[0082] When the parameter θ is between 0 and 1,
[0083] In order to verify the effectiveness of the algorithm under different parameters, n-fold cross validation is used, in which part of the data is used as a test set and the remaining data is used as a training set.
[0084] In this step, as an additional implementation method, the first-order network similarity is defined as:
[0085] For any traditional Chinese medicine compound F and indication D, assume that the target set perturbed by the traditional Chinese medicine compound F is Target f Target of indication D perturbation is Target d , then the formula expression of the first-order network similarity, that is, the original Jaccard similarity, is:
[0086]
[0087] In this step, as an additional implementation method, the N-order network similarity is defined as:
[0088] The N-order subnetwork of the Chinese herbal formula F or indication D is represented by F n or D n For any node i∈F n Or node j∈D n , meet dis min(i,j) ≤N, that is, the distance between points i and j is less than or equal to N, which becomes F n With D n The N-order similar subnetwork is denoted as F n ≈D n At this time, the formula for the N-order similarity between the Chinese medicine compound F and the indication D is:
[0089]
[0090] Among them, θ is a parameter, and its value is between 0 and 1.
[0091] For the PPI matrix A, dis min(i,j)≤N satisfies, existence And a ij >0, where
[0092]
[0093] Furthermore, based on the first-order network similarity and the N-order subnetwork similarity, the process of predicting the association between the traditional Chinese medicine compound and the indication includes:
[0094] The first-order network similarity and the N-order sub-network similarity are weighted and summed to obtain the comprehensive similarity between the Chinese medicine compound and its indications;
[0095] The association between the Chinese herbal compound and the indication is predicted based on the comprehensive similarity, and the comprehensive similarity scores are ranked. The higher the score, the greater the association between the Chinese herbal compound and the indication.
[0096] In this step, as an additional implementation method, the similarity between any Chinese herbal compound F and indication D is expressed as follows:
[0097]
[0098] That is, the similarity between TCM compound F and indication D is the sum of the similarities from order 1 to order n. The higher the DDASim score, the greater the correlation between TCM compound F and indication D.
[0099] Example 1
[0100] Relevant databases were used to obtain drug component targets and indication-related targets in the TCM compound. The obtained targets were integrated and deduplicated, and then imported into the database to obtain a protein interaction network F composed of TCM compound targets and a protein interaction network D composed of indication-related targets.
[0101] If the target set of the Chinese medicine compound network F is (a, b, c, d, h), and the target set of the indication network D is (b, c, e, f, g), then:
[0102] Calculate the first-order similarity between the two
[0103] Remove the first-order similarity nodes b and c, and generate the second-order sub-networks [(a, d), (a, h), (d, h)] and [(e, f), (e, g)]. Assuming that (a, d) and (e, f) are second-order similar in the PPI network, then
[0104] If the second-order similarity nodes a, d, e, and f are removed, the iteration stops because the remaining nodes h and g cannot form a third-order network.
[0105] Calculate the high-order network similarity as:
[0106]
[0107] Compared with traditional methods such as MNBDR, SCMFDD, and PREDICT, the HODDA model of this embodiment significantly improves prediction accuracy and precision. This demonstrates that the HODDA model of this embodiment can effectively predict the relationship between TCM compounds and their indications, thereby resolving issues such as the broad indications of TCM compounds and the unclear target population, providing methodological support for the accurate prediction of TCM compound indications.
[0108] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for predicting the indications of traditional Chinese medicine compound prescriptions based on a high-order network, characterized in that: include: Constructing a traditional Chinese medicine compound network and an indication network, and defining a first-order network and an N-order subnetwork in the traditional Chinese medicine compound network and the indication network; Calculating the first-order network similarity and the N-order subnetwork similarity in the traditional Chinese medicine compound network and the indication network; Based on the first-order network similarity and the N-order subnetwork similarity, the association between the traditional Chinese medicine compound and the indication is predicted.
2. The method according to claim 1, characterized in that The process of building a TCM compound network and indication network includes: Obtain targets corresponding to ingredients in traditional Chinese medicine compound prescriptions and targets related to indications; Using protein interaction network information, establish a traditional Chinese medicine compound network and indication network; The nodes in the TCM compound network represent the targets corresponding to the medicinal ingredients in the TCM compound, and the edges represent the associations between the targets. The nodes in the indication network represent indication-related targets, and the edges represent the associations between targets.
3. The method according to claim 2, characterized in that The process of obtaining targets corresponding to the ingredients of Chinese herbal compound and targets related to indications includes: Search for drug ingredient targets and indication-related targets in traditional Chinese medicine compounds based on the database; The acquired targets were integrated and deduplicated, and the TCM compound network and indication network were established using the STRING database.
4. The method according to claim 1, wherein The process of defining the first-order network and the N-order subnetwork in the traditional Chinese medicine compound network and the indication network includes: The target set of TCM compound perturbation and the target set of indication perturbation are respectively regarded as first-order networks; In the protein interaction network, for any traditional Chinese medicine compound or indication, there are N nodes. If any two nodes in the node are connected, then the N nodes constitute an N-order subnetwork of the traditional Chinese medicine compound or indication; Wherein, N is greater than 1, and the N-order subnetwork does not contain similar nodes of 1 to N-1 subnetworks.
5. The method according to claim 1, wherein The process of defining the N-order sub-network further includes: In the process of determining an N-order subnetwork, if the current-order subnetwork contains nodes similar to those of a lower-order subnetwork, the similar nodes are removed and a higher-order subnetwork is determined until no higher-order subnetwork can be constructed.
6. The method according to claim 1, characterized in that The process of calculating the first-order network similarity includes: For any traditional Chinese medicine compound and indication, determine the set of perturbed targets respectively; The Jaccard similarity formula was used to calculate the first-order network similarity between the target sets.
7. The method according to claim 1, characterized in that The process of calculating the similarity of the N-order subnetwork includes: Determine the N-order subnetwork of traditional Chinese medicine compound prescriptions and indications; For any two nodes, if the distance between them in the protein interaction network is less than or equal to N, the corresponding subnetwork is regarded as the N-order similarity subnetwork; The N-order similarity between the Chinese herbal compound and its indications is calculated based on the preset parameters and the HODDA algorithm.
8. The method according to claim 7, characterized in that The process of calculating the N-order similarity between the Chinese herbal compound and the indication according to the preset parameters and formula includes: The parameters of the HODDA algorithm are optimized and selected using an n-fold cross-validation method.
9. The method according to claim 7, characterized in that The value range of the preset parameter is between 0 and 1.
10. The method according to claim 1, characterized in that Based on the first-order network similarity and the N-order subnetwork similarity, the process of predicting the association between the traditional Chinese medicine compound and the indication includes: Performing weighted summation on the first-order network similarity and the N-order sub-network similarity to obtain a comprehensive similarity between the Chinese medicine compound and the indication; The association between the Chinese herbal compound and the indication is predicted based on the comprehensive similarity, and the comprehensive similarity scores are ranked. The higher the score, the greater the association between the Chinese herbal compound and the indication.