Method for determining core substances and core substance groups during traditional Chinese medicine metabolism
Through the method of combining CCM algorithm and PisCES algorithm, a dynamic substance-matter-related network was constructed, and the core substances and core substance groups in the metabolism of traditional Chinese medicine were discovered, which solved the problem of failure to explore dynamic correlations in the existing technology and provided a more comprehensive basis for new drug research and development.
Patent Information
- Application Number
- CN202310147520.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-22
AI Technical Summary
The existing technology has failed to effectively explore the dynamic correlation between substances during the study of traditional Chinese medicine metabolism, resulting in a lack of comprehensive guidance on the research and development of new drugs.
The CCM algorithm is used to calculate the local and global causal relationships between various substances during the metabolism of traditional Chinese medicine, and a dynamic substance-matter correlation network is constructed based on sliding window technology and weighted Pearson correlation. The PisCES algorithm is used to detect rigid substance clusters in the metabolism process, and the core substance and core substance groups are determined through the analysis of the dynamic topological attributes of the complex network.
It has achieved a comprehensive exploration of the dynamic correlation between substances during the metabolism of traditional Chinese medicine, and provided more accurate guidance on the research and development of new drugs.
Smart Images

Figure CN116153436B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for determining core substances and core substance groups in the metabolism process of traditional Chinese medicine, and belongs to the field of computer-aided drug design. Background Art
[0002] As a traditional medicine in China, traditional Chinese medicine plays a unique advantage in the treatment of diseases such as cancer and cardiovascular and cerebrovascular diseases. Traditional Chinese medicine has the characteristics of a multi-component and multi-target synergistic action mechanism. Drug combination therapy has more advantages than single therapy. For traditional Chinese medicine, this advantage is usually attributed to the interaction of various active ingredients of herbs. However, there are few research reports on the key active substances of drug pairs in the treatment of diseases. The relevant research work mainly focuses on mining the correlation between traditional Chinese medicine components or between traditional Chinese medicine components and in-vivo components with static data. However, the treatment process of diseases involves the dynamic and complex metabolic process of traditional Chinese medicine components and in-vivo substances, and the interaction relationship between substances evolves continuously over time. Therefore, the dynamic interaction relationship between substances needs to be considered in the research process, so as to more comprehensively explore the dynamic correlation between substances in the metabolism process of traditional Chinese medicine, and provide a new perspective and method for new drug research and development.
[0003] At present, more and more methods such as network pharmacology, pharmacodynamics, pharmacokinetics, and metabolomics are being applied to the research of the correlation of active ingredients in traditional Chinese medicine. For example, Cheng ("Network-based prediction of drug combinations". Nature communications, 2019.10(1):1-11) proposed a network-based method to determine drug-drug correlations by quantifying the correlations between drug targets and disease proteins, and thereby identify effective drug combinations. Wang ("Network-based modeling of herb combinations in traditional Chinese medicine". Briefings in bioinformatics, 2021.22(5):bbab106) proposed a framework of network pharmacology on this basis to quantify the interaction magnitude of herb pairs. Wicha S G ("A general pharmacodynamic interaction model identifies perpetrators and victims in drug interactions". Nature communications, 2017.8(1):1-11) proposed a general pharmacodynamic interaction model to quantitatively determine the drug interaction magnitude in a directed manner. Butterweck V ("Potential of pharmacokinetic profiling for detecting herbal interactions with drugs". Clinical pharmacokinetics, 2008.47(6):383-397) used pharmacokinetic methods to predict Herb-Drug interactions and pointed out the clinical significance of drug interactions during metabolism. However, the above methods only mine the interaction relationships between substances during metabolism based on static data and do not involve dynamic relationships.
[0004] Wang X J (《Chinmedomics: Newer Theory and Application》. Chinese Herbal Medicines, 2016.8(4):299-307; 《An integrated chinmedomics strategy for discovery of effective constituents from traditional herbal medicine》. Scientific reports, 2016.6(1):1-12; 《Chinmedomics: A Powerful Approach Integrating Metabolomics with Serum Pharmacochemistry to Evaluate the Efficacy of Traditional Chinese Medicine》, Engineering, 2019.5(1):60-68; 《Chinmedomics, a new strategy for evaluating the therapeutic efficacy of herbal medicines》, Pharmacology & therapeutics, 2020.216:107680) proposed the concept of "Chinmedomics", combining the dynamic changes of traditional Chinese medicine ingredients, metabolites and endogenous substances. Further integrating the methods of traditional Chinese medicine metabolomics and serum pharmacology, Pearson correlation analysis was performed on the static metabolic data, and then the material basis and mechanism of action were explained. Although this method involves the dynamic change law of substances, it only conducts qualitative research on it and cannot obtain the dynamic correlation results of all substances in the whole metabolic process. If only relying on it to determine the core substances and core substance groups in the metabolic process, the results will be relatively one-sided. Summary of the Invention
[0005] In order to more comprehensively explore the dynamic correlations among substances during the metabolism of compound traditional Chinese medicine, the present invention uses the Convergent Cross Mapping (CCM) algorithm, which can calculate the causal relationship scores between substances using time series information, to discover the local and global causal relationships among substances during the metabolism of compound traditional Chinese medicine. A dynamic substance-substance correlation network is constructed through the sliding window technique combined with weighted Pearson correlation. The Global spectral clustering indynamic networks (PisCES) algorithm is used to detect the rigid substance clusters during the metabolism process. Furthermore, the importance of substances is analyzed using the dynamic topological properties of complex networks to determine the core substances and core substance groups during the metabolism process. The method of the present application provides a new perspective and method for new drug research and development by comprehensively exploring the dynamic correlations among substances during the metabolism of traditional Chinese medicine.
[0006] A method for determining core substances and core substance groups during the metabolism of traditional Chinese medicine, the method comprising:
[0007] Step 1, collecting time series data of traditional Chinese medicine metabolism;
[0008] Step 2, for the collected time series data of traditional Chinese medicine metabolism, using the CCM algorithm to obtain the local and global causal relationships among substances during the metabolism of traditional Chinese medicine;
[0009] Step 3, according to the obtained local and global causal relationships among substances, constructing a dynamic substance-substance correlation network through the sliding window technique combined with weighted Pearson correlation;
[0010] Step 4, using the PisCES algorithm to detect the rigid substance clusters during the metabolism process; the rigid substance clusters refer to substances that always maintain a close relationship during the metabolism process;
[0011] Step 5, analyzing the importance of each substance during the metabolism of traditional Chinese medicine according to the dynamic topological properties of the constructed dynamic substance-substance correlation network to determine the core substances and core substance groups during the metabolism process.
[0012] Optionally, the time series data of traditional Chinese medicine metabolism collected in Step 1 is the relative content of each substance generated during the metabolism of traditional Chinese medicine. For any two substances u and v, their time series data is recorded as two time series of length T and where each parameter represents the relative content of the corresponding substance at each moment within a time period of length T.
[0013] Optionally, Step 2 includes:
[0014] Step 2.1, respectively for and Construct a shadow manifold. After performing E-dimensional lag embedding on the time series data of substance u, we obtain τ represents the time lag; the set composed of these vectors is denoted as the shadow manifold Obtained in the same way;
[0015] Step 2.2, by calculating the Euclidean distance Find the E + 1 nearest neighbors at time t, t1,..., t i ,..., t E+1 Respectively represent x u (t)'s time subscript, with the distances from near to far;
[0016] Step 2.3, calculate the weight w of the i-th nearest neighbor i , where d i = exp{-dis x u (t), x u (t i )] / dis x u (t), x u (t1)]};
[0017] Step 2.4, using The time subscripts corresponding to the E + 1 nearest neighbors found, through these E + 1 time subscripts corresponding Values to calculate the estimated value Denoted as
[0018] Step 2.5, calculate and The Pearson correlation coefficient ρ between uv :
[0019]
[0020] Use c uv To represent the influence of cause u on result v, that is, the one-way causal score from substance u to substance v:
[0021]
[0022] That is, obtain the local causal relationships between substances in the traditional Chinese medicine metabolism process;
[0023] Step 2.6, using each substance as a node, the causal relationship score c uvConstruct a weighted directed causal network with the weight being the directed edge from node u to v; perform topological analysis on the causal network, with the out-degree of the nodes in the causal network in-degree respectively representing the influence intensity of substances as causes and results in the metabolic system; the degree centrality c of the node u = c u: + c :u As an importance index of substances, that is, to obtain the global causal relationships among substances in the traditional Chinese medicine metabolism process.
[0024] Optionally, the step 3 includes:
[0025] Step 3.1: Use the sliding window technique to divide the time series of length T into H time windows of length W;
[0026] Step 3.2: Calculate the substance-substance correlation ρ within each time window using the weighted Pearson correlation coefficient u,v ;
[0027] Step 3.3: Use substances as nodes and the substance-substance correlation ρ in step 3.2 u,v as the weight of the edge between nodes u and v to construct a weighted undirected substance-substance correlation network, with each network being a network snapshot;
[0028] Step 3.4: As the time window slides backward, obtain a series of network snapshots on the time axis, jointly constituting a dynamic substance-substance correlation network.
[0029] Optionally, the step 4 includes:
[0030] Step 4.1: For any network snapshot h, determine its optimal number of clusters K h ;
[0031] Step 4.2: Calculate the eigenvectors of the Laplacian matrix L of the network snapshot h h and obtain U h = V h (V h )', V h represents the matrix composed of the first K h eigenvectors of L h (V h )'V h = I, where I represents the identity matrix;
[0032] Step 4.3: Perform Laplacian smoothing on U h through iteration, where represents the number of iterations and α is the smoothing coefficient;
[0033] Step 4.4: Perform k-means clustering on the Laplacian-transformed matrix ;
[0034] Step 4.5: Use the five-fold cross-validation method to measure the reliability of the clustering result with modularity as the evaluation index, and find the optimal smoothing coefficient α;
[0035] Step 4.6: Set α to the optimal value obtained in Step 4.5, repeat Steps 4.1 - 4.4 for N times, and count the number of times two nodes are assigned to the same class in the N results to represent the probability that nodes u and v belong to the same class;
[0036] Step 4.7: Use as the adjacency matrix to obtain a new network, perform k-means clustering on the new network, and use the clustering result as the final clustering result.
[0037] Optionally, the said Step 5 includes:
[0038] Calculate three centrality indexes of Temporal Degree, Temporal Closeness, and Temporal Betweenness in the dynamic substance-substance correlation network:
[0039] (1) Temporal Degree, TemD
[0040]
[0041] where D h (v) represents the degree of node u in the h-th network snapshot;
[0042] (2) Temporal Closeness, TemC
[0043]
[0044] Δ [h,H] (v,u) represents the Temporal shortest path distance from v to u in the network snapshots [h, H].
[0045] (3) Temporal Betweenness, TemB
[0046]
[0047] where g [h,H] (q,u) represents the number of Temporal shortest paths from q to u in the network snapshots [h, H], g [h,H](q, u, v) represents the number of Temporal shortest paths from q to u passing through v.
[0048] Optionally, in step 4.3, the smoothing coefficient α < 0.13.
[0049] Optionally, the method further includes normalizing TemD and TemC by dividing them by H(|V| - 1), and normalizing TemB by dividing it by H(|V| - 2)(|V| - 1) / 2.
[0050] This application also provides an application of the method for determining the core substances and the core substance group in the process of traditional Chinese medicine metabolism in the field of computer-aided drug design.
[0051] The beneficial effects of the present invention are:
[0052] The PisCES algorithm used in this application is a global dynamic network spectral clustering algorithm, which is improved from the spectral clustering algorithm in the static network and has been widely applied in biomedical fields such as proteins, brain functions, and cancers. However, for the first time, this application proposes to use this algorithm to utilize the time series data of substance contents in the metabolic process to mine the correlations between substances, so as to obtain the dynamic correlations between substances in important metabolic processes. After the present invention mines the global and local causal relationships between substances in the metabolic process based on the CCM algorithm, it further constructs a dynamic substance-substance correlation network by combining a sliding window and weighted Pearson correlation. Based on the PisCES algorithm, rigid substance clusters in the metabolic process are discovered, and then the dynamic correlations between substances or multiple substances in the metabolic system are explored. Based on the dynamic correlations between substances in the metabolic process, it can more accurately and comprehensively reflect the dynamic and complex metabolic process of traditional Chinese medicine components and substances in the body, thus providing a more accurate guiding basis for new drug research and development. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is a matrix diagram of the local causal analysis results between 76 substances involved in the process of treating ischemic stroke with Xiangdan injection obtained by the method of the present invention.
[0055] Figure 2 It is a causal network diagram between 76 substances involved in the process of treating ischemic stroke with Xiangdan injection obtained by the method of the present invention.
[0056] Figure 3 It is a dynamic evolution diagram of the substance cluster during the treatment of ischemic stroke with Xiangdan injection obtained by using the method of the present invention. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0058] Example 1:
[0059] This example provides a method for determining core substances and core substance groups during the metabolism of traditional Chinese medicine. The method includes:
[0060] Step 1, collect time-series data of traditional Chinese medicine metabolism;
[0061] In practical applications, the time-series data of traditional Chinese medicine metabolism can be obtained by regularly collecting blood, urine, etc. to detect the types and relative contents of substances during the metabolism process.
[0062] Step 2, for the collected time-series data of traditional Chinese medicine metabolism, use the CCM algorithm to obtain the local and global causal relationships between substances during the metabolism of traditional Chinese medicine;
[0063] Step 3, according to the obtained local and global causal relationships between substances, construct a dynamic substance-substance correlation network by combining the sliding window technique with weighted Pearson correlation;
[0064] Step 4, use the PisCES algorithm to detect rigid substance clusters during the metabolism process; the rigid substance clusters refer to substances that always maintain a close relationship during the metabolism process; according to different actual application scenarios, the substances defined by the rigid substance clusters are different, and drug researchers can determine the rigid substance clusters in the metabolism of traditional Chinese medicine.
[0065] Step 5, analyze the importance of each substance during the metabolism of traditional Chinese medicine according to the dynamic topological attributes of the constructed dynamic substance-substance correlation network, so as to determine the core substances and core substance groups during the metabolism process.
[0066] Example 2
[0067] This embodiment provides a method for determining core substances and core substance groups in the process of traditional Chinese medicine metabolism. Taking the treatment of ischemic stroke with Xiangdan injection as an example, during the determination process, the CCM algorithm is used to analyze the local and global causal relationships among various substances in the metabolism of compound traditional Chinese medicine, and substance pairs with strong causal relationships and key substances are found. A dynamic substance-substance correlation network is constructed by combining a sliding window with weighted Pearson correlation technology, and the PisCES algorithm is used to detect rigid substance clusters in the metabolic process. Further, the importance of substances is analyzed using the dynamic topological properties of complex networks to determine the core substances and core substance groups in the metabolic process.
[0068] This method first requires collecting time series data of traditional Chinese medicine metabolism. After injecting Xiangdan injection into mice, mouse blood is collected at regular intervals, and the types and relative contents of substances in it are detected to obtain the time series data of traditional Chinese medicine metabolism. Through detection, it is determined that a total of 76 substances are produced during the treatment of ischemic stroke with Xiangdan injection. The following details how to analyze the local and global causal relationships among these 76 substances, how to construct a dynamic substance-substance correlation network of 76 substances by combining a sliding window with weighted Pearson correlation technology, and how to use the PisCES algorithm to detect rigid substance clusters in the metabolic process. Further, the importance of substances is analyzed using the dynamic topological properties of complex networks to determine the core substances and core substance groups in the metabolic process.
[0069] Assume that in the time series data of traditional Chinese medicine metabolism collected, the time series data of any two substances u and v are recorded as two time series of length T and where each parameter represents the relative content of the corresponding substance at each moment within a time period of length T. This method includes:
[0070] S1: Calculate the causal relationship scores between pairwise substances in traditional Chinese medicine metabolism based on the CCM algorithm for local causal relationship analysis.
[0071] For two time series of length T, and The specific algorithm steps are as follows:
[0072] Step 1: Respectively establish shadow manifolds for and After performing E-dimensional lag embedding on the time series of substance u, we get where τ represents time lag. The set composed of these vectors is denoted as the shadow manifold is obtained in the same way.
[0073] Step 2: Find the E + 1 nearest neighbors at time t by calculating the Euclidean distance t1,...,ti ,...,t E+1 respectively represent (from the nearest to the farthest) x u the time subscript of (t).
[0074] Step 3: Calculate the weight w of the i-th nearest neighbor i , where d i = exp{-dis x u (t), x u (t i )] / dis x u (t), x u (t1)]}.
[0075] Step 4: Using the time subscripts corresponding to the E + 1 nearest neighbors found, through these E + 1 time subscripts corresponding values to calculate the estimated value denoted as
[0076] Step 5: Calculate the Pearson correlation coefficient ρ and between uv , Use c uv to represent the influence of cause u on result v, that is, the unidirectional causal score from substance u to substance v, represents the reciprocal causal score between substances u and v.
[0077] Combined with the data characteristics, τ is taken as 1, and cross-validation is used to optimize E. Finally, E = 7 is selected to calculate the causal relationships among 76 substances in the metabolic process. The results are as Figure 1 shown. Respectively rank c uv and , and obtain the top 10 substance pairs in terms of scores, as shown in Table 1. These substance pairs have a strong causal relationship in the metabolic process.
[0078] Table 1: The top 10 substance pairs of c uv and
[0079]
[0080] S2: Construct a causal network and conduct a global causal relationship analysis.
[0081] Using substances as nodes and the causal relationship score c uv As the weight of the directed edge from node u to v, a weighted directed causal network is constructed. Topological analysis is performed on the causal network. At this time, the out-degree of the nodes in the network In-degree respectively represent the influence intensity of substances as causes and results in the metabolic system; the degree centrality c of the nodes u = c u: + c :u is used as an importance index of substances.
[0082] Taking 76 substances as nodes, the causal relationships between pairwise substances as edges, and the causal relationship scores as the weights of the edges, a causal relationship network is constructed, as shown in Figure 2 . The top 10 substances in terms of c u: , c :u and c u are respectively called causative substance, effective substance, and CE substance, as shown in Table 2.
[0083] Table 2: The top 10 substances in terms of scores
[0084]
[0085] S3: Using the sliding window technique combined with the weighted Pearson correlation coefficient to calculate the correlation between substance-substance, and constructing a dynamic substance-substance correlation network.
[0086] Step 1: Using the sliding window technique, a time series of length T is segmented into time windows of length W, and a total of H time windows are obtained;
[0087] Step 2: In each time window, the correlation ρ u,v between substance-substance is calculated using the weighted Pearson correlation coefficient;
[0088] Step 3: Taking substances as nodes and the ρ u,v in Step 2 as the weight of the edge between nodes u and v, a weighted undirected substance-substance correlation network is constructed, and each network is used as a network snapshot;
[0089] Step 4: As the time window slides backward, a series of network snapshots on the time axis are obtained, and they jointly constitute a dynamic substance-substance correlation network.
[0090] For a time series of length 721, the Hamming sliding window method is used to segment the time series, and the weights of the Hamming window are defined as follows:
[0091]
[0092] Set the window length W = 7 and the step size to 1, and a total of H = 715 time windows are obtained.
[0093] S4: Cluster the dynamic network obtained in S4 based on the PisCES algorithm to obtain rigid substance clusters, and analyze the substances that always maintain a close relationship during the metabolic process.
[0094] Mine the dynamic relationships between substances during the metabolic process, and transform the problem into identifying communities in the dynamic network, that is, dynamic network clustering. Use the PisCES algorithm to mine the rigid substance clusters during the metabolic process, including the following steps:
[0095] Step 1: Determine the optimal number of clusters K in the network snapshot h h ;
[0096] Step 2: Calculate the eigenvectors of the Laplacian matrix L of the network snapshot h h and obtain U h = V h (V h )', V h represents the matrix composed of the first K h eigenvectors of L h (V h )'V h = I, where I represents the identity matrix;
[0097] Step 3: Perform Laplacian smoothing on U h through iteration, represents the number of iterations, and α is the smoothing coefficient;
[0098] Step 4: Perform k-means clustering on the Laplacian-transformed matrix ;
[0099] Step 5: Use the five-fold cross-validation method to measure the reliability of the clustering results with modularity as the evaluation index, and find the optimal smoothing coefficient α;
[0100] Step 6: Set α to the optimal value obtained in Step 5, repeat Steps 1 - 4 100 times, and count the number of times that two nodes are assigned to the same class in the 100 results represents the probability that nodes u and v belong to the same class.
[0101] Step 7: Use as the adjacency matrix of the network, perform k-means clustering on the network, and the clustering results are used as the final clustering results.
[0102] Optimize α within the range of [0, 0.13] using the cross-validation method, and finally determine the optimal smoothing coefficient α = 0.09. The dynamic evolution of the substance clusters is as Figure 3 shown. Nine rigid substance clusters rSC were found during the metabolic process. Substances within the same rigid substance cluster do not change at all from s1 to s 701 at all, which indicates that substances within the same rigid substance cluster always maintain a close correlation during the metabolic process. And at the initial stage of metabolism, the relationships between substances change complexly, and the substance clusters change violently, which indicates that the interactions between substances are constantly occurring during this period. And from s 301 onwards, the substance clusters gradually tend to a stable state, and each substance tends to reach a stable state.
[0103] It should be noted that in this embodiment, for the process of treating ischemic stroke with Xiangdan injection, substances within the same rigid substance cluster do not change at all from s1 to s 701 at all, which is considered to always maintain a close correlation during the metabolic process, that is, it is regarded as a rigid substance cluster.
[0104] Figure 3 In, the abscissa represents the final clustering results of a series of network snapshots on the time axis.
[0105] S5: Calculate the three dynamic network centrality indicators of Temporal Degree, Temporal Closeness, and Temporal Betweenness of each node in the dynamic network obtained in S4, which are used to measure the importance of each substance during the metabolic process.
[0106] Temporal Degree, Temporal Closeness, and Temporal Betweenness are respectively improved from the three centrality indicators of Degree, Closeness, and Betweenness in the static network.
[0107] (1) Temporal Degree (TemD)
[0108]
[0109] where D h (v) represents the degree of node u in the h-th network snapshot.
[0110] (2) Temporal Closeness (TemC)
[0111]
[0112] Δ [h,H](v, u) represents the Temporal shortest path distance from v to u in the network snapshot [h, H].
[0113] (3) Temporal Betweenness (TemB)
[0114]
[0115] where g [h,H] (q, u) represents the number of Temporal shortest paths from q to u on the network snapshot [h, H], and g [h,H] (q, u, v) represents the number of Temporal shortest paths from q to u passing through v.
[0116] Calculate the three dynamic centrality metrics for 76 nodes. The top 10 substances and their scores are shown in Table 3.
[0117] Table 3: Nodes and scores ranked top 10 in TemD, TemC, and TemB
[0118]
[0119] In this embodiment, it is considered that the top 10 substances with the scores shown in Table 3 are the core substances and core substance groups in the process of treating ischemic stroke with Xiangdan injection.
[0120] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining core substances and core substance groups during the metabolism of traditional Chinese medicine, characterized in that, The method includes: Step 1, collecting time series data of traditional Chinese medicine metabolism; Step 2, for the collected time series data of traditional Chinese medicine metabolism, using the CCM algorithm to obtain the local and global causal relationships among various substances in the traditional Chinese medicine metabolism process; Step 3, according to the obtained local and global causal relationships among various substances, constructing a dynamic substance-substance correlation network by combining the sliding window technique with weighted Pearson correlation; Step 4, using the PisCES algorithm to detect rigid substance clusters in the metabolism process; the rigid substance clusters refer to substances that always maintain a close relationship in the metabolism process; Step 5, analyzing the importance of various substances in the traditional Chinese medicine metabolism process according to the dynamic topological attributes of the constructed dynamic substance-substance correlation network, so as to determine the core substances and core substance groups in the metabolism process; The said Step 2 includes: Step 2.1, respectively for and establish a shadow manifold. After performing E-dimensional lag embedding on the time series data of substance u, we obtain where τ represents the time lag; the set composed of these vectors is denoted as the shadow manifold obtained in the same way; Step 2.2, by calculating the Euclidean distance find the E+1 nearest neighbors at time t, t1,..., t i ,..., t E+1 respectively represent the time subscripts of x u (t), with the distances from near to far; Step 2.3, calculate the weight w of the i-th nearest neighbor i , where d i = exp{-dis[x u (t), x u (t i )] / dis[x u (t), x u (t1)]}; Step 2.4, using the time subscripts corresponding to the E + 1 nearest neighbors found, and corresponding to these E + 1 time subscripts calculate the estimated value using the values denoted as Step 2.5, calculate and the Pearson correlation coefficient ρ uv : Use c uv represents the influence of cause u on result v, that is, the unidirectional causal score from substance u to substance v: That is, obtaining the local causal relationships among various substances in the traditional Chinese medicine metabolism process; Step 2.6, using each substance as a node and the causal relationship score c uv as the weight of the directed edge pointing from node u to v, construct a weighted directed causal network; perform topological analysis on the causal network, and use the out-degree in-degree of the nodes in the causal network to represent the influence intensity of the substance as a cause and an effect in the metabolic system respectively; the degree centrality c u = c u: + c :u is used as an importance index of the substance, that is, to obtain the global causal relationship among substances in the traditional Chinese medicine metabolism process.
2. The method according to claim 1, wherein The time series data of traditional Chinese medicine metabolism collected in the step 1 are the relative contents of various substances generated during the traditional Chinese medicine metabolism process. For any two substances u and v, their time series data are recorded as two time series with a length of T. and where each parameter represents the relative content of the corresponding substance at each moment within the time period with a length of T.
3. The method according to claim 2, wherein The said Step 3 includes: Step 3.1: Using the sliding window technique to divide the time series of length T into H time windows of length W; Step 3.2: Calculate the substance-substance correlation ρ using the weighted Pearson correlation coefficient within each time window u,v ; Step 3.3: Using substances as nodes, construct a weighted undirected substance-substance correlation network with the substance-substance correlation ρ in Step 3.2 as the weight of the edge between nodes u and v, and each network is used as a network snapshot; u,v Step 3.4: As the time window slides backward, a series of network snapshots on the time axis are obtained, which together constitute a dynamic substance-substance correlation network.
4. The method according to claim 3, characterized in that, The said Step 4 includes: Step 4.1: For any network snapshot h, determine its optimal number of clusters K h ; Step 4.2: Calculate the Laplacian matrix L of the network snapshot h h for its eigenvectors, and obtain U h = V h (V h )', where V h represents the matrix composed of the first K h eigenvectors of L h (V h )'V h = I, and I represents the identity matrix; Step 4.3: Perform Laplace smoothing on U h where denotes the number of iterations and α is the smoothing coefficient; Step 4.4: Perform k-means clustering on the Laplacian-transformed matrix ; Step 4.5: Using the five-fold cross-validation method to measure the reliability of the clustering results with modularity as the evaluation index, and finding the optimal smoothing coefficient α; Step 4.6: Set α to the optimal value obtained in Step 4.5, repeat Steps 4.1 - 4.4 for N times, and count the number of times that the two nodes are assigned to the same class in the N results represents the probability that nodes u and v belong to the same class; Step 4.7: Taking as the adjacency matrix to obtain a new network, performing k-means clustering on the new network, and taking the clustering result as the final clustering result.
5. The method according to claim 4, characterized in that, The said Step 5 includes: Calculating three centrality indexes of Temporal Degree, Temporal Closeness, and Temporal Betweenness in the dynamic substance-substance correlation network: (1) Temporal Degree, TemD Among which D h (v) represents the degree of node u in the h-th network snapshot; (2) Temporal Closeness, TemC Δ [h,H] (v, u) represents the Temporal shortest path distance from v to u in the network snapshot [h, H]; (3) Temporal Betweenness, TemB where g [h,H] (q, u) represents the number of Temporal shortest paths from q to u on the network snapshot [h, H], and g [h,H] (q, u, v) represents the number of Temporal shortest paths from q to u passing through v.
6. The method according to claim 4, wherein In the said Step 4.3, the smoothing coefficient α < 0.
13.
7. The method according to claim 4, characterized in that, The method further includes normalizing TemD and TemC by dividing them by H(|V| - 1), and normalizing TemB by dividing it by H(|V| - 2)(|V| - 1) / 2.
8. Application of the method for determining the core substances and core substance groups in the traditional Chinese medicine metabolism process according to any one of claims 1-7 in the field of computer-aided drug design.
Citation Information
Patent Citations
Method for evaluating influences of drugs on inter-module relations in biomolecule network
CN106709231A
Method for the preparation of biosynthetic device and their uses in diagnostics
CN109074422A