Traditional Chinese medicine direct target spot prediction method and system based on gene regulatory network propagation

By constructing a high-throughput experimental database of traditional Chinese medicine and tissue-specific protein-gene regulation network, combining pseudo-stable state hypothesis calculation and network transmission model, the direct target of traditional Chinese medicine is identified, and the problems of high noise rate and database differences in traditional Chinese medicine target research are solved, and more accurate target prediction is achieved.

CN120340595APending Publication Date: 2025-07-18INSTITUTE OF CHINESE MATERIA MEDICA CHINA ACADEMY OF CHINESE MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510402695.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The research on traditional Chinese medicine targets in the prior art has high noise rates and large database differences, making it difficult to accurately identify the direct targets of traditional Chinese medicine. Traditional methods rely on chemical composition information to lead to research limitations.

Method used

Using a method based on gene regulation network transmission, a high-throughput experimental database of traditional Chinese medicine was constructed, and a tissue-specific protein-gene regulation network was established. Based on pseudo-stable state hypothesis, an iterative calculation was used to identify the direct target of traditional Chinese medicine.

Benefits of technology

The direct targets of traditional Chinese medicine can be accurately identified without prior chemical structural information, broadening the scope of research on the mechanism of action of traditional Chinese medicine and improving the reliability and application value of research results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340595A_ABST
    Figure CN120340595A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine direct target prediction method and system based on gene regulatory network propagation, and relates to the technical field of traditional Chinese medicine research, first, a traditional Chinese medicine high-throughput experimental database is constructed, then a tissue-specific protein-gene regulatory network is established, then unexplained gene expression disturbance is calculated based on pseudo-homeostasis hypothesis, and the traditional Chinese medicine direct target prediction method and system based on gene regulatory network propagation are obtained. And then iterative calculation is carried out by adopting a network propagation model, and finally, a prediction result is obtained by sequencing the protein target spot influence score vectors. The method breaks through the traditional research limitation depending on traditional Chinese medicine chemical components, predicts the target spot by using traditional Chinese medicine induced gene expression change in combination with a multilevel biological network and a mathematical model, does not need priori chemical structure information, widens the traditional Chinese medicine action mechanism research range, and effectively solves the problems of high noise rate, large database difference and the like in the existing target spot research. Direct targets of the traditional Chinese medicine can be identified more accurately, the reliability and application value of research results are improved, and the method is of great significance to traditional Chinese medicine research which is difficult to analyze by a traditional method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traditional Chinese medicine research, and more specifically to a method and system for predicting direct targets of traditional Chinese medicine based on gene regulatory network propagation. Background Art

[0002] Currently, how to overcome the complexity of traditional Chinese medicine chemical components and reveal its multi-scale mechanisms of treating diseases is the core issue in the current research field of direct targets of traditional Chinese medicine. Currently, network pharmacology is considered an effective tool for elaborating the mechanisms of traditional Chinese medicine. This research method, namely the network construction method of "traditional Chinese medicine - chemical components - action targets", relies on the chemical components of traditional Chinese medicine. However, the targets of traditional Chinese medicine obtained based on component associations often contain a large amount of noise, resulting in a high noise rate. At the same time, for the component-based target prediction method, there are problems such as a huge difference in databases and serious homogenization, including quercetin, etc. In addition, the diversity of components contained in traditional Chinese medicine, processed products, and their prescriptions, as well as the complexity of the interaction process in the human body, bring difficulties and challenges to revealing the direct target prediction of traditional Chinese medicine.

[0003] With the emergence of high-throughput sequencing and the development of systems biology, the research on drug target prediction algorithms has been promoted by integrating multi-modal data and advanced computational methods. Differential gene analysis is currently the most commonly used method to elaborate the mechanisms of traditional Chinese medicine from the gene expression level. In addition, some studies have begun to consider correlation analysis based on expression profiles to predict targets, using machine learning algorithms such as pattern matching to construct complex relationships between drug interventions and gene expression changes. Such algorithms can reveal the biological associations between interfering molecules and gene expression by quantifying the similarities between different interfering entities, and have the potential to discover new drug targets and drug repositioning. Recently, it has also been used in analyzing the complex systems of traditional Chinese medicine, including KORE-Map, ITCM, and Herb-CMap, etc. However, this method can only be used for drugs with comprehensive reference gene expression profiles by comparing the similarities between drug gene expression profiles, and is not suitable for finding de novo targets. Therefore, it is particularly necessary to construct a direct target prediction framework that combines high-throughput gene expression profile data and is interpretable without relying on component information. This method is expected to break through the limitations of existing analytical models and provide important theoretical support and application basis for the precise mechanism research and multi-target prediction of traditional Chinese medicine.

[0004] In addition, the prediction framework based on molecular phenotype data can better reflect the actual effects of traditional Chinese medicine in biological systems and overcome the deficiencies of traditional component-based methods in reflecting the synergistic effects of multiple components and dose compatibility relationships. Through a systematic method that gets rid of the limitation of component information, the new prediction framework will be able to more accurately capture the complex action mechanisms of traditional Chinese medicine compounds and improve the reliability and application value of research results.

[0005] In view of the above-mentioned practical difficulties in traditional Chinese medicine (TCM) target research and the technical omissions in target prediction research, how to provide a method to utilize the extensive gene expression changes induced by TCM, combined with multi-level biological networks and mathematical models, to identify the direct targets of TCM, so as to achieve target prediction without prior chemical structure information is an urgent problem for those skilled in the art to solve. Summary of the Invention

[0006] In view of this, the present invention provides a method and system for predicting direct targets of traditional Chinese medicine based on gene regulatory network propagation, adopting a novel multi-level network fusion framework for predicting TCM targets without components. By integrating high-throughput gene expression profile data related to TCM and artificial intelligence algorithms, a new approach for predicting the mechanism of action of TCM is provided. Different from traditional methods based on TCM chemical components, the framework of the present invention utilizes the extensive gene expression changes induced by TCM, combined with multi-level biological networks and mathematical models, to identify the direct targets of TCM, thus achieving target prediction without prior chemical structure information. This not only broadens the scope of research on predicting the mechanism of action of available TCM, but is also particularly important for TCM that is difficult to analyze by traditional methods.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for predicting direct targets of traditional Chinese medicine based on gene regulatory network propagation, comprising:

[0009] Constructing a high-throughput experimental database of traditional Chinese medicine;

[0010] According to the high-throughput experimental database of traditional Chinese medicine, establishing a tissue-specific protein-gene regulatory network;

[0011] Calculating the unexplained gene expression perturbation based on the pseudo-steady state hypothesis;

[0012] Adopting a network propagation model, and propagating the unexplained gene expression perturbation in the tissue-specific protein-gene regulatory network through iterative calculation to obtain a protein target influence score vector;

[0013] Sorting the protein target influence score vector in descending order to obtain the sorting result of the predicted direct targets of traditional Chinese medicine.

[0014] Optionally, establishing a tissue-specific protein-gene regulatory network specifically includes:

[0015] Edges in tissue-specific protein-gene regulatory networks describe the regulation of gene expression by transcription factors and their associated proteins. Protein-gene regulatory networks are constructed by combining two types of networks, namely transcription factor-gene and protein-protein interactions. For each transcription factor, its associated proteins are identified in the protein interaction network, defined as proteins with a network distance of no more than 2 from the transcription factor in the network. Through the above method, by combining transcription factor-gene regulatory relationships with protein interaction information, tissue-specific protein-gene regulatory networks are constructed.

[0016] Optionally, it is characterized in that calculating the unexplained gene expression perturbation based on the pseudo-steady state hypothesis is specifically:

[0017] There is a known regulatory network matrix A, where A ji represents the regulatory weight of gene j on gene i; the expression change amount ΔE of gene i i is expressed as the sum of its own regulatory factors and regulatory influences from upstream genes; based on the ODE model of the gene transcription process, the specific calculation formula is as follows:

[0018]

[0019] Among them, G i refers to the ratio between the normal transcription rate constant and the degradation rate constant of mRNA, and E i represents the mRNA level of a certain gene in the sample; E j represents the mRNA level of a certain direct or indirect regulatory protein j that regulates gene i, and n represents the total number of gene expressions detected in the sample; when processing gene expression data, the gene expression level of the treatment group is usually processed as the ratio relative to the control group, that is, the logFC value;

[0020] The calculation formula of the ODE model based on the gene transcription process is simplified and transformed to obtain the unexplained gene expression perturbation.

[0021] Optionally, the simplification and transformation process is:

[0022] Divide both sides of equation (1) by the mRNA level of the control group to obtain the model of gene expression ratio:

[0023]

[0024] Among them, E ci represents the mRNA level of a certain gene in the control group sample, and G ci is the ratio between the normal transcription rate constant and the degradation rate constant of a certain gene mRNA in the control sample, and numerically is the same as G i ;

[0025] The model construction in Equation (2) implies an assumption that drug treatment only affects the mRNA transcription and / or degradation rate constants without causing any changes in the gene regulatory network; taking the logarithm of both sides of Equation (2) gives the following linear expression, simplifying the above model to:

[0026]

[0027] where P i represents the differential part of the change in gene i expression explained by the known regulatory network A, ΔE i represents the change in gene i expression, and A ji represents the regulatory weight of gene j on gene i in the regulatory network; transforming Equation (3) into matrix form gives the unexplained gene expression perturbation.

[0028] Optionally, the gene expression perturbation P unexplained by the regulatory network is expressed as:

[0029] P = (I - A)ΔE (4)

[0030] where I is an n×n identity matrix, ΔE is defined as the vector of gene expression changes [ΔE1, ΔE2, …, ΔE i T , A is defined as the protein-gene biological regulatory network matrix, and P is defined as the perturbation vector [P1, P2, …, P i T .

[0031] Optionally, the traditional Chinese medicine high-throughput experimental database includes traditional Chinese medicine high-throughput experimental data, tissue-specific protein-DNA / protein-protein interaction databases, and tissue-derived gene expression data.

[0032] Optionally, the core of the network propagation model is to propagate perturbation information from the gene level to the protein target level through the structure of the tissue-specific protein-gene regulatory network; the specific calculation formula is as follows:

[0033] S = (I - αA T ) -1 P (5)

[0034] where S represents the vector of influence scores of protein targets, reflecting the contribution degree of each protein to the gene expression change; I is the identity matrix used to maintain the basic structure of the system; α is the attenuation coefficient used to control the range of information propagation; A T is the transpose matrix of the regulatory network A used to implement the reverse propagation calculation of information; P is the calculated unexplained perturbation vector.

[0035] ​​A traditional Chinese medicine direct target prediction system based on gene regulatory network propagation, comprising:

[0036] A database construction module for constructing a high-throughput experiment database of traditional Chinese medicine;

[0037] A network modeling module for establishing a tissue-specific protein-gene regulatory network according to the high-throughput experiment database of traditional Chinese medicine;

[0038] A gene expression data analysis module for calculating the unexplained gene expression perturbation based on the pseudo-steady state hypothesis;

[0039] An iterative calculation module for propagating the unexplained gene expression perturbation in the tissue-specific protein-gene regulatory network by iterative calculation using a network propagation model to obtain a protein target influence score vector;

[0040] A result output module for sorting the protein target influence score vector in descending order to obtain the sorting result of the predicted direct targets of traditional Chinese medicine.

[0041] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a method and system for predicting direct targets of traditional Chinese medicine based on gene regulatory network propagation. First, a high-throughput experiment database of traditional Chinese medicine is constructed, then a tissue-specific protein-gene regulatory network is established, then the unexplained gene expression perturbation is calculated based on the pseudo-steady state hypothesis, then iterative calculation is performed using a network propagation model, and finally the protein target influence score vector is sorted to obtain the prediction result. The present invention breaks through the research limitation of traditional reliance on the chemical components of traditional Chinese medicine, uses the gene expression changes induced by traditional Chinese medicine in combination with multi-level biological networks and mathematical models to predict targets, without prior chemical structure information, broadens the research scope of the action mechanism of traditional Chinese medicine, effectively solves problems such as high noise rate and large database differences in existing target research, can more accurately identify the direct targets of traditional Chinese medicine, improves the reliability and application value of research results, and is of great significance for the research of traditional Chinese medicine that is difficult to analyze by traditional methods. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0043] Figure 1 It is a schematic flow chart of the method provided by the present invention;

[0044] Figure 2 It is a comparison of the AUROC values of the IFTP-TCM and DE methods (logfc) under different parameters provided by the present invention. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] An embodiment of the present invention discloses a method for predicting direct targets of traditional Chinese medicine based on gene regulatory network propagation, as Figure 1 shown, including:

[0047] Construct a high-throughput experimental database of traditional Chinese medicine;

[0048] According to the high-throughput experimental database of traditional Chinese medicine, establish a tissue-specific protein-gene regulatory network;

[0049] Based on the pseudo-steady state hypothesis, calculate the unexplained gene expression perturbation;

[0050] Adopt a network propagation model, and through iterative calculation, propagate the unexplained gene expression perturbation in the tissue-specific protein-gene regulatory network to obtain a protein target impact score vector;

[0051] Arrange the protein target impact score vector in descending order to obtain the ranking result of the predicted direct targets of traditional Chinese medicine.

[0052] In a specific embodiment, establishing a tissue-specific protein-gene regulatory network specifically includes:

[0053] The edges in the tissue-specific protein-gene regulatory network describe the regulation of gene expression by transcription factors and their associated proteins. The protein-gene regulatory network is constructed by combining two types of networks, namely transcription factor-gene and protein-protein interaction; for each transcription factor, identify its associated proteins in the protein interaction network, which are defined as proteins with a network distance not exceeding 2 from the transcription factor in the network; through the above method, combine the transcription factor-gene regulatory relationship and the protein interaction information to construct a tissue-specific protein-gene regulatory network.

[0054] In a specific embodiment, in the gene expression regulatory network, the expression change of a gene can be explained by the expression change of its upstream regulatory factors and the connection weights in the regulatory network. Based on the pseudo-steady state hypothesis, the gene expression change amount ΔE i can be explained through the regulatory relationship between genes. Assume that there is a known regulatory network matrix A, where A jiDenote the regulatory weight of gene j on gene i; the expression change of gene i, ΔE i Is expressed as the sum of its own regulatory factors and the regulatory influences from upstream genes; the method (IFTP-TCM) in the present invention is based on the ODE model of the gene transcription process, and the specific calculation formula is as follows:

[0055]

[0056] Where, G i Refers to the ratio between the normal rate of mRNA transcription and the degradation rate constant, E i Represents the mRNA level of a certain gene in the sample; E j Represents the mRNA level of a certain direct or indirect regulatory protein j that regulates gene i, n represents the total number of gene expressions in the detected samples; when processing gene expression data, usually the gene expression level of the treatment group is processed as the ratio relative to the control group, that is, the logFC value; divide both sides of equation (1) by the mRNA level of the control group to obtain the model of gene expression ratio:

[0057]

[0058] Where, E ci Represents the mRNA level of a certain gene in the control group sample, G ci That is, the ratio between the normal rate of mRNA transcription and the degradation rate constant of a certain gene in the control sample, and numerically the same as G i The same;

[0059] The model construction in equation (2) implies an assumption that drug treatment only affects the mRNA transcription and / or degradation rate constant, and will not cause any changes in the gene regulatory network; take the logarithm of both sides of equation (2) to obtain the following linear expression, and simplify the above model to:

[0060]

[0061] Where P i Represents the differential part of the expression change of gene i explained by the known regulatory network A, ΔE i Represents the expression change of gene i (logFC), A ji Represents the regulatory weight of gene j on gene i in the regulatory network; transform equation (3) into matrix form, and the gene expression perturbation P not explained by the regulatory network can be expressed as:

[0062] P = (I - A)ΔE (4)

[0063] Where, I is the n×n identity matrix, and ΔE is defined as the vector of gene expression changes [ΔE1, ΔE2, …, ΔE i T ​, A is defined as the biological regulatory network matrix of protein - gene, and P is defined as the perturbation vector [P1, P2, …, P i T . In some studies, this value is used to identify true drug - perturbed genes. However, in fact, only perturbed genes are obtained here, rather than true protein targets. Therefore, in this embodiment, a network diffusion model is further introduced to identify upstream protein targets.

[0064] Calculating the contribution of upstream protein targets based on the network propagation model

[0065] Overview of the network propagation model

[0066] Using the network propagation model, the unexplained perturbation information is propagated in the regulatory network through iterative calculation, so as to quantify the impact of each upstream protein on gene expression changes. After identifying the unexplained perturbation P, the next step is to evaluate the contribution of upstream protein targets to these perturbations. The basic idea based on network propagation is to assume that the unexplained perturbation P needs to be propagated through the regulatory network A to quantify the contribution of each upstream protein to the perturbation. The network propagation model passes the perturbation information to each node (protein target) in the network through iterative calculation.

[0067] Mathematical model

[0068] The core of the network propagation model is to propagate the perturbation information from the gene level to the protein target level through the structure of the regulatory network. The specific calculation formula is as follows:

[0069] S = (I - αA T ) -1 P (5)

[0070] where S represents the influence score vector of protein targets, reflecting the contribution degree of each protein to gene expression changes; I is the identity matrix, used to maintain the basic structure of the system; α is the attenuation coefficient, used to control the range of information propagation; A T is the transposed matrix of the regulatory network A, used to implement the reverse propagation calculation of information; P is the calculated unexplained perturbation vector.

[0071] In a specific embodiment, constructing the tissue - specific gene regulatory network is specifically as follows:

[0072] ​Using the Bayesian evidence integration method and the Naive Bayes classifier, various gene expression data evidences can be effectively integrated to calculate the posterior probability of gene interaction, and then a tissue-specific gene interaction network can be constructed to improve the accuracy of gene interaction prediction. Specifically, the Bayesian network integrates various construction methods based on gene expression data, including information theory methods (ARACNe), machine learning techniques (GENIE3), and hybrid and integrated methods (PANDA). In addition, when conducting drug perturbation studies in A293 cells, a gene dataset related to lung tissue was selected to construct the gene regulatory network. Through this method, the original network of each tissue was refined, enabling the network analysis results to more accurately reflect the gene interactions in the tissue or cells under study. The original protein-protein interaction network was constructed from the STRING database, and only those protein-protein interactions verified by experiments or confirmed from high-quality curated databases were selected. To further improve the accuracy and reliability of the network, only PPI interactions with scores exceeding 700 were retained, and these high-confidence interactions were converted into bidirectional relationships.

[0073] The edges in the protein-gene regulatory network describe the regulation of gene expression by TFs and their associated proteins, which are the molecular targets of interest in the present invention. The protein-gene regulatory network is constructed by combining two types of networks, namely TF-Gene and PPI. Specifically, for each transcription factor, its associated proteins are identified in the PIN, defined as the proteins with a network distance of no more than 2 from the TF in the network. By this method, combining the TF-gene regulatory relationship and the protein-protein interaction information, a protein-gene regulatory network that comprehensively reflects the transcriptional regulatory mechanism is constructed.

[0074] Calculate the scores of upstream direct protein targets based on the quasi-steady state assumption and the network propagation model

[0075] In biological systems, the gene expression regulatory network is complex, and multiple genes interact with each other through protein-protein interactions and regulatory relationships. To simplify the model, the quasi-steady state assumption is introduced, that is, it is assumed that the system is in an approximately steady state. At this time, in the gene expression regulatory network, the expression changes of genes can be explained by the expression changes of their upstream regulatory factors and the way of regulating the regulatory relationships between genes in the weighted connection network. IFTP-TCM is based on the ODE model of the gene transcription process. In some studies, this value is used to identify the true drug perturbation genes, but here only the perturbation genes are obtained, rather than the true protein targets. Therefore, in this embodiment, the network diffusion model is further introduced to identify the upstream protein targets.

[0076] The network propagation model is used to spread the unexplained perturbation information in the regulatory network through iterative calculation, so as to quantify the influence of each upstream protein on gene expression changes. The basic idea based on network propagation is that the unexplained perturbed genes need to be spread through the gene regulatory network to quantify the contribution of each upstream protein to the perturbation. The network propagation model gradually transfers the perturbation information to the upstream targets (protein-level targets) in the network through iterative calculation. After calculating the scores of protein targets through the diffusion model and sorting them, the ranking results of the direct targets predicted by traditional Chinese medicine are obtained.

[0077] Verify IFTP-TCM using real datasets of traditional Chinese medicine and traditional Chinese medicine small molecules:

[0078] First, the accuracy of the direct targets of the holistic effects of traditional Chinese medicine inferred by IFTP-TCM on 10 kinds of THM qi-tonifying traditional Chinese medicines was evaluated using the perturbation dataset from Korea-map. The transcript expression information was derived from the water and ethanol extracts of THM and herbs, prepared at three different concentrations, and applied to four representative human cell lines (A549, HepG2, HT29, and SW1783). To ensure the specificity of network organization, gene expression data were systematically collected from the GTEx and TCGA databases in this example. For example, when conducting drug perturbation studies in A293 cells, the lung tissue datasets in the GTEx database and the TCGA database were selected to construct the gene regulatory network.

[0079] The AUROC score of the prediction results of each traditional Chinese medicine was measured using the protein score ranking list predicted by IFTP-TCM. IFTP-TCM was compared with the DE analysis method based on the Deseq2 package, which is most commonly used for high-throughput data, and the protein target prediction ranking list of each traditional Chinese medicine was compared with the reference traditional Chinese medicine targets constructed from traditional Chinese medicine-ingredient-targets in the HIT2.0 database. For the complete area (AUC) under the receiver operating characteristic (ROC) curve, IFTP-TCM was always superior to DE analysis, as Figure 2As shown, the AUROC of target prediction in different parameters of IFTP-TCM and DE analysis (deseq2) is summarized, showing that IFTP-TCM is significantly better than DE analysis in all four datasets from different tissue sources. The drug target prediction from DE analysis has the worst AUROC, and the average AUROC scores of action prediction results in different tissues are 0.521 (AUROC range in lung: 0.400 - 0.679), 0.500 (AUROC range in liver: 0.288 - 0.600), 0.513 (AUROC range in colon: 0.295 - 0.630), and 0.543 (AUROC range in brain: 0.322 - 0.687). Meanwhile, the target predictions of different parameters of IFTP-TCM are all better than DE analysis, and the average AUROC scores of the four datasets from different tissue sources are 0.638 (AUROC range in lung: 0.638 - 0.751), 0.677 (AUROC range in liver: 0.519 - 0.914), 0.671 (AUROC range in colon: 0.482 - 0.780), and 0.682 (AUROC range in brain: 0.458 - 0.877).

[0080] Although IFTP-TCM aims to predict the direct targets of the holistic effects of traditional Chinese medicine (without component information), its performance can only be systematically benchmarked with currently commonly used traditional Chinese medicine-component-target databases such as HIT2.0, etc., because there is currently no gold standard dataset for systematically evaluating the direct targets of traditional Chinese medicine. Therefore, to further verify this method, it was applied to the perturbation experiment of traditional Chinese medicine small molecules with known targets, and the experimental data came from the unified high-throughput experimental platform ITCM based on the active components of traditional Chinese medicine. It can be seen that after applying this method to rank the predicted targets of small molecules, Table 1 shows the analysis results, which predict the true traditional Chinese medicine small molecule targets, including 17 targets with rankings less than 50. Overall, these results support that this method has good performance in identifying the true direct targets of traditional Chinese medicine.

[0081] Table 1 Prediction of direct targets of traditional Chinese medicine small molecules with known targets based on IFTP-TCM

[0082]

[0083]

[0084] Compared with other existing technologies, this method has the following advantages:

[0085] 1. Compared with the existing computational methods that rely on the fold change of gene expression, such as differential gene analysis, to screen key genes, this method can transfer the changes at the gene expression level to the upstream protein level through network diffusion by combining the biological regulatory network, identify the direct targets of traditional Chinese medicine, and the prediction effect is more ideal.

[0086] 2. Compared with other methods for predicting traditional Chinese medicine targets, this method does not require the component information of traditional Chinese medicine, avoiding the difficulties in target research brought by the complex component information of traditional Chinese medicine. At the same time, it can reflect the overall effect of traditional Chinese medicine by means of gene expression profile information.

[0087] A traditional Chinese medicine direct target prediction system based on gene regulatory network propagation, including:

[0088] Database construction module, constructing a traditional Chinese medicine high-throughput experiment database;

[0089] Network modeling module, establishing a tissue-specific protein-gene regulatory network according to the traditional Chinese medicine high-throughput experiment database;

[0090] Gene expression data analysis module, calculating the unexplained gene expression perturbation based on the pseudo-steady state hypothesis;

[0091] Iterative calculation module, using the network propagation model, propagating the unexplained gene expression perturbation in the tissue-specific protein-gene regulatory network through iterative calculation to obtain a protein target influence score vector;

[0092] Result output module, sorting the protein target influence score vector in descending order to obtain the sorting result of the predicted direct targets of traditional Chinese medicine.

[0093] The specific structure and its operation method are as follows:

[0094] Execution subject: computer system (including the module for constructing tissue-specific protein-gene regulatory network, network modeling module, iterative calculation module)

[0095] Step 1: Construct a tissue-specific protein-gene regulatory network

[0096] Execution subject: network modeling module

[0097] Input data:

[0098] Tissue-specific protein-DNA / protein-protein interaction database

[0099] Gene expression data from tissue sources

[0100] Process description:

[0101] The gene transcription kinetics model is fitted by the ridge regression algorithm to determine the edge weight A of the regulatory network matrix A ji , representing the regulatory weight (intensity) of gene j on gene i (including positive and negative directions).

[0102] Construct a tissue-specific protein-gene regulatory network (PGRN), with nodes being genes / proteins and edge weights being A ji .

[0103] Key role: Quantify the direct / indirect regulatory relationships between genes and provide a topological structure basis for perturbation propagation.

[0104] Step 2: Calculate the unexplained gene expression perturbations

[0105] Execution entity: The gene expression data analysis module related to traditional Chinese medicine

[0106] Mathematical model: A system of linear equations based on the pseudo-steady state assumption

[0107] P = (I - A)ΔE

[0108] Parameter definitions:

[0109] P: The vector of unexplained perturbations (unknown)

[0110] I: Identity matrix

[0111] A: The known regulatory network matrix (A ji is a known quantity)

[0112] ΔE: The vector of gene expression changes, ΔE i = logFC

[0113] Calculation process:

[0114] Solve for P through matrix inversion operation, and screen out the abnormal perturbation genes beyond the network prediction range.

[0115] Key role: Identify the abnormal changes in gene expression caused by drug action and exclude the part that can be explained by the known regulatory network.

[0116] Step 3: Quantify the target contribution based on the network propagation model

[0117] Execution entity: Iterative calculation module

[0118] When the algorithm converges, the equation is satisfied: S = αA T S + P

[0119] Parameter definitions:

[0120] S: The vector of protein target influence score (unknown)

[0121] α: Decay coefficient (α ∈ (0, 1), default 0.9)

[0122] A T : Transpose of the regulatory network matrix

[0123] P: Unexplained perturbation vector

[0124] Iterative algorithm:

[0125] Initialize S = P

[0126] while not converged:

[0127] S_new = α * A^T @ S_old + P

[0128] S_old = S_new

[0129] Key role: Propagate gene-level perturbations to upstream proteins through a feedback loop mechanism, and quantify the contribution of targets to abnormal expression.

[0130] Step 4: Target scoring and ranking

[0131] Execution entity: Result output module of the computer system

[0132] Processing process: Sort the finally converged S vector in descending order to generate a target priority list. Set a threshold to screen significant targets (such as TopK S values).

[0133] Verification mechanism: Verify the overlap degree between predicted targets and known traditional Chinese medicine action targets.

[0134] Innovation: Break through the traditional method that relies on chemical components, and directly infer targets through gene expression perturbations.

[0135] Technical advantages:

[0136] The network propagation model solves the problem of target attribution for indirect regulatory relationships (such as feedback loops).

[0137] Matrix operations and iterative algorithms ensure the computational efficiency of large-scale networks.

[0138] An example of parameter settings is shown in Table 2:

[0139] Table 2 Parameter settings

[0140]

[0141]

[0142] All matrix operations in the formula need to satisfy the dimension compatibility condition (e.g., A is an n×n matrix, ΔE is an n×m vector), where n is the number of genes and m is the number of samples.

[0143] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the identical or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For related parts, reference can be made to the descriptions in the method section.

[0144] The above descriptions of the disclosed embodiments enable those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for predicting direct targets of traditional Chinese medicine based on gene regulatory network propagation, characterized in that Including: Constructing a high-throughput experimental database of traditional Chinese medicine; Establishing a tissue-specific protein-gene regulatory network based on the high-throughput experimental database of traditional Chinese medicine; Calculating the unexplained gene expression perturbation based on the pseudo-steady state hypothesis; Using a network propagation model to propagate the unexplained gene expression perturbation in the tissue-specific protein-gene regulatory network through iterative calculation to obtain a protein target impact score vector; Sorting the protein target impact score vector in descending order to obtain the sorting result of the predicted direct targets of traditional Chinese medicine.

2. The traditional Chinese medicine direct target prediction method based on gene regulatory network propagation according to claim 1, wherein The establishment of the tissue-specific protein-gene regulatory network is specifically as follows: The edges in the tissue-specific protein-gene regulatory network describe the regulation of gene expression by transcription factors and their associated proteins. The protein-gene regulatory network is constructed by combining two types of networks, namely transcription factor-gene and protein-protein interactions; for each transcription factor, its associated proteins are identified in the protein interaction network, which are defined as proteins with a network distance of no more than 2 from the transcription factor in the network; through the above methods, combining the transcription factor-gene regulatory relationship and the protein interaction information, a tissue-specific protein-gene regulatory network is constructed.

3. A traditional Chinese medicine direct target prediction method based on gene regulatory network propagation according to claim 1, characterized in that, The specific calculation of the unexplained gene expression perturbation based on the pseudo-steady state hypothesis is as follows: There is a known regulatory network matrix A, where A ji represents the regulatory weight of gene j on gene i; the change in expression level ΔE of gene i i is expressed as the sum of its own regulatory factors and the regulatory influences from upstream genes; based on the ODE model of the gene transcription process, the specific calculation formula is as follows: Among them, G i refers to the ratio between the normal rate of mRNA transcription and the degradation rate constant, and E i represents the mRNA level of a certain gene in the sample; E j represents the mRNA level of a certain direct or indirect regulatory protein j that regulates gene i, and n represents the total number of gene expressions in the tested samples; when processing gene expression data, the gene expression levels of the treatment group are processed into ratios relative to the control group, that is, logFC values; Simplifying and transforming the calculation formula of the ODE model based on the gene transcription process to obtain the unexplained gene expression perturbation.

4. A traditional Chinese medicine direct target prediction method based on gene regulatory network propagation according to claim 3, characterized in that, The simplification and transformation process is as follows: Dividing both sides of equation (1) by the mRNA level of the control group to obtain a model of gene expression ratio: Among them, E ci represents the mRNA level of a certain gene in the control group samples, and G ci that is, the ratio between the normal speed of the mRNA transcription rate and the degradation rate constant of a certain gene in the control samples, and numerically it is the same as G i ; The model construction in equation (2) implies an assumption that drug treatment only affects the mRNA transcription and / or degradation rate constant, without causing any changes in the gene regulatory network; taking the logarithm of both sides of equation (2) gives the following linear expression, simplifying the above model to: where P i represents the differential part of the expression change of gene i explained by the known regulatory network A, and ΔE i represents the expression change amount of gene i, and A ji represents the regulatory weight of gene j on gene i in the regulatory network; transform equation (3) into matrix form to obtain the unexplained gene expression perturbation.

5. The traditional Chinese medicine direct target prediction method based on gene regulatory network propagation according to claim 4, wherein The gene expression perturbation P unexplained by the regulatory network is expressed as: P = (I - A)ΔE (4) where I is an n×n identity matrix, ΔE is defined as the vector of gene expression change amounts [ΔE1, ΔE2, …, ΔE i T , A is defined as the protein-gene biological regulatory network matrix, and P is defined as the perturbation vector [P1, P2, …, P i T .​​ 6. The traditional Chinese medicine direct target prediction method based on gene regulatory network propagation according to claim 1, wherein The high-throughput experimental database of traditional Chinese medicine includes high-throughput experimental data of traditional Chinese medicine, a tissue-specific protein-DNA / protein-protein interaction database, and gene expression data from tissue sources.

7. A traditional Chinese medicine direct target prediction method based on gene regulatory network propagation according to claim 1, characterized in that, The core of the network propagation model is to propagate the perturbation information from the gene level to the protein target level through the structure of the tissue-specific protein-gene regulatory network; the specific calculation formula is as follows: S = (I - αA T ) -1 P (5) Where S represents the protein target impact score vector, reflecting the contribution degree of each protein to the gene expression change; I is the identity matrix, which is used to maintain the basic structure of the system; α is the attenuation coefficient, which is used to control the range of information propagation; A T is the transposed matrix of the regulatory network A, which is used to implement the reverse propagation calculation of information; P is the calculated unexplained perturbation vector.

8. A traditional Chinese medicine direct target prediction system based on gene regulatory network propagation, characterized in that, Applying a method for predicting direct targets of traditional Chinese medicine based on gene regulatory network propagation according to any one of claims 1-7, including: A database construction module for constructing a high-throughput experimental database of traditional Chinese medicine; A network modeling module for establishing a tissue-specific protein-gene regulatory network based on the high-throughput experimental database of traditional Chinese medicine; A gene expression data analysis module for calculating the unexplained gene expression perturbation based on the pseudo-steady state hypothesis; An iterative calculation module for using a network propagation model to propagate the unexplained gene expression perturbation in the tissue-specific protein-gene regulatory network through iterative calculation to obtain a protein target impact score vector; The result output module sorts the protein target impact score vectors in descending order to obtain the sorting result of the directly predicted traditional Chinese medicine targets.