Virtual screening method for high-importance compounds of traditional Chinese medicine compound
By combining the UNIQ system, KEGG enrichment analysis, disease-related modular biological networks, and CHM-FIEFP software, the problems of low accuracy and efficiency in screening active ingredients of traditional Chinese medicine compound formulas were solved, achieving efficient and accurate compound screening with a significantly improved validation rate.
Patent Information
- Application Number
- CN202511045357.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-14
AI Technical Summary
The screening of active ingredients in traditional Chinese medicine compound formulas suffers from low accuracy, low efficiency, and low experimental feasibility. Existing technologies lack systematic screening methods, making it difficult to accurately assess the contribution of each component and their interactions.
The UNIQ system was used to screen compounds of principal herbs in traditional Chinese medicine compound prescriptions. Combined with KEGG enrichment analysis and disease-related modular biological networks, the QED system was used to evaluate the drug properties of the compounds. The entropy weight method was used to calculate the entropy weights of the compounds using CHM-FIEFP software to screen out compounds of high importance.
It significantly improves the accuracy and efficiency of screening active ingredients in traditional Chinese medicine compound preparations. The screened compounds have high verifiability and universality, and can complete in a short time what would take years to do with traditional methods. The verification rate is higher than that of traditional random screening methods.
Smart Images

Figure BDA0005522302770000071 
Figure BDA0005522302770000072 
Figure BDA0005522302770000073
Abstract
Description
Technical Field
[0001] This invention belongs to the field of pharmacology, and in particular relates to a virtual screening method for highly important compounds in traditional Chinese medicine compound preparations. Background Technology
[0002] Traditional Chinese medicine (TCM) compound formulas are an important component of TCM, and their unique compatibility theories and rich clinical experience provide valuable resources for modern drug development. In recent years, advancements in analytical techniques and computational methods, particularly the development of artificial intelligence and computational models, have provided new technical means for in-depth research into the material basis and mechanisms of action of TCM compound formulas. However, the screening of active ingredients in TCM compound formulas still faces significant challenges. The primary problem is that compound formulas contain a large number of components with varying compositions and concentrations, making the identification of key active ingredients extremely difficult. Secondly, the mechanisms of action of each component in a compound formula are complex, often involving the synergistic effects of multiple molecular targets and signaling pathways. Furthermore, existing technologies lack systematic screening methods, making it difficult to accurately assess the contributions of each component and their interactions, which severely restricts the modernization of TCM compound formula research.
[0003] Currently, three main methods are used to screen active ingredients from traditional Chinese medicine (TCM) compound prescriptions: The first is the experience-based method, which selects active ingredients based on traditional medication experience and analyzes them according to the theory of prescription compatibility. While this method inherits traditional wisdom, it lacks scientific systematicity and is inefficient. The second is the stepwise screening method, which gradually separates the compound into monomeric compounds and then tests the activity of each compound. Although this method is intuitive, it is time-consuming, costly, and prone to overlooking synergistic effects between compounds. The third is the network pharmacology method, which predicts possible mechanisms of action by constructing compound-target networks. However, this method lacks sufficient experimental data support, and the accuracy of the prediction results is limited.
[0004] CN114113364A discloses a predictive method for network pharmacology of traditional Chinese medicine (TCM). Through sample preparation, HPLC fingerprinting, screening of active ingredients and potential targets, network construction, enrichment analysis, and virtual validation, this method can effectively predict the effective targets of TCM components in target diseases, with results consistent with in vivo experiments. However, this method has several key problems: First, it does not consider the in vivo detectability of compounds, potentially leading to the screening of compounds that may not be effectively detected and tracked in vivo; second, the reliability of compound information is difficult to guarantee due to the lack of cross-validation across multiple databases; third, it lacks a scientific quantitative scoring system, making it difficult to objectively evaluate the screening results; and finally, it insufficiently assesses the matching effect between the compound and its components, meaning the screened components cannot fully reflect the overall effect of the compound.
[0005] In view of the shortcomings of existing technologies, there is an urgent need to develop a computer-aided screening method that is efficient and disease-specific in the field of screening active ingredients of traditional Chinese medicine compound prescriptions. Summary of the Invention
[0006] To address the aforementioned technical problems, the present invention aims to provide a virtual screening method for highly important compounds in traditional Chinese medicine compound formulas. This method successfully solves the problems of low accuracy and low experimental feasibility in the screening of active ingredients in traditional Chinese medicine compound formulas, and provides a new technical path for modern research on traditional Chinese medicine compound formulas.
[0007] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0008] A virtual screening method for highly important compounds in traditional Chinese medicine compound preparations includes the following steps:
[0009] S1. Identify the components of traditional Chinese medicine compound formulas to obtain initial compounds, and annotate the initial compounds; use the UNIQ system to screen the compounds of the principal drug in the traditional Chinese medicine compound formulas to obtain high-confidence compounds;
[0010] S2. Use the UNIQ system to perform KEGG enrichment analysis on the high-confidence compounds to evaluate their potential pathways of action; compare the enriched pathways of action with the established disease-related modular biological network to screen out potential regulatory compounds.
[0011] S3. The potential regulatory compounds were evaluated using the QED system, and the compounds with the highest QED scores were selected as high-drug-like compounds.
[0012] S4. The high-potency compounds were evaluated using CHM-FIEFP software. The fit score for each high-potency compound was calculated using the entropy weight method. Compounds with high scores were considered high-importance compounds.
[0013] Preferably, in step S1, a Q-Exactive HFX mass spectrometer is used to identify the components of the traditional Chinese medicine compound, and the initial compounds are annotated using the PubChem database.
[0014] Preferably, in step S1, the HERB v2.0 and HerbioMap database integrated in the UNIQ system are used to perform key screening based on compounds in the principal drug.
[0015] Preferably, in step S2, the R package clusterProfilerv4.9.2.2 is used to perform KEGG signaling pathway enrichment analysis on the target corresponding to the high-confidence compound, and the screening condition is significant pathways with an adjusted p value < 0.05.
[0016] More preferably, the target point is predicted using the DrugCIPHER algorithm.
[0017] Preferably, in step S2, the construction of the disease-related modular biological network includes the following steps:
[0018] (a) Obtain joint cohort transcriptome data from the TCGA, TARGET, and GTEx databases from the UCSC Xena platform;
[0019] (b) Differential gene expression analysis was performed on the normalized RNA-seq data using DESeq2, and differentially expressed genes with a p-value <0.05 after Benjamini-Hochberg adjustment were defined.
[0020] (c) The Wilcoxon signed-rank exact test was used to analyze the proteomic data and screen for differentially expressed proteins.
[0021] Preferably, in step S2, compounds with a high overlap rate between the enrichment analysis results corresponding to high-confidence compounds and the enrichment analysis results related to the obtained disease modular biological network are selected.
[0022] Preferably, in step S4, the high-potency drug compounds and the targets corresponding to the diseases treated by traditional Chinese medicine compound prescriptions are subjected to GO-BP and KEGG enrichment analyses, respectively, and the results are used for screening and evaluation.
[0023] Preferably, in step S4, the formula for calculating the fitting score using the entropy weight method is: F = W' Ri ×R i +W' ri ×r i ;
[0024] In the formula, F is the fitting score; R i and r i These are the Term similarity score and Target similarity score for high-potency drug compounds, respectively; W' Ri and W' ri These represent the weights of the Term similarity score and the Target similarity score for high-drug-activity compounds, respectively.
[0025] This invention also provides the application of highly important compounds obtained by the above method in the preparation of drugs for treating diseases.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] (1) The screening method of the present invention significantly improves the accuracy of screening active ingredients in traditional Chinese medicine compound prescriptions. In the initial screening stage, the present invention focuses on the principal ingredient by combining the theory of monarch, minister, assistant and guide, which is in line with the principle of the principal ingredient in traditional Chinese medicine theory, and theoretically increases the probability of selecting target compounds. Furthermore, KEGG enrichment analysis is used to associate compounds with disease pathways to ensure that the screened compounds have therapeutic relevance. This mechanism-based screening method is more scientific than traditional empirical screening.
[0028] (2) The screening method of this invention significantly improves the efficiency of screening active ingredients in traditional Chinese medicine compound prescriptions. This invention employs the tandem application of multiple modern analytical technology platforms to achieve comprehensive evaluation from chemical composition and biological effects to pharmacological properties. High-throughput component identification is achieved through Q-Exactive HFX mass spectrometry, while the application of the UNIQ system and CHM-FIEFP software greatly accelerates data analysis and prediction. This multi-platform integration approach compresses the screening work that traditionally takes many years to complete into a shorter time.
[0029] (3) The screening method of this invention has high verifiability and universality. Taking Dang Gui Si Ni Tang as an example, this invention showed significant antitumor activity in 3 out of 5 highly important compounds, with a verification rate of 60%, which is far higher than the success rate of traditional random screening methods. This high verification rate stems from the rigorous screening by the multi-dimensional evaluation system in the early stage. In addition, the component identification, pathway analysis, pharmacological evaluation, and efficacy prediction of the screening method of this invention are all based on universal evaluation criteria, so it can be extended to the research of other traditional Chinese medicine compound prescriptions. Attached Figure Description
[0030] Figure 1 This invention presents a phased target screening and quantitative evaluation system for the components of Dang Gui Si Ni Tang compound. In the figure, A represents the main target screening process; B represents the computational processing layer; and C represents the entropy weight method optimization.
[0031] Figure 2 The figure shows the virtual screening process of the compound ingredients of Dang Gui Si Ni Tang and the pharmacodynamic test results of the highly important compounds obtained by the screening. In the figure, A is the flowchart of the screening framework, which shows the screening process from 1000 initial compounds to the final 5 highly important compounds; B and C show the dose-dependent inhibitory curves of three highly important compounds on HGC27 and AGS cells and the pharmacodynamic test results in mice, respectively. Detailed Implementation
[0032] This invention provides a virtual screening method for highly important compounds in traditional Chinese medicine compound formulas, comprising the following steps:
[0033] S1. High-quality screening of medicinal herb components to obtain highly reliable compounds:
[0034] (1) Identify the components of traditional Chinese medicine compound, obtain the initial compounds, and annotate the initial compounds.
[0035] In this invention, a Q-Exactive HFX mass spectrometer is preferably used to identify the components of traditional Chinese medicine compound. The Q-Exactive HFX mass spectrometer, with its advantages of ultra-high resolution and sensitivity, can maximize the accuracy of initial component identification, laying a solid foundation for subsequent screening. While other mass spectrometry instruments can also perform component analysis, they cannot achieve the same level of resolution and sensitivity.
[0036] In this invention, it is preferred to use the PubChem database (organic small molecule bioactivity data) to annotate the initial compound. More preferably, the annotation includes, but is not limited to, the molecular formula and molecular weight, bioactivity, hepatotoxicity, structures and solubility of the initial compound.
[0037] (2) Based on the principle of monarch, minister, assistant and guide in traditional Chinese medicine theory, the UNIQ system was used to screen the compounds in the monarch drug of traditional Chinese medicine compound prescriptions to obtain highly reliable compounds.
[0038] The UNIQ system (Unified Network Integration and Query System) is a comprehensive platform for bioinformatics analysis. This system integrates multi-level information on compounds, targets, and pathways for network analysis and data mining. In traditional Chinese medicine research, the UNIQ system is primarily used for predicting the biological functions of compounds, pathway enrichment analysis, and efficacy evaluation. A key feature of this system is its ability to simultaneously process multi-dimensional biological data and provide visualized analytical results.
[0039] In this invention, it is preferable to use HERB v2.0 and the HerbioMap database (which have been pre-collected by the system) in the UNIQ system to identify the compound components from the principal drug in traditional Chinese medicine compound prescriptions.
[0040] S2. High-precision prediction of biological effects to obtain potential regulatory compounds:
[0041] (1) KEGG enrichment analysis of the high-confidence compounds was performed using the UNIQ system to evaluate their potential pathways of action.
[0042] In this invention, it is preferred to use the R package clusterProfilerv4.9.2.2 to perform enrichment analysis of relevant KEGG signaling pathways on the targets corresponding to the above-mentioned high-confidence compounds; further preferred is that the corresponding targets are obtained based on the algorithm DrugCIPHER (Network-based relating pharmacological and genomic spaces for drug target identification).
[0043] In this invention, it is preferable to evaluate the potential pathways of action based on the obtained significance adj.p-value of the relevant entries.
[0044] (2) The enriched pathways were compared with the established disease-related modular biological networks to screen out potential regulatory compounds.
[0045] In this invention, the preferred process for constructing disease-related modular biological networks is as follows: Disease-related transcriptome data is obtained through the UCSC Xena platform—specifically, from a joint cohort of the TCGA, TARGET, and GTEx databases (https: / / xena.ucsc.edu / ). The analysis of disease-related modular biological network construction preferably uses normalized data recalculated and processed by UCSC TOILRNA-seq. Differential gene expression analysis is then performed using the R package DESeq2, with differentially expressed genes (DEGs) defined as a Benjamini-Hochberg (BH) adjusted p-value < 0.05. Preferably, proteomics-level data analysis employs the Wilcoxon signed-rank exact test to identify differentially expressed proteins (DEPs). Molecules that are simultaneously upregulated or downregulated are classified as differentially expressed molecules (DEMs) for further investigation.
[0046] In this invention, the preferred method for screening potential regulatory compounds is to select compounds whose enrichment analysis results are highly correlated with the enrichment analysis results related to the obtained disease modular biological network, and whose two enrichment results show a high overlap rate. Further preferred methods involve screening compounds whose pathway-level enrichment analysis results corresponding to the predicted targets of high-confidence compounds have a high degree of matching with the enrichment analysis results of key molecules in the obtained disease modular biological network, and whose enrichment results suggest potential intervention for the disease. As one possible implementation method, taking Dang Gui Si Ni Tang (Angelica Sinensis Decoction) for gastric cancer as an example, compounds with at least one predicted target appearing in the "Gastric Cancer" pathway of KEGG are screened.
[0047] S3. High-throughput chemical property identification to obtain highly drug-like compounds:
[0048] The potential regulatory compounds were systematically evaluated using QED (Qualitative Evaluation and Diagnosis), and ranked based on the scoring results. Compounds with the highest QED scores were selected as high-potency drug-like compounds, with the top 10 compounds receiving the highest QED scores being preferentially selected as high-potency drug-like compounds. This step ensured that the screened compounds possessed favorable pharmaceutical properties.
[0049] The Quantitative Drug Evaluator (QED) score is a comprehensive scoring method for assessing the drug efficacy of compounds. This system calculates a comprehensive score between 0 and 1 by weighting multiple physicochemical parameters such as molecular weight, lipid-water partition coefficient, number of hydrogen bond donors, and number of hydrogen bond acceptors. Compared to the traditional Lipinski rule, the QED score system is more suitable for assessing the drug efficacy of natural products because its scoring criteria are based on the characteristics of marketed drugs and optimized for the specific characteristics of natural products.
[0050] S4. Highly efficient component fitting and prediction to obtain highly important compounds:
[0051] The high-potency compounds were evaluated using CHM-FIEFP software. The fit score for each compound was calculated using the entropy weight method. Compounds with high scores were identified as high-importance compounds, and the top 5 compounds with the highest scores were selected as high-importance compounds.
[0052] CHM-FIEFP software (Traditional Chinese Medicine Compound-Component Effect Fitting Prediction Software) is a prediction software specifically developed for traditional Chinese medicine compound prescriptions. This software integrates compound structure information, target data, and disease pathway information through machine learning algorithms to predict and score the therapeutic efficacy of chemical components in traditional Chinese medicine compound prescriptions. The software is characterized by its strong targeting, and the prediction model has been trained and validated with a large amount of traditional Chinese medicine data.
[0053] In this invention, it is preferable to perform GO-BP and KEGG enrichment analyses on the high-drug-activity compounds obtained above and the targets corresponding to the diseases treated by traditional Chinese medicine compound prescriptions, respectively, and screen the relevant results to obtain highly important compounds.
[0054] The development principle of this invention is as follows:
[0055] The results obtained after enrichment analysis (Biological Process and Pathway) (Note: the enrichment analysis results corresponding to the potential core targets related to the disease treated by the prescription are in the master table, and the enrichment analysis results corresponding to the potential core targets of each potential component in the prescription are in the target sub-tables) are organized into three columns: the first column is the term name, the second column is the target name, and the third column is the p-value. The tables are then accessed using the CHM-FIEFP model.
[0056] This model first removes terms with p-values greater than 0.05 from each table, and then calculates the similarity score R between each term in the sub-table and the overall table. This process filters out terms in each sub-table that are identical to those in the overall table, allowing them to proceed to the next round of calculation. Next, for each filtered target sub-table, the similarity score r for the target is calculated. Finally, the weight W between R and r is obtained using the "entropy weight method". R W r Furthermore, the weight score of each sub-table is calculated using the formula: F = W Ri ×R i +W ri ×r i Finally, the efficacy fitting score (F) for each sub-table (i.e., each potential component in the prescription) is calculated.
[0057] The specific method for evaluating high-drug-potency compounds and obtaining highly important compounds using CHM-FIEFP software in this invention is as follows:
[0058] This system was developed using a scoring model, assigning scores to the Term and Target attributes of each target sub-table based on their fit to the overall table. Then, using the entropy weighting method, the system calculates the weights of each target sub-table with respect to the Term and Target attributes, ultimately calculating the total score for each target sub-table. The specific methodology is detailed in [link to documentation]. Figure 1 .
[0059] (1) Remove all entries with p-value > 0.05 from the master table and each sub-table, and then update the master table and each sub-table;
[0060] (2) Calculate the evaluation index R of the relevant attribute Term of each sub-table. The calculation formula of R is as shown in (3-1) and (3-2).
[0061] N' i =N i ∩N f (3-1)
[0062]
[0063] Where, N i N represents the number of terms corresponding to the i-th component in the prescription. f This indicates the number of terms in a prescription, where f is the representative prescription, and N' is the number of terms in the prescription. i N represents i With N f The number of Term values after taking the intersection.
[0064] (3) Remove all entries in the master table and each sub-table that are inconsistent with the attribute Term, and then update the master table and each sub-table again;
[0065] (4) Calculate the evaluation index r of the relevant attribute Term of each sub-table. The calculation formula of r is as shown in (3-3) and (3-4).
[0066] n' i (T i ) = n i (T i )∩n(T i (3-3)
[0067]
[0068] Where, n i (T i ) represents the number of Targets corresponding to the i-th term of the i-th component in the formula, n(T) i ) represents the total number of Targets corresponding to the i-th term, n' i (T i ) represents n i (T i ) and n(T i The number of Target values after taking the intersection, n f (T i ) represents the total number of targets corresponding to the i-th term in the prescription.
[0069] (5) Calculate the weights W of attributes Term and Target in the evaluation process of each sub-table based on the "entropy weight method". Ri and W ri The calculation method is as follows:
[0070] a) First, calculate the R values for each type of sample in each target sub-table. i and r i Standardize them separately:
[0071] Assume that each subtable is related to R i and r i The results can be formed into vectors such as (3-5) and (3-6):
[0072] R = {R1, R2, ..., R} m} (3-5)
[0073] r = {r1, r2, ..., r} m} (3-6)
[0074] For R i With r iStandardization yields R' i With r' i The calculation formulas are as follows (3-7) and (3-8);
[0075]
[0076] b) Calculate R' i With r' i The information entropy E is as shown in (3-9);
[0077] E R’i,r 'i=-P i logP i (3-9)
[0078] in,
[0079]
[0080] c) Calculate R' i With r' i The weights W are as shown in (3-11);
[0081]
[0082] (6) Finally, the efficacy fitting scores (F) of each target sub-table can be obtained, describing the verification between each target sub-table and the total table. The calculation formula is as shown in (3-12).
[0083] F = W' Ri ×R i +W' ri ×r i (3-12).
[0084] The present invention also provides the application of the highly important compounds obtained by the above method in the preparation of drugs for treating diseases; the preferred disease is gastric cancer, and the highly important compounds are one or more of cinnamic acid, ferulic acid, 2-hydroxyphenylpropionic acid, 2-hydroxycinnamic acid and vanillin, and are screened in Dang Gui Si Ni Tang.
[0085] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0086] Example 1
[0087] Taking the Dang Gui Si Ni Tang compound (Angelica sinensis, Cinnamomum cassia, Paeonia lactiflora, Asarum heterotropoides, etc.) as an example, the virtual screening method of this invention is used to screen for highly important compounds. The specific steps are as follows:
[0088] S1. High-quality screening of medicinal herb components to obtain highly reliable compounds:
[0089] The components of Danggui Sini Decoction were identified using a Q-Exactive HFX mass spectrometer to obtain initial compound information; subsequently, all detected compounds were annotated using the PubChem database.
[0090] Based on the principles of principal, assistant, adjuvant, and guide herbs in Traditional Chinese Medicine (TCM), the UNIQ system was used to screen compounds from the principal herbs in Dang Gui Si Ni Tang (Angelica sinensis and Cinnamomum cassia) [the UNIQ system's HERB v2.0 and HerbioMap database (which pre-collected relevant components) were used to identify compounds from the principal herbs]. From an initial pool of 1000 compounds, 54 high-confidence compounds from the principal herbs were selected. These are: palmitic acid, curcumin, benzoic acid, myristic acid, pentadecanoic acid, sebacic acid, uridine, 4-hydroxy-3-methoxycinnamic acid, nonanoic acid, 2-hydroxyphenylpropionic acid, adenine, linoleic acid, succinic acid, heptanoic acid, azelaic acid, eugenol, 3,4-dihydroxybenzoic acid, sucrose, 4-vinylphenol, phthalic acid, 3-hydroxybenzaldehyde, guanosine, protocatechuic aldehyde, vanillin, ferulic acid, cytidine, oleic acid, diethyl phthalate, and 4-hydroxybenzoic acid. Acids, adenosine, choline, trigonelline, coumarin, 2-phenylethyl formate, styrene, (-)-linalool, (+)-camphene, phthalic anhydride, p-cymene, p-methoxycinnamicaldehyde, uracil, anethole, cis-cinnamicaldehyde, benzaldehyde, N-trans-feruloyltyramine, (E)-o-methoxycinnamicaldehyde, (2E)-2-hydroxycinnamic acid, cinnamic acid, (-)-α-phellandrene, cinnamicaldehyde, ligustilide, nicotinic acid, γ-terpinene, and phenylacetaldehyde.
[0091] S2. High-precision prediction of biological effects to obtain potential regulatory compounds:
[0092] KEGG enrichment analysis was performed on 54 high-confidence compounds selected using the UNIQ system [the R package clusterProfilerv4.9.2.2 was used to perform enrichment analysis on the relevant KEGG signaling pathways corresponding to the targets of the above compounds (obtained based on the DrugCIPHER algorithm)] to evaluate their potential pathways of action (the evaluation and screening were based on the significance of the relevant entries, adj.p-value).
[0093] The enriched pathways were compared with the established disease-related modular biological network to screen for compounds with disease treatment relevance. This involved screening compounds whose enrichment analysis results matched the enrichment analysis results of the disease-related modular biological network, with a high overlap between the two sets of enrichment results. High-confidence compounds were selected based on the high degree of match between the enrichment analysis results of the predicted target pathways and the enrichment analysis results of key molecules in the disease-related modular biological network, and which showed potential for intervention in gastric cancer (at least one predicted target appeared in the "Gastric Cancer" pathway of KEGG). Through this step, 36 compounds with potential regulatory potential were finally identified. Specifically, it contains curcumin, benzoic acid, 4-hydroxy-3-methoxycinnamicaldehyde, 2-hydroxyphenylpropionic acid, linoleic acid, eugenol, 3,4-dihydroxybenzoic acid, sucrose, 4-vinylphenol, phthalic acid, 3-hydroxybenzaldehyde, protocatechuic aldehyde, vanillin, ferulic acid, oleic acid, diethyl phthalate, 4-hydroxybenzoic acid, adenosine, choline, trigonelline, coumarin, 2-phenylethyl formate, styrene, (-)-linalool, (+)-camphene, phthalic anhydride, p-methoxycinnamicaldehyde, anethole, N-trans-feruloyltyramine, (E)-o-methoxycinnamicaldehyde, (2E)-2-hydroxycinnamic acid, cinnamic acid, (-)-α-phellandrene, ligustilide, nicotinic acid, and γ-terpinene.
[0094] S3. High-throughput chemical property identification to obtain highly drug-like compounds:
[0095] The 36 potential regulatory compounds were systematically evaluated using a quantitative drug efficacy score (QED), resulting in 36 scores. Based on the scores, the 10 compounds with the highest QED scores were selected as high-potency compounds. This step ensured that the screened compounds possessed good pharmaceutical properties, namely: cinnamic acid, ferulic acid, 2-hydroxyphenylpropionic acid, (2E)-2-hydroxycinnamic acid, vanillin, anethole, phthalic acid, eugenol, diethyl phthalate, and N-trans-feruloyltyramine.
[0096] S4. Highly efficient component fitting and prediction to obtain highly important compounds:
[0097] The CHM-FIEFP software was used to further evaluate 10 high-potency drug compounds. The software employed entropy weighting to calculate the fit score for each compound (GO-BP and KEGG enrichment analyses were performed on the 10 high-potency drug compounds and the target sites corresponding to the diseases treated by Dang Gui Si Ni Tang, and the relevant results were further screened). Finally, the five compounds with the highest scores were selected as highly important: cinnamic acid, ferulic acid, 2-hydroxyphenylpropionic acid, 2-hydroxycinnamic acid, and vanillin.
[0098] To verify the reliability of the screening results, this invention used gastric cancer cell lines HGC27 and AGS for in vitro pharmacodynamic experiments. The specific process is as follows:
[0099] Human gastric cancer cell lines HGC-27 and AGS were purchased from China Infrastructure of Cell Line Resources (China). They were then cultured in RPMI-1640 and F12K media supplemented with 10% (v / v) FBS and 100 U / ml streptomycin / penicillin, respectively, at 37°C and 5% CO2. After trypsinization and resuspending the cells from the culture dishes, 10 μL of the cell suspension (1:1 dilution) was used for cell counting using a cell counting chamber, and the counts were performed at 5 × 10⁻⁶ cells / mL. 3 Cells were seeded at a density of [number] cells / well in 96-well cell culture plates and then placed in a cell culture incubator to allow them to adhere. After adhesion the following day, cells were treated with different concentrations of candidate compounds for 24 hours, and cell viability was assessed using the CCK8 assay. The results showed that ferulic acid, cinnamic acid, and 2-hydroxycinnamic acid treatment for 24 hours had significant dose-dependent inhibitory effects on the aforementioned gastric cancer cells. Figure 2 As shown in B.
[0100] Further in vivo pharmacodynamic evaluation was conducted using a human gastric cancer PDX mouse (NCG) model, divided into a blank control group, a ferulic acid monotherapy group, a cinnamic acid monotherapy group, and a 2-hydroxycinnamic acid monotherapy group, to evaluate its anti-gastric cancer effect. The specific process is as follows:
[0101] Four-week-old male NCG mice were purchased from IDMO. All procedures for the animal experiments were performed in accordance with the regulations of the Ethics Committee of Tianjin Medical University Cancer Hospital. All 12 mice were randomly divided into 4 groups of 3 mice each. When tumors were visible to the naked eye, the mice were treated with 200 μL of physiological saline (blank control group) and three compound solutions (300 mg / kg body weight), respectively, once daily during the experiment. On day 14, all mice were sacrificed at the end of the experiment (no mice died unexpectedly during the entire experiment), and the tumor specimens were photographed. The results showed that, compared with the blank control group, ferulic acid, cinnamic acid, and 2-hydroxycinnamic acid all inhibited the growth of human gastric cancer xenografts in mice to varying degrees (compared to the blank control group). Figure 2 As shown in C.
[0102] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A virtual screening method for highly important compounds in traditional Chinese medicine compound formulas, characterized in that, Includes the following steps: S1. Identify the components of traditional Chinese medicine compound formulas to obtain initial compounds, and annotate the initial compounds; use the UNIQ system to screen the compounds of the principal drug in the traditional Chinese medicine compound formulas to obtain high-confidence compounds; S2. KEGG enrichment analysis of the high-confidence compounds was performed using the UNIQ system to evaluate their potential pathways of action; The enriched pathways were compared and analyzed with established disease-related modular biological networks to screen for potential regulatory compounds. S3. The potential regulatory compounds were evaluated using the QED system, and the compounds with the highest QED scores were selected as high-drug-like compounds. S4. The high-potency compounds were evaluated using CHM-FIEFP software. The fit score for each high-potency compound was calculated using the entropy weight method. Compounds with high scores were considered high-importance compounds.
2. The virtual screening method according to claim 1, characterized in that, In step S1, a Q-Exactive HFX mass spectrometer is used to identify the components of the traditional Chinese medicine compound, and the initial compounds are annotated using the PubChem database.
3. The virtual screening method according to claim 1, characterized in that, In step S1, the HERB v2.0 and HerbioMap database integrated in the UNIQ system are used to perform key screening based on compounds in the principal drug.
4. The virtual screening method according to claim 1, characterized in that, In step S2, the R package clusterProfilerv4.9.2.2 is used to perform KEGG signaling pathway enrichment analysis on the target sites corresponding to the high-confidence compounds. The screening criteria are significant pathways with an adjusted p-value < 0.
05.
5. The virtual screening method according to claim 4, characterized in that, The target points were predicted using the DrugCIPHER algorithm.
6. The virtual screening method according to claim 1, characterized in that, In step S2, the construction of the disease-related modular biological network includes the following steps: (a) Obtain joint cohort transcriptome data from the TCGA, TARGET, and GTEx databases from the UCSC Xena platform; (b) Differential gene expression analysis was performed on the normalized RNA-seq data using DESeq2, and differentially expressed genes with a p-value <0.05 after Benjamini-Hochberg adjustment were defined. (c) The Wilcoxon signed-rank exact test was used to analyze the proteomic data and screen for differentially expressed proteins.
7. The virtual screening method according to claim 1, characterized in that, In step S2, compounds with a high overlap rate between the enrichment analysis results corresponding to high-confidence compounds and the enrichment analysis results related to the obtained disease modular biological network are screened.
8. The virtual screening method according to claim 1, characterized in that, In step S4, the high-potency drug compounds and the targets corresponding to the diseases treated by traditional Chinese medicine compound prescriptions are subjected to GO-BP and KEGG enrichment analyses, and the results are used for screening and evaluation.
9. The virtual screening method according to claim 1, characterized in that, In S4, the formula for calculating the fitting score using the entropy weight method is: F = W' Ri ×R i +W' ri ×r i ; In the formula, F is the fitting score; R i and r i These are the Term similarity score and Target similarity score for high-potency drug compounds, respectively; W' Ri and W' ri These represent the weights of the Term similarity score and the Target similarity score for high-drug-activity compounds, respectively.
10. The use of highly important compounds screened based on the method of claim 1 in the preparation of drugs for treating diseases.