Screening method and application of anti-breast cancer key targets of traditional Chinese medicine composition for treating triple negative breast cancer pulmonary metastasis

By integrating Mendelian randomization and multi-omics data, a multi-target network of XLCF was screened, which solved the problems of single and systematic target validation in the treatment of lung metastases of triple-negative breast cancer by traditional Chinese medicine composition. This achieved highly accurate target screening and assessment of potential side effects, and reduced the risk of research and development.

CN121366641APending Publication Date: 2026-01-20SHANXI CANCER HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511482854.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

In the existing technology, the target verification method of the traditional Chinese medicine composition Xianling Cifang (XLCF) in the treatment of lung metastasis of triple-negative breast cancer is single, lacks systematicity and integration, makes it difficult to reflect the synergistic effect of "multi-component-multi-target" in traditional Chinese medicine, and the experimental results are easily affected by specific experimental conditions.

Method used

We employed a method integrating Mendelian randomization and multi-omics data to screen traditional Chinese medicine components and targets using the TCMSP and HERB databases. We then constructed a network using Cytoscape, conducted Mendelian randomization analysis and SMR/colocalization analysis, validated the results using clinical databases, and finally performed molecular docking to achieve target screening supported by multi-level evidence.

Benefits of technology

It effectively reduced the false positive rate of network pharmacology prediction, improved the accuracy and robustness of target screening, reduced the blindness of experiments, reduced the risk of failure in the later stages of research and development, and provided assessment support for pleiotropic effects and potential adverse reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366641A_ABST
    Figure CN121366641A_ABST
Patent Text Reader

Abstract

The invention relates to a screening method of anti-breast cancer key targets of a traditional Chinese medicine composition for treating triple-negative breast cancer pulmonary metastasis, which sequentially comprises the following steps: component-target preliminary screening, two sample Mendel randomization, SMR and co-localization, full phenotype association and side effect evaluation, pathway enrichment, clinical expression verification, molecular docking verification and the like. After Mendel randomization (MR) and SMR are placed in network pharmacology, population genetic variation is used as a natural randomization tool, evidence of causal relationship between exposure (target level) and outcome (breast cancer risk) is provided on a public GWAS level, traditional association analysis mixing and reverse causal risk are reduced, and the false positive rate of network pharmacology prediction is greatly reduced; pheWAS-MR is introduced to perform system identification on phenotypes of various diseases, so that a potential off-target effect can be identified in advance, the multiple effects and possible adverse reactions of potential treatment targets in traditional Chinese medicine compound targets can be evaluated, powerful data support is provided for subsequent clinical test design and drug alert strategies, and the risk of later failure of research and development is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a screening method for key breast cancer target points of a traditional Chinese medicine composition for treating triple-negative breast cancer lung metastasis by integrating Mendelian randomization and multi-omics data, and application thereof. BACKGROUND

[0002] Breast cancer ranks first in the global female malignant tumor incidence rate. A traditional Chinese medicine composition for treating triple-negative breast cancer lung metastasis disclosed in Chinese patent CN119633091A, also referred to as Xianling Cifang (XLCF) hereinafter, is composed of five Chinese medicines, i.e. mountain cymbidium (Cremastra Appendiculata), honeycomb (Nidus Vespae), Epimedium brevicornum, Curcuma zedoaria and Paris polyphylla. XLCF is clinically used for the adjuvant treatment of breast cancer and has effects such as delaying lung metastasis and prolonging survival period. However, the material basis of XLCF is not clear, the action target points are not clear, and sufficient causal evidence is lacking.

[0003] The previous research of the research group found that XLCF can significantly prolong the survival period of triple-negative breast cancer patients with positive axillary lymph nodes, and the mechanism is mainly achieved by inhibiting angiogenesis and regulating polarization of tumor-associated macrophages. Some existing researches have discussed the effects of certain single ingredients and single target points in XLCF on the treatment of breast cancer. For example, the polymeric saponin component Polyphyllin III of Paris polyphylla mainly inhibits the proliferation of breast cancer cells MBA-MB-231 through the iron death pathway mediated by ACSL4; the XLCF extract may reduce M2 macrophages in tumors and lungs by reducing STAT6, thereby showing an anti-lung metastasis function; the XLCF formula inhibits breast cancer metastasis through the microRNA-134-SLUG axis.

[0004] However, the current researches mainly focus on the corresponding relationship between single ingredients and single target points, and it is difficult to reflect the overall view of the “multi-ingredient-multi-target” synergistic action of traditional Chinese medicine. In addition, the complex interaction network between target points is ignored. At the same time, the target point verification method is single, and the experimental results are easily affected by specific experimental conditions. The overall research paradigm is still limited to the reductionist framework of “single ingredient-single target point”, and lacks systematicness and integration. SUMMARY

[0005] In view of the above problems of the prior art, according to the embodiments of the present application, a Xianling Cifang target point screening method integrating Mendelian randomization and multi-omics data is provided to solve the above technical problems.

[0006] In order to achieve the above-mentioned purpose, the application provides a screening method of anti-breast cancer key target points of a traditional Chinese medicine composition for treating triple-negative breast cancer lung metastasis, which comprises the following steps in sequence:

[0007] 1, Xianlingcifang target point gene network construction

[0008] Screen the chemical components of XLCF five medicinal materials from the TCMSP database (https: / / tcmsp-e.com / tcmsp.php) (screening conditions: drug similarity DL≥0.18, oral bioavailability OB≥30%), and construct a candidate compound set; map the predicted target points of the candidate compounds to standardized gene symbols through the UniProtKB database, and integrate to form a first target point set;

[0009] Screen the chemical components and predicted target points of XLCF five medicinal materials from the HERB database (http: / / www.uniprot.org) to form a second target point set;

[0010] Use Cytoscape to construct a "traditional Chinese medicine-component-target point" network, and realize topological visualization;

[0011] Two target point sets are combined and de-redundant to form an XLCF target point gene set A.

[0012] 2, Two-sample Mendelian randomization analysis

[0013] Instrumental variable (IV) screening: take the XLCF target point gene set A as a query object, and extract the cis-pQTL of the target point gene from the deCODE public pQTL database; the standard for screening the instrumental variable cis-pQTL is: 1) P<5e-08; 2) exclude SNPs in the human major histocompatibility complex (MHC) region; 3) SNPs 500KB upstream and downstream of the gene; 4) remove linkage disequilibrium r 2 (LD-r 2 )<0.1. The selected data set is of European ethnic background.

[0014] Outcome data acquisition: obtain the GWAS ID (ebi-a-GCST004988 and ieu-b-4810) of breast cancer from the MRC IEU OpenGWAS database, and use the standardized association summary statistics as the outcome by using the R package TwoSampleMR.

[0015] Causal estimation: The breast cancer GWAS summary data was used as the outcome dataset. The Wald ratio method was used to evaluate the Mendelian randomization results of exposure with only one SNP, and the inverse variance weighted (IVW) method was used to evaluate the Mendelian randomization results of exposure with two or more SNPs. The adj.p < 0.05 was used as the screening criterion for significant causality.

[0016] Quality control: Heterogeneity tests, pleiotropy tests, and leave-one-out sensitivity analyses were performed on each exposure-outcome pair to ensure the robustness of the instrumental variables.

[0017] A causal candidate gene set B is formed by targeting quality control.

[0018] 3. SMR and colocalization analysis

[0019] SMR (Summary-data-based Mendelian Randomization) analysis: Using pQTLs corresponding to the causal candidate gene set as exposures and breast cancer GWAS as the outcome, the SMR test was performed, and proteins with P<0.05 in the SMR analysis results were included as key genes in subsequent analyses.

[0020] Bayesian colocalization analysis: For genes that pass SMR, SNPs in the ±50kb region are taken, and the posterior probability PP.H4 is calculated using the Bayesian colocalization framework; PP.H4 > 0.05 is considered as evidence of colocalization between GWAS and pQTL.

[0021] Genes supported by both SMR and colocalization constitute the key target set C.

[0022] 4. Full phenotypic association analysis and side effect assessment

[0023] The key target set C was analyzed using the AstraZeneca PheWAS Portal (https: / / azphewas.com / ) and PheWebDatabase (https: / / pheweb.org / ).

[0024] 5. Enrichment analysis of GO and KEGG pathways

[0025] We used the R package clusterProfiler to perform gene ontology function (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis on the key target set C, with a selection criterion of P<0.05.

[0026] 6. Expression validation in clinical databases

[0027] Data integration: The Breast Invasive Carcinoma (BRCA) dataset (TCGA-BRCA) was downloaded from the Cancer Genome Atlas (TCGA) database (https: / / portal.gdc.cancer.gov / ) using the R package TCGAbiolinks and used as a validation set for analysis. Simultaneously, it was normalized to FPKM (Fragments Per Kilobaseper Million) format and analyzed using the UCSC Xena database.

[28] Clinical data were obtained from (https: / / xena.ucsc.edu / ), including 1183 breast cancer samples and 110 normal controls. Data were standardized to FPKM format and matched with clinical information. GSE29044 (73 breast cancer samples and 36 control samples) and GSE29431 (54 breast cancer samples and 12 control samples) were downloaded from the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ) using the R package GEOquery. Batch effects were removed using the sva package, and the limma package was used for standardization to construct an integrated GEO dataset (127 breast cancer samples and 48 control samples).

[0028] Differential expression analysis: The expression differences between the breast cancer (BRCA) group and the control (Normal) group in the breast cancer dataset (TCGA-BRCA) and the integrated GEO dataset were analyzed using the R packages DESeq2 and limma, respectively.

[0029] 7. Molecular docking verification

[0030] XLCF Active Ingredient and Target Structure Acquisition: Active ingredients corresponding to drug targets were obtained from TCMSP / HERB, sorted by OB and DL scores, and the best representative compounds were selected as ligands; crystal structures of target proteins were downloaded from the PDB database, and key target sets were used as receptors; for cases where a target corresponds to multiple ligands, the OB score was used as the primary criterion. For target proteins without experimental structures, their predicted structures were downloaded from the AlphaFoldDB database.

[0031] Molecular docking: The AutoDock vina program from the CB-Dock2 website was used to perform blind docking and visualization between the target of traditional Chinese medicine compound and the corresponding small molecule compound.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] 1. The target points predicted by network pharmacology are usually based on correlation. By placing Mendelian randomization (MR) and SMR behind network pharmacology, using population genetic variation as a natural randomization tool, providing evidence of causal relationship between exposure (target point level) and outcome (breast cancer risk) at the public GWAS level, reducing the risk of traditional association analysis confounding and reverse causality, and greatly reducing the false positive rate of network pharmacology prediction.

[0034] 2. The five-layer progressive analysis process from traditional Chinese medicine-component-target network screening MR causal inference SMR / co-localization analysis to identify functional variation clinical transcriptome consistency verification molecular docking structure confirmation, each layer is based on independent data sources (pQTL / eQTL, GWAS, transcriptome data, protein crystal structure), realizing multi-source evidence support step by step, and effectively reducing the systematic bias caused by single database or specific experimental conditions.

[0035] 3. The whole process relies on public databases (TCMSP, HERB, deCODE, GWAS catalog, TCGA, GEO, PDB, etc.) and open source software (TwoSampleMR, coloc, AutoDock Vina, CB-Dock2, etc.), which can lock the target points and corresponding components worthy of further study without experiments, reducing blind experiments.

[0036] 4. By introducing PheWAS-MR, multiple disease phenotypes can be systematically identified, potential off-target effects can be identified in advance, the pleiotropic effects and possible adverse reactions of potential therapeutic targets in traditional Chinese medicine compound targets can be evaluated, and strong data support can be provided for subsequent clinical trial design and drug warning strategies, thereby effectively reducing the risk of failure in the later stage of research and development. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other embodiments according to these drawings without creative labor.

[0038] Figure 1 Flow chart for Xianlinggubao target point screening technology.

[0039] Figure 2 Traditional Chinese medicine-component-target network (A: TCMSP; B: HERB).

[0040] Figure 3 Scatter plot of different model effect estimates of Mendelian randomization.

[0041] Figure 4 Schematic diagram for GO and KEGG enrichment analysis.

[0042] Figure 5 Verification for drug target clinical data set.

[0043] Figure 6 Schematic diagram for molecular docking. DETAILED DESCRIPTION

[0044] The present application is further explained by the following specific examples, which should not be construed as limiting. The present application can be carried out by one of ordinary skill in the art and by one skilled in the art, based on the description given herein. The various modifications and variations which can be made to the present application will be apparent to those skilled in the art, without departing from the spirit and scope of the application. It should be noted that the following examples and features in the examples can be combined with each other, without conflict. It should also be noted that the terms used in the examples of the present application are intended to describe specific embodiments, and are not intended to limit the scope of protection of the present application. The test methods in the following examples, unless otherwise specified, are generally carried out under conventional conditions, or under the conditions recommended by the manufacturer.

[0045] When the examples give numerical ranges, it is understood that, unless the present application indicates otherwise, each numerical range can be selected from both ends of the range and any value between the two ends. Unless otherwise defined, all technical and scientific terms used in the present application are consistent with the understanding of the prior art by those skilled in the art and the description of the present application, and any method, equipment and material of the prior art similar or equivalent to the method, equipment and material of the present application can be used to realize the present application.

[0046] It should be noted that the terms such as "up", "down", "left", "right", "middle" and "one" used in the specification are only for the convenience of clear description, and are not intended to limit the scope of the present application, and the change or adjustment of the relative relationship without substantial change of the technical content is also regarded as the scope of the present application.

[0047] As shown in Figure 1 The preferred embodiment of the present application provides a screening method for key target points of anti-breast cancer Chinese medicine composition for treating lung metastasis of triple-negative breast cancer, which comprises the following steps in turn:

[0048] 1. Construction of Xianlingci Fang target point gene network

[0049] Chemical composition acquisition: Chemical components of five traditional Chinese medicines were retrieved from the TCMSP and HERB databases. The ADME parameter thresholds were set as DL≥0.18 and OB≥30%. The TCMSP database predicted 27 components and their corresponding 229 component targets. The HERB database obtained 351 components and 1100 targets of the five traditional Chinese medicines. After merging and deduplication, a total of 1249 traditional Chinese medicine compound targets were obtained, forming XLCF target gene set A.

[0050] A network of relationships between traditional Chinese medicine, its components, and targets was constructed using Cytoscape software. Figure 2 ).

[0051] 2. MR analysis

[0052] (1) Instrumental variable (IV) screening: cis-pQTLs (±100kb, P<5×10) of 1,249 XLCF targets were obtained from the deCODE database. -8 LD-r 2 <0.01), and finally 157 proteins entered the MR of the two samples.

[0053] (2) Causal estimation: Using cis-pQTLs associated with XLCF target genes as instrumental variables, causal inference analysis was performed in conjunction with breast cancer GWAS data (ebi-a-GCST004988, ieu-b-4810). The results showed that 11 proteins had significant causal associations with breast cancer, namely CASP8, CCL2, CNTN3, CDKN1A, HMCN2, CD200R1, MMP1, ST13, PFKFB2, YOD1, and GRK5. These 11 target genes formed the causal candidate gene set B. In the outcome ebi-a-GCST004988, CCL2 and CNTN3 were positively correlated with breast cancer risk, while CASP8, CDKN1A, HMCN2, CD200R1, and MMP1 were negatively correlated. In outcome IEU-B-4810, GRK5 was positively associated with breast cancer risk, while CASP8, PFKFB2, CDKN1A, HMCN2, ST13, and YOD1 were negatively associated. Since ST13 contains only one SNP, we only plotted MR scatter plots of CASP8, CCL2, CNTN3, CDKN1A, HMCN2, CD200R1, MMP1, PFKFB2, YOD1, and GRK5 with breast cancer risk. Figure 3 The intercept of the IVW model approaches 0.

[0054] (3) Quality control: heterogeneity test (Table 1), pleiotropy test (Table 2) and sensitivity analysis (Table 3).

[0055] Table 1. Mendelian randomization analysis of proteins on breast cancer heterogeneity test

[0056]

[0057]

[0058] Table 2. Mendelian randomization analysis of proteins on breast cancer pleiotropy test

[0059]

[0060] Table 3. Mendelian randomization analysis of proteins on breast cancer steiger directionality test

[0061]

[0062]

[0063] 3. SMR and colocalization analysis

[0064] SMR (Table 4): cis-pQTL as exposure, breast cancer GWAS as outcome, excluding pleiotropy; 8 proteins (CASP8, CCL2, CNTN3, CDKN1A, HMCN2, PFKFB2, ST13, YOD1) pSMR <0.05, suggesting causal robustness.

[0065] Colocalization (Table 5): Bayesian colocalization was performed on 8 genes ± 50 kb region, the top SNP of CCL2 gene was rs12601658, and the PP.H4 in colocalization analysis was 0.999; the top SNP of CNTN3 gene was rs549560, and the PP.H4 in colocalization analysis was 1.000; the top SNP of CDKN1A gene was rs3176345, and the PP.H4 in colocalization analysis was 0.823; the top SNP of YOD1 gene was rs1044145, and the PP.H4 in colocalization analysis was 0.968. The 8 genes formed the key target set C.

[0066] Table 4. SMR analysis results of proteins on breast cancer

[0067]

[0068] Table 5. Coloc colocalization analysis results of proteins and breast cancer

[0069]

[0070] 4. Full phenotype association analysis and side effect evaluation (Table 6)

[0071] CASP8, CNTN3, HMCN2 and CDKN1A were significantly associated with 2 to 3 non-breast cancer phenotypes (P < 5 x 10 -8 ) in the PheWeb database; however, these proteins did not show any significant phenotype association (P > 5 x 10 -8 ) in the PheWAS Portal database. The existing literature only supports the association between inguinal hernia and breast cancer risk, and there is no evidence that the remaining phenotypes are associated with breast cancer. Therefore, the risk of adverse drug reactions or unexpected levels of pleiotropic effects is low when targeting these 4 proteins for breast cancer treatment.

[0072] Table 6. Results of protein and breast cancer whole phenomenon association analysis

[0073]

[0074]

[0075] 5. GO and KEGG pathway enrichment analysis.

[0076] GO and KEGG analysis were performed on the 8 key genes. Inhibition of cell cycle G1 / S phase progression, inhibition of mitotic cell cycle G1 / S phase progression and other BPs were the main areas of enrichment of key genes; cell division site, cell cycle protein-dependent protein kinase holoenzyme complex, cleavage furrow, etc. CC; carbohydrate phosphatase activity, cysteine-type peptidase activity, ubiquitin-like protein ligase binding, sugar-phosphatase activity, ubiquitin protein ligase binding, etc. MF. There was also enrichment in the p53 signaling pathway, human cytomegalovirus infection, platinum resistance, IL-17 signaling pathway, South American trypanosomiasis, etc. biological pathways (KEGG). The enrichment analysis results were visualized using a bubble chart ( Figure 4 A). MF, CC, BP and KEGG network diagrams were generated ( Figure 4 B- Figure 4 E).

[0077] 6. Clinical database expression verification Figure 5

[0078] The TCGA-BRCA dataset was downloaded by R package TCGAbiolinks and analyzed as a verification set. After excluding data samples lacking clinical information, a total of 1183 breast cancer samples with clinical information and 110 control samples in Counts format sequencing data were obtained, and the corresponding clinical information is shown in Table 7. After removing the batches of GSE29044 and GSE29431 using the sva package, they were combined into an integrated GEO dataset ( Figure 5 A- Figure 5 D), and the corresponding GEO data information is shown in Table 8. ​

[0079] In the integrated GEO dataset, Figure 5 E- Figure 5 F), the expression difference of CNTN3 and ST13 between normal tissues and breast cancer tissues also showed high significance (P < 0.001); the expression difference of CDKN1A between the two groups was significant (P < 0.01); and the expression difference of CCL2 was also statistically significant (P < 0.05).

[0080] In the TCGA-BRCA dataset, Figure 4 G-H), the expression difference of five key genes CCL2, CNTN3, HMCN2, PFKFB2 and ST13 was highly statistically significant (P < 0.001); the expression difference of CASP8 and CDKN1A between the control group and the breast cancer group was also significant (P < 0.01); and the expression difference of YOD1 between the two groups was statistically significant (P < 0.05).

[0081] Table 7. TCGA-BRCA patient baseline characteristics table

[0082]

[0083]

[0084] Table 8. GEO microarray chip information

[0085]

[0086] 7. Molecular docking Figure 6 )

[0087] The XLCF target points CASP8, CCL2 and CDKN1A are the action target points of Quercetin. The XLCF target points CNTN3, HMCN2, PFKFB2, ST13 and YOD1 are subjected to molecular docking, i.e. Tyrosol. The docking results of the eight target points with Quercetin and Tyrosol are shown in Table 6. Figure 6

[0088] The above examples are only illustrative of the principles and effects of the present application, and are not intended to limit the present application. Any person skilled in the art can modify or change the above examples without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical thought disclosed by the present application should be covered by the claims of the present application.​

Claims

1. A screening method for key target points of Chinese medicine composition for treating triple-negative breast cancer lung metastasis against breast cancer, characterized in that, Comprise the following steps in turn: Step 1-Ingredient-target preliminary screening Extract XLCF five medicinal ingredients in TCMSP according to DL≥0.18, OB≥30%, and standardized target genes according to UniProtKB, to obtain the first target set; extract the same medicinal ingredients and targets in HERB to obtain the second target set; combine the two sets and remove redundancy to form the XLCF target gene set; Step 2-Two-sample Mendelian randomization cis-pQTLs were extracted from deCODE database with step 1 gene set as query, P<5x10 -8 , excluding MHC region, LD-r 2 <0.1 screening tool variables; using breast cancer GWAS summary data as outcome, Wald ratio for single SNP, IVW for multiple SNPs to estimate causality, adj. p<0.05 as significant; quality control by heterogeneity test, pleiotropy test, leave-one-out method to get causal candidate gene set; Step 3-SMR and colocalization Take the pQTL of the candidate gene set in step 2 as exposure, and breast cancer GWAS as outcome, perform SMR test, and retain genes with P<0.05; perform Bayesian colocalization on the ±50kb region, and select genes with PP.H4>0.05 to form the key target set; Step 4-Whole phenotype association and side effect evaluation Input the key target set into AstraZeneca PheWAS Portal and PheWeb Database for whole phenotype association analysis; Step 5-Pathway enrichment Use clusterProfiler to perform GO and KEGG analysis on the key target set, and screen functions and pathways with P<0.05; Step 6-Clinical expression verification Download TCGA-BRCA dataset of 1183 cases of breast cancer and 110 cases of normal control, and integrated GEO dataset of 127 cases of breast cancer and 48 cases of normal control after sva batch removal and limma standardization in GSE29044, GSE29431; compare the expression difference between breast cancer and normal tissue using DESeq2 and limma respectively to verify the key target set; Step 7-Molecular docking verification Select the representative compound with the highest OB and DL score from TCMSP / HERB as the ligand; obtain the protein structure of the key target set from PDB or AlphaFoldDB as the receptor; use CB-Dock2 AutoDock Vina for blind docking, output the binding mode and binding energy, and confirm the direct binding ability of active ingredients and key targets.

2. The use of the method of claim 1 in screening anti-breast cancer key targets of traditional Chinese medicine compositions for treating triple-negative breast cancer lung metastasis.

Citation Information

Patent Citations

  • Traditional Chinese medicine composition for treating triple-negative breast cancer pulmonary metastasis and preparation method and application thereof

    CN119633091A