Method for treating liver cancer by AI-assisted network pharmacology exploration of qi-tonifying and middle-jiao-strengthening decoction

Through AI-assisted network pharmacology methods, multi-database resources and artificial intelligence technology are integrated to screen the active ingredients of Buqi Jianzhong Decoction and the core targets of liver cancer, solving the problems of complex Chinese medicine ingredients and difficult to identify liver cancer targets, and realizing the identification of the precise treatment mechanism of liver cancer and screening of active ingredients.

CN119993256AInactive Publication Date: 2025-05-13ANHUI POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510054627.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Because the ingredients of Buqi Jianzhong Decoction are complex, it is difficult to identify its effective ingredients and functions through traditional methods, and the causes of liver cancer are complex, making it difficult to find effective therapeutic targets.

Method used

AI-assisted network pharmacology methods are adopted to integrate multi-database resources and artificial intelligence technology, systematically screen potential active ingredients, predict targets, build disease-related gene networks, identify key core target genes, and screen candidate compounds that are highly related to the target through AI algorithms.

Benefits of technology

The precise identification of the mechanism and active ingredients of Buqi Jianzhong Decoction for treating liver cancer has been achieved, and the accuracy of online pharmacological analysis and prediction has been improved, providing a new perspective for the modernization of traditional Chinese medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993256A_ABST
    Figure CN119993256A_ABST
Patent Text Reader

Abstract

The invention provides an AI-assisted network pharmacology exploration method for treating liver cancer by Qi-tonifying and Zhongzhong decoction, which is used for revealing a molecular mechanism of the Qi-tonifying and Zhongzhong decoction for treating liver cancer. The research belongs to the field of computer-aided drug design, and aims to systematically promote key links of drug development through a mode of combining multi-database resource integration with an AI technology. Firstly, a plurality of authoritative databases are comprehensively utilized to screen potential active components and target genes thereof, gene expression data in a GEO database are combined, a disease-related gene network is constructed through an AI algorithm, and core target genes for liver cancer treatment are accurately recognized. Secondly, on the basis of a ChEMBL database and an AI auxiliary screening method, candidate compounds highly related to core target genes are mined; afterwards, through a full-automatic molecular docking technology and molecular dynamics simulation, the binding performance and stability of the candidate compound and the core target spot on the molecular level are deeply evaluated, and therefore the stable compound with the excellent binding characteristic is screened out. Finally, the efficacy and safety of the Qi-tonifying and middle-jiao-strengthening decoction are verified through in-vitro experiments, and a theoretical basis and experimental support are further provided for a molecular mechanism of the Qi-tonifying and middle-jiao-strengthening decoction for treating liver cancer. According to the research, on the basis of multi-database resource integration, artificial intelligence and an advanced calculation simulation technology are combined, a systematic process from drug discovery to verification is established, and an innovative method and a practical path are provided for modernization and precise tumor treatment of traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for exploring the Buqi Jianzhong Decoction for treating liver cancer by using AI-assisted network pharmacology, and belongs to the field of computer-aided drug design. Background Art

[0002] Liver cancer is characterized by high morbidity and mortality, complex and diverse molecular mechanisms, and difficult treatment. Although modern therapies have improved, challenges such as drug resistance and toxic side effects still exist. Traditional Chinese medicine has become an important supplement to tumor treatment with its multi-component, multi-target, and low toxicity characteristics. According to TCM theory, the pathogenesis of liver cancer is mainly related to "liver depression", "spleen deficiency", and "qi stagnation and blood stasis". Buqi Jianzhong Decoction is a classic Chinese medicine prescription. It is widely regarded as a comprehensive adjuvant therapy for the treatment of liver cancer because of its unique effects of tonifying the spleen and replenishing qi, soothing the liver and relieving depression. However, due to its complex composition, it is challenging to reveal its mechanism through traditional methods.

[0003] Buqi Jianzhong Decoction (BJD) originated from the Southern Song Dynasty's Jisheng Fang and is an important prescription known for soothing the liver and relieving depression. It is composed of nine Chinese herbs: Atractylodes (15g), Atractylodes (15g), Poria (15g), Tangerine Peel (10g), Magnolia Bark (10g), Orientis (10g), Ginseng (15g), Ophiopogon (15g), and Scutellaria (10g). It is known as a good medicine for promoting blood circulation, relieving pain, and improving liver function. However, due to its complex composition, which can act on multiple targets, it is difficult for us to identify the active ingredients and their functions. In addition, the etiology of liver cancer is complex, it is difficult to find effective therapeutic targets, and it is challenging to reveal its mechanism through traditional methods. Therefore, this study combined network pharmacology, artificial intelligence, molecular docking and simulation technology, and in vitro experimental verification to study the mechanism and active ingredients of BJD in the treatment of HCC, thereby improving the accuracy of network pharmacology analysis and prediction, and providing a new perspective for the modernization of traditional Chinese medicine. Summary of the invention

[0004] The present invention proposes an AI-assisted network pharmacology method for exploring the Buqi Jianzhong Decoction for the treatment of liver cancer, which systematically completes the key steps of drug development by integrating multiple database resources and artificial intelligence (AI) technology. It is characterized in that the method includes the following technical solutions:

[0005] S1: Screen potential active ingredients in the TCMSP and ECTM2.0 databases. After the screening is completed, the SMILES structure of the obtained compound is imported into the Swiss Target Prediction website to predict the corresponding targets of each chemical component. Combined with the GeneCards database (https: / / www.genecards.org / ) and the OMIM database (https: / / www.omim.org / ) to screen genes closely related to the disease, and integrate the gene expression data of the GEO database, use advanced AI algorithms to build a disease-related gene network, and accurately identify key core target genes.

[0006] In one embodiment of the present invention, S1 is: importing the intersection targets into the STRING database (https: / / string-db.org / ) to construct a PPI network, screening out interactions with a confidence level > 0.4, and limiting them to the "Homosapiens" range.

[0007] In one embodiment of the present invention, S1 is: using Cytoscape 3.10.3 (https: / / cytoscape.org / ) to perform analysis to evaluate key topological properties such as degree centrality, betweenness centrality and closeness centrality.

[0008] In one embodiment of the present invention, S1 is: on the Cytoscape 3.10.3 platform, with the help of CytoHubba and CytoNCA plug-ins, the median value is used as the screening threshold to identify potential targets. To further improve the target list, the identified nodes are cross-compared with the differentially expressed genes (DEGs) in the GEO dataset (GSE16476) and the gene set enrichment analysis (GSEA) results.

[0009] In one embodiment of the present invention, S1 is: using machine learning (including random forest, support vector machine, LASSO regression) and deep learning (artificial neural network) algorithms to evaluate the importance of the target. Cross-validation screening is used to obtain core targets with high specificity and stability.

[0010] In one embodiment of the present invention, S1 is: evaluating the diagnostic performance of core targets through ROC (Receiver Operating Characteristic) curve analysis.

[0011] In one embodiment of the present invention, S1 is: performing KEGG functional enrichment analysis on core targets with the help of R packages (clusterProfiler and org.Hs.eg.db).

[0012] In one embodiment of the present invention, S1 is: using the OmicShare tool (https: / / www.omicshare.com / tools / home / report / koenrich.html) to visualize the KEGG analysis results.

[0013] S2: With the help of ChEMBL database, AI algorithm is used to screen candidate compounds that are highly relevant to core targets.

[0014] In one embodiment of the present invention, S2 is: select candidate compounds with the help of AI algorithm, and the biological activity data of small molecule compounds corresponding to the core target protein are derived from the ChEMBL database (https: / / www.ebi.ac.uk / ChEMBL / ). The core target protein and its matching small molecules are determined by querying the UniProt ID website (https: / / www.uniprot.org / ). In terms of biological activity screening criteria, IC 50 The value was converted into pIC 50 In order to standardize.

[0015] In one embodiment of the present invention, S2 is: in the biological assay type, compounds of the binding assay category are given priority.

[0016] In one embodiment of the present invention, S2 is: after completing the preliminary screening of the compounds, evaluate their ADME (absorption, distribution, metabolism and excretion) characteristics to determine the similarity between them. Then, the compounds are screened again according to the Ro5 rule. Compounds that meet these conditions are considered to be more potential drug candidates.

[0017] In one embodiment of the present invention, S2 is: studying the filtering strategy of interfering compounds (PAINS, Pan-Assay Interference Compounds) introduced into drug screening to improve the accuracy and efficiency of compound screening.

[0018] In one embodiment of the present invention, S2 is: to further improve the screening process, the molecular fingerprint (MACCS, Morgan2 / 3) and the active label (pIC 50 ≥7.0) as input features to efficiently screen potential candidate compounds.

[0019] In one embodiment of the present invention, S2 is: using machine learning algorithms (RF and SVM) and deep learning methods (ANN and convolutional neural network: CNN) to carry out model prediction and screen out compounds with higher affinity (pIC 50≥9.0) for further research.

[0020] In one embodiment of the present invention, S2 is: In order to select the most ideal core target structure, the initial step is to screen high-quality structures that meet specific standards from the PDB database. These standards are set as: Homo sapiens, molecular weight (≥100Da), resolution and single-chain structures, ensuring the quality and suitability of the selected structure.

[0021] In one embodiment of the present invention, S2 is: using OpenBabel API programming to perform preprocessing operations on the protein structure. By writing custom code, water molecules and non-essential ligands are removed from the structure. At the same time, the SMILES string of each ligand is converted into PDBQT format. The generated PDBQT file is optimized according to the SMILES string of the ligand and the preset pH value (7.4).

[0022] S3: Molecular docking and molecular dynamics simulation techniques are used to evaluate the binding ability and stability of candidate compounds with core targets, and to screen for stable candidate complexes.

[0023] In one embodiment of the present invention, S3 is: carrying out a detailed molecular docking simulation with the help of the Smina tool based on the AutoDockVina algorithm. Smina performs a comprehensive sampling of the potential binding modes of the ligand within the defined binding pocket, and then calculates the binding energy of each conformation;

[0024] In one embodiment of the present invention, S3 is: for those complexes with binding energy (≤-13.0 kcal / moL), further analysis and visualization using PyMOL and Discovery Studio;

[0025] In one embodiment of the present invention, S3 is: a molecular dynamics simulation verification step, using GROMACS software to implement molecular dynamics simulation (MD).

[0026] S4: Verify the efficacy and safety of candidate compounds through in vitro cell experiments, and provide theoretical basis and experimental basis for new drug research and development.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] The present invention integrates authoritative databases such as TCMSP, GeneCards, GEO, ChEMBL, etc., constructs a disease-related gene network through an AI algorithm, accurately identifies the core target genes of liver cancer, and combines AI technology to screen candidate compounds that are highly correlated with the target. Utilize fully automatic molecular docking and molecular dynamics simulation to deeply evaluate the binding performance and stability of compounds and targets, and optimize the drug screening process. The study also verified the efficacy and safety of Buqi Jianzhong Decoction through in vitro experiments to ensure reliable results. This method realizes the full process innovation of data integration, AI screening and experimental verification, and provides a scientific basis and practical path for the study of the molecular mechanism of Buqi Jianzhong Decoction in the treatment of liver cancer. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 and Figure 2 :Intersecting targets and potential targets of BJD and HCC.

[0030] Figure 3 :DEG S Volcano plot (A) and DEG S Heat map (B); ridge distribution diagram (C); GSEA results of significantly up-regulated genes (D) and down-regulated genes (E).

[0031] Figure 4 : Potential core targets.

[0032] Figure 5 :Identification of core genes using a comprehensive artificial intelligence algorithm. RF analysis of the top 20 important targets (A); SVM-based ranking of candidate target feature importance (B); LASSO regression analysis and coefficient values ​​of candidate targets (C); gene importance ranking based on artificial neural network analysis (D); Venn diagram of core targets identified by four methods (E).

[0033] Figure 6 : KEGG enrichment analysis of core targets.

[0034] Figure 7 : ROC curves for diagnostic evaluation of AKT1 and CTNNB1.

[0035] Figure 8 :Key signaling pathways related to AKT1 / AKT2 and CTNNB1 in HCC (https: / / www.cellsignal.com / pathways / ); PI3K / Akt signaling pathway (A); Wnt / β-Catenin signaling pathway (B).

[0036] Fig. 9: Molecular filtering and substructure analysis of AKT1, AKT2 and CTNNB1 compounds; Molecular features of compounds that meet the Ro5 standard (A); Radar chart of molecular features that meet the Ro5 standard (B); Molecular features of compounds that do not meet the Ro5 standard (C); Substructure frequency analysis of filtered compounds (D).

[0037] Fig.10 :Screening of active compounds for AKT1, AKT2 and CTNNB1 based on machine learning and deep learning. Performance evaluation of machine learning models (A); top 20 highly active compounds (B).

[0038] Fig.11 : Structural and ligand analysis of kinase-related PDB entries; PDB: 3O96 and ligand IQO (A); PDB: 4EJN and ligand OR4 (B); PDB: 7AFW and ligand IQO (C); PDB: 9C1W and ligand XOO (D); PDB: 8Q61 and ligand K06 (E); PDB: 3QKM and ligand SM9 (F).

[0039] Fig.12 : Binding energy distribution of 80 complexes.

[0040] Fig.13 :Visualization of molecular docking results (TOP-6).

[0041] Fig.14 : Molecular dynamics simulation of AKT1-CHEMBL3701766; RMSD curve (A); 3D Gibbs free energy diagram (B); 2D Gibbs free energy diagram (C); RMSF curve (D); Rg curve (E); H bond number (F).

[0042] Fig.15 : Molecular dynamics simulation of AKT2-CHEMBL3701763; RMSD curve (A); 3D Gibbs free energy diagram (B); 2D Gibbs free energy diagram (C); RMSF curve (D); Rg curve (E); H bond number (F).

[0043] Fig.16 :Effects of BJD on proliferation of HepG2 cells.

[0044] Fig.17 :Effects of BJD on the migration and invasion ability of HepG2 cells (Bar=100 μm); scratch test results (A); cell migration rate (B); invasion test results (C); number of invaded cells (D). DETAILED DESCRIPTION

[0045] Example 1: Based on the combination of gene expression data in the GEO database, AI algorithms are used to construct disease-related gene networks to accurately identify the core target genes for BJD treatment of HCC

[0046] The specific steps are as follows:

[0047] 1. Methods for screening cross-targets between BJD and HCC

[0048] (1): The active ingredients of BJD in this example are from TCMSP (https: / / tcmspw.com / index.php). The screening criteria include oral bioavailability (OB) ≥ 30%, drug similarity (DL) ≥ 0.18, and compliance with Lipinski's five rules (Ro5) to ensure that the selected compounds have good pharmacokinetic properties. First, the TCMSP database was used to screen 35 active ingredients of BJD (Table 1). The SMILES structures of these active ingredients were then input into the Swiss TargetPrediction target prediction database to predict their targets (786). In order to improve the accuracy of target identification, an additional 292 targets were retrieved from the ETCM2.0 database. After merging and removing duplicates from the targets of the two databases, a total of 1,077 targets were identified. By querying the GeneCards database (https: / / www.genecards.org / ) and the OMIM database (https: / / www.omim.org / ), 1,273 and 521 HCC-related genes were obtained, respectively. After deduplication, 1,669 HCC-related genes were identified. The intersection of BJD-related genes and HCC-related genes was identified using Venn analysis (https: / / www.bioinformatics.com.cn / static / others / jvenn / ), resulting in 297 cross-targets that may be associated with the therapeutic effect of BJD in HCC treatment ( Figure 1 ).

[0049] Table 1. Potential active ingredients for HCC

[0050]

[0051]

[0052] (2): These intersection targets were imported into the STRING database (https: / / string-db.org / ) to construct a PPI network. Interactions were screened with a confidence score > 0.4 and the interactions were limited to "Homo sapiens". The network was exported and analyzed using Cytoscape 3.10.3 (https: / / cytoscape.org / ) to assess key topological properties, including degree, betweenness centrality, and closeness centrality. The results of the two algorithms were intersected, and 43 potential action targets were identified ( Figure 2 and Table 2 ).

[0053] Table 2. Potential targets of HCC in the treatment of BJD

[0054]

[0055]

[0056] 2. Identify potential core targets through differentially expressed genes (DEGs) analysis in GEO datasets

[0057] (1): In Cytoscape 3.10.3 software, CytoHubba and CytoNCA plug-ins were used to identify potential targets. The median value was used as the screening threshold. In order to further optimize the target list, these identified nodes were cross-compared with the differentially expressed genes (DEGs) and gene set enrichment analysis (GSEA) results in the GEO dataset (GSE16476). The specific operation was to extract DEGs from the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ), and then use R (limma) to normalize and analyze the extracted data.

[0058] (2): If Figure 3 A and Figure 3 As shown in B, by analyzing the DEGs of the liver cancer control group (C) and the experimental group (P), it was found that there were 147 significantly changed genes (75 up-regulated genes and 72 down-regulated genes) in the GEO dataset (GSE164760). In addition, Figure 3 C illustrates the distribution trend of enriched pathways, further emphasizing the key role of metabolic dysregulation (such as drug metabolism and foreign compound metabolism) and immune signaling in liver cancer progression. GSEA analysis further discovered significantly enriched molecular pathways. The upregulated genes were mainly related to the TGF-β signaling pathway, antigen processing and presentation, and complement and coagulation cascades, indicating the presence of pro-inflammatory and immune-related activities in the tumor microenvironment ( Figure 3D). In contrast, downregulated genes were mainly enriched in pathways related to metal ion transport, focal adhesion, and extracellular matrix-receptor interactions, indicating dysfunctional hepatocytes and impaired structural integrity within the tumor microenvironment ( Figure 3 E).

[0059] (3): By intersecting the PPI network nodes with the previously obtained DEGs, potential core targets were finally identified. That is, 43 potential active targets and 147 DEGs S The Venn online tool was imported to identify 17 potential core targets ( Figure 4 ).

[0060] 3. Screening core targets through multimodal AI algorithms

[0061] like Figure 5 As shown in A, 17 potential core targets were ranked according to their feature importance in RF analysis, among which HIF1A, AKT1, and CTNNB1 scored the highest, indicating that HIF1A, AKT1, and CTNNB1 play a key role in the regulatory network affecting HCC progression. In SVM-RFE analysis, the top genes according to their contribution to the classification model were AKT1, CASP3, HIF1A, and CTNNB1 ( Figure 5 B). These genes were identified as key to distinguishing HCC from non-HCC samples, further strengthening their relevance in the disease context. LASSO regression showed significant differences in the coefficients between the core targets. The coefficient of AKT1 (-8.89E-4) indicated that its gene expression was negatively correlated with target activity and played an inhibitory regulatory role in the relevant pathway; while the coefficient of CTNNB1 was 6.72E-4, suggesting that inhibition of CTNNB1 may have some effect but the effect is weak ( Figure 5 C). In the ANN analysis, CTNNB1 had the highest weight (0.134), further highlighting its significant contribution to the HCC-related gene network ( Figure 5 D). The prominence of these genes in multiple modalities suggests that they have the potential to serve as HCC biomarkers or therapeutic targets. Finally, the four-model cross-validation identified AKT1 (PDB: 3O96) and CTNNB1 (PDB: 7AFW) as core targets ( Figure 5 E). The results suggest that these two genes are key targets of the molecular pathways driving HCC progression and may serve as candidate genes for further study and therapeutic development.

[0062] 4. KEGG enrichment analysis

[0063] In order to reveal the biological effects and related pathways of the core targets, functional enrichment analysis was performed on them. On this basis, the Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis was further implemented to clarify the key signaling pathways associated with the core targets. This enrichment analysis was performed using the R package (clusterProfiler and org.Hs.eg.db), and the results were then visualized using the OmicShare tool (https: / / www.omicshare.com / tools / home / report / koenrich.html). KEGG pathway enrichment analysis showed that CTNNB1 and AKT1 genes were significantly enriched in multiple liver cancer-related signaling pathways, including the PI3K / Akt signaling pathway, the Wnt / β-catenin signaling pathway, and the HIF-1 signaling pathway ( Figure 6 ). This suggests that CTNNB1 may promote the progression of liver cancer by regulating tumor stem cell properties and angiogenesis, while AKT1 may regulate the survival and proliferation of tumor cells by participating in metabolic disorders and hypoxia adaptation, indicating that they play a key role in the formation and progression of liver cancer.

[0064] 5. Diagnostic column analysis

[0065] like Figure 7 As shown in the figure, the area under the receiver operating characteristic (ROC) curve (AUC) of AKT1 and CTNNB1 were 0.836 and 0.899, respectively. The higher AUC values ​​indicated that these two genes had higher sensitivity and specificity, highlighting their potential diagnostic biomarker value in identifying HCC patients.

[0066] 6. Pathway and kinase analysis of core targets

[0067] AKT1 is a key kinase in the PI3K / Akt pathway that regulates cell survival and metabolism in liver cancer ( Figure 8 A), which is activated by PIP3 (UniProt ID: P31749). The homolog of AKT1, AKT2, plays a similar role, and the two kinases may compensate for or synergize with each other under certain conditions. The close relationship between AKT1 and AKT2 emphasizes the need to consider both in understanding liver cancer progression and therapeutic intervention. CTNNB1 encodes β-Catenin, a key mediator in the Wnt / β-Catenin pathway that is frequently dysregulated in liver cancer ( Figure 8 B). Upon activation, β-catenin accumulates in the cytoplasm and translocates to the nucleus, driving oncogene expression and promoting tumor progression. In addition, CTNNB1 (UniProt ID P35222) demonstrated a regulatory relationship between kinase activity and β-catenin function, highlighting CTNNB1 as a potential therapeutic target.

[0068] Example 2: Combining the ChEMBL database with AI-assisted screening to obtain candidate compounds highly relevant to core targets

[0069] The specific steps are as follows:

[0070] 1. Using AI algorithms to screen candidate compounds

[0071] (1): The biological activity data of small molecule compounds related to core target proteins are all derived from the ChEMBL database (https: / / www.ebi.ac.uk / ChEMBL / ). The core target proteins and their corresponding small molecules are identified by querying their UniProt IDs (https: / / www.uniprot.org / ). In terms of biological activity screening criteria, IC 50 The value is taken into account and converted into pIC 50 In addition, compounds whose bioassay type was classified as "binding assay" were given priority. Finally, 3995 compounds with SMILES characterization and pIC 50 Value compound.

[0072] (2): After the initial screening of the compounds, their similarity was evaluated by evaluating their ADME (absorption, distribution, metabolism and excretion) properties. The compounds were further screened based on Ro5. The criteria of Ro5 include molecular weight (≤500Da), hydrogen bond donor (HBD≤5), hydrogen bond acceptor (HBA≤10) and LogP (≤5). Compounds that meet these criteria are considered more likely to become successful drug candidates. Finally, 2,521 compounds were obtained. Fig. 9 As shown in A, the average molecular weight of 2,521 compounds is 428Da, and the standard deviation is small, indicating that the molecules are more stable and smaller. The average values ​​of HBA and HBD counts are 5.83 and 1.95, respectively, indicating that there are fewer hydrogen bonding characteristics. The average LogP value is 3.65, indicating that the candidate compounds have moderate lipophilicity ( Fig. 9 B). In summary, 2,521 retained compounds exhibited desirable pharmacokinetic properties and were distributed within the acceptable range specified by Ro5.

[0073] (3): To improve the accuracy and efficiency of compound screening, the study then used interference compound in drug screening (PAINS) filtering to minimize the possibility of nonspecific binding and false positive interactions. In addition, compounds with irrelevant or unwanted substructures were excluded ( Fig. 9 C). Structural frequency analysis of the remaining 2,079 compounds ( Fig. 9D) Common substructures such as oxygen-nitrogen single bonds, aliphatic long chains, and carbonyl groups are highlighted. These substructures are considered to have strong drug-like properties and are very suitable for further development. These compounds also show good ADME properties and structural specificity, laying a solid foundation for high-throughput screening and candidate drug development.

[0074] 2. Screening of active compounds using machine learning and deep learning

[0075] (1): In order to further optimize the screening process, a classification model based on compound fingerprints was constructed and validated, using molecular fingerprints (SMILES) and activity tags (pIC 50 ≥7.0) as the screening index, 1141 active compounds and 938 inactive compounds were classified from the compound data set to efficiently screen potential candidate compounds. Then RF, SVM, ANN and CNN were used for model prediction to screen out compounds with higher affinity (pIC 50 ≥9.0) for further study. Fig.10 As shown in A, among these models (RF model: AUC = 0.85; SVM model: AUC = 0.86; ANN model: AUC = 0.82; CNN model: AUC = 0.82), the SVM model had the highest AUC, indicating that the model performed well in identifying active compounds. Finally, 82 highly active compounds (pIC 50 ≥9.0). These compounds showed excellent activity characteristics and had high potential for drug development. 50 The top 20 most active compounds were subjected to molecular docking analysis ( Fig.10 B) provides a basis for optimizing lead compounds.

[0076] 3. Screen structures from the PDB database based on core targets (select protein targets based on kinases from the PDB database)

[0077] (1): To obtain the optimal core target structure, the optimal structure was first screened from the PDB database. The screening criteria were: Homo sapiens, molecular weight (≥100Da), resolution and single-chain structures to ensure the quality and applicability of the selected structures. The protein structure was then preprocessed using the OpenBabel API, and custom code was written to remove water molecules and non-essential ligands. The SMILES strings of each ligand were then converted to PDBQT format. The generated PDBQT file was optimized based on the SMILES string of the ligand and the preset pH value (7.4) to ensure that it was both reasonable and stable in the subsequent molecular docking process.

[0078] (2): Based on the identified core kinases (CTNNB1 is P35222, AKT1 is P31749, and AKT2 is P31751), high-resolution protein structures were systematically retrieved and screened from the PDB. All selected structures were selected according to strict criteria, including high-resolution To allow for precise structural analysis, and the presence of ligands (≥100 Da) relevant to small molecule drug discovery. Structure of AKT1 3O96 ( Fig.11 A) and 4EJN ( Fig.11 B) was selected to analyze key ligand interactions in the PI3K / Akt signaling pathway. For AKT2, we selected the structure 3QKM ( Fig.11 C) and 8Q61 ( Fig.11 D), highlighting its unique ligand binding kinetics and functional specificity. Finally, the representative structure selected for CTNNB1 is 7AFW ( Fig.11 E) and 9C1W ( Fig.11 F), both structures with bound ligand, revealing a key interaction site in the Wnt / β-Catenin signaling pathway.

[0079] Example 3: Using fully automated molecular docking and molecular dynamics simulation technology, the binding performance and stability of candidate compounds at the molecular level were evaluated, and stable candidate complexes were screened

[0080] The specific steps are as follows:

[0081] 1. Molecular docking and target-ligand interaction analysis

[0082] (1): In order to achieve efficient molecular docking, the present invention uses the Smina tool based on the AutoDockVina algorithm to carry out comprehensive molecular docking simulations. Smina can automatically parse the PDBQT files of PDB binding proteins and ligands, and then automatically identify the position of the binding pocket and calculate the spatial size of the docking site to ensure the accuracy and efficiency of the docking calculation. By comprehensively sampling the potential binding modes of the ligand in the defined binding pocket, Smina is able to calculate the binding energy of each conformation. The docking results of 80 pairs of compound-targets showed that the average binding free energy was -12.08 kcal / moL, indicating that the active compound has a strong binding affinity with the target ( Fig.12). The binding free energy values ​​of highly active compounds were significantly lower than the average, indicating stronger target-ligand interactions. These compounds showed different interaction modes at the key binding sites of CTNNB1, AKT1, and AKT2.

[0083] (2): Among the complexes with binding energy ≤-13.0 kcal / moL, the top six target-ligand interactions were further visualized using PyMOL and Discovery Studio software to ensure accurate interpretation of the molecular interactions. Among them, the CHEMBL3701766-3O96 complex exhibited the highest binding affinity (-13.8 kcal / moL). The interaction mainly involved hydrogen bonds with ASN-53, SER-205, TYR-272, and ASP-274 residues, as well as hydrophobic interactions with THR-82, LEU-210, LEU-264, LYS-268, TYR-272, ASP-274, and ASP-292 residues (Table 3). Other target protein active groups also exhibited higher free binding energy values, all lower than -13.0 kcal / moL ( Fig.13 A to F and Table 3).

[0084] Table 3. Interaction analysis of the complexes

[0085]

[0086] 2. Molecular dynamics simulation and free energy decomposition analysis

[0087] (1) Molecular dynamics simulations (MD) were performed using GROMACS to evaluate the stability and dynamics of the compound-target complexes CHEMBL3701766-AKT1 and CHEMBL3701763-AKT2, and MD and binding free energy decomposition analysis were performed ( Fig.14 and Fig.15 The simulations were run for 100 ns in explicit solvent and the system was equilibrated under NPT conditions. Trajectory analysis included root mean square deviation (RMSD), root mean square fluctuation (RMSF), radius of gyration (Rg), Gibbs free energy, and hydrogen bonding analysis.

[0088] (2): The RMSD curve is an important indicator for evaluating the stability of the protein-ligand complex. For the CHEMBL3701766-AKT1 complex, the RMSD curve tends to be stable after 10 ns, with an average RMSD of 0.40 nm ( Fig.14 In contrast, the CHEMBL3701763-AKT2 complex reached stability after 44 ns, with an average RMSD of 0.6 nm ( Fig.15A). In addition, the Gibbs free energy diagram shows a clear minimum energy region, indicating that the binding conformation is stable ( Fig.14 B and C; Fig.15 B and C).

[0089] The RMSF curve indicates the fluctuation degree of amino acid residues during MD. RMSF analysis shows that the AKT1 amino acid residues in the range of 287 to 316 show greater flexibility ( Fig.14 D), while AKT2 amino acid residues in the range of 37-52, 103-148, and 289-385 showed greater flexibility ( Fig.15 D). The Rg curve is used to characterize the compactness and stability of the structure. The results show that the complex maintains equilibrium throughout the process and no obvious structural changes are observed ( Fig.14 E and Fig.15 E). In addition, the number of hydrogen bonds in the MD process ranges from 0 to 3 ( Fig.14 F) and 1~6( Fig.15 F) changes between the two groups, further demonstrating the stability of the binding.

[0090] (3): In order to further investigate the binding stability of the complex, the MM / GBSA method was used to calculate the binding free energy based on the final 10ns stable RMSD trajectory. As shown in Table 4, the total binding free energy of the AKT1-CHEMBL3701766 complex is -32.40 kcal / moL, of which the van der Waals force contribution is -49.14 kcal / moL, the electrostatic force contribution is -9.67 kcal / moL, and the gas phase energy contribution is -58.82 kcal / moL. These results indicate that the total binding energy is beneficial to the stability of the AKT1-CHEMBL3701766 complex system. In addition, the total binding free energy of the AKT2-CHEMBL3701763 complex was determined to be -55.42 kcal / moL, of which the van der Waals force is -64.53 kcal / moL, the electrostatic force is -34.80 kcal / moL, and the gas phase energy is -99.33 kcal / moL. These values ​​also support the stability of the AKT2-CHEMBL3701763 complex system.

[0091] Table 4. Energy of the composite MMPB / GBSA (kcal / moL)

[0092]

[0093] Example 4: Verification of the efficacy and safety of BJD by in vitro experiments

[0094] The specific steps are as follows:

[0095] 1. Verification through cell experiments

[0096] (1) In this experiment, the first step was to explore the effects of different concentrations of BJD on HepG2 cells. The specific operation was as follows: BJD (100 g) was soaked for 1 hour, decocted twice (40 min), centrifuged, and the supernatant was concentrated to 1 g per mL of the original medicinal material and stored at -20°C. HepG2 cells were inoculated in DMEM medium (Gibco) containing 10% FBS, penicillin (100 U / mL), and streptomycin (100 U / mL) at 37°C and 5% CO 2 The cells were cultured in a moist environment. In this experiment, cells from passages 4 to 8 were used.

[0097] Cell counting kit-8 (CCK8; Biyuntian, Shanghai) was used to analyze cell viability. HepG2 cells in the logarithmic growth phase were selected and seeded in a 96-well cell culture plate at a density of 5×105 cells / well. After the cells adhered to the wall, different concentrations of BJD (0, 12.5, 25.0, 50.0, 75.0, 100.0, 125.0, 150.0, 200.0 μg / mL) were added and cultured for 24 hours. After the incubation, 100 μL of CCK-8 reagent (10%) was added to each well and incubated for 1 hour. A blank group containing only CCK-8 reagent but no cells was added was set as a control. The optical density (OD) value was measured at 450 nm, and the cell viability was calculated using the following formula:

[0098]

[0099] It was found that different concentrations of BJD showed a dose-dependent inhibitory effect on HepG2 cells. The calculated half inhibitory concentration (IC 50 ) was 185.4 μg / mL( Fig.16 ). Based on this result, the concentrations of BJD in subsequent experiments were set to 0, 50, 100, and 200 μg / mL (C, L, M, and H).

[0100] 2. Exploring the effect of BJD on HepG2 cells through scratch repair assay and transwell assay

[0101] (1) To evaluate cell migration ability, a wound healing experiment was performed. HepG2 cells were seeded onto six-well plates at a density of 5×105 cells per well. When the cells grew to confluence, a line was drawn in the center of the cell monolayer using a 200 μL pipette tip, perpendicular to the plate surface. The plate was then washed with PBS to remove any cell debris. After washing three times with PBS, different concentrations of BJD solution (0, 50, 100, 200 μg / mL) were added. The scratch area was imaged at 0 and 24 hours using an inverted microscope, and the cell migration rate was analyzed using ImageJ software. The healing rate was calculated using the following formula:

[0102]

[0103] HepG2 cells were seeded in 24-well plates and cultured overnight. Different concentrations of BJD (0, 50, 100, 200 μg / mL) were added to each well according to the settings of different experimental groups. Then, each group of cells was digested with trypsin and prepared into a concentration of 5×10 3 cells / mL of cell suspension. Add 100 μL of cell suspension to the upper chamber of the Transwell chamber coated with Matrigel, and add DMEM medium containing 10% FBS to the lower chamber. After 48 hours of culture, wash the cells three times with PBS and gently wipe off the remaining cells in the upper chamber with a cotton swab. For cells that migrated to the lower chamber, fix them with 4% paraformaldehyde and stain them with 0.1% crystal violet. Repeat the operation three times for each experimental group. For each repetition, randomly select 5 high-power fields under an inverted microscope to count the number of migrated cells. The entire experiment needs to be repeated three times.

[0104] The results are as follows Fig.17 As shown in A, the scratch spacing of each group of cells was roughly the same at 0h. Compared with the blank control group, the sample significantly inhibited the proliferation rate of HepG2 after 24h treatment, and the inhibitory effect increased with the increase of concentration (P<0.05). After 24h treatment with 200μg / mL, the cell repair rate was 19.03±1.72% ( Fig.17 B). The results showed that after 24 hours of BJD treatment, the invasion ability of cells was significantly inhibited, and the number of Transwell migrated cells was significantly reduced compared with the blank group (P<0.05) ( Fig.17 C); when the concentration of BJD was 200 μg / mL, the number of Transwell cells was 83.67±5.51 ( Fig.17 D).

[0105] Unless otherwise defined, all terms, symbols and other scientific terms used herein are intended to have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. In some cases, terms with conventionally understood meanings are defined herein for the purpose of clarification or ease of reference, and such definitions herein should not be construed as indicating significant differences from conventional understandings in the art. The technical methods described or cited herein are generally well understood by those skilled in the art and are adopted by conventional methods. Unless otherwise stated, the use of commercially available kits, reagents and instruments is carried out in accordance with the protocols and parameters given by the manufacturer.

[0106] Although the disclosure is disclosed as above, the protection scope of the disclosure is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the disclosure, and these changes and modifications will fall within the protection scope of the present invention.

Claims

1. AI-assisted network pharmacology explores the method of Buqi Jianzhong Decoction for treating liver cancer, characterized in that: The method comprises the following steps: S1: Screen potential active ingredients in the TCMSP and ECTM2.0 databases. After the screening is completed, the SMILES structure of the obtained compound is imported into the Swiss Target Prediction website to predict the corresponding targets of each chemical component. Then the GeneCards database (https: / / www.genecards.org / ) and the OMIM database (https: / / www.omim.org / ) are used to determine the genes associated with the disease. At the same time, combined with the gene expression information included in the GEO database, the gene network associated with the disease is constructed with the help of advanced AI algorithms, so as to accurately identify the key core target genes; S2: Obtain candidate compounds that are highly relevant to core targets through the ChEMBL database combined with AI-assisted screening. S3: Using fully automatic molecular docking and molecular dynamics simulation technology, the binding performance and stability of candidate compounds are evaluated at the molecular level, and finally stable candidate complexes are screened; S4: Verify the efficacy and safety of Buqi Jianzhong Decoction through in vitro experiments to provide theoretical basis and experimental support for drug development.

2. The method according to claim 1, characterized in that In the S1, the effective ingredients of BJD for treating HCC are screened from the TCMSP database according to OB≥30%, DL≥0.18 and in accordance with the Ro5 rule.

3. The method according to claim 1, characterized in that The intersection targets in S1 were imported into the STRING database to construct a PPI network, and the interactions were screened with a confidence score > 0.4 and limited to "Homo sapiens".

4. The method according to claim 1, characterized in that In the S1, the CytoHubba and CytoNCA plug-ins in Cytoscape 3.10.3 were used to identify potential targets with the median as the threshold, and the target list was improved by intersecting with the DEGs and GSEA results of the GEO dataset (GSE16476). DEGs were extracted from GEO and normalized by R (limma) and differential analysis, and intersected with the PPI network nodes to determine potential core targets.

5. The method according to claim 1, characterized in that In S1, machine learning (RF, SVM, LASSO) and deep learning (ANN) were used to evaluate the importance of targets, Venn analysis was used to integrate the results to determine the core targets, and ROC curve analysis was used to evaluate the diagnostic performance.

6. The method according to claim 1, characterized in that In S1, functional enrichment and KEGG pathway analysis were performed on the core targets using R packages (clusterProfiler and org.Hs.eg.db), and the results were visualized using the OmicShare tool.

7. The method according to claim 1, characterized in that In S2, an AI algorithm is used to screen candidate compounds from the ChEMBL database, and core target proteins and small molecules are identified by UniProt ID. 50 Values ​​converted to pIC 50 As the criterion, "binding assay" type compounds were given priority; ADME properties of the initial screening compounds were evaluated and further screening was performed based on Ro5.

8. The method according to claim 1, characterized in that In the S2, PAINS filtering is used to improve the accuracy and efficiency of screening; molecular fingerprints (MACCS, Morgan2 / 3) and active tags (pIC 50 ) as the feature, and use machine learning (RF, SVM) and deep learning (ANN, CNN) to screen for high affinity (pIC 50 ≥9.0).

9. The method according to claim 1, characterized in that In S2, the PDB database was used to identify the genes with molecular weight ≥ 100 Da and resolution ≤ The optimal core target structure was screened based on the single-chain structure. Then, the Python-coded API OpenBabel was used for preprocessing to remove water molecules and non-essential ligands, convert the ligand SMILES string into PDBQT format and optimize it.

10. The method according to claim 1, characterized in that In the S3, the Smina tool based on the AutoDockVina algorithm was used to perform molecular docking simulation, parse PDB and SMILES, identify the binding pocket position and calculate the docking site size, calculate the binding energy of each conformation, and analyze and visualize the complexes with binding energy ≤-13.0 kcal / moL using PyMOL and Discovery Studio.