Method and system for automatically screening traditional Chinese medicine efficacy substances based on large model
By constructing a deep neural network model based on automated data collection and deep learning methods using large models, the problem of low efficiency in identifying active substances in traditional Chinese medicine was solved, achieving efficient and accurate screening of active substances in Chinese medicine and improving the scientific rigor and systematic nature of the research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN UNIV OF TRADITIONAL CHINESE MEDICINE
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional methods are insufficient for efficiently and systematically identifying the active substances in traditional Chinese medicine, resulting in a lack of scientific rigor and systematicity in the study of the efficacy of traditional Chinese medicine.
By employing automated data acquisition and deep learning methods based on large models, and through semantic understanding and knowledge reasoning, a deep neural network model is constructed to achieve automated screening of active substances in traditional Chinese medicine.
This improves the efficiency and accuracy of screening active substances in traditional Chinese medicine, reduces the blind spots of traditional experiments, and provides scientific and efficient technical support for the research of active substances in traditional Chinese medicine.
Smart Images

Figure CN121963955A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of traditional Chinese medicine informatics and artificial intelligence technology, specifically relating to an automated screening method and system for active substances in traditional Chinese medicine based on a large model. Background Technology
[0002] The traditional Chinese medicine (TCM) industry is currently at a critical stage of transformation, shifting from traditional experience-based guidance to modern scientific understanding. In TCM research and quality control, the identification and screening of pharmacodynamic substances is a crucial step in building a modern TCM research system. Pharmacodynamic substances in TCM typically refer to chemical components that can directly act on disease-related targets and produce actual biological effects, forming the material basis for the clinical efficacy of TCM. However, due to the complex composition, multiple targets, and intricate interactions of TCM, traditional methods relying on experience and individual experiments are insufficient for efficiently and systematically identifying pharmacodynamic substances. Therefore, there is an urgent need for an efficient screening method that integrates modern data mining and intelligent algorithms to improve the scientific rigor and systematic nature of TCM pharmacodynamic substance research.
[0003] With the development of artificial intelligence, especially technologies such as Large Language Models (LLM), deep learning, and graph neural networks, models have shown significant advantages in drug activity prediction, target association reasoning, and multimodal data fusion, enabling the automatic extraction of high-dimensional nonlinear relationships in molecular structures, pharmacological features, and target networks. Meanwhile, multi-source databases of traditional Chinese medicine (TCM) chemical components, target effects, and pharmacological activities are becoming increasingly comprehensive, providing a solid data foundation for training high-quality intelligent models. Against this backdrop, there is an urgent need for a TCM pharmacodynamic substance screening technology that combines large models to achieve intelligent processing from data acquisition, model construction, automated training to predictive output, thereby improving the systematicness, scientific rigor, and efficiency of TCM pharmacodynamic substance research.
[0004] In view of this, the present invention is hereby proposed. Summary of the Invention
[0005] To address the aforementioned technical problems in existing technologies, this invention provides an automated screening method and system for active substances in traditional Chinese medicine based on a large model. By automating the collection and standardized processing of multi-source traditional Chinese medicine data, and utilizing a large model for deep learning training and knowledge reasoning, it achieves efficient and accurate screening of active substances, reduces the blindness of traditional experiments, improves prediction accuracy, and provides scientific and efficient technical support for the discovery and development of active substances in traditional Chinese medicine.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows: Firstly, an automated screening method for active substances in traditional Chinese medicine based on a large model includes: S1. Determine the target attributes for drug efficacy screening, and obtain drug efficacy-related targets based on the target attributes using the semantic understanding and knowledge reasoning capabilities of the large model; S2. Based on the pharmacodynamic-related targets, call the large model-driven data acquisition module to automatically collect the target-related active compound dataset; S3. Perform automated structure and property standardization processing on the target-related active compound dataset to construct a standard model training dataset; S4. Identify the target Chinese medicinal materials or Chinese medicinal formulas, and use the knowledge enhancement function of the large model to automatically extract and integrate the compound dataset contained in the Chinese medicinal materials or formulas; S5. Standardize the dataset of Chinese herbal compounds to construct a standard dataset of Chinese herbal compounds. S6. Input the standard model training dataset into the deep neural network model to train the model and obtain the target efficacy and target target pharmacological substance screening model. S7. Input the standard Chinese herbal compound dataset into the pharmacodynamic substance screening model, combine the inference mechanism of the large model to obtain the prediction results, and output them in a visual format.
[0007] Furthermore, the pharmacodynamic-related targets were identified through automated literature mining and structured database retrieval from... and Database acquisition, and leveraging the named entity recognition and relation extraction capabilities of large models to improve retrieval accuracy.
[0008] Furthermore, step S2 specifically includes: Based on the pharmacological target, the automated data acquisition and structured extraction technology driven by a large model is used to obtain data on active compounds that have an active relationship with the pharmacological target from public databases. The active compound data are screened, and invalid or unqualified compound structures are eliminated using structural verification tools. Large-scale models are used for semantic deduplication and validity assessment of compounds, combined with structure verification tools to eliminate invalid or unacceptable compound structures; the screened compound structures are then used... The structure is converted into a unified standard format to verify its legality and to construct a dataset of target-related active compounds.
[0009] Furthermore, the data includes compound structure, target identifier, activity value, activity level, and structure-target association confidence labels generated by a large model.
[0010] Furthermore, step S3 specifically includes: Based on the target-related active compound dataset, the target-related active compound dataset is subjected to structure normalization. The Python script is used in conjunction with the semantic parsing function of the large model to extract the SMILES structure and its label information in batches as input to the deep neural network model, and exported as CSV, JSON or NumPy format.
[0011] Furthermore, step S4 specifically includes: Based on the preset efficacy target, the large model is invoked to screen Chinese medicinal materials or prescriptions that are highly correlated with the target efficacy from the knowledge base and database; the names, structures and CAS numbers of Chinese medicinal compounds are obtained and saved as standardized format files.
[0012] Furthermore, step S5 specifically includes: The structure of the Chinese herbal medicine compound dataset is standardized, the SMILES format is unified, and redundant information is eliminated. The compound attributes are semantically labeled using a large model to generate a standard Chinese herbal medicine compound dataset.
[0013] Furthermore, step S6 specifically includes: Utilize large models for automated feature engineering and feature importance analysis; construct a deep neural network model structure that combines knowledge enhancement from large models to automatically learn the complex relationship between compound features and drug targets; Automatically divides the training, validation and test sets, and suggests the selection of loss function and optimizer based on the large model. The training process is automatically monitored and saved, and an early stopping mechanism is automatically triggered when overfitting or performance stagnation occurs. After training, the model performance is evaluated using the validation and test sets, and the AUC, F1-score, and confusion matrix evaluation metrics are automatically output.
[0014] Furthermore, step S7 specifically includes: The standardized Chinese herbal compound dataset was preprocessed using a feature method consistent with the model input. The feature matrix is input into a drug efficacy screening model that combines large model knowledge enhancement, and the predicted results of each compound on the target efficacy or target are output. The prediction results are interpreted and visualized using a large model, highly active compounds are screened as candidate components, and an interactive display interface is generated.
[0015] Secondly, an automated screening system for active substances in traditional Chinese medicine based on a large model includes: Data acquisition module: Combines large-scale model semantic retrieval with database crawling to obtain drug efficacy-related target and compound data; The structure standardization module performs standardization and semantic tagging on datasets of target-related active compounds and traditional Chinese medicine compounds. Model training module: Input the standard model training dataset into the deep neural network model that combines knowledge from large models for training; Predictive reasoning module: Uses the trained model to predict, interpret, and visualize Chinese herbal compounds.
[0016] Furthermore, the data acquisition module includes: Database interface unit: Used to establish a connection with the database of Chinese medicine compounds, efficacy and target through program scripts, and automatically batch retrieve structure, efficacy and target information under the guidance of large model semantic parsing; Keyword retrieval unit: Used for joint queries across multiple databases based on target efficacy attributes and synonyms generated by a large model, automatically generating preliminary candidate datasets.
[0017] Furthermore, the structural standardization module includes: Format processing unit: Used to call chemical information processing tools and perform unified format conversion on the captured compound structures according to the structure analysis rules generated by the large model; Redundant information removal unit: Used to combine the anomaly detection capabilities of rule engine and large model to clean invalid, missing, duplicate or structurally conflicting data.
[0018] Furthermore, it also includes a feature extraction module, which is used to automatically extract features from the standardized compound dataset, specifically including: Molecular fingerprint generation unit: used to convert the molecular structure of a compound into numerical fingerprint features using molecular fingerprints; Graph structure coding unit: used to extract the topological features of compounds through a graph structure encoder; Physicochemical property calculation unit: used to calculate molecular weight and LogP basic properties, and to assist the model in learning the relationship between compound characteristics and drug targets.
[0019] Furthermore, the model training module includes: Neural network architecture configuration unit: used to support users in selecting fully connected network, graph convolutional neural network, and graph attention network model architecture; Training parameter optimization unit: used to support hyperparameter setting and automatic search; Training monitoring unit: Used to record training metrics such as loss value, accuracy, and AUC, and supports early stopping and model checkpoint saving.
[0020] Furthermore, the prediction inference module includes: Standard data import unit: used to import standardized Chinese herbal compound datasets and complete predictions under the guidance of large model inference; Activity score output unit: Used to output the activity score or probability value of each Chinese medicine compound on a specific target or target efficacy attribute, and supports threshold screening and sorting.
[0021] Furthermore, it also includes a result display module, which is used to display the inference results output by the prediction inference module, specifically including: Visualization output unit: Used to display the model prediction results of candidate pharmacodynamic substances in the form of charts, rankings, etc., and generate visualization reports or interactive display interfaces; File Export Unit: Used to support outputting results in PDF or Excel format.
[0022] Compared with existing technologies, the present invention provides an automated screening method and system for active substances in traditional Chinese medicine (TCM) based on a large model. The method determines the target pharmacodynamic attributes and acquires relevant target points, collects data on target-related active compounds and compounds from the target TCM herbs or formulas, constructs training and prediction sets after standardization, and trains a deep neural network model to obtain a screening model for active substances, ultimately achieving the prediction of specific active substances in TCM. The system includes modules for data acquisition, structural standardization, model training, and predictive inference, supporting automated data acquisition, standardization processing, model optimization, and result visualization output. This invention can improve the efficiency and accuracy of screening active substances in TCM, reduce the blind spots of traditional experiments, and provide scientific and technological support for the discovery and development of active ingredients in TCM. Attached Figure Description
[0023] Figure 1 A flowchart of an automated screening method for active substances in traditional Chinese medicine provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the automated screening system for active substances in traditional Chinese medicine provided in an embodiment of the present invention. Detailed Implementation
[0024] The technical solution of the present invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are not all embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0025] It should be noted that, unless otherwise specifically stated, the relative arrangement and numerical expressions of the components and steps described in these embodiments should not be construed as limiting the scope of the invention.
[0026] The following description of exemplary embodiments is merely illustrative and is not intended to limit the invention or its application or use in any way. Techniques, methods, and apparatus known to those skilled in the art may not be discussed in detail herein, but where applicable, such techniques, methods, and apparatus should be considered part of this specification.
[0027] Example 1 See Figure 1 , Figure 1 This is a flowchart of an automated screening method for active pharmaceutical ingredients (APIs) in traditional Chinese medicine (TCM) based on a large-scale model, proposed in this invention. This method sequentially completes steps such as target API identification, data collection and standardization, deep neural network model training, and API prediction. It accurately screens for specific active pharmaceutical ingredients in target TCM herbs or prescriptions based on their target efficacy attributes. Standardization improves data quality, and large-scale model training enhances prediction reliability, reducing the blind spots of traditional experimental screening and providing efficient scientific methodological support for the discovery of active ingredients in TCM, TCM research and development, and pharmacological mechanism studies. Specific steps may include: S1. Determining the target and acquiring target points: Determining the target attributes for drug efficacy screening, and acquiring drug efficacy-related target points based on the target attributes; specifically, this may include: S11. Based on the clinical application background, disease type, or efficacy requirements, determine the target efficacy attributes (such as anti-inflammatory, anti-tumor, and other pharmacological functions) through manual setting or based on a question-and-answer system (such as a query interface built on a medical ontology), and standardize them into a set of keywords for automated calling in subsequent steps.
[0028] S12. Employing automated literature mining and structured database retrieval technologies, data is collected from TCMSPs, drug banks, and other sources. ), public chemistry databases ( ), Bioactivity Database ( Databases such as [database name missing] are used to obtain pharmacodynamic target information related to the target drug's efficacy attributes. Specific methods include: using biological Python libraries (…). ) or the Entrez module calls the PubMed application programming interface of the National Center for Biotechnology Information (NCBI). This allows for batch keyword retrieval of documents. Using Selenium tools , Web-based crawling is used to extract data from the database's web interface. Additionally, for applications that support RESTful APIs (...),... The target database can use the requests library to automatically submit keywords and parse the returned target data in JSON (JavaScript Object Notation) or XML (Extensible Markup Language) format.
[0029] S13. Use Python scripts to clean, organize, deduplicate, and functionally categorize the acquired target information. Use pandas to process the crawled or exported target data (such as gene names, etc.). The target function is formatted; for targets with synonyms or inconsistent naming in different databases, a formatting process is adopted. and Achieve unified and standardized target annotation. Based on target-specific GO annotation... Pathway information and other data are used to classify the biological functions of targets. This allows for the screening of core targets that are highly correlated with the efficacy of the target drug and have a clear mechanism of action, providing a basis for subsequent model construction.
[0030] S2. Collect target-related active compound data: Based on the pharmacodynamic-related targets, collect a dataset of target-related active compounds; S21. Based on the pharmacodynamic target determined in step S1, acquire data on active compounds that have an active relationship with the pharmacodynamic target from a public database using automated data acquisition technology; specifically, using web crawling technology, by calling the bioactive database application programming interface (API) ) or use Network resource client ( This Python package downloads compound data in batches by target identifier (target ID). The information obtained includes compound structure (SMILES, simplified molecular input specification), target ID, activity values (such as IC50, half-maximal inhibitory concentration; EC50, half-maximal effective concentration; Ki, inhibition constant, etc.), and activity level, as well as structure-target association confidence labels generated by a large model.
[0031] S22. The active compound data is screened, and invalid or unqualified compound structures are removed using structure verification tools; null values, missing structures, and duplicate records are removed using the pandas library; preliminary screening is performed by setting numerical thresholds (e.g., IC50 < 10 μM is considered active); and the structure verification is performed using the ChemInformatics Development Kit (CHEDK). Conduct structural verification to eliminate invalid SMILES or unqualified compound structures.
[0032] S23. Use the screened compound structures The structure is converted into a unified standard format to verify its legitimacy and to construct a dataset of target-related active compounds. The function verifies the validity of compound structures; outputs standard formats such as SMILES, International Compound Identifier (InChI), and Molecular Structure File (Mol); and uses comma-separated value files (CSV) or JavaScript object representation. ) Construct structured data storage to provide usable data for subsequent model training.
[0033] S3. Construct a standard training set: Standardize the dataset of active compounds related to the target and construct a standard model training dataset. S31. Based on the target-related active compound dataset collected in S2, perform structure normalization on the target-related active compound dataset; convert the activity values of active compounds into binary labels according to efficacy, define classification thresholds (e.g., set IC50 < 10 μM as activity = 1, otherwise 0), and sort the data according to [" (Simplified molecular input specifications) Save in the format of "(Active Tag)"; Through numerical computation library ( ) and data analysis library ( Write scripts to implement batch processing, and at the same time check whether the label distribution is balanced to ensure the rationality of the training data.
[0034] S32. A Python script is used to extract the SMILES structures and their corresponding activity tag information in batches, which are then used as input data for a deep neural network model.
[0035] pass Or deep learning chemistry toolkit ( ) Analyze the molecular characterization of SMILES; the processed dataset can be exported as CSV, JSON, or This format is used for subsequent model training.
[0036] S4. Collect data on Chinese medicinal compounds: Identify the target Chinese medicinal materials or Chinese medicinal formulas, and collect a dataset of Chinese medicinal compounds contained in the target Chinese medicinal materials or formulas. S41. First, based on the preset efficacy targets (such as anti-inflammatory, anti-tumor, antiviral, etc.), and combining traditional medical knowledge with modern research results, select Chinese medicinal materials or prescriptions that are strongly related to the target efficacy. The screening process can be implemented by constructing a Python automated script with keyword matching capabilities. This script, combined with positive rules (such as "treats ×× disease syndrome" or "belongs to ×× meridian"), performs text parsing on textual materials such as the *Shennong Bencao Jing*, *Dictionary of Chinese Medicine Prescriptions*, and *Dictionary of Chinese Materia Medica*, or online open databases such as the *China Traditional Chinese Medicine Information Database*.
[0037] In text data processing, calling Word segmentation tools are used to perform word frequency analysis and topic keyword extraction, assisting in cluster analysis to determine the matching degree between the efficacy direction and traditional Chinese medicine prescriptions. For modern research findings, automated data collection programs based on web crawlers (such as those based on...) are constructed. or Automatically retrieves lists of prescriptions and medicinal materials from traditional Chinese medicine prescription databases or literature retrieval systems such as PubMed, CNKI, and Wanfang Database, and then... The mapping standard library for Chinese and English synonyms, such as mapping, completes the standardization of medicinal material names.
[0038] S42. After identifying the target Chinese medicinal materials or prescriptions, use Python to write batch query and crawling scripts to automatically call... By using the public APIs or web-based data interfaces of traditional Chinese medicine databases, and simulating POST or GET requests, the corresponding chemical component data can be crawled by passing in keywords of the name of the medicinal material / prescription.
[0039] For scenarios with complex webpage structures or content dynamically rendered via JavaScript, Selenium is used to automatically control the browser to parse the rendered page, and combined with... or Achieve accurate information extraction. The crawled content includes the names, structures (such as SMILES and InChI, whose format must be consistent with the compound format in step S3), CAS, etc., of Chinese herbal compounds, and saves this data in .csv (comma-separated value) format to ensure that it can be directly used for large model training and screening of pharmacodynamic substances.
[0040] S5. Construct a standard prediction set: Standardize the traditional Chinese medicine compound dataset to construct a standard traditional Chinese medicine compound dataset. S51. Perform structural standardization processing on the traditional Chinese medicine compound dataset collected in step S4, including unifying SMILES / InChI representation, verifying the validity of chemical structures, removing invalid or incomplete structures, standardizing atom and bond types, standardizing stereochemical information, eliminating redundant isomers and duplicate records, and ensuring that the structure of each compound is unique and correct in the dataset. S52. A Python script is used to batch extract the SMILES structure and its label information as a standard dataset of traditional Chinese medicine compounds, which serves as input data for prediction by a deep neural network model. During the extraction process, regular expression matching is used to validate the standardized structure information fields (such as the SMILES string format); [Further details about the process are needed for accurate translation.] The database undergoes data cleaning and reconstruction, specifically including: deleting unstructured data, removing duplicate records, standardizing naming formats (such as standardizing capitalization and merging aliases), and removing components without clear biological function annotations. The final output is a dataset of traditional Chinese medicine compounds with a unified structure and clearly defined fields, which can be saved as... , or This format provides standardized input for subsequent model training and prediction tasks.
[0041] S6. Training the screening model: Input the standard model training dataset into the deep neural network model to train the model and obtain the screening model for the target efficacy and the pharmacological substances of the target target. S61. Perform automated feature extraction processing on the standard model training dataset in step S3; automatically learn the complex relationship between compound features and drug efficacy targets; the extracted features include molecular structure information, physicochemical properties, topological structure information, etc., and convert the compounds into numerical representations that can be used for model training through molecular fingerprints (such as ECFP, extended connection fingerprint), graph structure encoders (such as convolutional neural networks GCN), or other feature vector methods.
[0042] S62. Construct a deep neural network model, including an input layer, several hidden layers, and an output layer. The hidden layers can contain graph neural network layers, attention mechanism layers, and fully connected layers, used to automatically capture the complex nonlinear relationship between compound structural features and drug targets. Basic parameter settings are as follows: input dimension based on feature vector length; number of hidden layer nodes typically 128–512; ReLU activation function; output layer set to sigmoid (multi-label classification) or softmax (single-label classification) depending on the task type; batch size 32–128.
[0043] S63. Use DataLoader to implement batch loading and multi-threading acceleration, and divide the dataset into training, validation, and test sets. Training employs supervised learning, with the loss function automatically selected based on the task: Cross-Entropy Loss is used for classification tasks, and Mean Squared Error (MSE) is used for regression tasks. The optimizer is automatically matched (e.g., Adam or SGD), and learning rate scheduling (cosine annealing or dynamic adjustment) is supported. During training, metrics such as loss, accuracy, and AUC are automatically recorded, model checkpoints are saved, and an early stopping mechanism is triggered to prevent overfitting. Automatically matches optimizers (such as Adam and SGD) and supports learning rate scheduling strategies (such as cosine annealing and dynamic adjustment). Automated monitoring and saving are implemented during training: metrics such as loss, accuracy, and AUC (area under the curve) are automatically recorded; model checkpoints and log files are saved at each training epoch; and an early stopping mechanism is automatically triggered when overfitting or performance stagnation occurs.
[0044] S64. After training, the model performance is evaluated using the validation and test sets, and the AUC, F1-score, and confusion matrix are automatically output as evaluation metrics. The model structure or parameters can be dynamically adjusted based on prediction performance, ultimately forming a target pharmacodynamic substance screening model with high prediction accuracy and generalization ability. The AUC, F1 score, confusion matrix, and other evaluation metrics are automatically output via scikit-learn, and visualization charts (such as ROC curves, PR curves, and learning curves) are automatically generated and exported as PDF or PNG files.
[0045] S7. Predicting active ingredients: Input the standard Chinese herbal compound dataset into the active ingredient screening model to obtain the prediction results.
[0046] S71. Preprocess the standardized Chinese herbal compound dataset from step S5 using the same feature extraction method as the model input; all features are automatically stored as... Matrix or This format ensures that the input feature matrix perfectly matches the input interface of the training model.
[0047] S72. Input the feature matrix into the trained drug efficacy substance screening model. The model outputs the prediction results of each Chinese medicine compound on the target efficacy or target based on the learned drug efficacy-target-structure correlation law. S73. Screen and sort the prediction results, set thresholds based on activity scores or probability values, and extract highly active Chinese herbal compounds as candidate components, which are then used as candidate components for the target pharmacodynamic substances. S74. Compare and analyze the candidate pharmacodynamic substances with the target pharmacodynamic attributes, generate a visualization report or interactive display interface, and output the prediction screening result set.
[0048] By calling Python such as or Read the database automatically from authoritative databases (such as...) The script retrieves relevant pharmacodynamic literature, bioactivity data, or known target information from sources such as [list of sources], and performs intelligent comparative analysis using keyword matching, semantic recognition, and other methods. It then cross-validates the effectiveness and rationality of candidate compounds by combining model predictions with external database data, thereby achieving data-supported automated assisted judgment.
[0049] Ultimately, Python is used to automatically generate visual reports or interactive display interfaces, outputting a set of predicted screening results, thus forming a standardized and reusable process for identifying pharmacodynamic substances.
[0050] Example 2 See Figure 2 , Figure 2 This is a schematic diagram of an automated screening system for active ingredients in traditional Chinese medicine (TCM) based on a large model, as proposed in this invention. This system accurately screens for substances with specific medicinal effects in target TCM herbs or prescriptions based on target efficacy attributes, reducing the blind spots of traditional experiments and improving the efficiency and accuracy of discovering active ingredients in TCM, thus providing technical support for TCM research and development and pharmacological mechanism studies. Figure 2 As shown, the screening and prediction system for active substances in traditional Chinese medicine may specifically include: M1, Data Acquisition Module: Used to acquire drug efficacy-related target information based on the target's efficacy attributes, and to collect datasets of active compounds related to the target and datasets of Chinese medicinal compounds contained in the target Chinese medicinal materials or Chinese medicinal formulas, providing a basic dataset for large-scale model training. Specific functions include: M11, Database Interface Unit: Used to establish a connection with the Chinese medicine database through program scripts, and supports automated batch retrieval of compound structure, efficacy, and target information; M12, Keyword Retrieval Unit: Used to perform multi-database joint queries based on the input target attributes and automatically generate a preliminary candidate dataset.
[0051] M2, Structure Normalization Module: Used to normalize the target-related active compound dataset to construct a standard model training dataset, and to normalize the traditional Chinese medicine compound dataset to construct a standard traditional Chinese medicine compound dataset. Specific functions include: M21, Format Processing Unit: Used to call chemical information processing tools to perform unified format conversion on the captured compound structures; M22, Redundancy Removal Unit: Used to clean invalid, missing, duplicate, or structurally conflicting data.
[0052] M3, Model Training Module: Used to input the standard model training dataset into the deep neural network model for training, thereby obtaining a screening model for the target drug efficacy and the pharmacodynamic substances of the target target. Specific functions include: M31, Neural Network Structure Configuration Unit: Used to support users in selecting fully connected network, graph convolutional neural network, and graph attention network model structures; M32, Training Parameter Optimization Unit: Used to support hyperparameter setting and automatic search; M33, Training Monitoring Unit: Used to record loss values, accuracy, and AUC training metrics, and supports... With model save.
[0053] M4, Prediction and Inference Module: Used to input the standard Chinese herbal compound dataset into the pharmacodynamic substance screening model and obtain prediction results. Specific functions include: M41, Standard Data Import Unit: Supports batch import of standardized Chinese herbal compound data and completes large model prediction inference; M42, Activity Score Output Unit: Used to output the activity score or probability value of each Chinese medicine compound on a specific target or target efficacy attribute, and supports threshold screening and sorting.
[0054] M5, Feature Extraction Module: This module is used for automated feature extraction from the standardized compound dataset. Specific functions include: M51, Molecular fingerprint generation unit: used to convert the molecular structure of a compound into numerical fingerprint features using molecular fingerprints; M52, Graph Structure Encoding Unit: Used to extract the topological features of compounds through a graph structure encoder; M53, Physicochemical Property Calculation Unit: Used to calculate molecular weight and LogP basic properties, and to assist the model in learning the relationship between compound characteristics and drug targets.
[0055] M6. Result Display Module: Used to display the inference results output by the predictive inference module. Specific functions include: M61, Visualization Output Unit: Used to display the model prediction results of candidate pharmacodynamic substances in the form of charts, rankings, etc., and generate visualization reports or interactive display interfaces; M62, File Export Unit: Used to support output of results in PDF or Excel format.
[0056] In summary, the present invention has the following advantages: 1. By automating data collection and batch processing, traditional manual retrieval and experimentation are replaced, significantly reducing the time cost of data acquisition and preprocessing; 2. Construct a drug efficacy screening system that combines large models to automatically learn the complex relationship between compounds and targets, achieve intelligent prediction, and reduce the random screening behavior of traditional experiments; 3. By standardizing data processing to unify data formats and cleaning redundant information, combined with model training monitoring and evaluation, we can ensure data quality and prediction accuracy. 4. To provide systematic scientific methods and technical tools for the discovery of active ingredients and the study of pharmacological mechanisms of traditional Chinese medicine, and to promote the transformation of traditional Chinese medicine research and development from experience-driven to data-driven.
[0057] The above specific embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An automated screening method for active substances in traditional Chinese medicine based on a large model, characterized in that, include: S1. Determine the target attributes for drug efficacy screening, and obtain drug efficacy-related targets based on the target attributes using the semantic understanding and knowledge reasoning capabilities of the large model; S2. Based on the pharmacodynamic-related targets, call the large model-driven data acquisition module to automatically collect the target-related active compound dataset; S3. Perform automated structure and property standardization processing on the target-related active compound dataset to construct a standard model training dataset; S4. Identify the target Chinese medicinal materials or Chinese medicinal formulas, and use the knowledge enhancement function of the large model to automatically extract and integrate the compound dataset contained in the Chinese medicinal materials or formulas; S5. Standardize the dataset of Chinese herbal compounds to construct a standard dataset of Chinese herbal compounds. S6. Input the standard model training dataset into the deep neural network model to train the model and obtain the target efficacy and target target pharmacological substance screening model. S7. Input the standard Chinese herbal compound dataset into the pharmacodynamic substance screening model, combine the inference mechanism of the large model to obtain the prediction results, and output them in a visual format.
2. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 1, characterized in that, The pharmacodynamic targets were obtained from TCMSP, DrugBank, PubChem, and ChEMBL databases through automated literature mining and structured database retrieval, and the retrieval accuracy was improved by leveraging the named entity recognition and relation extraction capabilities of large models.
3. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 1, characterized in that, Step S2 specifically includes: Based on the pharmacological target, the automated data acquisition and structured extraction technology driven by a large model is used to obtain data on active compounds that have an active relationship with the pharmacological target from public databases. The active compound data are screened, and invalid or unqualified compound structures are eliminated using structural verification tools. Large-scale models are used for semantic deduplication and validity assessment of compounds, combined with structure verification tools to eliminate invalid or unacceptable compound structures; the screened compound structures are then used... The structure is converted into a unified standard format to verify its legality and to construct a dataset of target-related active compounds.
4. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 3, characterized in that, The data includes compound structure, target identifier, activity value, activity level, and structure-target association confidence labels generated by a large model.
5. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 1, characterized in that, Step S3 specifically includes: Based on the target-related active compound dataset, the target-related active compound dataset is subjected to structure normalization. The Python script is used in conjunction with the semantic parsing function of the large model to extract the SMILES structure and its label information in batches as input to the deep neural network model, and exported as CSV, JSON or NumPy format.
6. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 1, characterized in that, Step S4 specifically includes: Based on the preset efficacy target, the large model is invoked to screen Chinese medicinal materials or prescriptions that are highly correlated with the target efficacy from the knowledge base and database; the names, structures and CAS numbers of Chinese medicinal compounds are obtained and saved as standardized format files.
7. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 1, characterized in that, Step S5 specifically includes: The structure of the Chinese herbal medicine compound dataset is standardized, the SMILES format is unified, and redundant information is eliminated. The compound attributes are semantically labeled using a large model to generate a standard Chinese herbal medicine compound dataset.
8. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 1, characterized in that, Step S6 specifically includes: Utilize large models for automated feature engineering and feature importance analysis; construct a deep neural network model structure that combines knowledge enhancement from large models to automatically learn the complex relationship between compound features and drug targets; Automatically divides the training, validation and test sets, and suggests the selection of loss function and optimizer based on the large model. The training process is automatically monitored and saved, and an early stopping mechanism is automatically triggered when overfitting or performance stagnation occurs. After training, the model performance is evaluated using the validation and test sets, and the AUC, F1-score, and confusion matrix evaluation metrics are automatically output.
9. The automated screening method for active substances in traditional Chinese medicine based on a large model according to claim 1, characterized in that, Step S7 specifically includes: The standardized Chinese herbal compound dataset was preprocessed using a feature method consistent with the model input. The feature matrix is input into a drug efficacy screening model that combines large model knowledge enhancement, and the predicted results of each compound on the target efficacy or target are output. The prediction results are interpreted and visualized using a large model, highly active compounds are screened as candidate components, and an interactive display interface is generated.
10. An automated screening system for active substances in traditional Chinese medicine based on a large model, characterized in that, include: Data acquisition module: Combines large-scale model semantic retrieval with database crawling to obtain drug efficacy-related target and compound data; The structure standardization module performs standardization and semantic tagging on datasets of target-related active compounds and traditional Chinese medicine compounds. Model training module: Input the standard model training dataset into the deep neural network model that combines knowledge from large models for training; Predictive reasoning module: Uses the trained model to predict, interpret, and visualize Chinese herbal compounds.
11. The automated screening system for active substances in traditional Chinese medicine based on a large model according to claim 10, characterized in that, The data acquisition module includes: Database interface unit: Used to establish a connection with the database of Chinese medicine compounds, efficacy and target through program scripts, and automatically batch retrieve structure, efficacy and target information under the guidance of large model semantic parsing; Keyword retrieval unit: Used for joint queries across multiple databases based on target efficacy attributes and synonyms generated by a large model, automatically generating preliminary candidate datasets.
12. The automated screening system for active substances in traditional Chinese medicine based on a large model according to claim 10, characterized in that, The structural standardization module includes: Format processing unit: Used to call chemical information processing tools and perform unified format conversion on the captured compound structures according to the structure analysis rules generated by the large model; Redundant information removal unit: Used to combine the anomaly detection capabilities of rule engine and large model to clean invalid, missing, duplicate or structurally conflicting data.
13. The automated screening system for active substances in traditional Chinese medicine based on a large model according to claim 10, characterized in that, It also includes a feature extraction module, which is used to automatically extract features from the standardized compound dataset, specifically including: Molecular fingerprint generation unit: used to convert the molecular structure of a compound into numerical fingerprint features using molecular fingerprints; Graph structure coding unit: used to extract the topological features of compounds through a graph structure encoder; Physicochemical property calculation unit: used to calculate molecular weight and LogP basic properties, and to assist the model in learning the relationship between compound characteristics and drug targets.
14. The automated screening system for active substances in traditional Chinese medicine based on a large model according to claim 10, characterized in that, The model training module includes: Neural network architecture configuration unit: used to support users in selecting fully connected network, graph convolutional neural network, and graph attention network model architecture; Training parameter optimization unit: used to support hyperparameter setting and automatic search; Training monitoring unit: Used to record training metrics such as loss value, accuracy, and AUC, and supports early stopping and model checkpoint saving.
15. The automated screening system for active substances in traditional Chinese medicine based on a large model according to claim 10, characterized in that, The predictive reasoning module includes: Standard data import unit: used to import standardized Chinese herbal compound datasets and complete predictions under the guidance of large model inference; Activity score output unit: Used to output the activity score or probability value of each Chinese medicine compound on a specific target or target efficacy attribute, and supports threshold screening and sorting.
16. The automated screening system for active substances in traditional Chinese medicine based on a large model according to claim 10, characterized in that, It also includes a results display module, which is used to display the inference results output by the prediction inference module, specifically including: Visualization output unit: Used to display the model prediction results of candidate pharmacodynamic substances in the form of charts, rankings, etc., and generate visualization reports or interactive display interfaces; File Export Unit: Used to support outputting results in PDF or Excel format.