Intelligent extraction method and system for traditional Chinese medicine clinical scientific research knowledge

Automatically identify and analyze traditional Chinese medicine terms through natural language processing and deep learning technology, and construct a knowledge graph with graph database technology, solving the problem of low efficiency in extracting knowledge of traditional Chinese medicine clinical research, and achieving efficient knowledge extraction and potential law discovery.

CN120072336AInactive Publication Date: 2025-05-30ZHEJIANG CHINESE MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510136362.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively process unstructured or semi-structured data in the field of traditional Chinese medicine, resulting in low efficiency in extracting clinical research knowledge in traditional Chinese medicine and difficulty in discovering potential laws and correlations.

Method used

Natural language processing technology and deep learning algorithms are used to automatically identify and label traditional Chinese medicine terms, explore the relationship between terms through correlation analysis, build a knowledge graph in the field of traditional Chinese medicine, and use graph database technology to perform efficient storage and reasoning analysis.

Benefits of technology

It realizes intelligent extraction of clinical research knowledge of traditional Chinese medicine, significantly improves the efficiency of scientific researchers to extract effective information from a large amount of data, avoids subjective deviations and time costs in manual analysis, and supports the formulation of personalized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072336A_ABST
    Figure CN120072336A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data mining, in particular to a traditional Chinese medicine clinical scientific research knowledge intelligent extraction method and system, and the method comprises the following steps: S1, collecting traditional Chinese medicine clinical data through multiple channels; s2, cleaning, denoising and formatting the traditional Chinese medicine clinical data collected in the S1; s3, analyzing the traditional Chinese medicine clinical data processed in the S2, and automatically identifying and marking traditional Chinese medicine terms; s4, mining the relation between the terms through correlation analysis, and constructing a knowledge graph of the traditional Chinese medicine field; s5, scientific research knowledge is extracted from the processed traditional Chinese medicine clinical data, and a traditional Chinese medicine clinical research report or knowledge abstract is generated. By automatically applying natural language processing and deep learning technologies, traditional Chinese medicine terms are accurately extracted and associated from traditional Chinese medicine clinical data, the traditional Chinese medicine knowledge graph is effectively constructed, the scientific research efficiency is remarkably improved, and accurate knowledge support and decision basis are provided for traditional Chinese medicine clinical research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data mining technology, and in particular to a method and system for intelligently extracting clinical scientific research knowledge from traditional Chinese medicine. Background Art

[0002] With the continuous development of clinical research in Traditional Chinese Medicine (TCM), a large amount of TCM clinical data, medical literature, and scientific research papers have been generated and stored in various databases and information sources. TCM clinical data covers a wealth of patient medical history information, the relationship between symptoms and diseases, the effects of drugs and treatment methods, etc. Although these data have extremely high scientific research value, due to the complex data format, non-uniform information, and most of them are stored in unstructured or semi-structured form, it is difficult to use them directly for effective scientific research analysis. In addition, existing TCM clinical data mostly rely on manual analysis and experience summary, lacking systematic and automated processing methods, resulting in low knowledge extraction efficiency and difficulty in discovering potential patterns and associations.

[0003] While some existing methods for analyzing TCM data based on natural language processing (NLP) exist, most rely on manually designed rules and lack the ability to deeply learn TCM terminology and its relationships. Traditional knowledge graph construction methods are unable to effectively handle the complex relationships and terminology unique to TCM, and fail to fully utilize graph database technology for efficient storage and reasoning analysis of these relationships. Therefore, there is an urgent need for a method and system for intelligently extracting TCM clinical research knowledge to address these issues. Summary of the Invention

[0004] Based on the above objectives, the present invention provides a method and system for intelligent extraction of clinical scientific research knowledge of traditional Chinese medicine.

[0005] The intelligent extraction method of TCM clinical research knowledge includes the following steps:

[0006] S1: Collect TCM clinical data through various channels, including patient medical records, medical literature, and scientific research papers;

[0007] S2: Clean, denoise, and format the TCM clinical data collected in S1 to remove irrelevant information and standardize the data format;

[0008] S3: Use natural language processing technology to analyze the TCM clinical data processed by S2, automatically identifying and annotating TCM terms, including diseases, symptoms, drugs, and treatment methods;

[0009] S4: Based on the TCM terms annotated in S3, we use association analysis to explore the relationships between terms, including the relationship between diseases and symptoms, and between drugs and efficacy, and build a knowledge graph in the field of TCM.

[0010] S5: Based on the knowledge graph constructed in S4 and the mined term relationships, scientific research knowledge is extracted from the processed TCM clinical data, and a TCM clinical research report or knowledge summary is generated.

[0011] Optionally, the S1 specifically includes:

[0012] S11: Automatically extract medical record data containing patient basic information, medical history, symptoms, diagnosis, and treatment plan by accessing the electronic medical record management system of a hospital or clinic;

[0013] S12: Access medical literature databases, including CNKI or PubMed, to collect information on TCM theories, treatment methods, and drug applications from scientific research literature, clinical research reports, and medical monographs related to TCM;

[0014] S13: By accessing academic paper databases, including CNKI, Wanfang Data, or Web of Science, we collected scientific research paper data in the field of traditional Chinese medicine, including paper titles, abstracts, keywords, research methods, results analysis, and conclusions.

[0015] Optionally, the S2 specifically includes:

[0016] S21: Use data cleaning algorithms to process the collected TCM clinical data, delete invalid or duplicate data records, remove missing values ​​and data with incorrect formats; for records with missing values, use interpolation or data filling techniques to repair them to ensure data integrity;

[0017] S22: Use noise filtering technology to denoise the data. For text data, use a stop word removal algorithm based on word frequency statistics to remove irrelevant common words. For numerical data, use mean filtering or median filtering to remove outliers and noise data.

[0018] S23: Standardize the denoised data and convert data of different formats and scales into a standard format. For numerical data, standard deviation standardization or minimum-maximum standardization is used to ensure that the data are within the same range. For text data, a unified encoding format is used for text encoding.

[0019] S24: Reorganize and format the data, convert data from different sources into a unified database format, including SQL database or NoSQL database format, and store the processed TCM clinical data in a standard database for future use.

[0020] Optionally, the S3 specifically includes:

[0021] S31: Use a dictionary-based word segmentation algorithm to segment the pre-processed TCM clinical data, dividing the continuous text data into separate word units, ensuring that the word segmentation results are consistent with the standard dictionary of TCM terminology;

[0022] S32: Using a part-of-speech tagging tool to perform part-of-speech tagging on the text after word segmentation, identifying the part-of-speech category of each word;

[0023] S33: Apply named entity recognition models to automatically identify and classify entities related to diseases, symptoms, drugs, and treatments in TCM clinical data based on pre-trained TCM terminology recognition models;

[0024] S34: Annotate the identified TCM terms to generate structured data with term category labels.

[0025] Optionally, the S33 specifically includes:

[0026] S331: Select a model based on a bidirectional encoder representation converter and pre-train the selected model using a corpus containing annotated Chinese medicine terms. During the training process, the masked language model task is used to optimize the model parameters θ by minimizing the loss function L for predicting masked words. The expression is:

[0027] Among them, L is the loss function, N is the number of training samples, and w i is the i-th word, m is the masking window size, P is the conditional probability, and θ is the model parameter;

[0028] S332: Input the TCM clinical data after part-of-speech tagging in S32 into the pre-trained TCM term recognition model, generate a context representation vector for each word by encoding the input text, and use the conditional random field layer to predict the entity label of each word;

[0029] S333: Based on the entity labels predicted in S332, the entities are classified into four categories: diseases, symptoms, drugs, and treatment methods.

[0030] Optionally, the S4 specifically includes:

[0031] S41: Analyze the TCM terminology data identified and annotated in S3 to identify frequent co-occurrence patterns between terms. Specifically, the Apriori algorithm is used to scan the TCM clinical dataset to determine frequent item sets, and association rules are generated based on the frequent item sets to identify the associations between diseases and symptoms, and between drugs and efficacy terms.

[0032] S42: For the association rules mined in S41, the weight of each association relationship is calculated using the support and confidence indicators; the formula is: Weight(A→B)=Support(A→B)×Confidence(A→B), where Weight(A→B) is the weight of the association rule A→B, Support(A→B) is the support of the rule, and Confidence(A→B) is the confidence of the rule;

[0033] S43: Based on the weights of the association relationships calculated in S42, a knowledge graph in the field of traditional Chinese medicine is constructed using graph database technology. The knowledge graph consists of nodes and edges, where nodes represent traditional Chinese medicine terms and edges represent the association relationships between terms.

[0034] S44: Use node embedding algorithm to train the knowledge graph and generate vector representations of TCM terms.

[0035] Optionally, the S43 specifically includes:

[0036] S431: Create a corresponding node for each TCM term. The node attributes include the term name and term category. The specific steps include:

[0037] S4311: Term name nodeization, each TCM term identified in S42 is added to the graph database as a node;

[0038] S4312: Term category labeling: assign a category label to each node, whose value is disease, symptom, drug or treatment method, which is used to reflect the category attribute of the term;

[0039] S432: Based on the association weights calculated in S42, edges are created between the terms. The edge attributes include association type and weight value. The specific steps include:

[0040] S4321: Definition of association type, defining the type of association between terms, including concomitant, therapeutic, or impact;

[0041] S4322: assigning weight values, assigning the association weights calculated in S42 to the corresponding edges to indicate the strength of the associations;

[0042] S4323: Edge creation: Create an edge for each pair of related terms in the graph database, whose attributes include association type and weight value;

[0043] S433: Optimize the structure of the constructed knowledge graph; the specific steps include:

[0044] S4331: Create indexes for node attributes and edge attributes;

[0045] S4332: Partition the graph based on term categories and association types to reduce query path length;

[0046] S4333: Delete redundant edges in the graph.

[0047] Optionally, the S44 specifically includes:

[0048] S441: Select node embedding algorithm, including Node2Vec algorithm and TransE algorithm;

[0049] S442: Input the constructed TCM knowledge graph G = (V, E) into the selected node embedding algorithm, where V is the node set and E is the edge set; each node v∈V represents a TCM term, and each edge e∈E represents the association relationship between terms;

[0050] S443: Use the selected node embedding algorithm to learn the vector representation of each node v∈V Where d is the dimension of the vector; the node embedding algorithm maximizes the similarity between nodes by optimizing the objective function. The specific objective function is defined as: Among them, L is the loss function, σ is the Sigmoid function, z u and z v are the vector representations of node u and node v respectively, and T represents the transpose of the vector.

[0051] Optionally, the S5 specifically includes:

[0052] S51: Based on the knowledge graph of Traditional Chinese Medicine (TCM) constructed in S4, we use graph query technology to reason about the relationships between TCM terms. We use the graph database query language to query disease, symptom, drug, and treatment method nodes, and mine potential relationships in TCM clinical data.

[0053] S52: Based on the term association information retrieved in S51, combined with the TCM terminology annotated in S3 and the specific medical records in TCM clinical data, scientific research knowledge is extracted. The extracted content includes the combination of diseases and symptoms, the relationship between drugs and efficacy, and the effectiveness evaluation of treatment plans;

[0054] S53: Through preset templates, the extracted scientific research knowledge is organized into a complete TCM clinical research report or knowledge summary. The report content includes disease analysis, symptom description, drug efficacy evaluation and treatment recommendations.

[0055] The intelligent extraction system for TCM clinical research knowledge is used to implement the above-mentioned intelligent extraction method for TCM clinical research knowledge, and includes the following modules:

[0056] Data collection module: collects TCM clinical data through various channels, including patient medical records, medical literature and scientific research papers;

[0057] Data preprocessing module: used to clean, denoise and format the TCM clinical data collected by the data acquisition module to remove irrelevant information and standardize the data format;

[0058] Natural Language Processing Module: Utilizes natural language processing technology to analyze TCM clinical data processed by the data preprocessing module, automatically identifying and annotating TCM terms, including diseases, symptoms, drugs, and treatment methods;

[0059] Relationship mining module: Based on the TCM terms annotated by the natural language processing module, it conducts association analysis to mine the relationships between terms, including the relationship between diseases and symptoms, and between drugs and efficacy;

[0060] Knowledge graph construction module: Based on the term relationships generated by the relationship mining module, a knowledge graph in the field of traditional Chinese medicine is constructed using graph database technology;

[0061] Scientific research knowledge extraction module: Based on the knowledge graph, it automatically extracts scientific research knowledge from processed TCM clinical data and generates TCM clinical research reports or knowledge summaries.

[0062] Beneficial effects of the present invention:

[0063] The present invention, by adopting natural language processing technology and deep learning algorithms, can automatically analyze TCM clinical data, accurately identify and label TCM terms such as diseases, symptoms, drugs and treatment methods; by mining and analyzing the relationships between these terms, the present invention can effectively discover the potential laws and knowledge in TCM clinical data, and generate a knowledge graph in the field of TCM with high scientific research value; this automated knowledge extraction process significantly improves the efficiency of scientific researchers in extracting effective information from large amounts of data, avoiding the subjective bias and time cost in manual analysis.

[0064] The present invention uses graph database technology to efficiently store and query terms in the field of traditional Chinese medicine and their associations, making the generated traditional Chinese medicine clinical research reports or knowledge summaries more accurate and comprehensive; through association analysis, the relationship between drugs and efficacy, diseases and symptoms, etc. is mined. This invention can not only support the formulation of personalized treatment plans, but also provide important data support for traditional Chinese medicine clinical research. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0066] Figure 1 Schematic diagram of the knowledge intelligent extraction method according to an embodiment of the present invention;

[0067] Figure 2 Schematic diagram of a knowledge intelligent extraction system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0068] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It is also noted that, to provide a more detailed description, the following embodiments are best and preferred embodiments, and those skilled in the art may employ alternative methods for implementing certain known technologies. Furthermore, the accompanying drawings are intended only to provide a more detailed description of the embodiments and are not intended to limit the present invention.

[0069] It should be noted that references in the specification to "one embodiment," "an embodiment," "an exemplary embodiment," "some embodiments," etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not every embodiment necessarily includes such specific features, structures, or characteristics. In addition, when specific features, structures, or characteristics are described in conjunction with an embodiment, it is within the knowledge of persons skilled in the relevant art to implement such features, structures, or characteristics in conjunction with other embodiments (whether or not explicitly described).

[0070] In general, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending at least in part on the context, allow for the presence of other factors that are not necessarily explicitly described.

[0071] like Figure 1 As shown, the intelligent extraction method of TCM clinical research knowledge includes the following steps:

[0072] S1: Collect TCM clinical data through various channels, including patient medical records, medical literature, and scientific research papers;

[0073] S2: Clean, denoise, and format the TCM clinical data collected in S1 to remove irrelevant information and standardize the data format to ensure data availability and consistency;

[0074] S3: Use natural language processing technology to analyze the TCM clinical data processed by S2, automatically identifying and annotating TCM terms, including diseases, symptoms, drugs, and treatment methods;

[0075] S4: Based on the traditional Chinese medicine terms marked in S3, mine the relationships between terms through association analysis, including the associations between diseases and symptoms, and between drugs and curative effects, and construct a knowledge graph in the field of traditional Chinese medicine;

[0076] S5: Based on the knowledge graph constructed in S4 and the mined term relationships, extract valuable scientific research knowledge from the processed traditional Chinese medicine clinical data, and generate a traditional Chinese medicine clinical research report or knowledge summary.

[0077] S1 specifically includes:

[0078] S11: Automatically extract medical record data containing patients' basic information, medical history, symptoms, diagnosis, and treatment plan content by accessing the electronic medical record management system of hospitals or clinics;

[0079] S12: Collect information on traditional Chinese medicine theories, treatment methods, and drug applications in scientific research literature, clinical research reports, and medical monographs in the field of traditional Chinese medicine by accessing medical literature databases, including CNKI or PubMed;

[0080] S13: Collect scientific research paper data in the field of traditional Chinese medicine by accessing academic paper databases, including CNKI, Wanfang Data, or Web of Science, including paper titles, abstracts, keywords, research methods, result analysis, and conclusions, covering traditional Chinese medicine clinical research and curative effect analysis; The specific steps for collecting traditional Chinese medicine clinical data through the electronic medical record system, medical literature database, and scientific research paper database are described, and the technical means for each step are refined to ensure the accuracy and integrity of data collection, providing a reliable data basis for subsequent extraction of traditional Chinese medicine clinical scientific research knowledge.

[0081] S2 specifically includes:

[0082] S21: Process the collected traditional Chinese medicine clinical data using data cleaning algorithms, delete invalid or duplicate data records, remove missing values and data with format errors; for records with missing values, use interpolation or data filling techniques for repair to ensure data integrity;

[0083] S22: Perform denoising processing on the data using noise filtering techniques; for text data, use a stop word removal algorithm based on word frequency statistics to剔除 irrelevant common words (such as "de", "shi", etc.); for numerical data, use mean filtering or median filtering methods to remove outliers and noise data;

[0084] S23: Standardize the denoised data and convert data of different formats and scales into a standard format. For numerical data, standard deviation standardization or minimum-maximum standardization is used to ensure that the data are within the same range. For text data, a unified encoding format (such as UTF-8) is used for text encoding.

[0085] S24: Reorganize and format the data, convert data from different sources into a unified database format, including SQL database or NoSQL database format, and store the processed TCM clinical data in a standard database for backup; the above steps ensure the high quality and consistency of the collected TCM clinical data through data cleaning, denoising, standardization and formatting, providing a reliable data foundation for subsequent TCM terminology recognition and knowledge extraction.

[0086] S3 specifically includes:

[0087] S31: Use a dictionary-based word segmentation algorithm to segment the pre-processed TCM clinical data, dividing the continuous text data into separate word units, ensuring that the word segmentation results are consistent with the standard dictionary of TCM terminology;

[0088] S32: Use part-of-speech tagging tools to tag the text after word segmentation, identify the part-of-speech category of each word, and assist in subsequent entity recognition;

[0089] S33: Apply named entity recognition (NER) models to automatically identify and classify entities of diseases, symptoms, drugs, and treatments in TCM clinical data based on pre-trained TCM terminology recognition models;

[0090] S34: Label the identified TCM terms and generate structured data with term category labels to facilitate subsequent knowledge relationship mining.

[0091] S33 specifically includes:

[0092] S331: A model based on Bidirectional Encoder Representations from Transformers (BERT) is selected and pre-trained using a corpus containing annotated TCM terms to enhance the model's understanding of TCM terminology. During training, a Masked Language Model task is used to optimize the model parameters θ by minimizing the loss function L for predicting masked words, expressed as:

[0093] Among them, L is the loss function, N is the number of training samples, and w iis the i-th word, m is the masking window size, P is the conditional probability, and θ is the model parameter;

[0094] S332: Input the TCM clinical data after part-of-speech tagging in S32 into the pre-trained TCM term recognition model, and generate the context representation vector v for each word by encoding the input text. i , and use the conditional random field layer to assign entity labels y to each word i The prediction of entity recognition probability is expressed as: Among them, Y is the entity label sequence, X is the input sequence, and f k is the characteristic function, λ k is the weight parameter, N is the sequence length, and K is the number of feature functions;

[0095] S333: Based on the entity labels predicted in S332, the entities are classified into four categories: disease, symptom, drug, and treatment method. The specific classification is achieved by maximizing the probability that each entity belongs to a certain category. The classification function is defined as: C(e i )=argmax c∈{疾病,症状,药物,治疗方法} P(c|e i ), where C(e i ) is entity e i The classification result, c is the potential category, P(c|e i ) is entity e i The conditional probability of belonging to category c; through the above steps, TCM terms can be accurately identified and classified, ensuring the semantic accuracy and structuring of TCM clinical data. By using the pre-trained natural language processing model and conditional random field algorithm, the accuracy of entity recognition and classification is significantly improved, providing a reliable data foundation for subsequent knowledge relationship mining and scientific research knowledge extraction, and enhancing the intelligent level of TCM clinical research.

[0096] S4 specifically includes:

[0097] S41: Analyze the TCM terminology data identified and annotated in S3 to identify frequent co-occurrence patterns between terms. Specifically, the Apriori algorithm is used to scan the TCM clinical data set to determine frequent item sets, and generate association rules based on the frequent item sets to identify the associations between diseases and symptoms, and drugs and efficacy terms. First, the Apriori algorithm iteratively expands the frequent item sets, calculates the support (A) of all individual items, and selects frequent item sets whose support is not less than a predetermined threshold σ. Then, these frequent item sets are used to generate association rules, ensuring that the support and confidence of each rule are not less than a predetermined threshold γ. The formula is as follows:

[0098]

[0099] Among them, A and B are TCM term sets, σ ​​is the support threshold, and γ is the confidence threshold;

[0100] S42: For the association rules mined in S41, the weight of each association relationship is calculated using the support and confidence indicators; the formula is: Weight(A→B)=Support(A→B)×Confidence(A→B), where Weight(A→B) is the weight of the association rule A→B, Support(A→B) is the support of the rule, and Confidence(A→B) is the confidence of the rule;

[0101] S43: Based on the weights of the association relationships calculated in S42, a knowledge graph in the field of traditional Chinese medicine is constructed using graph database technology. The knowledge graph consists of nodes and edges, where nodes represent traditional Chinese medicine terms (e.g., disease, symptom, drug, treatment method), and edges represent associations between terms (e.g., "accompanying," "treatment," "influence," etc.).

[0102] S44: A node embedding algorithm is used to train the knowledge graph, generate vector representations of TCM terms, and enhance the model's ability to understand and express the complex relationships between terms. Through the above steps, the associations between TCM terms can be systematically mined to construct an accurate and comprehensive TCM knowledge graph. Association rule mining and graph database technology are used to effectively identify key relationships such as diseases and symptoms, drugs and efficacy, and enhance the expressive power and application value of the knowledge graph. The optimization of the knowledge graph ensures the scientific nature and practicality of the model, and provides a solid data foundation for subsequent scientific research knowledge extraction.

[0103] S43 specifically includes:

[0104] S431: Create a corresponding node for each TCM term. The node attributes include the term name and term category. The specific steps include:

[0105] S4311: Term name nodeization, each TCM term identified in S42 is added to the graph database as a node;

[0106] S4312: Term category labeling: assign a category label to each node, whose value is disease, symptom, drug or treatment method, which is used to reflect the category attribute of the term;

[0107] S432: Based on the association weights calculated in S42, edges are created between the terms. The edge attributes include association type and weight value. The specific steps include:

[0108] S4321: Definition of association type, defining the type of association between terms, including concomitant, therapeutic, or impact;

[0109] S4322: assigning weight values, assigning the association weights calculated in S42 to the corresponding edges to indicate the strength of the associations;

[0110] S4323: Edge creation: Create an edge for each pair of related terms in the graph database, whose attributes include association type and weight value;

[0111] S433: Optimize the structure of the constructed knowledge graph to improve query efficiency and data association accuracy. Specific steps include:

[0112] S4331: Create indexes for node and edge attributes to speed up queries;

[0113] S4332: Partition the graph based on term categories and association types to optimize the graph structure layout and reduce the query path length;

[0114] S4333: Delete redundant edges in the graph to ensure the simplicity and effectiveness of the knowledge graph; through the above steps, a knowledge graph in the field of traditional Chinese medicine can be systematically constructed to accurately reflect the relationship between traditional Chinese medicine terms, providing a solid data foundation for subsequent scientific research knowledge extraction and clinical research.

[0115] S44 specifically includes:

[0116] S441: Select a node embedding algorithm suitable for the TCM terminology knowledge graph, including the Node2Vec algorithm and the TransE algorithm. The selected algorithm must be able to effectively capture the semantic relationships and structural information of the nodes in the knowledge graph.

[0117] S442: Input the constructed TCM knowledge graph G = (V, E) into the selected node embedding algorithm, where V is the node set and E is the edge set; each node v∈V represents a TCM term, and each edge e∈E represents the association relationship between terms;

[0118] S443: Use the selected node embedding algorithm to learn the vector representation of each node v∈V Where d is the dimension of the vector; the node embedding algorithm maximizes the similarity between nodes by optimizing the objective function. The specific objective function is defined as: Among them, L is the loss function, σ is the Sigmoid function, z u and z v are the vector representations of node u and node v respectively, and T represents the transpose of the vector; according to the optimized model parameters, the final vector representation z of all nodes v∈V is generated vThese vectors are used to represent the semantic and structural features of TCM terms, which facilitates subsequent knowledge extraction and analysis tasks. Through the above steps, TCM terms can be effectively converted into high-dimensional vector representations, preserving the semantic and structural relationships between terms.

[0119] S5 specifically includes:

[0120] S51: Based on the TCM knowledge graph constructed in S4, graph query technology is used to reason about the relationships between TCM terms. Graph database query languages ​​(such as Cypher or SPARQL) are used to query disease, symptom, drug, and treatment method nodes, thereby mining potential relationships in TCM clinical data. Specifically, queries involve the association between diseases and symptoms, and the relationship between drugs and efficacy.

[0121] S52: Based on the terminology information retrieved in S51, combined with the TCM terminology annotated in S3 and the specific medical records in TCM clinical data, valuable scientific research knowledge is extracted. The extracted content includes the combination of diseases and symptoms, the relationship between drugs and efficacy, and the effectiveness evaluation of treatment plans;

[0122] S53: Through preset templates, the extracted scientific research knowledge is organized into a complete TCM clinical research report or knowledge summary. The report content includes disease analysis, symptom description, drug efficacy evaluation and treatment recommendations, which is convenient for reference by scientific researchers and clinicians, improves scientific research efficiency and reduces manual intervention.

[0123] like Figure 2 As shown, the intelligent extraction system of TCM clinical research knowledge is used to implement the above-mentioned intelligent extraction method of TCM clinical research knowledge, and includes the following modules:

[0124] Data collection module: collects TCM clinical data through various channels, including patient medical records, medical literature and scientific research papers;

[0125] Data preprocessing module: used to clean, denoise and format the TCM clinical data collected by the data acquisition module to remove irrelevant information and standardize the data format;

[0126] Natural Language Processing Module: Utilizes natural language processing technology to analyze TCM clinical data processed by the data preprocessing module, automatically identifying and annotating TCM terms, including diseases, symptoms, drugs, and treatment methods;

[0127] Relationship mining module: Based on the TCM terms annotated by the natural language processing module, it conducts association analysis to mine the relationships between terms, including the relationship between diseases and symptoms, and between drugs and efficacy;

[0128] Knowledge graph construction module: Based on the term relationships generated by the relationship mining module, a knowledge graph in the field of traditional Chinese medicine is constructed using graph database technology;

[0129] Scientific research knowledge extraction module: Based on the knowledge graph, it automatically extracts scientific research knowledge from processed TCM clinical data and generates TCM clinical research reports or knowledge summaries.

[0130] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.

[0131] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. An intelligent extraction method for clinical scientific research knowledge of traditional Chinese medicine, characterized in that: The following steps are involved: S1: Collect TCM clinical data through various channels, including patient medical records, medical literature, and scientific research papers; S2: Clean, denoise and format the TCM clinical data collected in S1 to remove irrelevant information and standardize the data format; S3: Use natural language processing technology to analyze the TCM clinical data processed by S2, automatically identify and annotate TCM terms, including diseases, symptoms, drugs, and treatment methods; S4: Based on the TCM terms annotated in S3, the relationships between terms are mined through association analysis, including the relationships between diseases and symptoms, and between drugs and efficacy, and a knowledge graph in the field of TCM is constructed; S5: Based on the knowledge graph constructed in S4 and the mined terminology relationships, scientific research knowledge is extracted from the processed TCM clinical data, and a TCM clinical research report or knowledge summary is generated.

2. The intelligent extraction method of clinical scientific research knowledge of traditional Chinese medicine according to claim 1 is characterized in that: The S1 specifically includes: S11: Automatically extract medical record data containing basic patient information, medical history, symptoms, diagnosis, and treatment plan by accessing the electronic medical record management system of a hospital or clinic; S12: By accessing medical literature databases, including CNKI or PubMed, collect information on TCM theories, treatment methods, and drug applications from scientific research literature, clinical research reports, and medical monographs related to TCM; S13: By accessing academic paper databases, including CNKI, Wanfang Data, or Web of Science, we collected scientific research paper data in the field of traditional Chinese medicine, including paper titles, abstracts, keywords, research methods, result analysis, and conclusions.

3. The intelligent extraction method of TCM clinical research knowledge according to claim 1 is characterized in that: The S2 specifically includes: S21: Use data cleaning algorithms to process the collected TCM clinical data, delete invalid or duplicate data records, and remove missing values ​​and data with incorrect formats; for records with missing values, use interpolation or data filling technology to repair them to ensure data integrity; S22: Use noise filtering technology to denoise the data; for text data, use a stop word removal algorithm based on word frequency statistics to remove irrelevant common words; for numerical data, use mean filtering or median filtering to remove outliers and noise data; S23: Standardize the denoised data and convert data of different formats and scales into a standard format. For numerical data, standard deviation standardization or minimum-maximum standardization method is used to ensure that the data is within the same range. For text data, a unified encoding format is used for text encoding. S24: Reorganize and format the data, convert data from different sources into a unified database format, including SQL database or NoSQL database format, and store the processed TCM clinical data in a standard database for backup.

4. The intelligent extraction method of TCM clinical research knowledge according to claim 1 is characterized in that: The S3 specifically includes: S31: Use a dictionary-based word segmentation algorithm to segment the preprocessed TCM clinical data, dividing the continuous text data into separate word units to ensure that the word segmentation results are consistent with the standard dictionary of TCM terminology; S32: using a part-of-speech tagging tool to perform part-of-speech tagging on the text after word segmentation, and identifying the part-of-speech category of each word; S33: Apply named entity recognition model to automatically identify and classify entities of diseases, symptoms, drugs and treatment methods in TCM clinical data based on pre-trained TCM terminology recognition model; S34: Annotate the identified TCM terms to generate structured data with term category labels.

5. The intelligent extraction method of clinical scientific research knowledge of traditional Chinese medicine according to claim 4 is characterized in that: The S33 specifically includes: S331: Select a model based on a bidirectional encoder representation converter, and use a corpus containing TCM terminology annotations to pre-train the selected model. During the training process, the masked language model task is used to optimize the model parameter θ by minimizing the loss function L of predicting masked words. The expression is: Among them, L is the loss function, N is the number of training samples, and w i is the i-th word, m is the masking window size, P is the conditional probability, and θ is the model parameter; S332: Input the TCM clinical data after part-of-speech tagging in S32 into a pre-trained TCM term recognition model, generate a context representation vector for each word by encoding the input text, and predict the entity label of each word using a conditional random field layer; S333: Based on the entity labels predicted in S332, the entities are classified into four categories: diseases, symptoms, drugs, and treatment methods.

6. The intelligent extraction method of TCM clinical research knowledge according to claim 1 is characterized in that: The S4 specifically includes: S41: Analyze the TCM terminology data identified and annotated in S3 to identify the frequent co-occurrence patterns between terms; specifically, use the Apriori algorithm to scan the TCM clinical data set, determine frequent item sets, and generate association rules based on frequent item sets to identify the association between diseases and symptoms, and drugs and efficacy terms; S42: For the association rules mined in S41, the weight of each association relationship is calculated using the indicators of support and confidence; the formula is: Weight(A→B)=Support(A→B)×Confidence(A→B), where Weight(A→B) is the weight of the association rule A→B, Support(A→B) is the support of the rule, and Confidence(A→B) is the confidence of the rule; S43: Based on the weight of the association relationship calculated in S42, a knowledge graph in the field of traditional Chinese medicine is constructed using graph database technology; the knowledge graph is composed of nodes and edges, the nodes represent traditional Chinese medicine terms, and the edges represent the association relationship between the terms; S44: Use node embedding algorithm to train the knowledge graph and generate vector representations of TCM terms.

7. The intelligent extraction method of TCM clinical research knowledge according to claim 6 is characterized in that: The S43 specifically includes: S431: Create a corresponding node for each TCM term, where the node attributes include the term name and term category. The specific steps include: S4311: Nodeization of term names, adding each TCM term identified in S42 as a node to the graph database; S4312: Term category labeling, assigning a category label to each node, whose value is disease, symptom, drug or treatment method, to reflect the category attribute of the term; S432: creating edges between terms according to the association relationship weights calculated in S42, where the edge attributes include association types and weight values; the specific steps include: S4321: Definition of association type, defining the type of association between terms, including concomitant, therapeutic, or impact; S4322: assigning a weight value, assigning the association relationship weight calculated in S42 to the corresponding edge to indicate the strength of the association relationship; S4323: edge creation, creating an edge for each pair of related terms in the graph database, whose attributes include association type and weight value; S433: Structural optimization of the constructed knowledge graph; the specific steps include: S4331: Create indexes for node attributes and edge attributes; S4332: Partition the graph according to term categories and association types to reduce query path length; S4333: Delete redundant edges in the graph.

8. The intelligent extraction method of TCM clinical research knowledge according to claim 7 is characterized in that: The S44 specifically includes: S441: Select a node embedding algorithm, including the Node2Vec algorithm and the TransE algorithm; S442: Input the constructed TCM knowledge graph G = (V, E) into the selected node embedding algorithm, where V is a node set and E is an edge set; each node v∈V represents a TCM term, and each edge e∈E represents an association relationship between terms; S443: Use the selected node embedding algorithm to learn the vector representation of each node v∈V Where d is the dimension of the vector; the node embedding algorithm maximizes the similarity between nodes by optimizing the objective function. The specific objective function is defined as: Among them, L is the loss function, σ is the Sigmoid function, z u and z v are the vector representations of node u and node v respectively, and T represents the transpose of the vector.

9. The intelligent extraction method of TCM clinical research knowledge according to claim 1 is characterized in that: The S5 specifically includes: S51: Based on the knowledge graph of TCM constructed by S4, the graph query technology is used to infer the relationship between TCM terms. Through the graph database query language, the nodes of diseases, symptoms, drugs, and treatment methods are queried to mine the potential relationships in TCM clinical data. S52: Based on the term association information queried in S51, combined with the TCM terms annotated in S3 and the specific medical records in TCM clinical data, scientific research knowledge is extracted. The extracted content includes the combination of diseases and symptoms, the relationship between drugs and efficacy, and the effect evaluation of treatment plans; S53: Through preset templates, the extracted scientific research knowledge is organized into a complete TCM clinical research report or knowledge summary. The report content includes disease analysis, symptom description, drug efficacy evaluation and treatment recommendations.

10. A system for intelligent extraction of clinical scientific research knowledge of traditional Chinese medicine, used to implement the method for intelligent extraction of clinical scientific research knowledge of traditional Chinese medicine as claimed in any one of claims 1 to 9, characterized in that: Includes the following modules: Data collection module: collect TCM clinical data through various channels, including patient medical records, medical literature and scientific research papers; Data preprocessing module: used to clean, denoise and format the TCM clinical data collected by the data acquisition module to remove irrelevant information and standardize the data format; Natural language processing module: Use natural language processing technology to analyze the TCM clinical data processed by the data preprocessing module, automatically identify and annotate TCM terms, including diseases, symptoms, drugs and treatment methods; Relationship mining module: Based on the TCM terms annotated by the natural language processing module, the module conducts association analysis to mine the relationships between terms, including the relationships between diseases and symptoms, and between drugs and efficacy; Knowledge graph construction module: based on the terminology relationships generated by the relationship mining module, a knowledge graph in the field of traditional Chinese medicine is constructed using graph database technology; Scientific research knowledge extraction module: Based on the knowledge graph, it automatically extracts scientific research knowledge from the processed TCM clinical data and generates TCM clinical research reports or knowledge summaries.