Construction method, system and device of credible traditional Chinese medicine evidence-based diagnosis and treatment agent

By constructing a large language model in the field of traditional Chinese medicine and a multi-dimensional diagnosis and treatment quality assessment strategy, and optimizing the intelligent agent for evidence-based diagnosis and treatment in traditional Chinese medicine, the problems of low accuracy and uninterpretability in the process of syndrome differentiation and treatment of existing systems have been solved, and high precision and reliability of traditional Chinese medicine-assisted diagnosis and treatment have been achieved.

CN120878092APending Publication Date: 2025-10-31UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510998861.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing TCM auxiliary diagnosis and treatment systems suffer from problems such as limited diagnostic capabilities of deep learning models, poor generalization ability of BERT-like models, and insufficient clinical reliability of large language models, making them difficult to apply in real TCM diagnosis and treatment scenarios. In particular, they have low accuracy and are uninterpretable in the process of syndrome differentiation and treatment.

Method used

By collecting textual corpora in the field of traditional Chinese medicine, using large language models for preprocessing and knowledge processing, we construct pre-training data and evidence-based diagnosis and treatment knowledge bases in the field of traditional Chinese medicine. Combining multi-dimensional diagnosis and treatment quality assessment strategies and online reinforcement learning methods, we optimize the intelligent agent for evidence-based diagnosis and treatment in traditional Chinese medicine.

Benefits of technology

It improves the accuracy and reliability of TCM evidence-based diagnosis and treatment intelligence, and can provide clinically feasible treatment plans, solving the problems of insufficient syndrome differentiation ability and poor interpretability of existing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878092A_ABST
    Figure CN120878092A_ABST
Patent Text Reader

Abstract

The invention provides a credible traditional Chinese medicine evidence-based diagnosis and treatment intelligent agent construction method, system and device, and relates to the technical field of traditional Chinese medicine evidence-based diagnosis and treatment intelligent agents.The method comprises the steps that text corpora in the traditional Chinese medicine field are collected and preprocessed through a natural language processing method; based on a large language model, performing knowledge processing on the text corpus; based on the pre-training data in the traditional Chinese medicine field, completing pre-training on the base large language model to obtain a large language model in the traditional Chinese medicine field; based on the traditional Chinese medicine field evidence-based diagnosis and treatment data, performing supervision fine tuning on the traditional Chinese medicine field big language model, and constructing a traditional Chinese medicine evidence-based diagnosis and treatment agent; based on a multi-dimensional diagnosis and treatment quality evaluation strategy, through an online reinforcement learning method, a traditional Chinese medicine evidence-based diagnosis and treatment agent is continuously optimized, and a credible traditional Chinese medicine evidence-based diagnosis and treatment agent is constructed. According to the scheme, the accuracy, the reliability and the continuous optimization capability of the traditional Chinese medicine evidence-based diagnosis and treatment intelligent agent can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of TCM evidence-based diagnosis and treatment intelligent agent technology, and in particular to a method, system and device for constructing a reliable TCM evidence-based diagnosis and treatment intelligent agent. Background Technology

[0002] With the rapid development of artificial intelligence technology, solving the problem of reliable evidence-based diagnosis and treatment decision-making in traditional Chinese medicine (TCM) has gradually become a core issue in the field of TCM evidence-based diagnosis and treatment intelligent agents. However, most existing TCM-assisted diagnosis and treatment systems suffer from problems such as poor data quality and limited knowledge base, as well as poor reliability of the reasoning process, making them difficult to apply in real-world TCM diagnosis and treatment scenarios. Specific problems include:

[0003] Problem 1: Deep learning models have limited diagnostic capabilities and are not interpretable. TCM auxiliary diagnosis and treatment systems based on deep learning generally suffer from insufficient model capacity, with the number of parameters typically in the millions, making it difficult to fully learn the complex nonlinear relationships of TCM syndrome differentiation and treatment. In clinical applications, these systems exhibit insufficient diagnostic accuracy, especially for complex syndromes such as concurrent syndromes and compound syndromes, where the accuracy rate drops significantly. More importantly, their "black box" nature means they cannot provide the diagnostic reasoning process, making it difficult for doctors to understand and verify the system's diagnostic conclusions, which severely restricts clinical acceptance.

[0004] Problem 2: BERT-based models have poor generalization ability and are limited to certain diseases. Although BERT-based models perform reasonably well in closed tests for specific diseases, they are essentially classification models and require a predefined fixed syndrome system. This makes it impossible for the system to handle the dynamic evolution of common syndromes in TCM clinical practice and novel combinations of syndromes that have not been seen before. In addition, general BERT has obvious biases in learning the representation of TCM professional terms, and its semantic capture of characteristic terms such as "pulse is wiry and slippery" is inaccurate, affecting the precision of syndrome differentiation. The system also has difficulty supporting the multi-round consultation process necessary for TCM diagnosis and treatment.

[0005] Problem 3: Large language models lack clinical reliability and deviate from actual needs; existing LLM (Large Language Model) driven TCM systems generally suffer from three major defects:

[0006] First, the problem of hallucinations is serious; large language models can generate false syndromes that do not conform to traditional Chinese medicine theory or clinical practice.

[0007] Second, there is poor consistency in decision-making; the same symptom input may lead to different diagnostic results due to slight adjustments in the prompt text.

[0008] Third, there is a functional positioning bias. Large language models are better at answering questions about TCM knowledge than providing actual diagnostic and treatment decision support, and cannot provide specific treatment plans that are clinically feasible.

[0009] These defects make it difficult for the system to meet the reliability requirements of real-world clinical scenarios.

[0010] Therefore, how to develop highly reliable and accurate TCM evidence-based diagnosis and treatment models, so as to fully integrate computer technology with TCM syndrome differentiation and treatment theory, and apply them to actual clinical applications, has become an important research direction in the industry. Summary of the Invention

[0011] The purpose of this invention is to provide a method, system, and device for constructing a reliable TCM evidence-based diagnostic and treatment intelligent agent, so as to solve at least one of the above-mentioned technical problems existing in the prior art.

[0012] Firstly, to address the aforementioned technical problems, this invention provides a method for constructing a reliable TCM evidence-based diagnostic and treatment intelligent agent, comprising the following steps:

[0013] Step 1: Collect text corpora in the field of Traditional Chinese Medicine and preprocess them using natural language processing methods;

[0014] Step 2: Based on the large language model, perform knowledge processing on the text corpus to obtain pre-training data in the field of traditional Chinese medicine, evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine, and evidence-based diagnosis and treatment data in the field of traditional Chinese medicine.

[0015] Step 3: Based on the pre-training data in the field of traditional Chinese medicine, complete the pre-training on the base large language model to obtain the large language model in the field of traditional Chinese medicine;

[0016] Step 4: Based on evidence-based diagnosis and treatment data in the field of traditional Chinese medicine, supervised fine-tuning is performed on the large language model in the field of traditional Chinese medicine to construct a traditional Chinese medicine evidence-based diagnosis and treatment agent;

[0017] Step 5: Based on a multidimensional diagnosis and treatment quality assessment strategy, continuously optimize the TCM evidence-based diagnosis and treatment agent through online reinforcement learning methods, and build a trustworthy TCM evidence-based diagnosis and treatment intelligent agent.

[0018] In this way, a trustworthy TCM evidence-based diagnosis and treatment intelligent agent driven by a large language model can be constructed, which makes full use of the characteristics of large language models, such as good interpretability and strong generalization ability, while avoiding the problems of illusion and poor decision consistency, and can provide specific treatment plans with clinical operability.

[0019] In one feasible implementation, the text corpus includes standard guidelines, books and documents, web resources, clinical data, etc., and may involve various text types such as PDF, TXT, and EXCEL.

[0020] The aforementioned standard guidelines were downloaded from the National Standards Information Public Service Platform, the State Administration of Traditional Chinese Medicine platform, etc., and cover sub-fields of traditional Chinese medicine such as syndrome differentiation, treatment, and health preservation.

[0021] The books and documents mentioned were collected from TCM university libraries, TCM paper databases, and TCM laboratories, including diagnostic and treatment guidelines, TCM textbooks, TCM papers, TCM patents, and TCM classics.

[0022] The web page resources mentioned are open-source knowledge downloaded from web pages using web crawling technology, including TCM text corpora, TCM theoretical knowledge, etc.

[0023] The clinical data was collected from hospitals, renowned TCM doctors, and typical TCM medical case collections, including TCM clinical electronic medical records and TCM prescription data.

[0024] This ensures the accuracy, comprehensiveness, and authority of knowledge in the field of traditional Chinese medicine.

[0025] In one feasible implementation, the specific method of preprocessing includes:

[0026] Step a1: Based on the OCR (Optical Character Recognition) method, perform character recognition on PDF-type text corpora; based on Python tools, process non-TXT-type text corpora to construct a TXT-type TCM-related text corpus.

[0027] Step a2: Based on heuristic rules, perform data cleaning on the TCM text corpus to remove common problems such as character errors, garbled formats, and invalid information caused by OCR technology and Python tools, thereby improving the quality of the corpus.

[0028] In one feasible implementation, step 2 specifically includes:

[0029] Step 21: Based on the Locality Sensitive Hash algorithm, after deduplication of the text corpus, and then based on the large language model, improve the data quality of the text corpus, and then construct pre-training data for the field of traditional Chinese medicine.

[0030] Step 22: Construct an evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine by sequentially performing text semantic segmentation, text block vectorization, and multi-level indexing;

[0031] Step 23: Construct evidence-based diagnosis and treatment data in the field of traditional Chinese medicine through data quality classification and prompting engineering.

[0032] In one feasible implementation, step 3 specifically includes:

[0033] Step 31: Mix the pre-training data in the TCM domain with the pre-training data in the general domain according to a preset ratio to obtain mixed pre-training data;

[0034] Step 32: Based on the mixed pre-trained data, perform unsupervised training on the large language model to obtain a large language model for the field of traditional Chinese medicine;

[0035] In this way, by leveraging the superior Chinese long text comprehension capabilities of large language models, an ideal large language model for the field of traditional Chinese medicine can be constructed.

[0036] In one feasible implementation, step 4 specifically includes:

[0037] Step 41: Convert the format of the TCM evidence-based diagnosis and treatment data to obtain the TCM evidence-based diagnosis and treatment proxy dataset;

[0038] Step 42: Based on the TCM evidence-based diagnosis and treatment agent dataset, train a large language model for the TCM domain using the cross-entropy loss function to obtain the TCM evidence-based diagnosis and treatment agent.

[0039] In one feasible implementation, the multidimensional diagnostic and treatment quality assessment strategy includes accuracy indicators, reliability indicators, and safety indicators.

[0040] The accuracy index is used to evaluate the similarity between the diagnostic decision-making results of TCM evidence-based diagnosis and treatment agency and the actual syndrome type;

[0041] The credibility index is used to evaluate the reliability of TCM evidence-based diagnosis and treatment agents in citing diagnosis and treatment knowledge during the evidence-based reasoning process.

[0042] The aforementioned safety indicators are used to evaluate whether the treatment principles and prescriptions of TCM evidence-based diagnosis and treatment are applicable to actual situations.

[0043] In one feasible implementation, the online reinforcement learning method specifically includes:

[0044] Step b1: Based on the multidimensional diagnosis and treatment quality assessment strategy, calculate the reward score between the predicted results and the actual results of TCM evidence-based diagnosis and treatment proxy.

[0045] Step b2: The TCM evidence-based diagnosis and treatment agent generates prediction results in real time during online reinforcement learning;

[0046] Step b3: Based on the reward score, update the model parameters of the TCM evidence-based diagnosis and treatment agent in real time, thereby optimizing the evidence-based diagnosis and treatment capabilities.

[0047] Secondly, based on the same inventive concept, this application also provides a system for constructing a trustworthy TCM evidence-based diagnosis and treatment intelligent agent, including a data acquisition module, a data processing module and a result generation module;

[0048] The data acquisition module is used to collect text corpora in the field of traditional Chinese medicine;

[0049] The data processing module includes a preprocessing unit, a knowledge processing unit, a pre-training unit, a supervised fine-tuning unit, and a reinforcement learning unit.

[0050] The preprocessing unit preprocesses the text corpus using natural language processing methods.

[0051] The knowledge processing unit, based on a large language model, performs knowledge processing on the text corpus to obtain pre-training data in the field of traditional Chinese medicine, a knowledge base for evidence-based diagnosis and treatment in the field of traditional Chinese medicine, and evidence-based diagnosis and treatment data in the field of traditional Chinese medicine.

[0052] The pre-training unit, based on pre-training data in the field of traditional Chinese medicine, completes pre-training on the base large language model to obtain a large language model in the field of traditional Chinese medicine.

[0053] The supervised fine-tuning unit, based on evidence-based diagnosis and treatment data in the field of traditional Chinese medicine, performs supervised fine-tuning on a large language model in the field of traditional Chinese medicine to construct an evidence-based diagnosis and treatment agent for traditional Chinese medicine.

[0054] The reinforcement learning unit, based on a multidimensional diagnosis and treatment quality assessment strategy, continuously optimizes the TCM evidence-based diagnosis and treatment agent through online reinforcement learning methods, and constructs a trustworthy TCM evidence-based diagnosis and treatment intelligent agent.

[0055] The result generation module is used to generate credible TCM evidence-based diagnosis and treatment intelligence externally.

[0056] Thirdly, based on the same inventive concept, this application also provides a device for constructing a trustworthy TCM evidence-based diagnosis and treatment intelligent agent, including a processor, a memory, and a bus. The memory stores instructions and data read by the processor, and the processor is used to call the instructions and data in the memory to execute the construction method of the trustworthy TCM evidence-based diagnosis and treatment intelligent agent as described above. The bus connects the functional components for transmitting information.

[0057] By adopting the above technical solution, the present invention has the following beneficial effects:

[0058] This invention provides a method, system, and apparatus for constructing a trustworthy TCM evidence-based diagnostic and treatment intelligent agent. By integrating TCM standard guidelines, literature, and clinical data, a domain knowledge base is built and a large language model is fine-tuned. Dynamic optimization is achieved through multi-dimensional diagnostic and treatment quality assessment strategies and online reinforcement learning methods, significantly improving the accuracy, reliability, and continuous optimization capabilities of the TCM evidence-based diagnostic and treatment intelligent agent, providing necessary technical support for TCM-related auxiliary diagnosis and treatment. This solution combines natural language processing technology and a large language model to preprocess and process TCM-related text of various data types, forming TCM-related pre-training data, a TCM-related evidence-based diagnostic and treatment knowledge base, and TCM-related evidence-based diagnostic and treatment data. Through pre-training and supervised fine-tuning, a TCM evidence-based diagnostic and treatment agent is constructed, possessing a certain ability to assist in diagnostic and treatment decision-making. This solution continuously optimizes the TCM evidence-based diagnostic and treatment agent by constructing a multi-dimensional diagnostic and treatment quality assessment strategy and combining it with online reinforcement learning methods, thus constructing a trustworthy TCM evidence-based diagnostic and treatment intelligent agent to assist in actual clinical diagnosis and treatment. Attached Figure Description

[0059] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0060] Figure 1 A flowchart illustrating a method for constructing a trusted TCM evidence-based diagnostic and treatment intelligent agent, as provided in an embodiment of the present invention;

[0061] Figure 2 A flowchart of the preprocessing method provided in an embodiment of the present invention;

[0062] Figure 3 for Figure 1 Flowchart of the specific method for step 2;

[0063] Figure 4 for Figure 1 Flowchart of the specific method for step 3;

[0064] Figure 5 for Figure 1 Flowchart of the specific method for step 4;

[0065] Figure 6 A flowchart of the online reinforcement learning method provided in an embodiment of the present invention;

[0066] Figure 7 This is a system diagram illustrating the construction of a trusted TCM evidence-based diagnostic and treatment intelligent agent, as provided in an embodiment of the present invention. Detailed Implementation

[0067] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0068] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0069] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0070] In the description of this invention, words such as "exemplarily," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary" in this invention should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of this invention, the meaning expressed by "and / or" can be both, or either one.

[0071] In the description of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction, they convey the same meaning.

[0072] In the description of this invention, sometimes the subscript such as W1 may be mistakenly written as a non-subscript form such as W1. Without emphasizing the difference, the meaning they express is the same.

[0073] The present invention will be further explained below with reference to specific embodiments.

[0074] It should also be noted that the specific embodiments or implementation methods described below are a series of optimized settings listed by the present invention to further explain the specific content of the invention, and these settings can be combined or used in conjunction with each other.

[0075] Example 1:

[0076] like Figure 1 As shown in the figure, this embodiment provides a method for constructing a trustworthy TCM evidence-based diagnostic and treatment intelligent agent, which includes the following steps:

[0077] Step 1: Collect text corpora in the field of Traditional Chinese Medicine and preprocess them using natural language processing methods;

[0078] Step 2: Based on the large language model, perform knowledge processing on the text corpus to obtain pre-training data in the field of traditional Chinese medicine, evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine, and evidence-based diagnosis and treatment data in the field of traditional Chinese medicine.

[0079] Step 3: Based on the pre-training data in the field of traditional Chinese medicine, complete the pre-training on the base large language model to obtain the large language model in the field of traditional Chinese medicine;

[0080] Step 4: Based on evidence-based diagnosis and treatment data in the field of traditional Chinese medicine, supervised fine-tuning is performed on the large language model in the field of traditional Chinese medicine to construct a traditional Chinese medicine evidence-based diagnosis and treatment agent;

[0081] Step 5: Based on a multidimensional diagnosis and treatment quality assessment strategy, continuously optimize the TCM evidence-based diagnosis and treatment agent through online reinforcement learning methods, and build a trustworthy TCM evidence-based diagnosis and treatment intelligent agent.

[0082] In this way, a trustworthy TCM evidence-based diagnosis and treatment intelligent agent driven by a large language model can be constructed, which makes full use of the characteristics of large language models, such as good interpretability and strong generalization ability, while avoiding the problems of illusion and poor decision consistency, and can provide specific treatment plans with clinical operability.

[0083] Furthermore, the text corpus includes standard guidelines, books and documents, web resources, clinical data, etc., and may involve various text types such as PDF, TXT, and EXCEL;

[0084] The aforementioned standard guidelines were downloaded from the National Standards Information Public Service Platform, the State Administration of Traditional Chinese Medicine platform, etc., and cover sub-fields of traditional Chinese medicine such as syndrome differentiation, treatment, and health preservation.

[0085] The books and documents mentioned are collected from TCM university libraries, TCM paper databases, and TCM laboratories, including TCM textbooks, TCM papers, TCM patents, TCM classics, and other documents.

[0086] The web page resources mentioned are open-source knowledge downloaded from web pages using web crawling technology, including TCM text corpora, TCM theoretical knowledge, etc.

[0087] The clinical data was collected from hospitals, renowned TCM doctors, and typical TCM medical case collections, including TCM clinical electronic medical records and TCM prescription data.

[0088] This ensures the accuracy, comprehensiveness, and authority of knowledge in the field of traditional Chinese medicine.

[0089] Furthermore, such as Figure 2 As shown, the specific methods of preprocessing include:

[0090] Step a1: Based on conventional OCR technology (such as PaddleOCR), perform text recognition on PDF-type text corpora; based on Python tools, process non-TXT-type text corpora to construct a TXT-type TCM-related text corpus.

[0091] Preferably, data processing is performed on non-TXT type text corpora, including: for text corpora of text types such as EXCEL and JSON, connecting the field information of each part based on a preset splicing template to form a TXT file;

[0092] Step a2: Based on conventional heuristic rules, perform data cleaning on the TCM text corpus to remove common problems such as character errors, garbled formats, and invalid information caused by OCR technology and Python tools, thereby improving the quality of the corpus;

[0093] Preferably, the data cleaning includes: batch processing of a text corpus in the field of traditional Chinese medicine based on conventional heuristic rules and a preset regular expression cleaning strategy.

[0094] Furthermore, such as Figure 3 As shown, step 2 specifically includes:

[0095] Step 21: Based on the Locality Sensitive Hash algorithm, after deduplication of the text corpus, and then based on the large language model, improve the data quality of the text corpus, and then construct pre-training data for the field of traditional Chinese medicine.

[0096] Preferably, step 21 specifically includes:

[0097] Step 211: Use a local sensitive hashing algorithm based on cosine similarity to remove duplicates from the text corpus;

[0098] Preferably, step 211 specifically includes:

[0099] Step 2111: Perform sliding window segmentation on the text corpus according to the context length supported by the large language model (e.g., 4K) to obtain a set of pre-trained text blocks in the field of traditional Chinese medicine.

[0100] Step 2112: Semantically encode each TCM text in the pre-trained text block set in the TCM field using the BGE-large-zh vector model to obtain a vector representation v;

[0101] Step 2113: For each v, generate k random hyperplanes; calculate the hash value of each bit to obtain the bit signature (64 bits or 128 bits), the specific formula is as follows:

[0102] h i (v)=sign(α i ·v)∈0,1

[0103] Among them, h i (v) represents the i-th hash value of v; α i Represents the i-th random hyperplane; sign represents the sign function;

[0104] Step 2114: Calculate the Hamming distance of the bit signatures between each text and filter them: If the Hamming distance is less than the preset threshold, it is considered a duplicate text, and only one text in the duplicate text is kept.

[0105] Step 212: Through prompting engineering, the original text is corrected and reconstructed by leveraging the semantic understanding and generation capabilities of the large language model to improve the data quality of the text corpus; the specific prompting text descriptions are shown in Table 1.

[0106] Table 1

[0107]

[0108] Step 22: Construct an evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine by sequentially performing text semantic segmentation, text block vectorization, and multi-level indexing;

[0109] Preferably, step 22 includes:

[0110] Step 221: Select treatment guidelines, TCM textbooks, TCM classics, and TCM papers as the basic knowledge corpus for the evidence-based treatment knowledge base in the field of TCM;

[0111] Step 222: Based on the basic knowledge corpus, design prompt texts and perform semantic segmentation of the text using a large language model to obtain segmentation results. This is to exclude content that is not related to diagnosis and treatment information, thereby ensuring the professionalism and reliability of the evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine. The specific prompt texts are shown in Table 2.

[0112] Table 2

[0113]

[0114] Step 223: Semantically represent the titles and content in the segmented results using the BGE-large-zh vector model; construct multi-layer knowledge tags using the Leiden hierarchical clustering algorithm, combined with title information, to form a multi-layer index with both vector indexing and path navigation capabilities, so as to organize and manage the semantic content of diagnosis and treatment knowledge, which is beneficial for subsequent retrieval; thus obtaining the evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine.

[0115] Step 23: Construct evidence-based diagnosis and treatment data in the field of traditional Chinese medicine through data quality classification and prompting engineering;

[0116] Preferably, step 23 specifically includes:

[0117] Step 231: Based on the standard definitions and specifications of disease and syndrome classification in the evidence-based diagnosis and treatment knowledge base of the TCM field, reject the sampling of diagnosis and treatment data that does not fall within the corresponding scope of discussion; through a BERT-like language model, combined with manually annotated high-quality datasets and low-quality datasets, train a binary classifier to determine the quality of diagnosis and treatment data, remove low-quality diagnosis and treatment data, and retain high-quality diagnosis and treatment data.

[0118] Step 232: First, based on the retrieval enhancement method, combined with the conventional BM25 algorithm and BGE-large-zh vector model, a hybrid retrieval strategy is used to recall the syndrome differentiation knowledge related to the patient's condition, disease, and syndrome in the diagnosis and treatment data from the evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine, and to recall the treatment principles, treatment methods, prescriptions, and other content related to the patient in the diagnosis and treatment data; then, the prompt text for the synthesized evidence-based syndrome differentiation reasoning path is designed, as shown in Table 3.

[0119] Table 3

[0120]

[0121] The prompt text for the synthetic evidence-based reasoning path is redesigned, as shown in Table 4;

[0122] Table 4

[0123]

[0124] By using prompting engineering and combining it with a large language model, data synthesis is performed to obtain synthesized data;

[0125] Finally, the prompt text for the quality verification of the synthesized data was designed, as shown in Table 5.

[0126] Table 5

[0127]

[0128] By using prompting engineering and combining it with a large language model, the quality of the synthesized data was verified, and evidence-based diagnosis and treatment data in the field of traditional Chinese medicine were obtained.

[0129] Furthermore, such as Figure 4 As shown, step 3 specifically includes:

[0130] Step 31: Mix the pre-training data in the TCM domain with the pre-training data in the general domain (high-quality data obtained from open-source corpora) according to a preset ratio (e.g., 1:1, 1:5 or 1:10) to obtain mixed pre-training data.

[0131] Step 32: Based on the hybrid pre-trained data, perform unsupervised training on the Qwen3-8B-Base large language model to obtain a large language model for the field of traditional Chinese medicine.

[0132] Preferably, the specific process of unsupervised training includes: using the cross-entropy loss function as the training loss function; given a sequence x = (x1, x2, ..., xn) consisting of T tokens. T );

[0133] Based on the Qwen3-8B-Base large language model, the word unit at each position is predicted using an autoregressive approach; for the t-th position, the Qwen3-8B-Base large language model predicts the word unit based on the previous t-1 word units x. <t Calculate the current word element x t The conditional probability distribution P θ (x t |x <t ), where θ represents the model parameters;

[0134] The cross-entropy loss function The specific formula is:

[0135]

[0136] Perplexity PPL is used as the training evaluation metric, and the specific calculation formula is as follows:

[0137]

[0138] Furthermore, the lower the perplexity of a large language model in the field of traditional Chinese medicine on the validation dataset, the more accurate the model is in modeling the data language of the traditional Chinese medicine field; therefore, based on the perplexity, the best-performing large language model in the field of traditional Chinese medicine can be obtained.

[0139] In this way, by leveraging the superior Chinese long text understanding capabilities of the Qwen3-8B-Base large language model, an ideal large language model for the field of traditional Chinese medicine can be constructed.

[0140] Furthermore, such as Figure 5 As shown, step 4 specifically includes:

[0141] Step 41: Convert the format of the TCM evidence-based diagnosis and treatment data to obtain a supervised fine-tuning dataset;

[0142] Preferably, the specific method for format conversion in step 41 is as follows:

[0143] The original format for TCM evidence-based diagnosis and treatment data is: D = I, R, J, Z, F;

[0144] Where I represents the content of the illness; R represents evidence-based reasoning, i.e. the pathogenesis process; J represents the disease diagnosis and syndrome differentiation results; Z represents the treatment principles and methods; F represents the prescription recommendations; R, J, Z and F are all reasoning path data synthesized in step 232;

[0145] The data format of TCM evidence-based diagnosis and treatment was converted into the format of intelligent agent training dataset to obtain the TCM evidence-based diagnosis and treatment agent dataset, as shown in Table 6.

[0146] Table 6

[0147]

[0148] Wherein, Fs represents the treatment plan; FJ represents the solution; both Fs and FJ are results returned by the second search tool t2; the specific definitions of the first search tool t1 and the second search tool t2 are shown in Table 7;

[0149] Table 7

[0150]

[0151] Step 42: Based on the TCM evidence-based diagnosis and treatment agent dataset, train a large language model for TCM using the cross-entropy loss function to obtain a TCM evidence-based diagnosis and treatment agent.

[0152] Preferably, step 42 specifically includes:

[0153] Step 421: Decompose the TCM evidence-based diagnosis and treatment agent dataset into a multi-task supervised fine-tuning dataset, as shown in Table 8;

[0154] Table 8

[0155]

[0156] Step 422: Combine the multi-task supervised fine-tuning dataset and train a large language model in the field of traditional Chinese medicine through the cross-entropy loss function to obtain a TCM evidence-based diagnosis and treatment agent, thereby constructing a supervised fine-tuning agent.

[0157] Furthermore, the multidimensional diagnostic and treatment quality assessment strategy includes accuracy indicators, reliability indicators, and safety indicators;

[0158] The accuracy index is used to evaluate the similarity between the diagnostic decision-making results of TCM evidence-based diagnosis and treatment agency and the actual syndrome type; specifically, it includes:

[0159] The accuracy rate is used to evaluate the precision of TCM diagnostic decision-making. The specific calculation formula is as follows:

[0160]

[0161] Among them, acc J This represents the accuracy rate of disease diagnosis and syndrome differentiation results; N represents the total number of (real) TCM evidence-based diagnosis and treatment data; j represents the syndrome differentiation decision results output by the TCM evidence-based diagnosis and treatment agent.

[0162] The accuracy of the analysis of treatment principles in traditional Chinese medicine is evaluated using cosine similarity. The specific calculation formula is as follows:

[0163]

[0164] Among them, acc Z represents the accuracy rate of treatment principles and methods; z represents the treatment principle analysis results output by the TCM evidence-based diagnosis and treatment agent; sim(·) represents the cosine similarity; emb(·) represents the vector model;

[0165] A larger-scale language model is used as the judge to evaluate the consistency between the recommended medication and the patient's condition. The specific formula is as follows:

[0166]

[0167] Among them, acc med represents the accuracy rate of prescription recommendations; LLM(·) represents the consistency score between each prescription recommendation and the patient's condition given by the large language model, with a value range of 0-5; fj represents the prescription solution output by the TCM evidence-based diagnosis and treatment agent;

[0168] The credibility index is used to evaluate the reliability of TCM evidence-based diagnosis and treatment agents in citing diagnostic and treatment knowledge during evidence-based reasoning. Since this part consists entirely of unstructured natural language text, this embodiment evaluates the reliability of evidence-based reasoning and treatment principle analysis results from two levels: cosine similarity and large language model. Specific formulas include:

[0169]

[0170] Where Rel represents credibility; pre represents the content of the predicted evidence-based reasoning and governance analysis; and lab represents the content of the actual evidence-based reasoning and governance analysis.

[0171] The aforementioned safety index is used to evaluate whether the treatment principles and prescriptions of the TCM evidence-based diagnosis and treatment agency are applicable to the actual situation. In this embodiment, the safety index Safe of the prescriptions output by the TCM evidence-based diagnosis and treatment agency is evaluated from the aspects of drug contraindications, physical adaptability and drug practicality through a large language model.

[0172] Furthermore, such as Figure 6 As shown, the online reinforcement learning method specifically includes:

[0173] Step b1: Based on a multidimensional diagnosis and treatment quality assessment strategy, calculate the reward score between the predicted results and the actual results of TCM evidence-based diagnosis and treatment; the specific formula includes:

[0174] score=λ1*(acc J +acc Z +accmed )+λ2*Rel+λ3*Safe

[0175] Wherein, λ1, λ2 and λ3 are the weights of each indicator;

[0176] Step b2: The TCM evidence-based diagnosis and treatment agent generates prediction results in real time during online reinforcement learning;

[0177] Step b3: Based on the reward score, update the model parameters of the TCM evidence-based diagnosis and treatment agent in real time to optimize the evidence-based diagnosis and treatment capabilities; the specific calculation formula is as follows:

[0178]

[0179] Where, π θ This represents the model parameters for the updated TCM evidence-based diagnosis and treatment agent; πθ old This represents the model parameters of the old TCM evidence-based diagnosis and treatment proxy; A t The dominant function is represented by Q(s). t a t )-V(s t ), used to measure action a t Compared to the average performance within the group; Q(s) t a t ) indicates that in state s t Next, execute action a. t Expected return; V(s) t ) indicates that in state s t Below, the average reward for performing all actions; β represents the adjustment coefficient for the KL divergence;

[0180] In this way, by optimizing parameters in real time using online data, the diagnostic and treatment performance of TCM evidence-based diagnosis and treatment agents can be improved from three aspects: accuracy, reliability, and security, thereby realizing the construction of a trustworthy TCM evidence-based diagnosis and treatment intelligent agent.

[0181] Example 2:

[0182] like Figure 7 As shown, this embodiment provides a system for constructing a trustworthy TCM evidence-based diagnosis and treatment intelligent agent, including a data acquisition module, a data processing module, and a result generation module;

[0183] The data acquisition module is used to collect text corpora in the field of traditional Chinese medicine;

[0184] The data processing module includes a preprocessing unit, a knowledge processing unit, a pre-training unit, a supervised fine-tuning unit, and a reinforcement learning unit.

[0185] The preprocessing unit preprocesses the text corpus using natural language processing methods.

[0186] The knowledge processing unit, based on a large language model, performs knowledge processing on the text corpus to obtain pre-training data in the field of traditional Chinese medicine, a knowledge base for evidence-based diagnosis and treatment in the field of traditional Chinese medicine, and evidence-based diagnosis and treatment data in the field of traditional Chinese medicine.

[0187] The pre-training unit, based on pre-training data in the field of traditional Chinese medicine, completes pre-training on the base large language model to obtain a large language model in the field of traditional Chinese medicine.

[0188] The supervised fine-tuning unit, based on evidence-based diagnosis and treatment data in the field of traditional Chinese medicine, performs supervised fine-tuning on a large language model in the field of traditional Chinese medicine to construct an evidence-based diagnosis and treatment agent for traditional Chinese medicine.

[0189] The reinforcement learning unit, based on a multidimensional diagnosis and treatment quality assessment strategy, continuously optimizes the TCM evidence-based diagnosis and treatment agent through online reinforcement learning methods, and constructs a trustworthy TCM evidence-based diagnosis and treatment intelligent agent.

[0190] The result generation module is used to generate credible TCM evidence-based diagnosis and treatment intelligence externally.

[0191] Example 3:

[0192] This embodiment provides a device for constructing a trusted TCM evidence-based diagnosis and treatment intelligent agent, including a processor, a memory, and a bus. The memory stores instructions and data read by the processor, and the processor is used to call the instructions and data in the memory to execute the construction method of the trusted TCM evidence-based diagnosis and treatment intelligent agent as described above. The bus connects the various functional components for transmitting information.

[0193] In another embodiment, this solution can also be implemented using an integrated device, which may include corresponding modules that perform one or more steps in the various embodiments described above. A module may be one or more hardware modules specifically configured to perform the corresponding step, or implemented by a processor configured to perform the corresponding step, or stored in a computer-readable medium for implementation by a processor, or implemented through some combination thereof.

[0194] The processor executes the various methods and processes described above. For example, the method implementations in this scheme can be implemented as software programs tangibly contained in a machine-readable medium, such as memory. In some implementations, part or all of the software program can be loaded and / or installed via memory and / or a communication interface. When the software program is loaded into memory and executed by the processor, one or more steps of the methods described above can be performed. Alternatively, in other implementations, the processor can be configured to execute one of the methods described above by any other suitable means (e.g., by means of firmware).

[0195] This device can be implemented using a bus architecture. A bus architecture can include any number of interconnect buses and bridges, depending on the specific application of the hardware and overall design constraints. The bus connects various circuits, including one or more processors, memory, and / or hardware modules. The bus can also connect various other circuits such as peripherals, voltage regulators, power management circuitry, external antennas, etc.

[0196] Buses can be Industry Standard Architecture (ISA) buses, Peripheral Component Interconnect (PCI) buses, or Extended Industry Standard Component (EISA) buses, etc. Buses can be divided into address buses, data buses, control buses, etc.

[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a reliable TCM evidence-based diagnostic and treatment intelligent agent, characterized in that, include: Step 1: Collect text corpora in the field of Traditional Chinese Medicine and preprocess them using natural language processing methods; Step 2: Based on the large language model, perform knowledge processing on the text corpus to obtain pre-training data in the field of traditional Chinese medicine, evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine, and evidence-based diagnosis and treatment data in the field of traditional Chinese medicine. Step 3: Based on the pre-training data in the field of traditional Chinese medicine, complete the pre-training on the base large language model to obtain the large language model in the field of traditional Chinese medicine; Step 4: Based on evidence-based diagnosis and treatment data in the field of traditional Chinese medicine, supervised fine-tuning is performed on the large language model in the field of traditional Chinese medicine to construct a traditional Chinese medicine evidence-based diagnosis and treatment agent; Step 5: Based on a multidimensional diagnosis and treatment quality assessment strategy, continuously optimize the TCM evidence-based diagnosis and treatment agent through online reinforcement learning methods, and build a trustworthy TCM evidence-based diagnosis and treatment intelligent agent.

2. The construction method according to claim 1, characterized in that, Specific preprocessing methods include: Step a1: Based on the OCR method, perform text recognition on PDF-type text corpora; based on Python tools, process non-TXT-type text corpora to construct a TXT-type TCM-related text corpus. Step a2: Based on heuristic rules, perform data cleaning on the TCM text corpus.

3. The construction method according to claim 1, characterized in that, Step 2 specifically includes: Step 21: Based on the Locality Sensitive Hash algorithm, after deduplication of the text corpus, and then based on the large language model, improve the data quality of the text corpus, and then construct pre-training data for the field of traditional Chinese medicine. Step 22: Construct an evidence-based diagnosis and treatment knowledge base in the field of traditional Chinese medicine by sequentially performing text semantic segmentation, text block vectorization, and multi-level indexing; Step 23: Construct evidence-based diagnosis and treatment data in the field of traditional Chinese medicine through data quality classification and prompting engineering.

4. The construction method according to claim 1, characterized in that, Step 3 specifically includes: Step 31: Mix the pre-training data in the TCM domain with the pre-training data in the general domain according to a preset ratio to obtain mixed pre-training data; Step 32: Based on the mixed pre-trained data, perform unsupervised training on the large language model to obtain a large language model for the field of traditional Chinese medicine.

5. The construction method according to claim 1, characterized in that, The specific process of unsupervised training in step 32 includes: The cross-entropy loss function is used as the training loss function; given a sequence x = (x1, x2, ..., xn) consisting of T words. T ); Based on a large language model, the word unit at each position is predicted using an autoregressive approach; for the t-th position, the large language model predicts the word unit based on the previous t-1 words x. <t Calculate the current word element x t The conditional probability distribution P θ (x t |x <t ), where θ represents the model parameters; The cross-entropy loss function The specific formula is: Perplexity PPL is used as the training evaluation metric, and the specific calculation formula is as follows: The lower the perplexity of a large language model in the field of Traditional Chinese Medicine on the validation dataset, the more accurate the model is in modeling the data language in the TCM field.

6. The construction method according to claim 1, characterized in that, Step 4 specifically includes: Step 41: Convert the format of the TCM evidence-based diagnosis and treatment data to obtain the TCM evidence-based diagnosis and treatment proxy dataset; Step 42: Based on the TCM evidence-based diagnosis and treatment agent dataset, train a large language model for the TCM domain using the cross-entropy loss function to obtain the TCM evidence-based diagnosis and treatment agent.

7. The construction method according to claim 1, characterized in that, A multidimensional approach to assessing the quality of medical care includes accuracy indicators, reliability indicators, and safety indicators. The accuracy index is used to evaluate the similarity between the diagnostic decision-making results of TCM evidence-based diagnosis and treatment agency and the actual syndrome type; specifically, it includes: The accuracy rate is used to evaluate the precision of TCM diagnostic decision-making. The specific calculation formula is as follows: Among them, acc J This represents the accuracy rate of disease diagnosis and syndrome differentiation results; N represents the total number of TCM evidence-based diagnosis and treatment data; j represents the syndrome differentiation decision results output by the TCM evidence-based diagnosis and treatment agent; J represents the disease diagnosis and syndrome differentiation results in the TCM evidence-based diagnosis and treatment data. The accuracy of the analysis of treatment principles in traditional Chinese medicine is evaluated using cosine similarity. The specific calculation formula is as follows: Among them, acc z represents the accuracy rate of treatment principles and methods; z represents the analysis results of treatment principles output by the TCM evidence-based diagnosis and treatment agent; sim(·) represents the cosine similarity; emb(·) represents the vector model; Z represents the treatment principles and methods in the TCM evidence-based diagnosis and treatment data; A larger-scale language model is used as the judge to evaluate the consistency between the recommended medication and the patient's condition. The specific formula is as follows: Among them, acc med The value represents the accuracy of the prescription recommendation; LLM(·) represents the consistency score between each prescription recommendation and the patient's condition information given by the large language model, with a value range of 0-5; fj represents the prescription solution output by the TCM evidence-based diagnosis and treatment agent; I represents the condition information. The credibility index is used to evaluate the reliability of TCM evidence-based diagnosis and treatment agents in citing clinical knowledge during evidence-based reasoning; the specific formula includes: Where Rel represents credibility; pre represents the content of the predicted evidence-based reasoning and governance analysis; and lab represents the content of the actual evidence-based reasoning and governance analysis. The aforementioned safety index is used to evaluate whether the treatment principles and prescriptions of TCM evidence-based diagnosis and treatment agents are applicable to actual situations. Specifically, it uses a large language model to evaluate the safety index Safe of the prescriptions and prescriptions output by TCM evidence-based diagnosis and treatment agents from the aspects of drug contraindications, physical adaptability, and drug practicality.

8. The construction method according to claim 7, characterized in that, The online reinforcement learning method specifically includes: Step b1: Based on a multidimensional diagnosis and treatment quality assessment strategy, calculate the reward score between the predicted results and the actual results of TCM evidence-based diagnosis and treatment; the specific formula includes: score=λ1*(acc J +acc Z +acc med )+λ2*Rel+λ3*Safe;e Wherein, λ1, λ2 and λ3 are the weights of each indicator; Step b2: The TCM evidence-based diagnosis and treatment agent generates prediction results in real time during online reinforcement learning; Step b3: Based on the reward score, update the model parameters of the TCM evidence-based diagnosis and treatment agent in real time; the specific calculation formula is as follows: Among them, v θ This represents the model parameters for the updated TCM evidence-based diagnosis and treatment agent; This represents the model parameters of the old TCM evidence-based diagnosis and treatment proxy; A t The dominant function is represented by Q(s). t a t )-V(s t ), used to measure action a t Compared to the average performance within the group; Q(s) t a t ) indicates that in state s t Next, execute action a. t Expected return; V(s) t ) indicates that in state s t Below, the average reward for performing all actions; β represents the adjustment coefficient of the KL divergence.

9. A system for constructing a trustworthy TCM evidence-based diagnostic and treatment intelligent agent, characterized in that, It includes a data acquisition module, a data processing module, and a result generation module; The data acquisition module is used to collect text corpora in the field of traditional Chinese medicine; The data processing module includes a preprocessing unit, a knowledge processing unit, a pre-training unit, a supervised fine-tuning unit, and a reinforcement learning unit. The preprocessing unit preprocesses the text corpus using natural language processing methods. The knowledge processing unit, based on a large language model, performs knowledge processing on the text corpus to obtain pre-training data in the field of traditional Chinese medicine, a knowledge base for evidence-based diagnosis and treatment in the field of traditional Chinese medicine, and evidence-based diagnosis and treatment data in the field of traditional Chinese medicine. The pre-training unit, based on pre-training data in the field of traditional Chinese medicine, completes pre-training on the base large language model to obtain a large language model in the field of traditional Chinese medicine. The supervised fine-tuning unit, based on evidence-based diagnosis and treatment data in the field of traditional Chinese medicine, performs supervised fine-tuning on a large language model in the field of traditional Chinese medicine to construct an evidence-based diagnosis and treatment agent for traditional Chinese medicine. The reinforcement learning unit, based on a multidimensional diagnosis and treatment quality assessment strategy, continuously optimizes the TCM evidence-based diagnosis and treatment agent through online reinforcement learning methods, and constructs a trustworthy TCM evidence-based diagnosis and treatment intelligent agent. The result generation module is used to generate credible TCM evidence-based diagnosis and treatment intelligence externally.

10. A device for constructing a trustworthy TCM evidence-based diagnostic and treatment intelligent agent, characterized in that, It includes a processor, a memory, and a bus. The memory stores instructions and data read by the processor. The processor is used to call the instructions and data in the memory to execute the construction method as described in any one of claims 1-8. The bus connects the functional components for transmitting information.