RAG large model functional gastrointestinal disease auxiliary diagnosis method, system and device based on rule engine and search engine

By adopting a RAG model based on rules engine and search engine in the assisted diagnosis of functional gastroenterology, combined with multi-engine collaborative matching and weighted summary methods, the problems of limited rule base, high maintenance cost and poor interpretation in the existing technology are solved, and a more efficient, flexible and transparent diagnostic effect is achieved.

CN120148822APending Publication Date: 2025-06-13INST OF INFORMATION ON TRADITIONAL CHINESE MEDICINE CACMS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510223700.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing auxiliary diagnosis methods for functional gastrointestinal diseases have problems such as limited rule base, high maintenance costs, lack of flexibility and poor interpretation, and it is difficult to effectively deal with complex or rare cases.

Method used

The RAG model based on the rules engine and search engine is adopted to obtain the user's diagnostic medical case text data, identify named entities and extract entities. Combined with the optimized rules engine, Elasticsearch search engine and FASS vector library, multi-engine collaborative matching and weighted summary are carried out to generate comprehensive search results, and finally assisted diagnosis of functional gastrointestinal diseases.

Benefits of technology

Improves diagnostic accuracy and flexibility, reduces maintenance costs, enhances interpretability and transparency, supports personalized treatment recommendations, and promotes continuous learning and development of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148822A_ABST
    Figure CN120148822A_ABST
Patent Text Reader

Abstract

The invention discloses an RAG large model functional gastrointestinal disease auxiliary diagnosis method, system and device based on a rule engine and a search engine, FASS vector library retrieval, Elasticsearch semantic retrieval and rule engine matching are combined, a gastrointestinal disease model obtained by performing fine adjustment based on a chatglm4-9b large model is combined, a threshold value is introduced to define a proportion, and an RAG large model functional gastrointestinal disease auxiliary diagnosis result is obtained. The gastrointestinal diseases are accurately diagnosed according to the medical case text information input by the user, the method is more intelligent, flexible and easy to understand, and the accuracy and efficiency of functional gastrointestinal disease diagnosis are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disease assisted diagnosis and treatment, and particularly relates to a method, system and device for assisting in the diagnosis of functional gastrointestinal diseases by using a RAG large model based on a rule engine and a search engine. Background Art

[0002] Functional gastrointestinal diseases (FGIDs) are a group of functional gastrointestinal diseases, which are digestive system diseases caused by the interaction of physiological, psychosocial and social factors. Patients with FGIDs often have extra-gastrointestinal symptoms, such as dyspnea, palpitation, chronic headache, myalgia, etc., and their incidence rate is relatively high.

[0003] In recent years, with the continuous development of computer technology, the methods for using computers to assist in the diagnosis of functional gastrointestinal diseases have also been evolving. Currently, there are mainly the following two technical directions for using computers to assist in the diagnosis of functional gastrointestinal diseases:

[0004] One is the functional gastrointestinal disease assisted diagnosis based on a rule base. The core of this method is to build a detailed rule set, covering gender, age, and diagnostic information, syndromes, treatment principles and prescription suggestions corresponding to various symptoms. Key information, such as the gender, age and specific symptoms of the patient, is extracted from the medical record text through natural language processing technology. These extracted information are matched with the preset rules to generate corresponding syndrome diagnosis results, and the final diagnosis opinions are presented to the user.

[0005] However, this method has certain limitations and disadvantages. For example, ① the finiteness of the rule set: The rule base is usually established based on existing medical knowledge and clinical experience, but these rules may not cover all situations. In the face of complex and changeable conditions, especially rare cases or newly emerging symptoms, the existing rules may not be sufficient to provide accurate diagnoses. ② High maintenance cost: With the progress of medical research and the continuous emergence of new discoveries, the rule base needs to be updated regularly to maintain its timeliness and accuracy, which increases the maintenance cost and technical difficulty of the system. ③ Lack of flexibility: The rule-driven method is often relatively rigid and difficult to handle ambiguous situations or borderline cases. When the input information does not exactly match the preset rules, the system may produce incorrect or uncertain results. ④ Over-reliance on expert knowledge: The formulation of rules highly depends on the knowledge level of experts in the field. If there are differences in expert opinions, it may lead to contradictory information in the rule base.

[0006] Another machine learning-based auxiliary diagnosis for functional gastrointestinal diseases. This method first uses a word embedding model to convert the original text in medical records (including patients' basic information and symptom descriptions) into numerical vector form. Then, a variety of machine learning algorithms, such as Naive Bayes, decision trees, and random forests, are used to build a diagnostic model for classification. Through learning a large amount of labeled data, an efficient classifier that can identify different disease patterns is trained. Finally, the trained model is used to predict and assist in the diagnosis of new cases, providing more accurate and personalized diagnosis and treatment suggestions. However, this method also has certain limitations and disadvantages. For example, ① high data quality requirements: The performance of machine learning models depends to a large extent on the quality of the training data. If the data set has biases, noise, or insufficient sample size, it will seriously affect the performance and generalization ability of the model. ② Poor interpretability: Although many complex machine learning algorithms (such as deep neural networks) can achieve high prediction accuracy, their working mechanisms are often not transparent, known as "black box" models. In this case, it is difficult for doctors to understand the reasons for the model to make specific decisions, affecting trust and acceptance. ③ Feature engineering challenges: For unstructured medical text data, how to effectively extract useful features is a difficult problem. Poor feature selection not only reduces the model performance but may also introduce unnecessary computational burdens. ④ Continuous learning requirements: Over time, new disease types and treatment methods emerge continuously, so the model needs to be retrained regularly to adapt to the latest medical progress, which also means long-term resource investment and technical support.

[0007] Therefore, it is necessary to improve the existing technologies and provide a method that can assist in the diagnosis of gastrointestinal diseases more intelligently and flexibly. Summary of the Invention

[0008] Therefore, based on the above background, in view of the defects existing in the existing computer-aided diagnosis technology for gastrointestinal diseases, the present invention provides a method, system, and device for auxiliary diagnosis of functional gastrointestinal diseases based on the RAG large model of a rule engine and a search engine, which are more intelligent and flexible.

[0009] The technical solution of the present invention is as follows:

[0010] A method for auxiliary diagnosis of functional gastrointestinal diseases based on the RAG large model of a rule engine and a search engine, the method comprising:

[0011] Obtain the diagnostic medical record text data of the user, where the diagnostic medical record text data includes gender, age, symptoms, and four diagnostic signs;

[0012] Perform named entity recognition on the diagnostic medical record text data, extract entity extraction, and then input the entity information into the optimized rule engine to match the user's medical record data, and output the rule engine matching result;

[0013] After filtering the diagnostic medical record text data, input it into the optimized Elasticsearch search engine for retrieval, and output the Elasticsearch retrieval results;

[0014] After converting the diagnostic medical record text data into word vectors through the bge-m3 model, search in the FASS vector library for similarity comparison, and output the FASS retrieval results;

[0015] Weight and summarize the rule engine matching results, Elasticsearch retrieval results, and FASS retrieval results according to the retrieval threshold in proportion to obtain the comprehensive retrieval results;

[0016] Input the comprehensive retrieval results into the functional gastrointestinal disorder diagnosis model to perform the auxiliary diagnosis of functional gastrointestinal disorders and output the diagnosis results.

[0017] Furthermore, the functional gastrointestinal disorder diagnosis model is fine-tuned and constructed based on the chatglm4-9b large model, and its construction steps are as follows:

[0018] S1: Collect the medical record text data of functional gastrointestinal disorders and non-functional gastrointestinal disorders;

[0019] The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions;

[0020] S2: Manually review and verify the medical record text data, and classify the data into positive sample identification data and negative sample identification data;

[0021] The manual work is done by professional doctors;

[0022] The positive sample identification data means that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal disorders;

[0023] The negative sample identification data refers to that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to non-functional gastrointestinal disorders;

[0024] S3: Adopt the instruction-following mode to fine-tune the chatglm4-9b large model, and define the input and output examples of input and output, and finally construct the functional gastrointestinal disorder diagnosis model;

[0025] Specifically, the steps to fine-tune the chatglm4-9b large model are as follows:

[0026] S3-1: Fill in the training data for the positive sample identification data in step 2, input: [medical record text of functional gastrointestinal disorders], output: [gender / age / symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions];

[0027] S3-2: Similar to step S3-1, perform training data filling on the negative sample identification data in step 2;

[0028] S3-3: Combine the positive and negative sample identification data, divide the validation and test sets, set the training environment, define the training parameters, and fine-tune the large training model;

[0029] S3-4: Perform cyclic training, evaluate the final performance of the model on the test set, and finally construct a functional gastrointestinal disorder diagnosis model.

[0030] Furthermore, the rule engine adopts the EasyRules rule engine.

[0031] Furthermore, the optimization of the rule engine includes the following steps:

[0032] 1. Collect medical record text data of functional gastrointestinal disorders and non-functional gastrointestinal disorders;

[0033] The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions;

[0034] 2. Manually review and verify the medical records and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data;

[0035] The positive sample identification data means that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal disorders;

[0036] The negative sample identification data refers to that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to non-functional gastrointestinal disorders;

[0037] 3. Clean the positive sample identification data in step 2 to obtain the correspondence between [gender / age / symptoms / four diagnostic signs] and [syndrome differentiation / treatment principles / prescriptions];

[0038] 4. Initialize the EasyRules rule engine, synchronize the correspondence between [gender / age / symptoms / four diagnostic signs] and [syndrome differentiation / treatment principles / prescriptions] obtained in step 3, manually review the threshold of each relationship, perform filling, and apply it to the EasyRules rule engine;

[0039] The rule engine threshold is equal to a number. The larger the number, the higher the priority of the result when the conditions are simultaneously met if it occurs.

[0040] Specifically, the threshold is a value determined by professional doctors according to clinical guidelines and experience.

[0041] 5. Construct an algorithm for the subsequent scoring mechanism of the rule engine output result, and the subsequent scoring mechanism algorithm satisfies:

[0042] ① Rule output: When one or more rules are satisfied simultaneously, all matching results can be output;

[0043] ② Threshold weight scoring: Set a weight threshold N for each rule. When the number of records with the same result is X, the score of this classification result is X * N;

[0044] ③ Result sorting and selection: Sort the results in descending order according to the calculated scores, and select the result with the highest score as the final output.

[0045] Furthermore, the optimization of the Elasticsearch search engine includes the following steps:

[0046] (1) Collect medical case text data of functional gastrointestinal diseases and non-functional gastrointestinal diseases;

[0047] The medical case text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions;

[0048] (2) Manually review and verify the medical cases and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data;

[0049] The positive sample identification data is that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases;

[0050] The negative sample identification data refers to that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to non-functional gastrointestinal diseases;

[0051] (3) Clean and filter the positive sample identification data in step (2) to form a list;

[0052] (4) Configure the Elasticsearch search engine, synchronize the list data in step (3), generate indexed data for retrieval, and build a functional gastrointestinal disease data retrieval library;

[0053] (5) Reconstruct the scoring and sorting algorithm;

[0054] The reconstructed scoring and sorting algorithm is:

[0055] For the documents in the full-text retrieval results of medical records, according to the hit words, a TF-IDF algorithm with the characteristics of functional gastrointestinal diseases is constructed based on the word threshold for relevance scoring. Specifically: Term Frequency TF(t,d) = the number of times the domain word t appears in document d * threshold / the total number of all words in document d, Inverse Document Frequency IDF(m,D) = log(total number of all medical record documents D / number of retrieved documents t), TF-IDF(t,d,m,D) = TF(t,d) * IDF(m,D). Accumulate the TF-IDF values of the domain words in the retrieved documents, and finally sort them from high to low to obtain the top three;

[0056] Furthermore, the steps for building the Fass vector library are as follows:

[0057] ① Collect medical record text data of functional gastrointestinal diseases and non-functional gastrointestinal diseases;

[0058] The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions;

[0059] ② Manually review and verify the medical record and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data;

[0060] Positive sample identification data means that symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases;

[0061] Negative sample identification data refers to symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all corresponding to non-functional gastrointestinal diseases;

[0062] ③ Clean and filter the positive sample identification data in step ② to form a list set;

[0063] ④ Configure the bge-m3 word vector model and Fass vector initialization configuration;

[0064] ⑤ For the list set generated in step ④, synchronize the vector information to the Fass vector library through the bge-m3 word vector model;

[0065] ⑥ Package the input and output interfaces of the Fass vector library to make it applicable.

[0066] Furthermore, the matching results of the rule engine, the retrieval results of Elasticsearch, and the retrieval results of FASS are weighted and summarized according to the ratio of 5:3:2.

[0067] Based on the same inventive concept, the present invention also provides a RAG large model functional gastrointestinal disease auxiliary diagnosis system based on a rule engine and a search engine, including:

[0068] A data acquisition module for the user to input medical record text data;

[0069] A retrieval and analysis module, which is used to input the medical record text data entered by the user into a rule engine, an Elasticsearch engine, and a FASS vector library respectively for matching or retrieval, and perform weighted aggregation calculation on the results output by the rule engine, the Elasticsearch engine, and the FASS vector library according to 5:3:2 to determine the comprehensive retrieval result;

[0070] A functional gastrointestinal disorder auxiliary diagnosis module, which diagnoses according to the comprehensive retrieval result data by using a functional gastrointestinal disorder model and outputs a diagnosis result.

[0071] Based on the same inventive concept, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. The processor executes the computer program to implement the above-mentioned method for auxiliary diagnosis of functional gastrointestinal disorders based on a RAG large model using a rule engine and a search engine.

[0072] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, it implements the above-mentioned method for auxiliary diagnosis of functional gastrointestinal disorders based on a RAG large model using a rule engine and a search engine.

[0073] The beneficial effects achieved by adopting the present invention are as follows:

[0074] 1. Improve diagnostic accuracy: By combining a rule engine, a search engine, and large model technology, it is possible to more comprehensively cover various possibilities of the condition; the rule engine provides structured logical reasoning, while the search engine and vector retrieval can capture more subtle data associations, thereby improving the ability to identify complex or rare cases.

[0075] 2. Enhance the flexibility and adaptability of the auxiliary diagnosis method and its system: The dynamic update feature of the rule engine enables the system to quickly respond to new diagnostic conditions without waiting for a large-scale software update cycle. At the same time, by utilizing the reinforcement learning ability of the machine learning model, the diagnostic algorithm can be continuously optimized to ensure that it remains up-to-date.

[0076] 3. Reduce maintenance costs: Compared with traditional fixed rule libraries, the dynamic rule engine reduces the need for frequent manual updates, reducing the maintenance costs in long-term operation. In addition, the automated feature extraction process also reduces the dependence on manual intervention and improves work efficiency.

[0077] 4. Improve interpretability and transparency: Through the comprehensive application of FASS vector library retrieval, Elasticsearch semantic retrieval, and rule engine matching, not only specific diagnostic bases are provided, but also the interpretability of the model decision-making process is enhanced. This is crucial for building trust between doctors and patients.

[0078] 5. Facilitate personalized treatment recommendations: The fine-tuned large model can better understand individual differences and generate more personalized diagnosis and treatment plans for each patient. This customized service helps improve treatment effects and increase patient satisfaction.

[0079] 6. Support continuous learning and development: The system design includes mechanisms such as positive and negative sample construction and iterative training, ensuring that the diagnostic tool is always in the best performance state over time and with technological progress. Such an architecture also reserves room for future possible technological upgrades.

[0080] In summary, the present invention can improve the diagnosis efficiency of functional gastrointestinal diseases while significantly enhancing the intelligence level and user experience, which is of great significance for promoting the digital transformation in the field of medical health. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Attached Figure 1 is the overall flowchart of the present invention.

[0082] Attached Figure 2 is the flowchart for constructing the large model of functional gastrointestinal diseases of the present invention.

[0083] Attached Figure 3 is the flowchart for optimizing the rule engine of the present invention.

[0084] Attached Figure 4 is the flowchart for optimizing the Elasticsearch search engine of the present invention.

[0085] Attached Figure 5 is the flowchart for building the Fass vector library of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with its embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0087] Embodiment 1: A method for assisting in the diagnosis of functional gastrointestinal diseases by an RAG large model based on a rule engine and a search engine, the method includes (as Figure 1 shown):

[0088] Obtain the diagnostic medical record text data of the user, where the diagnostic medical record text data includes gender, age, symptoms, and the four diagnostic signs;

[0089] The four diagnostic signs are obtained from the traditional Chinese medicine diagnostic methods of inspection, auscultation and olfaction, inquiry, and palpation, and the four diagnostic signs can be obtained by a doctor after the four diagnostic methods.

[0090] In a specific application, in one implementation manner, the four diagnostic signs can be mainly detected and obtained by a tongue and face image instrument and a pulse diagnosis instrument.

[0091] Perform named entity recognition on the medical record text data, extract entities, and then input the entity information into the optimized rule engine to match the user's medical record data, and output the rule engine matching result;

[0092] After filtering the medical record text data, input it into the optimized Elasticsearch search engine for retrieval, and output the Elasticsearch retrieval result;

[0093] The operations for filtering the medical record text data include standardizing terms, removing invalid, incorrect, sensitive, and irrelevant information or symbols, etc. Among them, standardizing terms can be performed by matching and comparing with a list of variant and correct terms. The format of the list of variant and correct terms is {variant term: correct term}, for example: An example of the list of variant and correct terms term_standardization = {

[0094] "Astragalus membranaceus":"Huangqi", #variant term: correct term

[0095] "Wedelia chinensis":"Fangfeng",

[0096] "Rhizome of Atractylodes macrocephala":"Baizhu",

[0097] "Fruit of Chinese date":"Dazao"

[0098] }. The method for converting variant terms in the text to correct terms: def standardize_terms(text):

[0099] for variant, standard in term_standardization.items():

[0100] text = text.replace(variant, standard)

[0101] return text.

[0102] The operations for removing invalid, incorrect, sensitive, and irrelevant information or symbols include:

[0103] 1: Remove invalid information, including expired and duplicate data, by setting a time range or performing uniqueness checks to remove such information.

[0104] 2: Incorrect information includes the above positive and variant name filtering and the lack of quantity units.

[0105] 3: Remove the word "sensitive".

[0106] 4: Irrelevant information or symbols such as extra spaces and unnecessary punctuation marks are represented by the regular expression [^\w\u4e00-\u9fff], which is the Unicode range of Chinese characters.

[0107] After converting the medical record text data into word vectors through the bge-m3 model, search in the FASS vector library for similarity comparison and output the FASS retrieval results;

[0108] Weightedly summarize the rule engine matching results, Elasticsearch retrieval results, and FASS retrieval results according to the retrieval threshold in the ratio of 5:3:2 to obtain the comprehensive retrieval results;

[0109] Input the comprehensive retrieval results into the functional gastrointestinal disease diagnosis model to perform the auxiliary diagnosis of functional gastrointestinal diseases and output the diagnosis results. The diagnosis results are disease / syndrome differentiation / treatment principle / suggested prescription text information.

[0110] Specific syndrome differentiations, for example: liver qi stagnation, or spleen and stomach weakness, or qi stagnation and blood stasis, etc.

[0111] The functional gastrointestinal disease diagnosis model is fine-tuned and constructed based on the chatglm4-9b large model, and its construction steps are as follows:

[0112] S1: Collect medical record text data of functional gastrointestinal diseases and non-functional gastrointestinal diseases;

[0113] The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principle, and prescription;

[0114] S2: Manually review and verify the medical record text data, and classify the data into positive sample labeled data and negative sample labeled data;

[0115] Positive sample labeled data means that the symptoms / four diagnostic signs / syndrome differentiation / treatment principle / prescription all correspond to functional gastrointestinal diseases;

[0116] Negative sample labeled data refers to the symptoms / four diagnostic signs / syndrome differentiation / treatment principle / prescription all corresponding to non-functional gastrointestinal diseases;

[0117] S3: In the instruction-following mode, fine-tune the ChatGLM4-9B large model, define input and output examples, and finally construct a functional gastrointestinal disorder model;

[0118] Specifically, the steps to fine-tune the ChatGLM4-9B large model are as follows (as Figure 2 shown):

[0119] S3-1: Fill in the training data for the positive sample identification data in step 2, input: [Medical records text of functional gastrointestinal disorders], output: [Gender / Age / Symptoms / Four diagnostic signs / Syndrome differentiation / Treatment principle / Prescription], for example {

[0120] input: "Feng, female, 41, first visit, diarrhea, watery stools, borborygmus, nausea, vomiting, acid regurgitation, belching, fatigue, chills, thin tongue with light purple color, scanty and moist tongue coating, deep pulse, thready pulse, diarrhea, functional gastrointestinal disorder, weakness of the spleen and qi deficiency, internal accumulation of damp-heat, invigorate the spleen and boost qi, raise yang and stop diarrhea, Atractylodes macrocephala 15g, Saposhnikovia divaricata 10g, Chinese yam 20g, lotus seeds 20g, Terminalia chebula 15g, Gorgon fruit 20g, Amomum villosum 20g, Platycodon grandiflorum 10g, Poria cocos 20g, Coix lacryma-jobi 15g,",

[0121] output: "female / 41 / diarrhea, watery stools, borborygmus, nausea, vomiting, acid regurgitation, belching, fatigue, chills, thin tongue with light purple color, scanty and moist tongue coating, deep pulse, thready pulse, diarrhea / functional gastrointestinal disorder / weakness of the spleen and qi deficiency, internal accumulation of damp-heat / invigorate the spleen and boost qi, raise yang and stop diarrhea / Atractylodes macrocephala 15g, Saposhnikovia divaricata 10g, Chinese yam 20g, lotus seeds 20g, Terminalia chebula 15g, Gorgon fruit 20g, Amomum villosum 20g, Platycodon grandiflorum 10g, Poria cocos 20g, Coix lacryma-jobi 15g}"

[0122] S3-2: Similar to step S3-1, fill in the training data for the negative sample identification data in step 2, for example:

[0123] {input: "Male, 35 years old, had diarrhea, watery stools, borborygmus, nausea, vomiting, acid regurgitation, belching, accompanied by fatigue and chills after eating unclean food. The tongue was light purple with scanty and moist coating, and the pulse was deep and thready.)",

[0124] output: "male / 35 / watery stools, borborygmus, nausea, vomiting, acid regurgitation, belching, accompanied by fatigue and chills, light purple tongue with scanty and moist coating, deep and thready pulse / diarrhea after eating unclean food / Treatment: Codonopsis pilosula 20g, Atractylodes macrocephala 1g, Poria cocos 15g, Glycyrrhiza uralensis 6g, Citrus reticulata 12g, Pinellia ternata 10g, Armeniaca vulgaris 10g, Aster tataricus 12g, Tussilago farfara 12g, Platycodon grandiflorum 10g, Fritillaria thunbergii 12g, Trichosanthes kirilowii 10g}"

[0125] S3-3: Merge positive and negative sample identification data, divide the validation and test sets (divide the training set and test set at a ratio of 8:2), set up the training environment, define training parameters, and fine-tune the large model;

[0126] S3-4: Perform cyclic training, evaluate the final performance of the model on the test set, and finally construct a functional gastrointestinal disease model.

[0127] The rule engine of the present invention is the EasyRules rule engine.

[0128] The optimization of the rule engine includes the following steps (as Figure 3 shown):

[0129] 1. Collect medical record text data of functional gastrointestinal diseases and non-functional gastrointestinal diseases;

[0130] The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions;

[0131] 2. Manually review and verify the medical record text data, and classify the data into positive sample identification data and negative sample identification data;

[0132] Positive sample identification data means that symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases;

[0133] Negative sample identification data refers to symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all corresponding to non-functional gastrointestinal diseases;

[0134] The data in the above step 1 and step 2 can adopt the data in steps S1 and S2 of the functional gastrointestinal disease model constructed by fine-tuning based on the chatglm4-9b large model.

[0135] 3. Clean the positive sample identification data in step 2 to obtain the corresponding relationship between [gender / age / symptoms / four diagnostic signs] and [syndrome differentiation / treatment principles / prescriptions];

[0136] 4. Initialize the EasyRules rule engine, synchronize the corresponding relationship between [gender / age / symptoms / four diagnostic signs] and [syndrome differentiation / treatment principles / prescriptions] obtained in step 3, manually review the threshold of each relationship, fill it, and apply it to the EasyRules rule engine;

[0137] 5. Construct an algorithm for the subsequent scoring mechanism of the rule engine output result, and the subsequent scoring mechanism algorithm satisfies:

[0138] ① Rule output: When one or more rules are satisfied simultaneously, all matching results can be output;

[0139] ②Threshold weight scoring: Set a weight threshold N for each rule. When the number of records with the same result is X, the score of this classification result is X * N;

[0140] ③Result sorting and selection: Sort the results in descending order according to the calculated scores, and select the result with the highest score as the final output.

[0141] The optimization of the Elasticsearch search engine includes the following steps (as Figure 4 shown):

[0142] (1) Collect medical record text data of functional gastrointestinal diseases and non-functional gastrointestinal diseases;

[0143] The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions;

[0144] (2) Manually review and verify the medical records and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data;

[0145] Positive sample identification data means that symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases;

[0146] Negative sample identification data means that symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to non-functional gastrointestinal diseases;

[0147] The data in steps (1) and (2) using the functional gastrointestinal disease diagnosis model are the data in steps S1 and S2 fine-tuned and constructed based on the chatglm4-9b large model.

[0148] (3) Clean and filter the positive sample identification data in step (2) to form a list;

[0149] (4) Configure the Elasticsearch search engine, synchronize the list data in step (3), generate indexed data for retrieval, and construct a functional gastrointestinal disease data retrieval library;

[0150] (5) Reconstruct the scoring and sorting algorithm;

[0151] The reconstructed scoring and sorting algorithm is:

[0152] For the documents in the full-text retrieval results of medical records, a TF-IDF algorithm with the characteristics of functional gastrointestinal diseases is constructed according to the hit words for relevance scoring. Specifically: Term Frequency TF(t,d) = the number of times the domain term t appears in the document d * threshold / the total number of all words in the document d, Inverse Document Frequency IDF(m,D) = log(total number of all medical record documents D / number of retrieved and hit documents t), TF-IDF(t,d,m,D) = TF(t,d) * IDF(m,D). The TF-IDF values of the domain terms in the hit documents are accumulated, and finally sorted from high to low to obtain the top three;

[0153] (6) Combine the retrieval strategies for functional gastrointestinal diseases.

[0154] The steps for building the Fass vector library are as follows (as Figure 5 shown):

[0155] ① Collect the medical record text data of functional gastrointestinal diseases and non-functional gastrointestinal diseases;

[0156] The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions;

[0157] ② Manually review and verify the medical record and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data;

[0158] Positive sample identification data means that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases;

[0159] Negative sample identification data refers to that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to non-functional gastrointestinal diseases;

[0160] The data in the above step ① and step ② can adopt the data in steps S1 and S2 of the functional gastrointestinal disease model constructed by fine-tuning based on the chatglm4-9b large model.

[0161] ③ Clean and filter the positive sample identification data in step ② to form a list set;

[0162] ④ Configure the bge-m3 word vector model and the Fass vector initialization configuration;

[0163] ⑤ For the list set generated in step ④, synchronize the vector information to the Fass vector library through the bge-m3 word vector model;

[0164] ⑥ Package the input and output interfaces of the Fass vector library to make it applicable.

[0165] The present invention can significantly overcome the defects and limitations of the existing auxiliary diagnosis methods for gastrointestinal diseases. Compared with the auxiliary diagnosis of functional gastrointestinal diseases based on a rule library, specifically:

[0166] 1. The present invention has greater applicability and lower maintenance costs. By utilizing the dynamic characteristics of the rule engine, the present invention can supplement and update rule scripts in real time to ensure that new rules can take effect immediately. In this way, without affecting business continuity, the system can continuously adapt to new medical discoveries and clinical practices.

[0167] 2. The present invention has greater flexibility: By optimizing the rule engine, the present invention enables it to output all eligible results when matching multiple rules simultaneously, thereby enhancing the flexibility of the system and its ability to handle complex situations.

[0168] 3. The present invention can avoid over-reliance on expert knowledge: To address this, the present invention introduces a rule weight threshold and a scoring mechanism in the rule engine to overcome the problem of over-reliance on expert knowledge. For example, a subsequent scoring mechanism algorithm for the output results of the rule engine is constructed, and the subsequent scoring mechanism algorithm can meet the following requirements:

[0169] ① Rule output: When one or more rules are satisfied simultaneously, all matching results can be output;

[0170] ② Threshold weight scoring: A weight threshold N is set for each rule. When the number of records with the same result is X, the score of this classification result is X * N;

[0171] ③ Result sorting and selection: The results are sorted in descending order according to the calculated scores, and the result with the highest score is selected as the final output.

[0172] Through the optimization measures of the search engine, not only the flexibility and adaptability of the rule engine are improved, but also the autonomous decision-making ability of the system is enhanced, the dependence on a single expert opinion is reduced, and thus more accurate and reliable diagnostic support is provided.

[0173] Moreover, compared with the auxiliary diagnosis of functional gastrointestinal diseases based on machine learning, the present invention makes breakthroughs and innovations through the construction of a large model of functional gastrointestinal diseases based on the fine-tuning of the chatglm4-9b large model and its RAG method.

[0174] Regarding the issues of data quality, feature engineering, and continuous learning, the present invention constructs positive and negative sample data for the data set to form a positive and negative sample fine-tuning training set; Positive samples: Positive samples refer to those patient cases that meet the criteria for functional gastrointestinal diseases. Negative samples: Negative samples refer to those patient cases whose symptoms and signs do not conform to the diagnosis of functional gastrointestinal diseases and the corresponding syndromes. These cases may show some similar symptoms but are ultimately diagnosed with other syndromes or diseases. Finally, the large model is fine-tuned through instruction supervision fine-tuning to generate a new model with the characteristics of functional gastrointestinal diseases, and the continuous learning requirements can be completed through the cyclic iteration of the large model fine-tuning.

[0175] The present invention uses the FASS vector library retrieval, Elasticsearch semantic retrieval, and rule engine matching, a three-in-one approach to explain and generate evidence for the "black box" problem of model prediction, so as to avoid the problem of poor interpretability.

[0176] Moreover, in the present invention, the vocabulary table and scoring mechanism of the Elasticsearch search engine are innovated. First, all medical records and diagnostic information data are indexed using Elasticsearch. Then, the professional term vocabulary of functional gastrointestinal diseases is embedded in the Elasticsearch word segmentation table and the type and weight are defined. Specifically, for the documents in the full-text retrieval results of medical records, according to the hit words, a TF-IDF algorithm with the characteristics of functional gastrointestinal diseases is constructed based on the vocabulary threshold for relevance scoring. Specifically: term frequency TF(t,d) = the number of times the domain term t appears in the document d * threshold / the total number of all words in the document d, IDF(m,D) = log(total number of all medical record documents D / number of retrieved and hit documents t), TF-IDF(t,d,m,D) = TF(t,d) * IDF(m,D). The TF-IDF values of the domain terms in the hit documents are accumulated. Sort from high to low and obtain the top three.

[0177] Regarding the FASS vector library retrieval in the present invention, after segmenting the medical records and diagnostic information through the bge-m3 model, word embeddings are formed into vectors and entered into the FASS vector library. Then, word vectors are generated for the content input by the user, and similarity comparison is performed with the FASS vector library to obtain the top three records with the highest similarity.

[0178] The rule engine matching is as follows: entity extraction is performed on the text input by the user through a large model, and then the entity information is put into the rule engine. After being adapted by the rule engine, the corresponding matching results are found.

[0179] And in the present invention, through threshold definition, rule engine matching > Elasticsearch search engine > FASS vector library retrieval, with a ratio of 5:3:2, result sets obtained by processing the user input through the rule engine, Elasticsearch engine, and FASS vector library respectively are used to assist in the diagnosis using the constructed large model of gastrointestinal diseases, and finally a diagnostic plan for functional gastrointestinal diseases is output. Such a design scheme not only effectively overcomes the main disadvantages of the existing two diagnostic methods, but also provides a more intelligent, flexible, and easy-to-understand auxiliary tool for clinicians, greatly improving the accuracy and efficiency of the diagnosis of functional gastrointestinal diseases.

[0180] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine, characterized in that: The method comprises: Obtaining the user's diagnosis medical record text data, wherein the diagnosis medical record text data includes gender, age, symptoms, and four diagnostic signs; Perform named entity recognition on the diagnosis medical case text data, extract entities, and then input the entity information into the optimized rule engine to match the user's medical case text data, and output the rule engine matching results; After filtering the diagnostic medical record text data, input it into the optimized Elasticsearch search engine for retrieval, and output the Elasticsearch retrieval results; After converting the diagnostic medical record text data into word vectors through the bge-m3 model, it is searched in the FASS vector library for similarity comparison and the FASS search results are output; The rule engine matching results, Elasticsearch search results, and FASS search results are weighted and summarized according to the search threshold and proportion to obtain a comprehensive search result; By inputting the comprehensive search results into the functional gastrointestinal disease diagnostic model, auxiliary diagnosis of functional gastrointestinal disease can be performed and the diagnostic results can be output.

2. According to claim 1, a RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine is characterized in that: The functional gastrointestinal disease model is constructed by fine-tuning the chatglm4-9b model, and the construction steps are as follows: S1: Collect medical records of functional gastrointestinal disorders and non-functional gastrointestinal disorders; The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions; S2: Manually review and verify the medical record text data, and classify the data into positive sample identification data and negative sample identification data; The positive sample identification data is that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases; Negative sample identification data refers to symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions that all correspond to non-functional gastrointestinal diseases; S3: Use the command-following mode to fine-tune the chatglm4-9b model, define input and output samples, and finally build a functional gastrointestinal disease diagnosis model; Specifically, the steps for fine-tuning the chatglm4-9b large model are as follows: S3-1: Fill the training data for the positive sample identification data in step 2, input: [functional gastrointestinal disease medical case text], output: [gender / age / symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescription]; S3-2: Same as step S3-1, filling the negative sample identification data in step 2 with training data; S3-3: Merge positive and negative sample identification data, divide the validation test set, set up the training environment, define training parameters, and fine-tune the training large model; S3-4: Cycle training, evaluate the final performance of the model on the test set, and finally build a functional gastrointestinal disease diagnosis model.

3. The RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine according to claim 1 is characterized in that: The rule engine adopts EasyRules rule engine.

4. The RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine according to claim 3 is characterized in that: The optimization of the rule engine includes the following steps:

1. Collect medical records of functional gastrointestinal diseases and non-functional gastrointestinal diseases; The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions; 2. Manually review and verify the medical records and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data; The positive sample identification data is that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases; Negative sample identification data refers to symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions that all correspond to non-functional gastrointestinal diseases; 3. Clean the positive sample identification data in step 2 to obtain the corresponding relationship between [gender / age / symptoms / four diagnostic signs] and [syndrome differentiation / treatment principles / prescriptions]; 4. Initialize the EasyRules rule engine, synchronize the corresponding relationship between [gender / age / symptoms / four diagnostic signs] and [syndrome differentiation / treatment principles / prescriptions] obtained in step 3, manually review the threshold of each relationship, fill it in, and apply it to the EasyRules rule engine; 5. Construct a scoring mechanism algorithm for the output results of the rule engine, and the scoring mechanism algorithm satisfies: ①Rule output: When one or more rules are met at the same time, all matching results can be output; ②Threshold weight scoring: Set a weight threshold N for each rule. When the number of records with the same result is X, the score of the classification result is X*N; ③ Result sorting and selection: Arrange the results in descending order according to the calculated scores, and select the result with the highest score as the final output.

5. The RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine according to claim 1 is characterized in that: The optimization of the Elasticsearch search engine includes the following steps: (1) Collect medical records of functional gastrointestinal disorders and non-functional gastrointestinal disorders; The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions; (2) Manually review and verify the medical records and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data; The positive sample identification data is that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases; Negative sample identification data refers to symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions that all correspond to non-functional gastrointestinal diseases; (3) cleaning and filtering the positive sample identification data in step (2) to form a list; (4) configuring the Elasticsearch search engine, synchronizing the list data of step (3), generating indexed data for retrieval, and constructing a functional gastrointestinal disease data retrieval library; (5) Reconstruct the scoring and sorting algorithm; The reconstructed scoring and sorting algorithm is: For the documents in the full-text retrieval results of medical records, a TF-IDF algorithm with functional gastrointestinal disease characteristics was constructed according to the hit words to perform relevance scoring, specifically: word frequency TF(t,d) = the number of times domain word t appears in document d * threshold / the total number of all words in document d, IDF(m,D) = log(the number of all medical record documents D / the number of documents hit in the retrieval t), TF-IDF(t,d,m,D) = TF(t,d)*IDF(m,D), the TF-IDF values ​​of the domain words in the hit documents are accumulated, and finally sorted from high to low to obtain the top three; (6) Combined functional gastrointestinal disease search strategy.

6. The RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine according to claim 1 is characterized in that: The steps to build the Fass vector library are as follows: ① Collect medical records of functional gastrointestinal diseases and non-functional gastrointestinal diseases; The medical record text data includes gender, age, symptoms, four diagnostic signs, syndrome differentiation, treatment principles, and prescriptions; ②Manually review and verify the medical records and diagnostic text data, and classify the data into positive sample identification data and negative sample identification data; The positive sample identification data is that the symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions all correspond to functional gastrointestinal diseases; Negative sample identification data refers to symptoms / four diagnostic signs / syndrome differentiation / treatment principles / prescriptions that all correspond to non-functional gastrointestinal diseases; ③Cleaning and filtering steps ②Positive sample identification data to form a list set; ④Configure the bge-m3 word vector model and Fass vector initialization configuration; ⑤ For the list set generated in step ④, synchronize the vector information to the Fass vector library through the bge-m3 word vector model; ⑥ Encapsulate the Fass vector library input and output interfaces to make them applicable.

7. The RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine according to claim 1 is characterized in that: The rule engine matching results, Elasticsearch search results, and FASS search results are weighted and summarized according to the ratio of 5:3:

2.

8. A RAG large model functional gastrointestinal disease auxiliary diagnosis system based on rule engine and search engine, characterized in that: include: A data collection module is used to collect the user's diagnosis medical record text data, and the diagnosis medical record text data includes gender, age, symptoms, and four diagnostic signs; The retrieval analysis module is used to input the medical case text data input by the user into the rule engine, Elasticsearch engine, and FASS vector library for matching or retrieval, and to perform weighted aggregation calculation on the results output by the rule engine, Elasticsearch engine, and FASS vector library according to 5:3:2 to determine the comprehensive retrieval results; The functional gastrointestinal disease auxiliary diagnosis module uses the functional gastrointestinal disease model to diagnose based on the comprehensive search result data and outputs the diagnosis results.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the RAG large model functional gastrointestinal disease auxiliary diagnosis method based on rule engine and search engine as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the RAG large model functional gastrointestinal disease auxiliary diagnosis method based on a rule engine and a search engine as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Medical reasoning large model construction method and device, electronic equipment and storage medium

    CN120832956A

  • Medical reasoning large model construction method and device, electronic equipment and storage medium

    CN120832956B

  • Data retrieval method, system and equipment based on hybrid storage architecture and medium

    CN121434272A