Intelligent coated tongue inquiry device based on multi-modal knowledge collaboration

Through the intelligent tongue coating consultation device integrating the tongue image acquisition module, edge computing module and expert diagnosis engine module, the data quality and multimodal information fusion of the traditional Chinese medicine tongue diagnosis instrument are solved, and the efficient and accurate diagnosis of multi-round dialogue and consultation in traditional Chinese medicine is achieved, and the objectivity of traditional Chinese medicine diagnosis and human-computer interaction efficiency is improved.

CN120388764APending Publication Date: 2025-07-29TIANJIN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510447720.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing traditional Chinese medicine tongue diagnosis instruments have challenges in insufficient data quality and labeling, problems with multimodal information fusion and interactive diagnostic logic construction, which is difficult to meet the complex needs of traditional Chinese medicine dialectical treatment, and large language models lack the mastery of refined diagnostic processes in multiple rounds of dialogue.

Method used

An intelligent tongue coating consultation device is designed, integrating high-precision tongue image acquisition module, edge computing module and expert diagnosis engine module. Through tongue image analysis and multiple rounds of consultation, the accurate correlation between tongue image and disease and real-time diagnostic logic update is achieved.

Benefits of technology

It significantly improves the accuracy and objectivity of traditional Chinese medicine diagnosis, improves human-computer interaction efficiency, and builds a multi-round dialogue and triage data set to simulate real scenarios, enhancing the reliability and data security of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388764A_ABST
    Figure CN120388764A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent tongue coating inquiry device based on multi-modal knowledge cooperation, and the device comprises a tongue picture collection module which is internally provided with color temperature calibration and is used for the standardized collection of tongue color and coating quality characteristics; the interactive display screen integrates a touch panel and voice interaction, dynamically displays a tongue picture analysis result and an inquiry progress, and guides the patient to complete standardized tongue placement and symptom feedback of multiple rounds of dialogues with the active dialogue guide module; the edge calculation module integrates a lightweight hospital guide model and a visual module for tongue coating image processing, and is internally provided with a knowledge graph (KG, Known Graph) system for storing a dynamic knowledge graph to realize local knowledge reasoning and preliminary construction; and the expert diagnosis engine module is connected with the device through a bidirectional encryption channel, integrates an expert model deployed at a cloud end, and outputs a final diagnosis result according to the tongue picture characteristics and the symptom information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medical diagnosis devices, and particularly to a multimodal traditional Chinese medicine diagnosis device that integrates tongue image visual analysis and active multi-round dialogue interaction, and more particularly to an intelligent tongue coating interrogation device that realizes the integration of "inspection-interrogation" through a dynamic knowledge graph to coordinate multimodal models. Background Art

[0002] With the rapid rise of large language models (LLMs) in artificial intelligence, large models such as ChatGPT and LLaMA have made remarkable progress in aspects such as language understanding, text generation, and knowledge reasoning. These models have achieved a qualitative leap in generalization ability and context understanding by leveraging large-scale pre-trained corpora and advanced training strategies. However, when applied to the highly specialized medical field, problems have emerged: their generality causes the models to lack mastery of refined diagnostic processes, and it is particularly difficult to fully extract implicit patient condition information in multi-round dialogue scenarios. To improve applicability in medical scenarios, researchers have continuously tried to introduce methods such as knowledge graphs or RAG (Retrieval-Augmented Generation) to assist the models in invoking domain knowledge, but there are still shortcomings in the dynamic interrogation of patients in real scenarios after multiple iterations. Therefore, how to design a technical solution that can support high-quality multi-round dialogue interrogation in traditional Chinese medicine has become an urgent problem in the modernization, digitization, and intelligentization of traditional Chinese medicine.

[0003] Existing traditional Chinese medicine large models, such as BianQue, HuatuoGPT, and Zhongjing, obtain traditional Chinese medicine professional knowledge by fine-tuning a large amount of traditional Chinese medicine texts, clinical literature, and professional Q&A data on the basis of basic large models. These models are relatively mature in answering basic pathology, formula usage, and common diseases. However, due to the lack of high-quality, multi-round traditional Chinese medicine dialogue datasets, they can still only give short-range answers to single questions.

[0004] On this basis, in recent years, tongue diagnosis instrument technology has gradually received attention and demonstrated unique advantages in traditional Chinese medicine diagnosis. Traditional Chinese medicine tongue diagnosis relies on physicians' subjective observation of tongue images, which is often limited by factors such as experience level and environmental light, resulting in poor objectivity and consistency of diagnostic results. Existing tongue diagnosis instruments also face many challenges in practical applications: insufficient data quality and annotation: high-quality tongue image and multi-round interactive interrogation data are scarce, restricting the training and optimization of the system; the problem of multimodal information fusion: how to efficiently fuse images, texts, and expert knowledge to achieve precise association between tongue images and diseases is still a technical bottleneck; the construction of interactive diagnosis logic: existing solutions are not yet mature in terms of active questioning and self-updating diagnosis logic, and it is difficult to meet the complex requirements of traditional Chinese medicine syndrome differentiation and treatment.

[0005] Therefore, combining the advanced large language model with the multi-modal interaction ability of the intelligent tongue diagnosis instrument, the research and development of an intelligent consultation device that can not only automatically collect tongue coating images, but also actively conduct multi-round consultations and update the diagnostic logic in real time has become an important direction to promote the digitalization and intelligent development of traditional Chinese medicine. Solving these challenges can not only improve the accuracy and objectivity of traditional Chinese medicine diagnosis, but also provide strong technical support for the modern transformation of traditional medicine. Summary of the Invention

[0006] The present invention provides an intelligent tongue coating consultation device based on multi-modal knowledge collaboration. By integrating a high-precision tongue image acquisition module, an edge computing module, and an expert diagnosis engine module, the present invention realizes the full-process hardwareization from tongue image analysis to disease diagnosis. The device supports improving the accuracy of disease recognition under the condition of limited dialogue data through tongue image photos, multi-round symptom inquiries, and dynamic updates of the knowledge graph, and is equipped with a multi-round dialogue dataset generation technology to support the implementation and performance improvement of the above functions. All calculation and storage processes are realized through the edge computing module and the cloud collaborative interface module built into the device to ensure real-time performance and data security. See the following description for details:

[0007] An intelligent tongue coating consultation device based on multi-modal knowledge collaboration, the device includes:

[0008] Tongue image acquisition module: The device is built-in with color temperature calibration for the standardized acquisition of tongue color and coating texture features;

[0009] Interactive display screen: Integrating a touch panel and voice interaction, dynamically displaying the tongue image analysis results and the consultation progress, and guiding the patient to complete the standardized tongue placement and symptom feedback for multi-round conversations with the active dialogue guidance module;

[0010] Edge computing module: Integrating a lightweight guidance model and a visual module for tongue coating image processing, and built-in with a KG system for storing dynamic knowledge graphs to realize local knowledge reasoning and preliminary construction;

[0011] Expert diagnosis engine module: Connecting to the device through a two-way encryption channel, integrating an expert model deployed in the cloud, and outputting the final diagnosis result according to the tongue image features and symptom information.

[0012] Among them, the edge computing module includes:

[0013] Construct a traditional Chinese medicine knowledge graph, which is constructed using data on traditional Chinese medicine disease types and their corresponding symptoms;

[0014] In multi-round conversations, identify the next set of symptoms to be asked, ask questions and parse user responses, update the symptom recorder, and decide whether to exit the loop; this process continues until the confidence in disease differentiation exceeds a predefined threshold;

[0015] For each unknown symptom j, calculate the importance score by weighting the edge weight w between symptom j and disease i ji and its corresponding influence factor S i , and calculate the importance score;

[0016] Sort according to the FinalScore(j) of all unknown symptoms, and select the two to three symptoms with the highest scores as candidate symptoms for inquiry;

[0017] After receiving the patient's response, update the symptom recorder, and determine whether to exit the inquiry loop. When the cosine similarity of the currently most likely candidate disease exceeds the predefined threshold, the appropriate disease has been identified, marking the end of the diagnostic dialogue. The collected user symptom information and diseases will be input into the expert model for final diagnosis;

[0018] The expert model will propose update suggestions to strengthen or weaken the weights between symptoms and symptoms, and between symptoms and diseases.

[0019] Among them, the nodes in the graph represent symptoms and disease types, and the edges and their weights W represent the degree of association. The initial weights are assigned based on the frequencies that appear in the dataset.

[0020] The device further includes: constructing an interactive diagnostic algorithm;

[0021] (1) Preliminary symptom extraction: From the patient's original query input Q, extract a preliminary set of symptoms through a graph encoder;

[0022] (2) Interactive inquiry and update process: In the preset inquiry rounds, for each current symptom: find the related diseases PD of this symptom in the knowledge graph G and sort them by importance; for each candidate disease D, find the related symptoms PS from the graph and also sort them by importance; according to the sorting results, select some key symptoms from them, and then ask the patient about these symptoms to get the patient's answer PR; update the current set of symptoms according to the patient's answer, and calculate the cosine similarity between the current symptom and the candidate disease; if the similarity reaches the preset diagnostic threshold ∈, then end the interaction process;

[0023] (3) Final diagnosis and knowledge graph update: Through the expert model, synthesize the current symptoms and candidate diseases to generate a final diagnosis answer, and further update the symptom information in the knowledge graph.

[0024] Among them, the generated dialogue covers the entire medical process, including the cycle of the patient describing symptoms, the doctor asking questions and the patient's response, as well as the doctor providing a final diagnosis, simulating a complete real scenario.

[0025] The beneficial effects of the technical solution provided by the present invention are:

[0026] 1. The present invention innovatively develops an intelligent tongue coating consultation device. By integrating a high-precision tongue image acquisition module and an edge computing module, it breaks through the limitations of single-modal diagnosis of traditional Chinese medicine equipment. The device has a consultation model, a vision model, and a KG system built into the edge computing module, which form a collaborative diagnosis link with the expert model in the expert diagnosis engine module. Without additional data fine-tuning, it can not only maintain the integrity of the professional knowledge base of the Chinese medicine large model but also achieve an intelligent consultation process driven by a dynamic knowledge graph, significantly improving the efficiency of human-computer interaction and the reliability of diagnosis.

[0027] 2. The present invention constructs a multi-round dialogue triage dataset, which approximates real triage situations through real cases, eliminating the difficulty of manually collecting triage data. In addition, a small dataset with low-validity information is extracted from the complete dataset to simulate the real situation of patient consultation to evaluate the robustness of the device model.

[0028] 3. The accuracy of the inquiry results of the consultation mechanism of the present invention on the dataset reaches 84.68%, significantly improving the communication ability between the model and the patient during the inquiry process. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a physical diagram of the device of the present invention, where the edge computing module, as the intelligent brain, is responsible for building a model for diagnosis;

[0030] Figure 2 is a schematic diagram of the dialogue between the present invention and the patient;

[0031] Figure 3 is a flowchart of the model working in the device of the present invention;

[0032] Figure 4 is a schematic diagram of comparing the models of the present invention with four models, namely DEEPSEEK-v3, CHATGPT-4o, HuatuoGPT, and BianQue, using LLM.

[0033] Table 1 is a schematic diagram of the comparative evaluation of the present invention and existing medical models in terms of inquiry metrics;

[0034] Table 2 is a schematic diagram of the comparative evaluation of the present invention and large-scale LLM in terms of inquiry metrics. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.

[0036] Example 1

[0037] This embodiment provides an intelligent tongue coating consultation device based on multi-modal knowledge collaboration, which integrates a high-precision tongue image acquisition module, an edge computing module, and an expert diagnosis engine module. This embodiment will introduce in detail the device composition of the present invention and the patient usage process:

[0038] The intelligent tongue coating consultation device adopts an integrated design, and the main body includes:

[0039] Tongue image acquisition module: It adopts an integrated optical design. The ring-shaped shadowless supplementary light group eliminates shadow interference, and cooperates with an autofocus macro camera to achieve high-definition imaging of the entire tongue area. The device is built-in with a color temperature calibration module, which can automatically adapt to different ambient light conditions to ensure the standardized acquisition of key features such as tongue color and coating texture.

[0040] Interactive display screen: It integrates a touch panel and a voice interaction module, provides a visual operation interface and a voice guidance function, can dynamically display the tongue image analysis results and the consultation progress, and guides the patient to complete the standardized tongue placement and symptom feedback for multi-round conversations with the active dialogue guidance module;

[0041] Edge computing module: It integrates a lightweight guidance model and a visual module for tongue coating image processing, and is built-in with a KG system for storing dynamic knowledge graphs to achieve local knowledge reasoning and preliminary construction. This module supports fast response and local data processing, has a data caching mechanism, can provide basic diagnosis functions and local processing of privacy data in an offline state, and provides support for the update of the subsequent knowledge graph.

[0042] Expert diagnosis engine module: It is connected to the device through a two-way encryption channel, and integrates an expert model deployed in the cloud, which can output the final diagnosis result according to the tongue image characteristics and symptom information, and return it to the device. The differential transmission technology is adopted to only upload key feature data to reduce the communication bandwidth requirements.

[0043] Patient usage process:

[0044] Tongue image acquisition stage: The patient gently places the lower jaw on the device bracket, naturally extends the tongue according to the voice prompt, automatically triggers the supplementary light system, and the camera takes pictures of the tongue surface, tongue side, and sublingual area from multiple angles, and simultaneously completes the structured extraction of features such as tongue color (light red / crimson red), coating texture (thin white / yellowish and greasy), and tongue shape (swollen / lobed). The visual module generates a tongue image feature vector and triggers the subsequent consultation process;

[0045] Active Inquiry Interaction Phase: First, the patient briefly describes the main complaint symptoms through the touch screen or voice (such as "often feel tired recently"). Based on the disease-symptom association rules of the knowledge graph and combined with the preset traditional Chinese medicine syndrome differentiation logic, the KG system automatically generates a multi-round inquiry strategy. The triage model conducts natural language interaction and preferentially asks about symptoms that have differential diagnostic value for the current main complaint (such as "Do you have chills or spontaneous sweating?"). Each answer from the patient is recognized and transformed into a structured symptom feature vector by the triage model. The KG system uses the cosine similarity algorithm to match the syndrome type nodes in the knowledge graph in real time and dynamically adjusts the subsequent questioning path until the completeness verification of the symptom set is completed;

[0046] Multi-modal Diagnosis Generation Phase: The expert diagnosis engine integrates the tongue image feature vector and symptom information, and the diagnosis result is presented in the form of a structured report, including the syndrome type conclusion, confidence score, recommended prescriptions, and personalized health preservation suggestions; The knowledge graph automatically optimizes the symptom-syndrome type association weight according to the diagnosis data of this time to achieve continuous iteration of the model.

[0047] This embodiment provides a complete hardware implementation solution for the intelligent tongue coating inquiry device. All algorithms and models (including the visual analysis model, triage dialogue model, expert diagnosis model, and dynamic knowledge graph) are deployed on this device through hardware-level optimization.

[0048] Embodiment 2

[0049] The following further introduces the models of the devices in Embodiment 1 in combination with specific examples and calculation formulas. See the following description for details:

[0050] I. Multi-round Dialogue Mechanism Based on the Knowledge Graph of the KG System (Edge Computing Module)

[0051] Details of the multi-round dialogue questioning based on the Knowledge Graph of the KG system are introduced. By mapping the initial description of the patient to the symptom nodes and combining the questioning steps of multiple iterations, the possible disease range is gradually narrowed to achieve efficient and accurate discrimination of the patient's condition.

[0052]

[0053] 1. Knowledge Graph Construction

[0054] First, a traditional Chinese medicine knowledge graph is constructed as the core of the multi-round dialogue consultation. This knowledge graph is constructed using the existing data of traditional Chinese medicine disease types and their corresponding symptoms. The nodes in the graph represent symptoms and disease types, and the edges and their weights W represent the degree of association. The initial weights are assigned based on the frequency of occurrence in the dataset.

[0055] For the initial description provided by the user, after being parsed by the triage model, the symptoms that can be mapped to the knowledge graph are recorded as the known symptom group, and the patient symptom vector P is constructed:

[0056]

[0057] In the KG system, by traversing all diseases in the graph, for each disease i, its symptom vector D is defined i :

[0058] D i =(d i1 ,d i2 ,...,d in ),d ik =w ik

[0059] where w ki represents the edge weight between symptom k and disease i, and the value ranges from 0 to 1. By calculating the cosine similarity S between the patient's current condition and disease i i :

[0060]

[0061] where S i is used as the importance score of disease i. The higher the score, the greater the likelihood that the patient has disease i.

[0062] 2. Active multi-round dialogue mechanism and Gaussian noise perturbation

[0063] In the multi-round dialogue, multiple modules are iteratively executed: identifying the next set of symptoms to be asked, asking questions and parsing the user response, updating the symptom recorder, and deciding whether to exit the loop. This process continues until the confidence in disease differentiation exceeds a predefined threshold.

[0064] For each unknown symptom j, by weighting the edge weight w between symptom j and disease i ji and its corresponding influence factor S i , its importance score Score(j) is calculated:

[0065]

[0066] To ensure the robustness of symptom discovery, a Gaussian noise perturbation mechanism based on the normal distribution is introduced. Specifically, a random perturbation term that follows the N(0, σ 2 ) distribution is superimposed on the original importance score of each unknown symptom:

[0067] FinalScore(j)=Score(j)+∈, ∈~N(0, σ 2 )

[0068] Sort according to the FinalScore(j) of all unknown symptoms, and select the two to three symptoms with the highest scores as candidate symptoms for inquiry.

[0069] 3. Symptom recording and knowledge graph update

[0070] After receiving the patient's response, update the symptom recorder. Subsequently, determine whether to exit the inquiry loop. When the cosine similarity of the currently most likely candidate disease exceeds a predefined threshold, it can be considered that the appropriate disease D has been identified. final Marks the end of the diagnostic dialogue. The user symptom information and disease D collected final Will be input into the expert model for final diagnosis.

[0071] Considering the dynamic path selection and weight adjustment mechanism in routing data transmission, as well as the parameter update characteristics of backpropagation through the chain rule in graph neural networks, a knowledge graph update mechanism is introduced. Based on the final diagnosis of the expert model and the information during the consultation process, the expert model will propose update suggestions to strengthen or weaken the weights between symptoms, symptoms and diseases. This update process is only carried out after the conversation with the patient ends, ensuring no additional delay in the conversation.

[0072] Based on the above method, the present invention proposes an interactive diagnosis algorithm, and the process of this algorithm is as follows:

[0073] (1) Preliminary symptom extraction: From the patient's original query input Q, extract a preliminary set of symptoms through a graph encoder.

[0074] (2) Interactive inquiry and update process: In the preset inquiry rounds, for each current symptom: Find the possible diseases PD related to this symptom in the knowledge graph G and sort them by importance; for each candidate disease D, find the possible related symptoms PS from the graph and also sort them by importance; according to the sorting results, select some key symptoms from them, and then ask the patient about these symptoms to get the patient's answer PR; Update the current set of symptoms according to the patient's answer, and calculate the cosine similarity between the current symptom and the candidate disease; if the similarity reaches the preset diagnosis threshold ∈, then end the interaction process.

[0075] (3) Final diagnosis and knowledge graph update: Through the expert model, synthesize the current symptoms and candidate diseases to generate a final diagnosis answer, and further update the symptom information in the knowledge graph.

[0076] II. Patient guidance model and visual model (edge computing module)

[0077] 1. Patient guidance model

[0078] The medical guidance model in the present invention uses a relatively small-parameter-scale LLM as the core, thereby significantly reducing the inference cost while ensuring the logic and accuracy of traditional Chinese medicine (TCM) interrogation. The model mainly plays a role in the following three aspects:

[0079] (1) Simulating the real interaction process between TCM physicians and patients: The model can have continuous multi-round conversations with patients, and based on the TCM syndrome differentiation thinking, ask patients whether they have specific symptoms or signs. This can dynamically collect and update the known symptom information of patients in a timely manner, providing a more comprehensive reference for the condition.

[0080] (2) Understanding and analyzing patients' answers: The guidance model analyzes the natural language input of users, automatically extracts or matches the sorted TCM symptom terms, and updates its symptom set in real time. This technology can combine a dictionary or knowledge base with the function of bidirectional mapping between TCM professional terms and common sayings to achieve both professionalism and popularity when asking questions and analyzing answers.

[0081] (3) Aligning and standardizing the language between doctors and patients: When the model asks questions to patients, it uses a more colloquial expression method with the help of natural language generation technology to prevent information understanding deviation; after the patient answers, the corresponding content is then mapped or compared to the professional symptom entries in the TCM knowledge graph to ensure the consistency and accuracy of information in the subsequent diagnosis process.

[0082] 2. Visual model

[0083] In the present invention, in addition to receiving information from the conversation with patients during the interrogation stage, tongue diagnosis is introduced as auxiliary information. Patients submit photos of their tongue coatings as additional input. Tongue diagnosis is an important technique in TCM diagnostics, which evaluates the health status and disease characteristics of the human body by observing the color, shape, and characteristics of the tongue coating. However, since the accuracy of tongue coating recognition cannot be fully guaranteed and is affected by various accidental factors, instead of using a multi-modal large model to regard the tongue coating as an additional visual modality. On the contrary, the visual model of the present invention first uses a ResNet convolutional neural network to obtain tongue image features from the tongue coating image, which, together with the symptom information, provides more comprehensive patient information for the expert model, thereby improving the effectiveness of diagnosis and treatment.

[0084] III. Expert model (expert diagnosis engine module)

[0085] The expert model in the expert diagnosis engine of the present invention uses the Sun Simiao TCM large model to provide the final diagnosis and suggestions. The model has been fine-tuned on a large number of high-quality TCM datasets and has rich TCM knowledge. It uses this model to receive the symptom information and disease syndrome differentiation information collected by the guidance model and the tongue image features extracted by the visual model, and gives a diagnosis result and treatment plan based on TCM professional knowledge.

[0086] After ending the conversation with the patient, the expert model will recommend updating the knowledge graph in the KG system. This includes strengthening the connections between key symptoms and diseases as well as between key symptoms themselves, while weakening the connections between misleading symptoms and diseases and between misleading symptoms and other symptoms.

[0087] IV. Construction of the multi-round dialogue dataset

[0088] Due to the scarcity of dialogue data, the present invention constructs a multi-round doctor-patient dialogue dataset using classical traditional Chinese medicine knowledge data. This process relies on the capabilities of large language models, and by assigning the roles of patients and doctors to the large language models, this method has shown superiority in previous studies.

[0089] First, the data is organized into multiple pairs of "disease + symptom list" as basic medical knowledge units. Subsequently, for each set of knowledge, it is assumed that the patient only has this specific disease and exhibits all the symptoms in the symptom list. On this basis, the qwen-plus model is called to construct doctor-patient dialogues.

[0090] The generated dialogues cover the entire medical process, including the cycle of the patient describing symptoms, the doctor asking questions and the patient's responses, as well as the doctor providing a final diagnosis, simulating a complete real scenario. Before each question, the doctor's questions about symptoms are calculated based on the known information in the knowledge graph, rather than randomly selected. This effectively simulates the logical thinking process of traditional Chinese medicine doctors based on relevant medical knowledge.

[0091] To ensure the robustness of natural language during the dialogue process, several basic requirements are defined for the large language models of patients and doctors:

[0092] (1) Common sense: To better conform to the actual situation, the patient's initial description should focus on the main and easily observable symptoms, and conduct an objective dialogue in line with the patient's role.

[0093] (2) Colloquialism: Assuming that the patient does not have extensive professional medical knowledge, the doctor should ask questions in an easy-to-understand language, and the patient's answers should not contain irrelevant information beyond the symptoms.

[0094] (3) Honesty: Assuming that the patient can accurately judge the symptoms asked by the doctor and will not report any non-existent symptoms. This ensures that the doctor receives correct information and reduces the uncertainty in the multi-round dialogue dataset.

[0095] In the present invention, the large language model is set to play the roles of a patient and a doctor. The patient model is responsible for describing symptoms, and the doctor model is responsible for asking questions based on the information in the knowledge graph and providing a diagnosis at the end of the conversation. This role setting not only simulates a real medical conversation scenario but also improves the authenticity and effectiveness of the conversation through the interaction of the models.

[0096] In summary, through the above-mentioned parts, the embodiments of the present invention can conduct continuous conversations while maintaining their professional knowledge and achieve self-update and iteration without relying on fine-tuning of multi-turn conversation datasets, which not only maintains flexibility but also improves recognition efficiency, and reduces the risk of misdiagnosis to a certain extent. Relevant experiments have verified the robustness and practicality of this method in a dynamic environment.

[0097] Embodiment 3

[0098] The following combines a table and specific experimental data to verify the feasibility of the model in Embodiment 2, as detailed in the following description:

[0099] The experiment was conducted on two NVIDIA A6000 GPUs with a total video memory of 96GB. Other models used in the experiment (except ChatGPT, DeepSeek, and Qwen2.5-Max, which used the provided APIs) were deployed through the Hugging Face with Transformer framework.

[0100] The present invention uses the previously mentioned multi-turn doctor-patient conversation dataset, which contains more than two thousand high-quality data. In the experiment, only the initial description containing partial patient symptoms was used as the input, and the medical model's diagnosis of the patient's disease was output through an automated test program. The following models were selected as the baselines for the experiment:

[0101] · Qwen2.5-Max is a general Chinese language model that supports long context and multi-turn conversations;

[0102] · Sunsimiao is a Chinese medical domain language model that does not have the ability of multi-turn conversations;

[0103] · HuatuoGPT is a Chinese language model that is fine-tuned with instructions using a medical domain knowledge dataset. Since the model weights have not been open-sourced, manual tests were conducted on the web page;

[0104] · BianQue is a Chinese medical domain model that focuses on enhancing the ability to ask questions;

[0105] · ChatGPT (ChatGPT-4o) is one of the most popular multi-modal LLMs, with excellent multi-turn conversation and logical reasoning abilities;

[0106] · DeepSeek (DeepSeek-v3) is an LLM based on the MOE architecture, with excellent efficiency and performance compared to other mainstream models.

[0107] I. Metric Evaluation

[0108] The present invention uses the following three metrics to measure the ability of a medical model to interrogate patients:

[0109] (1) Diagnostic accuracy rate, defined as the percentage of the number of correct diagnoses made by the medical model for patients out of the total number of patients, used to measure the effectiveness of the medical model in diagnosis:

[0110]

[0111] where P represents the number of all patients, and P c represents the number of patients whose diseases finally diagnosed by the model are consistent with the diseases they suffer from.

[0112] (2) Q&A ratio, used to measure the initiative of the model during the questioning process:

[0113]

[0114] where Q m represents the number of rounds of questions the model asks the patient during the consultation, A i m represents the number of rounds of the model answering the patient (or making a diagnosis) during the consultation, and n is the number of patients in the test set.

[0115] (3) Questioning distance, used to compare the difference in the initiative of the model and professional doctors in asking questions during the questioning process:

[0116]

[0117] where Q m and A m represent the number of rounds of the model asking and answering the patient during the questioning process, Q d and A d represent the number of rounds of professional doctors asking and answering patients in the dataset, and n is the number of patients in the test set.

[0118] The experimental results are as follows. The interrogation effectiveness of the present invention was compared with existing medical models of the same scale (as shown in Table 1), and also compared with large-scale LLMs (as shown in Table 2).

[0119] Table 1 Comparison of the present invention and existing medical models in interrogation metrics

[0120]

[0121] As shown in Table 1, compared with existing medical models of the same scale, the present invention has stronger multi-round dialogue capabilities, can obtain more information about patients through active questioning, thus making a more accurate diagnosis of the patients, and providing effective suggestions under the condition of equivalent existing medical knowledge volume.

[0122] Table 2 Comparison between the present invention and large-scale LLM in terms of interrogation metrics

[0123]

[0124] As shown in Table 2, the present invention also has advantages compared with large-scale LLM. Although these large-scale LLM have excellent multi-round dialogue and logical reasoning, their questions during the interrogation are not based on pre-existing medical knowledge, so they usually cannot obtain more meaningful diagnostic information through the interrogation. While the present invention establishes rules based on the connections between different symptoms and the relationships between symptoms and diseases, and asks questions according to these rules, thus being more likely to obtain additional information about the patient and make a more accurate diagnosis.

[0125] II. LLM Evaluation

[0126] LLM has been proven to be an effective natural language evaluator. Using LLM to conduct comparative evaluation on texts can effectively reflect the quality of the generated texts.

[0127] In this evaluation, two sets of doctor-patient dialogues and evaluation metrics were provided to LLM, guiding it to evaluate which doctor model performs better according to the specified criteria and clarify the reasons behind its judgment. In the case where the superior performer is distinguishable, the task of LLM is to determine the winner; on the contrary, if the gap between the two is negligible, a tie is declared. Acknowledging the randomness of LLM, each comparative evaluation was conducted five times, and the most common judgment was adopted as the final result.

[0128] The evaluation includes four key criteria:

[0129] · Knowledge ability: This refers to the amount of medical knowledge demonstrated in the dialogue, including the ability to answer questions related to diseases and make disease judgments;

[0130] Diagnostic doctrine: Evaluate whether the provided diagnostic suggestions are professional and comprehensive;

[0131] Fluency: Evaluate the naturalness and fluency of the dialogue to ensure it is very similar to real-world consultation scenarios;

[0132] Respect: This measures the degree of full respect of the dialogue for privacy and other sensitive issues, and whether the language used expresses sufficient concern for the patient.

[0133] The four models of DEEPSEEK-v3, CHATGPT-4o, HuatuoGPT, and BianQue were comparatively evaluated through this device. The evaluation results are as Figure 4 shown. In most aspects of the four evaluation criteria, the embodiments of the present invention outperform the baseline model, especially showing significant advantages in terms of knowledge and professionalism.

Claims

1. An intelligent tongue coating interrogation device based on multi-modal knowledge collaboration, characterized in that The device includes: Tongue image acquisition module: The device has built-in color temperature calibration for standardized acquisition of tongue color and coating characteristics. Interactive display screen: Integrating a touch panel and voice interaction, it dynamically displays the results of tongue image analysis and the progress of the medical interview, guiding the patient to complete the standardized tongue placement and symptom feedback in multiple rounds of dialogue with the active dialogue guidance module. Edge computing module: Integrating a lightweight guidance model and a visual module for tongue coating image processing, and built-in with a KG system for storing dynamic knowledge graphs to achieve local knowledge reasoning and preliminary construction. Expert diagnosis engine module: Connects to the device through a two-way encrypted channel, integrating an expert model deployed in the cloud, and outputs the final diagnosis result based on tongue image characteristics and symptom information.

2. The intelligent tongue coating interrogation device based on multi-modal knowledge collaboration according to claim 1, characterized in that The edge computing module includes: Construct a traditional Chinese medicine knowledge graph, which is constructed using data on traditional Chinese medicine disease types and their corresponding symptoms. In multiple rounds of dialogue, identify the set of symptoms to be asked next, ask questions and parse the user response, update the symptom recorder, and decide whether to exit the loop; this process continues until the confidence in differentiating diseases exceeds a predefined threshold. For each unknown symptom j, calculate the importance score Score(j) by weighting the edge weight w between symptom j and disease i ji and its corresponding influence factor S i , and add a Gaussian noise random perturbation term ∈ to it to obtain FinalScore(j); Sort according to the FinalScore(j) of all unknown symptoms, and select the two to three symptoms with the highest scores as candidate symptoms for questioning. After receiving the patient's response, update the symptom recorder, judge whether to exit the questioning loop. When the cosine similarity of the currently most likely candidate disease exceeds the predefined threshold, the appropriate disease has been identified, marking the end of the diagnostic dialogue. The user symptom information and disease collected will be input into the expert model for final diagnosis. The expert model will propose update suggestions to strengthen or weaken the weights between symptoms, symptoms and diseases.

3. The intelligent tongue coating consultation device based on multi-modal knowledge collaboration according to claim 1, characterized in that, The nodes in the graph represent symptoms and disease types, and the edges and their weights W represent the degree of association. The initial weights are assigned based on the frequencies appearing in the dataset.

4. An intelligent tongue coating interrogation device based on multi-modal knowledge collaboration according to claim 1, characterized in that, The device also includes: constructing an interactive diagnostic algorithm. (1) Preliminary symptom extraction: From the patient's original query input Q, extract the preliminary symptom set through a graph encoder. (2) Interactive questioning and updating process: In the preset questioning rounds, for each current symptom: find the diseases PD related to this symptom in the knowledge graph G and sort them by importance; for each candidate disease D, find the related symptoms PS from the graph and also sort them by importance; according to the sorting results, select some key symptoms from them and then ask the patient about these symptoms to get the patient's answer PR; update the current symptom set according to the patient's answer and calculate the cosine similarity between the current symptom and the candidate disease; if the similarity reaches the preset diagnostic threshold ∈, then end the interaction process. (3) Final diagnosis and knowledge graph update: Through the expert model, synthesize the current symptoms and candidate diseases to generate the final diagnosis answer and further update the symptom information in the knowledge graph.

5. An intelligent tongue coating inquiry device based on multi-modal knowledge collaboration according to claim 1, characterized in that, The generated dialogue covers the entire medical process, including the cycle of the patient describing symptoms, the doctor asking questions and the patient responding, as well as the doctor providing the final diagnosis, simulating a complete real scenario.

Citation Information

Cited By

  • Medical service system and method based on multi-granularity semantic analysis and knowledge reasoning engine

    CN120578745A