Intelligent doctor-patient dialogue term simplification method and system based on large language model
Through the intelligent doctor-patient dialogue system based on a large language model, the problems of information asymmetry and insufficient emotional regulation in online medical care are solved, personalized medical terminology simplification and emotionally adaptive interpretation are achieved, and the efficiency of doctor-patient communication and patient satisfaction are improved.
Patent Information
- Application Number
- CN202510950494.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-28
AI Technical Summary
In online medical services, information asymmetry is a prominent issue in doctor-patient communication. Patients have difficulty understanding professional medical terminology, lack personalized support, and have insufficient emotional regulation, which affects communication efficiency and trust.
The intelligent doctor-patient dialogue system based on a large language model uses named entity recognition, multimodal modules, and sentiment analysis to adjust the interpretation content according to the patient's background and emotions, providing personalized methods for simplifying medical terminology and combining the patient's personal qualities and emotional state for popular explanations.
It improves the efficiency of doctor-patient communication, enhances patients' understanding and satisfaction, reduces information asymmetry, increases doctor-patient trust, and ensures the professionalism and accuracy of the explanations.
Smart Images

Figure CN120853995A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart healthcare technology, and relates to methods and systems that integrate artificial intelligence and medical information, particularly to a method and system for simplifying intelligent doctor-patient dialogue terminology based on a large language model. Background Art
[0002] In today's era of digitalization and intelligentization, the medical field is undergoing unprecedented transformation, with online medical services gradually emerging as an indispensable and crucial component of the healthcare system. This emerging model brings numerous conveniences to patients: remote consultations break down geographical barriers, allowing patients to receive professional medical advice without long-distance travel; health management services leverage smart devices and data analysis to provide patients with personalized health monitoring and guidance; and post-operative guidance, delivered through online platforms, ensures that patients continue to receive care and support after discharge.
[0003] However, while online healthcare services have demonstrated enormous potential, they have also exposed some pressing issues. Deficiencies in information delivery, patient understanding, and user experience act as barriers, hindering their further development. Among these, the information asymmetry in doctor-patient communication is particularly prominent.
[0004] During online doctor-patient communication, doctors often use a large number of highly specialized medical terms to ensure the accuracy and professionalism of their diagnoses. These terms cover a wide range of aspects, including disease categories, drug names, symptoms and signs, and surgical procedures. While these terms are essential tools for daily communication for medical professionals, they can be as obscure and difficult to understand as a foreign language for the average patient. For example, when describing a rare disease, a doctor might use specialized pathological terminology that the patient may never have heard of, making it difficult to comprehend their meaning. This undoubtedly increases the difficulty for the patient to understand their condition and treatment plan.
[0005] The lack of personalized support further exacerbates the problem of information asymmetry. Patients vary significantly in age, education level, and health awareness, leading to different needs and comprehension abilities regarding medical information. Younger patients typically possess strong learning abilities and a high receptiveness to new information; they may be more willing to accept more specialized medical explanations and hope to better cooperate with treatment by gaining a deeper understanding of their condition. Older patients with lower levels of education, however, may prioritize clear and intuitive explanations; they require doctors to use simple and easy-to-understand language to explain their condition and treatment plans. However, current online medical services often fail to adequately consider these differences and cannot provide personalized information support for different patients.
[0006] Furthermore, patient emotional regulation is a crucial issue that cannot be ignored. When patients feel anxious or confused, they need more care and comfort from doctors. However, in existing online medical systems, doctors may easily overlook these emotional changes due to busy schedules. Moreover, existing systems lack emotion recognition and dynamic adjustment capabilities, failing to adjust communication methods and content in a timely manner based on the patient's emotional state. This severely limits the human touch in communication and affects trust and cooperation between doctors and patients. Summary of the Invention
[0007] In view of this, the purpose of this invention is to provide a method and system for simplifying medical terminology in intelligent doctor-patient dialogue based on a large language model. It first acquires patient consultation information, including text, voice, or image data. After converting the voice to text, the consultation information is divided into text and image parts, and stored according to communication units. A named entity recognition module reads the text information and extracts medical terms; a multimodal module analyzes medical terms or images not described in the doctor's text and transmits them to a large language model interpretation module for processing. This module combines patient background information, explains the terms and analyzes the images using colloquial language at different levels, and adjusts the tone according to the patient's emotions. Finally, it personalizes the explanations based on the patient's individual situation and consultation experience. The doctor reviews and improves the explanations through a feedback process to ensure accuracy and appropriateness.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model, the method comprising:
[0010] The system obtains consultation information and patient personal information from the doctor-patient communication platform, and performs voice conversion processing on the consultation information.
[0011] Based on the patient's consultation information and personal information, the first language model is used to assess the patient's personal qualities from the perspectives of education level, occupational relevance, age factor, and medical activity level.
[0012] The named entity recognition method is used to divide the patient's consultation information into several communication units, and the corresponding medical terms are extracted from each communication unit.
[0013] The second language model is used to match medical terms in each communication unit with their corresponding medical texts to determine whether there is a corresponding text explanation for all medical terms; medical terms that fail to match are passed to the third language model.
[0014] Sentiment analysis is performed on each segmented communication unit, and the sentiment analysis results are then passed to the third language model.
[0015] The third language model provides personalized interpretations of received, unmatched medical terms based on the patient's personal qualities and emotional changes.
[0016] Furthermore, for online consultation scenarios, doctors and patients interact through online consultation platforms and obtain consultation information and personal information generated during the interaction through web services. Consultation information includes text information, image information, and voice information; personal information includes the patient's name, gender, age, occupation, and historical consultation records, which cover past medical history and historical chat records.
[0017] Text data consists of text-based inquiries entered by patients or doctors; voice data consists of inquiries communicated by voice input by patients or doctors, which are then converted into text using voice recognition technology; image data includes medical examination reports or medical images.
[0018] All data is categorized and stored according to different data types.
[0019] Furthermore, the process by which the first language model calculates a patient's personal qualities based on their medical history and personal information is as follows:
[0020] First, we obtain the patient's four-dimensional data, including education level (E), occupational relevance (O), age factor (A), and healthcare activity level (M).
[0021] Then, the final weight W is obtained based on education level (E), occupational relevance (O), age factor (A), and healthcare activity level (M). sim :
[0022] W sim =w E ·E+w O ·O+w A ·A+w M ·(1-M)
[0023] In the formula, w E ,w O ,w A ,w M , respectively represent the preset weight ratios of education level E, occupational relevance O, age factor A, and medical activity level M;
[0024] Finally, based on the calculated W sim The popularization and professionalism of the explanations are controlled according to decision thresholds. Three decision thresholds are set: W1, W2, and W1 > W2. sim ≥W1 belongs to the category of highly popular and low-tech; W2≤W sim <W1 is a moderately popular and moderately professional category; W sim<W2 is a category with low popularity and high professionalism.
[0025] Furthermore, for the education level E of the patient, it is pre-assigned according to their academic qualifications, which are at least divided into four stages: primary school and below, junior high school education, high school education, undergraduate and above, and their assignments correspond to {E1, E2, E3, E4} respectively;
[0026] For the occupational relevance O of the patient, it first judges whether the patient is a medical practitioner. If so, O = O0; if not, it connects to the O*NET occupational database to query the "treatment and counseling" knowledge level rating table of the patient. If the "treatment and counseling" knowledge level rating result of the patient exists in the occupational database, the rating result is standardized: O = 1 - S / 100; if the "treatment and counseling" knowledge level rating result of the patient does not exist in the occupational database, it is assigned according to the occupational type, where the occupational type is divided into science and technology education, clerical, skilled labor, and others, and they are assigned O1, O2, O3, O4 respectively;
[0027] For the age factor A of the patient, it is standardized according to age:
[0028]
[0029] For the medical activity M of the patient, it is normalized according to the number of historical inquiry terms and the number of historical explanation terms, that is:
[0030]
[0031] The values of the above four-dimensional data are all between [0, 1], serving as the measurement thresholds for each dimension.
[0032] Furthermore, the consultation information is stored in blocks according to separate communication units. For all consultation information, it is divided into several communication units using the stop word division strategy;
[0033] Then, for each communication unit, medical term extraction is carried out by integrating named entity recognition and a dynamic term list. The unit text is sent into a fine-tuned general information extraction model. The general information extraction model relies on transfer learning for multi-specialty medical term recognition and forms a dynamic word library by connecting to an external medical database API to regularly update relevant terms; at the same time, doctors are allowed to customize and create a word list.
[0034] Furthermore, the second large language model determines whether a medical term or image is explained by the doctor's text through a multi-modal module:
[0035] First, both the term and the image are each divided into two parts;
[0036] Then, the analysis is carried out by matching the batch with the doctor's text. Using the medical terminology or images contained in the analysis and communication unit, the judgment work is carried out according to the following criteria: whether the doctor's text involves two of the following definitions or descriptions:
[0037] {Symptoms or characteristics, causes or risk factors, treatment or management recommendations, preventive measures}
[0038] If at least two definitions or descriptions relate to the medical term, the medical term is considered to be interpreted; otherwise, it is considered not interpreted and is passed to the third language model.
[0039] Furthermore, the patient's emotional state is detected through an emotion analysis module. The processing procedure of the emotion analysis module is as follows:
[0040] First, enter the text of the patient's current communication unit in the text input layer;
[0041] Then, key sentiment features are extracted using predefined rules in the feature extractor. The sentiment feature extractor calculates the number of interrogative sentences, exclamation marks, the proportion of negative words, the number of uncertain words, and the sentiment tendency value in real time; the output feature vector is:
[0042]
[0043] Finally, based on the dynamic cue word big language model analysis method, the big language model generates structured cue words based on feature values, which drives the big language model to output emotion labels and confidence scores. The emotion labels include anxiety, confusion, satisfaction, and calmness.
[0044] Furthermore, in the third language model, taking into account the patient's personal qualities and emotional state, unexplained medical terms are further explained.
[0045] For highly educated patients, use professionally termed language to explain;
[0046] For patients with a secondary education level, use simple and easy-to-understand language, and combine everyday examples with analogies or explanations;
[0047] For patients with low levels of education, use the most accessible language and metaphors, and explain using specific objects or scenarios;
[0048] When explaining to patients with anxiety, use a gentle and reassuring tone, and add encouraging words.
[0049] For patients experiencing confusion, a step-by-step analysis approach is used, breaking down complex medical terms into multiple simple parts for explanation;
[0050] For patients with a positive outlook, include positive prognostic information in the explanation to enhance their confidence;
[0051] For patients with calm emotions, use a neutral and professional tone to explain, ensuring the accuracy and completeness of the explanation.
[0052] On the other hand, a system for implementing the aforementioned intelligent doctor-patient dialogue terminology simplification method based on a large language model is also proposed. This system includes a data acquisition module, a storage module, a named entity recognition module, a large language model multimodal module, a large language model interpretation module, a sentiment analysis module, a personalization module, a feedback module, and a consultation interface.
[0053] The data acquisition module is used to collect patients' consultation information, including text, voice, or image data;
[0054] The storage module is used to store consultation information according to independent communication units;
[0055] The named entity recognition module is used to extract medical terms from consultation information; medical terms include disease names, body parts, clinical manifestations, microorganisms, viruses, drug names, and medical procedures.
[0056] The large language model multimodal module is used to identify medical terms or images that have not been interpreted by the doctor's text.
[0057] The large language model explanation module is used to generate popular explanations for unexplained medical terms and images. The explanations for medical terms include metaphors and examples; the explanations for images include annotations and text descriptions of abnormal parts in the images.
[0058] The sentiment analysis module is used to monitor the patient's emotional state and adjust the tone of the explanation.
[0059] The personalization module is used to adjust the explanation content based on the patient's personal information and historical consultation records;
[0060] The feedback module is used to provide the explanation to doctors for review and to receive their optimization feedback.
[0061] The consultation interface includes a chat window and a large language model explanation window.
[0062] Furthermore, the named entity recognition module includes a fine-tuned general information extraction model and a dynamic vocabulary; the UIE model uses transfer learning technology to adapt to the terminology recognition needs of different medical specialties; the dynamic vocabulary is automatically updated with the latest medical terms and definitions on a regular basis through an API connection to a medical database; the vocabulary also supports user-defined terms.
[0063] The large language model multimodal module divides medical terms and images in the unit into two categories, matches them with doctors' text in batches, and analyzes and judges whether doctors interpret the term or image. The judgment rules include whether at least two of the following are involved: definition or description, symptoms and characteristics, causes or risk factors, treatment methods or management recommendations, and preventive measures.
[0064] The sentiment analysis module is implemented using natural language processing technology, supporting real-time sentiment monitoring and feedback.
[0065] The beneficial effects of the present invention are:
[0066] First, it significantly improves the efficiency of doctor-patient communication. This invention intelligently identifies and simplifies medical terminology and images, effectively eliminating communication barriers between doctors and patients. This allows patients to understand doctors' diagnoses and recommendations more quickly and accurately, thereby accelerating the treatment process and improving overall communication efficiency.
[0067] Secondly, patients' understanding of medical knowledge has been significantly improved. This invention uses simplified explanations and image analysis to transform complex medical terms into language and images that are easy for patients to understand, helping them better grasp their condition and treatment plans, and enhancing their self-health management capabilities.
[0068] Furthermore, this invention effectively reduces information asymmetry between doctors and patients. By personalizing the interpretation of information, it ensures that patients with different educational backgrounds, ages, and health awareness levels can obtain information tailored to their individual needs, thereby eliminating misunderstandings and communication barriers caused by information differences and increasing trust between doctors and patients.
[0069] Furthermore, patient satisfaction and confidence have been enhanced. The emotion analysis module of this invention can monitor the patient's emotional state in real time and adjust the tone of explanation accordingly, making the patient feel more attentive and caring about medical services. This humanized communication approach helps alleviate patients' anxiety and tension, improves their satisfaction and confidence, and promotes harmonious doctor-patient relationships.
[0070] Finally, this invention significantly improves the professionalism and accuracy of the explanations. Doctors review and optimize the explanations through a feedback mechanism, ensuring they are both easy to understand and possess sufficient professionalism and reliability. This not only helps improve patients' medical knowledge but also provides doctors with more accurate and comprehensive patient information, supporting them in making more scientific and rational treatment decisions.
[0071] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0072] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0073] Figure 1 This is a schematic diagram of the overall process of the intelligent doctor-patient dialogue terminology simplification method based on a large language model according to an embodiment of the present invention;
[0074] Figure 2 This is a schematic diagram of the process for calculating a patient's personal qualities based on four-dimensional data according to an embodiment of the present invention;
[0075] Figure 3 This is a schematic diagram illustrating the calculation process of occupational relevance in an embodiment of the present invention;
[0076] Figure 4 This is an example of doctor-patient communication under the embodiments of the present invention;
[0077] Figure 5 This is a schematic diagram illustrating the matching process of whether medical terms are interpreted in an embodiment of the present invention;
[0078] Figure 6 This is a schematic diagram of the sentiment analysis process according to an embodiment of the present invention;
[0079] Figure 7 This is a schematic diagram of the structure of an intelligent doctor-patient dialogue terminology simplification system based on a large language model according to another embodiment of the present invention. Detailed Implementation
[0080] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0081] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0082] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0083] Please see Figures 1 to 7 This is a method and system for simplifying terminology in intelligent doctor-patient dialogue based on a large language model.
[0084] Example 1
[0085] This embodiment introduces a method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model, such as... Figure 1 As shown, it includes the following steps:
[0086] S1. Obtain consultation information and patient personal information from the doctor-patient communication platform, and perform voice conversion processing on the consultation information;
[0087] S2. Based on the patient's consultation information and personal information, the first language model is used to judge the patient's personal qualities from the perspectives of education level, occupational relevance, age factor, and medical activity level.
[0088] S3. The patient's consultation information is divided into several communication units using named entity recognition, and the corresponding medical terms are extracted from each communication unit.
[0089] S4. Use the second language model to match the medical terms in each communication unit with the corresponding medical text, and determine whether there is a corresponding text explanation for all medical terms; pass the medical terms that fail to match to the third language model;
[0090] S5. Perform sentiment analysis on each segmented communication unit and transmit the sentiment analysis results to the third language model;
[0091] S6, the third major language model, provides personalized interpretations of received medical terms that failed to match, based on the patient's personal qualities and emotional changes.
[0092] In step S1 of this embodiment, for online consultation scenarios, doctors and patients mainly interact through online consultation platforms. Therefore, consultation information and personal information generated during the interaction are obtained through Web services. Consultation information includes text information, image information, and voice information; personal information includes patient name, gender, age, occupation, and historical consultation records, etc. Historical consultation records cover past medical history and historical chat records.
[0093] Text data consists of written inquiries input by patients or doctors; voice data consists of inquiries input by patients or doctors via voice, which are then converted into text using speech recognition technology; image data includes medical examination reports or easily readable medical images. All data is categorized and stored according to its data type.
[0094] In step S2 of this embodiment, as Figure 2 As shown, the process by which the first language model calculates a patient's personal qualities based on their medical history and personal information is as follows:
[0095] First, we obtain the patient's four-dimensional data, including education level (E), occupational relevance (O), age factor (A), and healthcare activity level (M).
[0096] For patients’ education level E, a value is pre-assigned according to their educational level, which is divided into at least four stages: primary school and below, junior high school, high school, and undergraduate and above, with the values corresponding to {E1, E2, E3, E4} respectively.
[0097] Regarding the patient's occupational relevance O, such as Figure 3 As shown, it first determines whether the patient is a medical practitioner. If so, O = O0; otherwise, it connects to the O*NET occupational database and queries the patient's "treatment and consultation" knowledge level table. If the occupational database contains the patient's "treatment and consultation" knowledge level assessment result, the assessment result is standardized: O = 1 - S / 100; if the occupational database does not contain the patient's "treatment and consultation" knowledge level table assessment result, it assigns a value according to the occupational type, which is divided into science and technology education, clerical, skilled labor, and other categories, and is assigned the values O1, O2, O3, and O4 respectively.
[0098] For the patient's age factor A, it is standardized based on an age range of 18-65 years:
[0099]
[0100] For the medical activity M of a patient, it is normalized according to the historical number of inquiry terms and the historical number of explanatory terms, that is:
[0101]
[0102] Then, the final weight is obtained based on the education level E, occupational relevance O, age factor A, and medical activity M:
[0103] W sim = w E ·E + w O ·O + w A ·A + w M ·(1 - M)
[0104] In the formula, w E , w O , w A , w M respectively represent the preset weight ratios of the education level E, occupational relevance O, age factor A, and medical activity M.
[0105] Finally, according to the calculated W sim the popularization level and professionalism degree of the explanatory content are controlled according to the decision threshold; three-level decision thresholds are set, including W1, W2, where W1 > W2. The decision thresholds are: W sim ≥ W1 belongs to the high-popularization and low-professionalism category; W2 ≤ W sim < W1 is the medium-popularization and medium-professionalism category; W sim < W2 is the low-popularization and high-professionalism category.
[0106] In this embodiment, the following specific quantization criteria shown in Table 1 are given:
[0107] Table 1
[0108]
[0109] The final weight is output as W sim = 0.4E + 0.3O + 0.2A + 0.1(1 - M); the weight design principle is: dominated by education level (40%) and occupation (30%), and assisted by age (20%) and medical activity (10%). The W sim value controls the popularization level and professionalism degree of the explanatory content according to the decision threshold;
[0110] The decision threshold is: W sim ≥ 0.7 belongs to the high-popularization and low-professionalism category; 0.5 ≤ W sim < 0.7 is the medium-popularization and medium-professionalism category; W sim < 0.5 is the low-popularization and high-professionalism category.
[0111] In step S3 of this embodiment, the consultation information is stored in blocks according to individual communication units. Specifically, for all consultation information, a stop word segmentation strategy is used to divide it into several communication units. For example... Figure 4 The dialogue example shown follows these segmentation rules: If the dialogue contains stop words, such as "okay" or "um," then these words are used as the dividing lines. If there are no stop words, then the dialogue is segmented according to a fixed number of input entries, i.e., 5 entries for both the doctor and the patient. Figure 4 The example shown contains the following:
[0112] Patient: I've had stomach pain for the past few days, especially after eating.
[0113] Doctor: Hello, stomach pain may be caused by gastritis. I suggest you get tested for Helicobacter pylori.
[0114] Patient: Okay, what is Helicobacter pylori?
[0115] Doctor: Hmm, this is a type of bacteria that can cause stomach upset.
[0116] Patient: I've had stomach pain before, could it be an old problem?
[0117] Doctor: It may be a relapse. I suggest you pay attention to your diet and avoid spicy foods.
[0118] Patient: Okay, thank you, doctor. I will be careful.
[0119] In this embodiment, based on the segmentation rules (stop word segmentation or fixed 5 inputs), the system processes the dialogue as follows: Stop words include "okay" (item 3) and "um" (items 4 and 7), so these stop words are used as segmentation points to divide the dialogue into the following communication units: Unit 1 (items 1-2): Before the stop word "okay" appears. Unit 2 (items 3-4): From "okay" to "um". Unit 3 (items 5-7): From "um" to the end of the dialogue. If there are no stop words, the system will segment according to the 5 inputs (i.e., 5 messages) from both the doctor and patient. Because stop words exist in this example, segmentation is performed using stop words first. The storage module uses the patient ID as an index to save the patient's personal information, historical consultation information, and explanation content.
[0120] Then, for each communication unit, medical terms are extracted using a combination of Named Entity Recognition (NER) and a dynamic terminology lexicon. The unit text is fed into a fine-tuned Universal Information Extraction (UIE) model. This model relies on transfer learning to meet the recognition needs of medical terms from multiple specialties. Moreover, the system connects to an external medical database API to form a dynamic lexicon, which is updated regularly with relevant terms. Doctors are also allowed to create their own lexicons to ensure that the terms are both timely and professional. For example, when patient 001 asks about "stomach pain," the system identifies terms such as "stomach pain" (clinical manifestation), "gastritis" (disease), and "Helicobacter pylori" (microorganism). If any terms are missed or if doctors find them meaningful, they can be added to the custom lexicon.
[0121] In step S4 of this embodiment, the second language model uses a multimodal module to determine whether medical terms or images are interpreted by the doctor's text, as shown in the attached figure. Figure 5 As shown, the terminology and images are first divided into two parts, and then analyzed by matching the batch with the doctor's text. This module uses the medical terminology or images contained in the analysis and communication unit to make judgments according to the following criteria: whether the doctor's text involves at least two of the following items: definition or description, symptoms or characteristics, causes or risk factors, treatment or management recommendations, and preventive measures.
[0122] Taking patient 001 as an example, in communication unit 1, the doctor only said, "The stomach pain may be caused by gastritis." This only indicates a disease that may cause stomach pain, without providing a definition, symptoms, or treatment information. Therefore, the system categorizes it as "gastritis" without explanation. In unit 2, regarding "Helicobacter pylori," the doctor stated, "This is a bacterium that may cause stomach discomfort," including the definition (bacteria) and symptoms (stomach discomfort), meeting both criteria. The system judges it as explained. As for images, if the patient uploads a gastroscopy image, and the doctor only says, "It shows a little inflammation," without describing the inflammation or its cause in detail, the system will judge it as unexplained. Then, the large language model explanation module will provide a simplified explanation.
[0123] In this embodiment, taking user 001's inquiry as an example, the unexplained words are "gastritis" and gastroscopy image. The second language model interpretation module will describe "gastritis" as an inflammation of the stomach lining, just like skin becoming red and swollen after being irritated. The gastroscopy image will be described as this photo showing the internal condition of your stomach, with the red parts indicating an inflammatory reaction.
[0124] In step S5 of this embodiment, the patient's emotional state is detected by the emotion analysis module. The emotion analysis module is structured as follows: a text input layer (communication unit), an emotion feature extractor, dynamic prompts from a large language model, and an emotion classification layer from a large language model.
[0125] The sentiment feature extractor calculates the number of interrogative sentences, exclamation marks, negative word ratio, number of uncertain words, and sentiment tendency value in real time. The process includes:
[0126] Input: The text of the patient's current communication unit;
[0127] Processing: Extract key sentiment features using predefined rules;
[0128] Output feature vector:
[0129]
[0130] Finally, based on the dynamic cue word big language model analysis method, the big language model generates structured cue words based on feature values, driving the big language model to output emotion labels and confidence scores. The emotion labels include anxiety, confusion, satisfaction, and calmness. The specific process is as follows:
[0131] [System Command]
[0132] You are a professional medical emotion analyst. Please analyze a patient's emotions based on the following characteristics:
[0133] - Number of questions: {question_count}
[0134] - Number of exclamation marks: {exclamation_count}
[0135] - Negation ratio: {negation_ratio:.2f}
[0136] -Uncertainty word count:{uncertainty_words}
[0137] - Sentiment score: {sentiment_score: .2f}
[0138] Please select the best matching emotion label from [anxiety, confusion, satisfaction, calm] and output it in the following format:
[0139] {"emotion":"emotion label","confidence":confidence score 0.0-1.0}
[0140] In the above embodiments, the sentiment analysis module follows the appendix. Figure 6The process and personalization modules work together to further refine the interpretation. Taking user 001 as an example, the sentiment analysis module will calculate and classify three communication units:
[0141] For Unit 1, after processing by the feature extractor, the dynamic prompt words are adjusted as follows:
[0142] You are a professional medical emotion analyst. Please analyze a patient's emotions based on the following characteristics:
[0143] -Number of interrogative sentences: 0
[0144] - Number of exclamation marks: 0
[0145] -Negative word ratio: 0.00
[0146] -Number of uncertain words: 0
[0147] - Emotional tendency score: -0.32 ("Stomach pain" brings negative emotions)
[0148] The sentiment analysis result of the large language model is: {"emotion":"anxiety","confidence":0.85};
[0149] For Unit 2, after processing by the feature extractor, the dynamic prompt words are adjusted as follows:
[0150] You are a professional medical emotion analyst. Please analyze a patient's emotions based on the following characteristics:
[0151] Number of questions: 1
[0152] - Number of exclamation marks: 0
[0153] -Negative word ratio: 0.00
[0154] -Number of uncertain words: 0
[0155] - Sentiment score: -0.15
[0156] The sentiment analysis result of the large language model is: {"emotion":"confusion","confidence":0.92};
[0157] For Unit 3, after processing by the feature extractor, the dynamic prompt words of the large language model are adjusted as follows:
[0158] You are a professional medical emotion analyst. Please analyze a patient's emotions based on the following characteristics:
[0159] Number of questions: 1
[0160] - Number of exclamation marks: 0
[0161] -Negative word ratio: 0.00
[0162] - Number of uncertain words: 1
[0163] - Sentiment score: +0.41
[0164] The sentiment analysis result of the large language model is: {"emotion":"calm","confidence":0.78}.
[0165] In step S6 of this embodiment, unexplained medical terms are then explained and image analyses are generated using different levels of popularization methods based on the patient's background.
[0166] Adjust the tone of the explanation according to the patient's emotional state;
[0167] The explanation content will be personalized based on the patient's personal information and medical history.
[0168] Doctors review and refine their explanations through a feedback mechanism to ensure accuracy.
[0169] Emotional states include anxiety, confusion, satisfaction, or calmness. The system uses the information obtained from the test to improve the tone of the explanation, making it more in line with the patient's current psychological needs. When anxiety occurs, the explanation will include reassuring statements; when confusion occurs, the explanation will be adjusted to a step-by-step analysis, emphasizing key points; when satisfaction occurs, positive prognostic information will be added; and if the tone is calm, the explanation will maintain a neutral and professional tone.
[0170] Specifically, among them,
[0171] For highly educated patients, use professionally termed language to explain;
[0172] For patients with a secondary education level, use simple and easy-to-understand language, and combine everyday examples with analogies or explanations;
[0173] For patients with low levels of education, use the most accessible language and metaphors, and explain using specific objects or scenarios;
[0174] When explaining to patients with anxiety, use a gentle and reassuring tone, and add encouraging words.
[0175] For patients experiencing confusion, a step-by-step analysis approach is used, breaking down complex medical terms into multiple simple parts for explanation;
[0176] For patients with a positive outlook, include positive prognostic information in the explanation to enhance their confidence;
[0177] For patients with calm emotions, use a neutral and professional tone to explain, ensuring the accuracy and completeness of the explanation.
[0178] Based on this, this method significantly improves the efficiency of doctor-patient communication and patients' understanding of knowledge, reduces information asymmetry, and is widely used in online consultations and health education, especially for chronic disease management and postoperative recovery guidance.
[0179] Example 2
[0180] This embodiment introduces an intelligent doctor-patient dialogue terminology simplification system based on a large language model, such as... Figure 7 As shown, it includes a data acquisition module, a storage module, a named entity recognition module, a large language model multimodal module, a large language model interpretation module, a sentiment analysis module, a personalization module, a feedback module, and a consultation interface;
[0181] More specifically, the data acquisition module is used to collect patient consultation information, including text, voice, or image data;
[0182] The storage module is used to store consultation information according to independent communication units; the storage module stores user personal information, historical consultation records, and explanation content; all of the information is stored according to the patient ID.
[0183] The Named Entity Recognition (NAME) module is used to extract medical terms from consultation information. These medical terms include disease names, body parts, clinical manifestations, microorganisms, viruses, drug names, and medical procedures. The NAME module includes a fine-tuned Universal Information Extraction (UIE) model and a dynamic vocabulary. The UIE model employs transfer learning technology, enabling it to quickly adapt to the terminology recognition needs of different medical specialties. The dynamic vocabulary is automatically updated periodically with the latest medical terms and definitions via an API connection to a medical database. The vocabulary also supports user-defined terms, allowing doctors to manually add specialized terms from specific fields to improve recognition accuracy.
[0184] The large language model multimodal module is used to analyze medical terms and determine which medical terms or images are not explained by the doctor's text. The large language model multimodal module divides the medical terms and images in the unit into two categories, matches and combines them with the doctor's text in batches, and analyzes and determines whether the doctor explains the term or image. The determination rules include whether it involves two of the following five items: definition or description, symptoms and characteristics, causes or risk factors, treatment methods or management recommendations, and preventive measures.
[0185] The large language model explanation module is used to generate popular explanations for unexplained medical terms and images. The explanations for medical terms include metaphors and examples; the explanations for images include annotations and textual descriptions of abnormal parts in the images to help patients intuitively understand the doctor's diagnostic basis.
[0186] The sentiment analysis module is used to monitor the patient's emotional state and adjust the tone of the explanation; the sentiment analysis module is used to analyze the patient's emotional tendency during the interaction process; the modules are implemented through natural language processing technology and support real-time sentiment monitoring and feedback.
[0187] The personalization module is used to adjust the explanation content based on the patient's personal information and historical consultation records;
[0188] The feedback module is used to provide the explanations to doctors for review and receive their optimization feedback. The feedback mechanism involves the system presenting the explanations generated by LLMs to doctors, who then confirm, modify, or add to these explanations to improve their accuracy and professionalism. This feedback mechanism is accomplished through an online interactive interface, allowing for real-time adjustment of the explanations during doctor-patient consultations.
[0189] The consultation interface includes a chat window and a large language model explanation window.
[0190] By intelligently recognizing and simplifying medical terminology and images, this invention can effectively improve the efficiency of doctor-patient communication and patients' understanding of their condition. Furthermore, it implements personalized adjustments for different patients, adapting to their current emotional state, thereby further optimizing patient satisfaction and confidence. The doctor's feedback mechanism ensures that the relevant content maintains sufficient professionalism and reliability. This method and system provide a more efficient and practical solution for smart healthcare.
[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model, characterized in that: The method includes: The system obtains consultation information and patient personal information from the doctor-patient communication platform, and performs voice conversion processing on the consultation information. Based on the patient's consultation information and personal information, the first language model is used to assess the patient's personal qualities from the perspectives of education level, occupational relevance, age factor, and medical activity level. The named entity recognition method is used to divide the patient's consultation information into several communication units, and the corresponding medical terms are extracted from each communication unit. The second language model is used to match medical terms in each communication unit with their corresponding medical texts to determine whether there is a corresponding text explanation for all medical terms; medical terms that fail to match are passed to the third language model. Sentiment analysis is performed on each segmented communication unit, and the sentiment analysis results are then passed to the third language model. The third language model provides personalized interpretations of received, unmatched medical terms based on the patient's personal qualities and emotional changes.
2. The method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model according to claim 1, characterized in that: In the context of online consultations, doctors and patients interact through an online consultation platform and obtain consultation information and personal information generated during the interaction through web services. Consultation information includes text information, image information, and voice information; personal information includes the patient's name, gender, age, occupation, and historical consultation records, which cover past medical history and chat history. Text data consists of text-based inquiries entered by patients or doctors; voice data consists of inquiries communicated by voice input by patients or doctors, which are then converted into text using voice recognition technology; image data includes medical examination reports or medical images. All data is categorized and stored according to different data types.
3. The method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model according to claim 1, characterized in that: The process by which the first language model calculates a patient's personal qualities based on their medical history and personal information is as follows: First, we obtain the patient's four-dimensional data, including education level (E), occupational relevance (O), age factor (A), and healthcare activity level (M). Then, the final weight W is obtained based on education level (E), occupational relevance (O), age factor (A), and healthcare activity level (M). sim : W sim =w E ·E+w O ·O+w A ·A+w M ·(1-M) In the formula, w E ,w O ,w A ,w M , respectively represent the preset weight ratios of education level E, occupational relevance O, age factor A, and medical activity level M; Finally, according to the calculated W sim Control the popularization level and professionalism degree of the explanatory content according to the decision threshold. Among them, three-level decision thresholds are set respectively, including W1 and W2. Among them, the decision threshold of W1>W2 is: W sim 3W1 belongs to the category of high popularization and low professionalism; W2£W sim <W1 is the category of medium popularization and medium professionalism; W sim <W2 is the category of low popularization and high professionalism.
4. The method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model according to claim 3, characterized in that: For patients’ education level E, a value is pre-assigned according to their educational level, which is divided into at least four stages: primary school and below, junior high school, high school, and undergraduate and above, with the values corresponding to {E1, E2, E3, E4} respectively. Regarding the patient's occupational relevance O, it first determines whether the patient is a medical practitioner. If so, O = O0; otherwise, it connects to the O*NET occupational database and queries the patient's "treatment and consultation" knowledge level table. If the occupational database contains the patient's "treatment and consultation" knowledge level assessment result, the assessment result is standardized: O = 1 - S / 100; if the occupational database does not contain the patient's occupational communication assessment result, it is assigned a value according to the occupational type, which is divided into science and technology education, clerical, skilled labor, and other categories, and assigned values O1, O2, O3, and O4 respectively. For the patient's age factor A, it is standardized according to age: For a patient's medical activity level M, it is normalized based on the number of historical follow-up questions and the number of historical explanations of terms. The values of the above four-dimensional data are all between [0,1], serving as the measurement threshold for each dimension.
5. The method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model according to claim 1, characterized in that: The consultation information is stored in blocks according to individual communication units. For all consultation information, a stop word segmentation strategy is used to divide it into several communication units. Then, for each communication unit, medical terms are extracted by combining named entity recognition and dynamic terminology lists. The unit text is fed into a fine-tuned general information extraction model, which uses transfer learning to identify multi-specialty medical terms and forms a dynamic terminology library by connecting to an external medical database API, and updates the relevant terms regularly. At the same time, doctors are allowed to create their own terms.
6. The method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model according to claim 1, characterized in that: The second major language model uses a multimodal module to determine whether medical terms or images are interpreted by the doctor's text: First, divide the terminology and images into two separate sets; Then, the analysis is carried out by matching the batch with the doctor's text. Using the medical terminology or images contained in the analysis and communication unit, the judgment work is carried out according to the following criteria: whether the doctor's text involves two of the following definitions or descriptions: {Symptoms or characteristics, causes or risk factors, treatment or management recommendations, preventive measures} If at least two definitions or descriptions relate to the medical term, the medical term is considered to be interpreted; otherwise, it is considered not interpreted and is passed to the third language model.
7. The method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model according to claim 1, characterized in that: The patient's emotional state is detected through an emotion analysis module. The processing procedure of the emotion analysis module is as follows: First, enter the text of the patient's current communication unit in the text input layer; Then, key sentiment features are extracted using predefined rules in the feature extractor. The sentiment feature extractor calculates the number of interrogative sentences, exclamation marks, the proportion of negative words, the number of uncertain words, and the sentiment tendency value in real time; the output feature vector is: Finally, based on the dynamic cue word big language model analysis method, the big language model generates structured cue words based on feature values, which drives the big language model to output emotion labels and confidence scores. The emotion labels include anxiety, confusion, satisfaction, and calmness.
8. The method for simplifying terminology in intelligent doctor-patient dialogue based on a large language model according to claim 1, characterized in that: In the third language model, considering the patient's personal qualities and emotional state, further explanations are provided for unexplained medical terms. For highly educated patients, use professionally termed language to explain; For patients with a secondary education level, use simple and easy-to-understand language, and combine everyday examples with analogies or explanations; For patients with low levels of education, use the most accessible language and metaphors, and explain using specific objects or scenarios; When explaining to patients with anxiety, use a gentle and reassuring tone, and add encouraging words. For patients experiencing confusion, a step-by-step analysis approach is used, breaking down complex medical terms into multiple simple parts for explanation; For patients with a positive outlook, include positive prognostic information in the explanation to enhance their confidence; For patients with calm emotions, use a neutral and professional tone to explain, ensuring the accuracy and completeness of the explanation.
9. A system for implementing the intelligent doctor-patient dialogue terminology simplification method based on a large language model as described in any one of claims 1-8, characterized in that: It includes a data acquisition module, a storage module, a named entity recognition module, a large language model multimodal module, a large language model interpretation module, a sentiment analysis module, a personalization module, a feedback module, and a consultation interface; among them, The data acquisition module is used to collect patients' consultation information, including text, voice, or image data; The storage module is used to store consultation information according to independent communication units; The named entity recognition module is used to extract medical terms from consultation information; medical terms include disease names, body parts, clinical manifestations, microorganisms, viruses, drug names, and medical procedures. The large language model multimodal module is used to identify medical terms or images that have not been interpreted by the doctor's text. The large language model explanation module is used to generate popular explanations for unexplained medical terms and images. The explanations for medical terms include metaphors and examples; the explanations for images include annotations and text descriptions of abnormal parts in the images. The sentiment analysis module is used to monitor the patient's emotional state and adjust the tone of the explanation. The personalization module is used to adjust the explanation content based on the patient's personal information and historical consultation records; The feedback module is used to provide the explanation to doctors for review and to receive their optimization feedback. The consultation interface includes a chat window and a large language model explanation window.
10. A simplified terminology system for intelligent doctor-patient dialogue based on a large language model according to claim 9, characterized in that: The Named Entity Recognition (UIE) module includes a fine-tuned general information extraction model and a dynamic vocabulary; the UIE model employs transfer learning technology to adapt to the terminology recognition needs of different medical specialties; the dynamic vocabulary is automatically updated with the latest medical terms and definitions periodically through an API connection to a medical database; the vocabulary also supports user-defined terms. The large language model multimodal module divides medical terms and images in the unit into two categories, matches them with doctors' text in batches, and analyzes and judges whether doctors interpret the term or image. The judgment rules include whether at least two of the following are involved: definition or description, symptoms and characteristics, causes or risk factors, treatment methods or management recommendations, and preventive measures. The sentiment analysis module is implemented using natural language processing technology, supporting real-time sentiment monitoring and feedback.