Consultation shunting dialogue system based on small expert model
Through the consultation shunt dialogue system based on the small expert model, the probability coding technology guided by structural entropy and retrieval enhancement generation technology are used to achieve accurate identification and dynamic response to patient needs, solving the problems of low efficiency, inaccurate information and insufficient emotional support in traditional medical consultations, and improving the quality of consultation and patient experience.
Patent Information
- Application Number
- CN202510464280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing medical consultation system is difficult to accurately identify patients' diverse needs, cannot provide personalized information or emotional support, and insufficient response strategies during dynamic consultation, affecting the efficiency and quality of consultation.
The consultation diversion dialogue system based on the small expert model is adopted. The requirements judgment module uses the probability coding technology guided by structural entropy to analyze the patient's consultation content, and is diverted to the corresponding medical information consultation expert model or emotional support expert model, and uses retrieval enhancement generation technology to obtain information, combining large language models and diagnostic knowledge graphs to provide personalized responses.
It improves the efficiency and accuracy of medical consultation, provides personalized emotional support, improves patient satisfaction and treatment compliance, and optimizes the allocation of medical resources.
Smart Images

Figure CN120296135A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and knowledge graphs, and in particular to a consultation diversion dialogue system based on small expert models. Background Art
[0002] In modern large model dialogue systems, the consultation efficiency and quality directly affect the user experience. Especially when facing complex health problems, patients often need timely and accurate information support and emotional care. Traditional medical consultation methods often have problems such as poor information transmission, long response times, and insufficient attention to patients' emotional needs. With the rapid development of artificial intelligence (AI) technology, dialogue systems based on large language models (LLMs) have gradually become effective tools for solving medical consultation problems. These systems can understand patients' consultation content through natural language processing technology and generate corresponding responses. Especially the application of the Mixture of Experts (MoE) architecture enables the dialogue system to dynamically call different expert models according to different consultation types to provide more accurate services. However, existing MoE architectures often face the black box characteristic, and the process of expert selection is difficult to precisely control, resulting in insufficient information accuracy and emotional support.
[0003] Intelligent medical consultation systems face a complex technical challenge when dealing with patients' needs: how to accurately distinguish and efficiently meet the diverse needs of patients. Patients' consultation content often includes two major types of needs: information acquisition and emotional support. The processing methods and response strategies for these two types of needs are significantly different. However, traditional medical consultation systems often adopt a unified processing flow, making it difficult to accurately identify patients' true intentions and resulting in poor consultation effects. In addition, the individual differences and context dependence of patients' expression methods further increase the difficulty of need recognition. Even if the system can initially distinguish the type of need, how to provide personalized information or emotional support according to the specific situation of the patient is still a major challenge. During the dynamic consultation process, patients' needs may change, and the system needs to have the ability to adjust the response strategy in real time. At the same time, the professionalism and sensitivity of medical consultation require the system to consider the patient's mental state and acceptance ability while providing accurate information. This forms a technical problem: how to accurately identify and dynamically respond to patients' diverse needs while ensuring the professionalism and personalization of consultation services, and continuously optimize the system performance during this process to improve the overall intelligence level of medical consultation. Summary of the Invention
[0004] The present invention provides a consultation diversion dialogue system based on small expert models, mainly including: A demand judgment module, a medical information consultation expert model, and an emotional support expert model; Obtain the consultation content of the user through the demand judgment module, analyze the consultation content through the probability coding technology guided by structural entropy, judge the type of the consultation content, and divert the consultation content to the corresponding small expert model according to the type of the consultation content, and return the response content generated by the small expert model for the consultation content to the user; If the type of the consultation content is medical information, then divert the consultation content to the medical information expert model, obtain relevant information from the hierarchical diagnosis knowledge graph through the retrieval-enhanced generation technology, and generate a medical information response; If the type of the consultation content is emotional support, then divert the consultation content to the emotional support expert model, and generate an emotional support response according to the historical data context of the user.
[0005] Further, the analysis of the consultation content through the probability coding technology guided by structural entropy includes: Convert the consultation content into a vector representation , and use an encoder to encode the vector representation to obtain a latent variable , calculate the structural entropy of the latent variable by constructing a graph structure, and judge the type of the consultation content according to the structural entropy; The calculation of the structural entropy of the latent variable by constructing a graph structure includes: Construct a graph based on the embedding of the latent variable , take information type and emotion type as demand type labels for the graph to perform to obtain a three-layer coding tree 1; Calculate the structural entropy of the middle-layer nodes of the three-layer coding tree, and adjust the probability distribution of the latent variable according to the maximization of the structural entropy; The judgment of the type of the consultation content according to the structural entropy includes: Perform a linear transformation on the probability distribution of the latent variable to obtain a corresponding score vector , and perform a conversion on the score vector based on the Softmax classification function, and calculate the information type demand probability and the emotion type demand probability ; If the information type demand probability Greater than the probability of the emotional need , determine that the type of the consultation content is medical information , if the probability of the emotional need is greater than the probability of the information need , determine that the type of the consultation content is emotional support .
[0006] Further, the hierarchical diagnosis knowledge graph includes four layers, and the multi-hop path from the top layer to the bottom layer is represented as: , is the set of all disease names extracted from the electronic health record EHR database , represents the set of subcategories of is the set of categories aggregated from the set of subcategories is the set of disease manifestation characteristics, including and two subtypes represents the disease-specific characteristics enhanced by the large language model , represents the characteristics decomposed from the symptom manifestations extracted from the electronic health record EHR database ; represents the hierarchical or subordinate relationship is the characteristic relationship between the disease and the disease manifestation characteristics The relevant information is obtained from the hierarchical diagnosis knowledge graph through the retrieval-enhanced generation technology to generate a medical information response, including: Semantically segment the description of the patient's symptom manifestations in the consultation content, and decompose the patient's symptom manifestations into patient characteristics ; Match the patient characteristics with the nodes of the layer in the hierarchical diagnosis knowledge graph to obtain a set of clinical characteristic nodes, and the set of clinical characteristic nodes is the set of feature nodes in the layer with a matching degree greater than the threshold; For each node in the set of clinical characteristic nodes , determine the closest disease subcategory node in the layer by using the upward traversal method as the target subcategory node ; Using the target subcategory node as the parent node, and using the downward traversal method to reach Layer, retrieve nodes of all layers adjacent to the target subcategory node to form a node set ; then, nodes in the layer connected to the nodes in the node set are used to form a node set ; finally, the node set , the node set and the feature relationship are combined to form a set of diagnostic difference knowledge graphs ; The set of diagnostic difference knowledge graphs is used, together with the patient's symptom manifestations , the most relevant documents corresponding to the patient's symptom manifestations in the electronic health record EHR database and the prompt thereof to generate the medical response information A, that is ; Furthermore, the emotional support expert model includes a soft prompt group , a dense retriever, and a large language model with frozen parameters; The soft prompt group consists of
[0007] randomly initialized soft prompts. Each soft prompt has virtual tokens, where is the hidden dimension of the LLM and is the prompt length; The dense retriever is used to calculate a set of similarity scores between the context embedding corresponding to the historical data context and each soft prompt to select a suitable target soft prompt from the soft prompt group ; The large language model with frozen parameters is used to guide the input historical data context and the target soft prompt to generate the emotional support response. The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: First, the consultation triage dialogue system receives the patient's consultation content and analyzes it through natural language processing technology (probability coding guided by structural entropy) to intelligently judge the type of consultation, including medical information consultation or emotional support. According to the judgment result, the system diverts the consultation content to the corresponding expert model. For medical information consultation, the system calls the expert model specialized in medical information and uses retrieval-augmented generation (RAG) technology to obtain the latest and reliable information from authoritative medical knowledge bases and generate personalized medical advice. For emotional support consultation, the system calls the expert model specialized in emotional support, understands the patient's emotional needs based on the patient's historical communication data, and provides a warm and considerate response. Finally, the system returns the generated response to the patient to ensure that the patient obtains timely and accurate information or emotional support during the consultation process, thereby improving the patient's satisfaction and treatment compliance. This workflow not only improves the efficiency and quality of medical consultations but also provides a personalized experience for patients, promoting the intelligent development of medical services. Description of the Drawings
[0008] Figure 1 It is a flowchart of a consultation triage dialogue system based on small expert models in the present invention; Figure 2 It is a schematic structural diagram of a four-layer hierarchical diagnosis knowledge graph in the present invention ; Figure 3 It is a schematic diagram of the framework structure of a consultation triage dialogue system based on small expert models in the present invention. Specific Embodiments Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0009] To facilitate the understanding of the technical solutions in the present invention, the following explanations are given for the relevant technical terms in the present invention, specifically as follows: 1. Retrieval-Augmented Generation (RAG) is a natural language processing technology that combines retrieval and generation. Its core idea is to first retrieve relevant information and then use a generation model to generate answers, thereby improving the quality and accuracy of text generation. Specifically, RAG first retrieves information related to the input query from a large knowledge base or document collection. Given a query , through the retrieval model retrieves relevant documents from the knowledge base . Among them, is the retrieved document set, and is the retrieval algorithm.
[0010] Through a search algorithm, the document fragments most relevant to the user's question are found. Next, the retrieved information will be combined with the user's input for use by a generation model (such as GPT, etc.) to generate a richer and contextually relevant answer. The advantage of RAG is that it can dynamically retrieve the latest information and does not solely rely on the knowledge learned by the model during training. This method helps to improve the accuracy and relevance of the generated content as it can utilize external information sources to supplement the model's knowledge. In addition, RAG can generate diverse answers to adapt to different context requirements and thus performs well in areas such as question-and-answer systems, dialogue generation, and information summarization. This method effectively addresses the limitations that traditional generation models may encounter when dealing with specific knowledge.
[0011] 2. Structural Entropy and Encoding Tree: Structural entropy is a concept used to measure the complexity and information content of a system or structure, and is commonly applied in fields such as network science and information theory. Intuitively, the structural entropy method encodes the tree structure by characterizing the uncertainty of the hierarchical topology. Figure The structural entropy is defined as the total minimum number of bits required to determine the node encoding words in. Structural entropy has achieved success in fields such as information retrieval, traffic prediction, and reinforcement learning. By minimizing the structural entropy of a given graph , The hierarchical clustering results of the vertices in are retained by the associated encoding tree.
[0012] Encoding Tree: Let be an undirected weighted graph, where is the set of vertices, is the set of edges, and is the edge weight matrix. The encoding tree of the graph is defined as a hierarchical root tree as follows: 1) For each tree node , a vertex subset is associated with it.
[0013] 2) The root node of the tree is associated with the vertex set , i.e., .
[0014] 3) For each , its direct successor is labeled according to , increasing from left to right according to , and the direct predecessor is labeled as .
[0015] 4) For each with direct successors, the vertex subset are disjoint and .
[0016] 5) For each leaf node , contains only one vertex in
[0017] Structural entropy: Given any root-encoding tree of a graph , the structural entropy on measures the remaining complexity in after being reduced by . For each non-root node , its assigned structural entropy is defined as: , where is the cut point, i.e., the sum of the weights of the edges between the nodes in and the nodes not in , and and is the volume, i.e., the sum of the node degrees in and .
[0018] Given of the structural entropy is defined as: .
[0019] The introduction of structural entropy is to quantify the internal information and organization of complex structures. It can be regarded as a measure of the relationships between the various components in a system.
[0020] The present invention constructs a consultation diversion dialogue system based on a small expert model to improve the efficiency and quality of medical consultations, especially in the fields of medical information consultation and emotional support. The system uses advanced natural language processing technologies to intelligently judge the types of patients' consultations and divert them to the corresponding large medical information consultation models or emotional support large models, thereby providing personalized and accurate services. Specifically, the objectives of the present invention include the following aspects: 1. Improve the efficiency of medical consultation: In traditional medical consultation scenarios, patients often fall into the dilemma of long waiting times, and the responses they receive are also full of uncertainty. This situation not only greatly reduces the patient's medical experience, but may also cause the condition to worsen due to time delays. The system of the present invention uses structural entropy to capture the structural information of the data, and with the help of intelligent consultation content analysis technology, it can quickly and accurately identify the type of patient needs, and then accurately divert the consultation content to the corresponding expert model. This efficient diversion mechanism can greatly shorten the patient's waiting time and ensure that the patient obtains the required information support or emotional comfort in the shortest time, thereby comprehensively improving the efficiency of medical consultation. The system uses natural language processing technology to deeply analyze the patient's consultation content, constructs a graph structure through structural entropy to mine the potential information of the problem, and accurately infers the patient's intention by analyzing the key elements such as semantics and context in the text, and determines whether his needs belong to information support or emotional support. This intelligent diversion mechanism not only significantly improves the consultation response speed, but also optimizes the allocation of medical resources, allowing patients to get help more efficiently.
[0021] 2. Improve the accuracy and reliability of medical information: In the medical field, accurate information is crucial, especially when patient consultation involves disease management and treatment plans. The present invention uses retrieval-augmented generation (RAG) technology in combination with an authoritative medical knowledge base to ensure that the information provided is based on the latest medical guidelines and research results. This method effectively improves the accuracy of medical information, helps patients obtain personalized medical advice, and avoids misleading information that may appear in traditional consultations. Specifically, for information consultations, the system automatically calls an expert model specializing in medical information, and through the combination with an authoritative knowledge base, ensures that the information provided is not only accurate but also timely. This combination not only improves the reliability of medical information, but also generates personalized medical advice based on the patient's specific health status, historical medical records and other information. This precise information support not only helps patients better understand their own health status, but also provides a strong basis for their decision-making.
[0022] 3. Provide personalized emotional support: Emotional support is also an aspect that cannot be ignored in the medical process, especially when facing patients undergoing long-term treatment. The emotional support model of the present invention can understand the emotional needs of patients based on their historical communication data and provide warm and considerate responses. This personalized emotional support can not only improve patient satisfaction but also enhance patient treatment compliance, thereby improving the overall treatment effect. Specifically, when patients express emotions such as anxiety and worry, the system can give timely responses and provide effective emotional support to help patients relieve psychological pressure. In addition, the system will adjust its response strategy according to the emotional state of patients to better meet their emotional needs. This personalized emotional support not only improves the psychological comfort of patients but also enhances their confidence in the treatment plan.
[0023] In summary, the present invention constructs a consultation triage dialogue system based on a small expert model, aiming to solve various problems existing in traditional medical consultations, including poor information transmission, long response time, and unclear answer fields. This system not only improves the efficiency of medical consultations and the accuracy of information but also provides personalized emotional support, reduces the consumption of medical resources, promotes the accessibility of medical services, and provides patients with a safer, more reliable, and more user-friendly medical consultation experience.
[0024] Specifically, a consultation triage dialogue system based on a small expert model in the present invention includes a demand judgment module, a medical information consultation expert model, and an emotional support expert model, as Figure 1 shown, including: S101. Obtain the consultation content of the user through the demand judgment module, analyze the consultation content through the probability encoding technology guided by structural entropy, judge the type of the consultation content, divert the consultation content to the corresponding small expert model according to the type of the consultation content, and return the response content generated by the small expert model for the consultation content to the user.
[0025] Specifically, analyzing the consultation content through the probability encoding technology guided by structural entropy includes: converting the consultation content into a vector representation , using an encoder to encode the vector representation to obtain a latent variable , calculate the structural entropy of the latent variable by constructing a graph structure, and judge the type of the consultation content according to the structural entropy; Calculating the structural entropy of the latent variable by constructing a graph structure includes: constructing an adjacency matrix according to the embedding of the latent variable , and using a three-layer encoding tree for the adjacency matrix Perform hierarchical partitioning, calculate the structural entropy of the middle-layer nodes of the three-layer coding tree, and adjust the latent variables according to the maximization of the structural entropy of the probability distribution; Judge the type of the consultation content according to the structural entropy, including: performing a linear transformation on the sub-probability distribution of the latent variable to obtain the corresponding score vector , based on the Softmax classification function, transform the score vector and calculate the information class demand probability and the emotional class demand probability ; if the information class demand probability is greater than the emotional class demand probability , determine that the type of the consultation content is medical information , if the emotional class demand probability is greater than the information class demand probability , determine that the type of the consultation content is emotional support .
[0026] S102. If the type of the consultation content is medical information, then divert the consultation content to the medical information expert model, and obtain relevant information from the hierarchical diagnosis knowledge graph through retrieval-augmented generation technology to generate a medical information response.
[0027] Specifically, the hierarchical diagnosis knowledge graph includes four layers, and the multi-hop path from the top layer to the bottom layer is expressed as: , is the set of all disease names extracted from the electronic health record EHR database , represents the set of sub-categories of is the set of categories aggregated from the set of sub-categories is the set of disease manifestation characteristics, including and two subtypes represents the disease-specific features enhanced by the large language model , represents the features decomposed from the symptom manifestations extracted from the electronic health record EHR database ; represents the hierarchical or subordinate relationship is the feature relationship between the disease and the disease manifestation characteristics Obtain relevant information from the hierarchical diagnosis knowledge graph through retrieval-augmented generation technology to generate a medical information response, including: performing semantic segmentation on the description of the patient's symptom manifestations in the consultation content, and decomposing the patient's symptom manifestations into patient features ; Match the patient characteristics with the nodes in the layer of the hierarchical diagnosis knowledge graph to obtain a set of clinical characteristic nodes, and the set of clinical characteristic nodes is a set of feature nodes in the layer with a matching degree greater than the threshold; for each node in the set of clinical characteristic nodes, determine the closest disease subcategory node in the layer by upward traversal as the target subcategory node ; using the target subcategory node as the parent node, traverse downward to the layer, and retrieve all nodes in the layer adjacent to the target subcategory node to form a node set , and then form a node set by the nodes in the layer connected to the nodes in the node set , and finally combine the node set , the node set and the feature relationship to form a set of diagnostic difference knowledge graphs ; use a large language model to generate medical response information A according to the patient's symptom manifestations , the most relevant documents in the electronic health record EHR database corresponding to the patient's symptom manifestations , the diagnostic difference knowledge graph and its prompt , that is . .
[0028] S103. If the type of the consultation content is emotional support, then divert the consultation content to the emotional support expert model to generate an emotional support response according to the user's historical data context.
[0029] Specifically, the emotional support expert model includes a soft prompt group , a dense retriever, and a large language model with frozen parameters; the soft prompt group consists of randomly initialized soft prompts. Each soft prompt has virtual tokens, is the hidden dimension of the LLM, is the prompt length; the dense retriever is used to calculate the context embedding corresponding to the historical data context The similarity score sets between each soft prompt , select appropriate target soft prompts from the soft prompt group ; a large language model with parameter freezing is used to guide the input historical data context and the target soft prompt to generate an emotional support response.
[0030] The calculation process of the similarity score set is as follows: For each soft prompt in the soft prompt group , first obtain the context embedding from the historical data context through the embedding layer , and respectively pass through two linear layers and to obtain the similarity scores and , and calculate the average value of the similarity scores and dimension by dimension and , and calculate the original similarity score according to the following formula : ; then process through the Softplus activation function to obtain the normalized similarity score , that is , i is an integer greater than or equal to 1 and less than or equal to K; Form the similarity score set by aggregating the similarity scores corresponding to each soft prompt in the soft prompt group .
[0031] The following further elaborates on the data collection, demand judgment module, medical information consultation expert model, and emotional support expert model in the consultation triage dialogue system, as Figure 3 shown: Data collection and data sources are the basis for constructing the demand judgment module of the medical escort large model, and their quality and diversity directly affect the performance and generalization ability of the model. The following will elaborate from two aspects: data sources and data types.
[0032] Data sources include hospital information systems (HIS), online medical platforms, social media and health communities, and medical escort agencies.
[0033] Hospital Information System (HIS): A large amount of patient-related data is stored in the hospital's information system, including medical records, diagnosis results, treatment plans, etc. These data can provide important information for understanding the patient's basic health status and medical needs. For example, through medical records, the patient's disease history, allergy history, etc. can be understood, which helps to judge the background and possible needs of the patient's consultation content.
[0034] Online medical platforms: With the development of Internet-based healthcare, many online medical platforms have accumulated rich patient consultation data. These data include conversation records between patients and doctors, health consultation posts, etc. The data on online medical platforms are diverse and real-time, and can reflect the needs of different patients in different scenarios.
[0035] Social media and health communities: Social media and health communities are important platforms for patients to communicate about health problems and share experiences. On these platforms, patients will post their symptoms, concerns, and needs, and these data can provide valuable information for understanding the patient's emotional state and potential needs. For example, patients expressing fear and anxiety about diseases in health communities may indicate that they need emotional support.
[0036] Medical escort agencies: During the process of providing escort services, medical escort agencies will record the needs and feedback of patients. These data can provide direct information for understanding the needs of patients during actual escort, and help to optimize the performance of the demand judgment module.
[0037] The data types include text data, structured data, and time series data; text data mainly includes the patient's consultation content, case records, diagnosis reports, etc. Text data is the core data for demand judgment, and by analyzing the text, the type and specific content of the patient's needs can be understood. Structured data: such as the patient's basic information (age, gender, and medical history), vital sign data (body temperature, blood pressure, heart rate, etc.). Structured data can provide supplementary information for text data, helping to understand the patient's needs more comprehensively. Time series data: If there is data on multiple consultations or long-term health monitoring of patients, time series data can reflect the changing trends and periodicity of the patient's needs. For example, the symptoms of some chronic disease patients may worsen within a specific time period, and by analyzing time series data, the patient's needs can be predicted in advance.
[0038] Data preprocessing is a key step to ensure the performance of the model, covering multiple important aspects: cleaning of text data, data standardization, and normalization.
[0039] Cleaning of Text Data: First, input the original text dataset. In the dataset, duplicate records can lead to overfitting and bias in the model. By removing duplicates, ensure that each sample is unique, thereby improving the training effect of the model. Spelling mistakes, inconsistent formats, or other inaccurate entries often occur in text data. At the same time, missing values have a significant impact on model training. Missing records can be deleted, missing values can be filled with the mean or median, or more complex methods (such as interpolation or prediction models) can be used to handle missing data to ensure the integrity of the dataset. Additionally, clean special symbols, numbers, and extra spaces in the text to reduce data noise and improve the readability of the text.
[0040] Data Standardization and Normalization: When dealing with data of different dimensions and distributions, standardization and normalization are essential steps. The system inputs the cleaned text dataset and performs standardization and normalization processing on it respectively. Data standardization is to convert numerical data into a distribution with a mean of 0 and a standard deviation of 1. This process helps to eliminate the scale differences between features, making the model more efficient and stable during training; normalization is to scale the feature values to a specific range (such as 0 to 1). When the data has different dimensions, normalization can bring all features to the same scale, avoiding some features having too much influence on the model. Through the above steps, data preprocessing can significantly improve the training efficiency and final performance of the model, providing a solid foundation for subsequent analysis and output.
[0041] For the demand judgment module of the medical large model, through the probability encoding technology guided by structural entropy, it can more effectively analyze the patient's consultation content and accurately judge the type of their needs. This module is mainly implemented through the following steps: Data Preprocessing and Encoding: Vectorize the patient's consultation content and convert it into a format suitable for model input, denoted as. Use an encoder to encode the input and map it to a Gaussian distribution, where is the latent variable, and all the distributions form the embedding space of the latent variable. This process follows the encoder operation of the classical probability encoding model and obtains the probability representation of the input data in the latent space through learning. The formula is: .
[0042] Constructing the Graph and Calculating Structural Entropy: Based on the encoded latent variable embedding , construct the graph . Construct the adjacency matrix to represent the relationship between latent variables. The formula is: , where is the sigmoid activation function, which is used to ensure that the elements of the adjacency matrix are positive values.
[0043] Construct a three-layer coding tree with tags (requirement type: information or sentiment) as the optimal partition of the data 1. The middle layer nodes of the coding tree represent the categories of the classification tasks (information or sentiment), and each leaf node (i.e., the input data ) is assigned to the corresponding middle node according to its tag. Define the assignment matrix ( is the number of leaf nodes, is the number of middle nodes, where , corresponding to information and sentiment respectively), indicates that the -th leaf node belongs to the -th category.
[0044] Calculate the structural entropy of the coding tree. The structural entropy is defined in Figure and the coding tree 1 as: , ; where, is the sum of the weights of the edges connecting the internal points and external points; is the sum of the degrees of all data points in the graph 1; is the volume; is the parent node. For the structural entropy of the middle layer nodes of the three-layer coding tree, the formula is: , using the adjacency matrix and the assignment matrix , its regularized loss format is: , where, is the all-one matrix with shape .
[0045] By maximizing the structural entropy , the probability distribution of the latent variables can be constrained, so that the data belonging to different requirement types (information and sentiment) can be better separated in the latent space, thereby enhancing the discriminative ability of the model for different requirement types.
[0046] Training and optimization of the requirement judgment module: Use the encoder-only architecture for probability coding, and its overall loss is: , where, is the basic loss of probability coding, and an appropriate loss function can be selected according to the task type, such as the cross-entropy loss in the classification task; is a hyperparameter used to control the structural entropy regularization loss weight. During the training process, by adjusting the model parameters, the overall loss is minimized , enabling the model to learn feature representations related to the demand type and improving the accuracy of demand judgment.
[0047] The demand type judgment of the demand judgment module is as follows: 1. New input data encoding: After the demand judgment module is trained, when new patient consultation content is received, it is first input into the trained encoder . The encoder is a key component that has learned the data feature representation during the training phase. It maps the input consultation content to the latent space to obtain the latent variable . This process extracts and transforms the features of the input data, enabling the data to exist in a form more suitable for model processing and analysis, that is .
[0048] 2. Demand type judgment based on the Softmax classifier: The Softmax function (i.e., the Softmax classifier) is a commonly used multi-class activation function that can convert the output of the classifier into a probability distribution, making the sum of the probabilities of all classes equal to 1. In our demand judgment module, assume there are two categories: information demand and emotional demand, represented by for information and used for for emotion.
[0049] Let the score vector obtained after the classifier performs a linear transformation on the latent variable be , where is the score for information demand, is the score for emotional demand. The score vector is obtained through a linear layer ( ), where is the weight matrix, is the bias vector.
[0050] The Softmax function transforms the score vector to calculate the probability of each category as follows: , where is the probability that the consultation content belongs to information demand, is the probability that it belongs to emotional demand.
[0051] According to the probability distribution output by the Softmax function, select the category with the highest probability as the final demand judgment result, that is: , through the above demand judgment module based on the probability encoding technology guided by structural entropy, the patient's consultation content can be analyzed more accurately, providing a reliable basis for subsequent calling of the corresponding expert models (medical information consultation model or emotional support model), and realizing the efficient diversion and precision of medical escort services.
[0052] For the medical information consultation model of the medical large model, given an electronic health record (EHR) database and a large language model , the goal is to construct a four-layer hierarchical diagnostic knowledge graph , and the multi-hop path from the top layer to the bottom layer of the hierarchical diagnostic knowledge graph is represented as: , is the set of all disease names extracted from the electronic health record (EHR) database , represents the set of subcategories of is the set of categories aggregated from the set of subcategories, is the set of disease manifestation characteristics, including and two subtypes, represents the disease-specific features enhanced by the large language model , represents the features decomposed from the symptom manifestations extracted from the electronic health record (EHR) database ; represents the hierarchical or subordinate relationship, is the feature relationship between the disease and the disease manifestation characteristics. The hierarchical diagnostic knowledge graph is as Figure 2 shown.
[0053] For the given hierarchical diagnostic knowledge graph and the input patient manifestations , let , representing a certain subcategory determined from , and the goal is to extract the diagnostic difference knowledge graph related to from the hierarchical diagnostic knowledge graph .
[0054] 1. Construction of disease knowledge graph: The forms and representations of diseases in the electronic health record database are diverse. First, through disease clustering, the set of original disease descriptions is unified into , denoted as: , where represents the clustering model applied to , and is an embedding model.
[0055] Then, using the unified , a four - layer hierarchical disease knowledge graph is constructed through hierarchical aggregation. This graph integrates the relationships between diseases and their potential categories, and each disease is aggregated into a sub - category and a category. Define the disease knowledge graph as , composed of and the large language model aggregated as follows: .
[0056] In the retrieval stage, use the large language model to perform topic aggregation to extract the most relevant topics from to aggregate sub - categories. Then, these sub - category topics are further aggregated into higher - level categories, forming a hierarchical structure from sub - categories to broader categories. Next, apply hierarchical clustering to assign the diseases in to the aggregated sub - category topics, and then assign the sub - topics to the topics. This method utilizes the powerful semantic understanding and topic extraction capabilities of the large language model, making it possible to classify diseases more precisely in topic aggregation. By applying hierarchical clustering to the topics based on the large language model, the diseases in are aggregated into a hierarchical structure. Hierarchical aggregation introduces multiple levels of granularity for , ensuring that diseases with different symptom manifestations can be appropriately classified. To effectively utilize the historical diagnostic results from the electronic health record (EHR) database as an accurate representation of disease symptom manifestations, decompose the symptom manifestations of the diseases in into discrete features . Each single feature (such as symptom, location, or activity limitation) from each is created as a node . This final decomposition results in a comprehensive hierarchical disease knowledge graph , which captures both the disease category information obtained from hierarchical aggregation and the related features.
[0057] 2. Enhancement of symptom manifestations in the knowledge graph: For a given hierarchical diagnostic knowledge graph and the input patient manifestations , let represent a certain sub - category determined from , and the goal is to obtain from the hierarchical diagnostic knowledge graph Extract the diagnostic difference knowledge graph related to .
[0058] Hierarchical disease knowledge graph The knowledge in only contains information from the Electronic Health Record (EHR) database , which is not sufficient to accurately diagnose all diseases, especially when differentiating diseases with similar clinical manifestations. Therefore, integrating external knowledge is crucial. To supplement the diagnostic knowledge graph with key knowledge not existing in , add external knowledge to , which helps to differentiate diseases with similar symptom manifestations. Traverse all diseases , and use prompts specially designed for searching and generating disease nuances on the large language model . As shown in the following formula: , where and represent the large language model and its prompt for enhancing disease symptom manifestations, respectively.
[0059] Each generated diagnostic key difference node is then connected to its corresponding through the relationship . In this way, a chain is obtained. For example, generate a symptom manifestation and relationship with the disease node "lumbar hyperosteogeny" and form a chain: <lumbar_spondylosis, has_symptom, stiffness_or_pain_in_the_lower_back>. Finally, the hierarchical disease knowledge graph is formed by integrating and of .
[0060] 3. Diagnostic difference knowledge graph search: 3.1. Patient symptom decomposition For the given patient symptom manifestation , perform sentence segmentation on and decompose it into more detailed patient characteristics , denoted as . Define a mapping function to describe this process as follows: .
[0061] Furthermore, calculate the patient characteristics and calculate the semantic similarity score 𝑠𝑖𝑚 between them as shown in the following formula: , where is the similarity model, is the embedding model applied to and before calculating the similarity.
[0062] 3.2. Clinical Feature Matching For each patient characteristic , retrieve the top 𝑚 most similar , where 𝑚 represents the number of closest matches selected. Overall, the system retrieves 𝑛×𝑚 matching nodes in the knowledge graph .
[0063] To address the situation where there are no close matches in in , we introduce an indicator function to filter out irrelevant matches: , and, .
[0064] Where represents the set of nodes that satisfy the condition . The indicator function ensures that only those with a similarity score higher than the threshold are selected into .
[0065] Through clinical feature matching, we have successfully matched the query with the most relevant clinical feature nodes in the hierarchical diagnosis knowledge graph .
[0066] 3.3. Upward Traversal To accurately match the most relevant for the patient, an upward traversal method is adopted. By aggregating votes based on the shortest path distance between and in the graph, the closest disease subcategory is determined.
[0067] For , calculate its shortest path to each disease subcategory by upward traversing in the graph. Denote the shortest path distance from to as . If Indicates the current closest disease sub - category node, then the vote count of is incremented by one. Then, during the reverse process, the votes of each are accumulated, and the node with the highest number of votes is determined as . This voting process is formally defined by the indicator function Using this as the parent node, we traverse down to , retrieve all adjacent to and their adjacent . Given , let represent the set of disease nodes belonging to . Similarly, define: , representing the set of feature nodes connected to the disease nodes in .
[0068] We connect all triples , where and , to form a set of diagnostic difference knowledge graphs: , where, represents the diagnostic difference knowledge graph used for large - language model reasoning next.
[0069] 3.4 Active Diagnostic Questioning Mechanism Inaccurate diagnoses often stem from insufficient or incomplete patient descriptions. To address this issue, an active diagnostic questioning mechanism is proposed here. When the initial input lacks some key information required for a doctor or large - language model to make a more precise diagnostic decision, this mechanism acts as a co - pilot and poses targeted follow - up questions.
[0070] In the diagnostic knowledge graph , a feature may be connected to multiple disease nodes , and the discriminability of each varies. For example, some features are more common, such as "waist pain", while other features represent more unique features, such as "pain intensifies when walking". Here, the discriminability score of is defined as the reciprocal of its degree centrality in the knowledge graph : , where represents the total number in the knowledge graph in . Calculate the discrimination scores for each feature node , and select those nodes with the highest discrimination scores, as follows: , , where represents the selected features with the highest discrimination scores, which are used to actively guide follow-up questions for a clear diagnosis.
[0071] The inference retrieval enhanced generation model triggered by the knowledge graph is the core component of the medical retrieval enhanced generation model, which uses a large language model (LLM) to generate diagnostic results, personalized treatment plans, and medication suggestions. In addition, the system will actively provide doctors with suggestions for follow-up questions to clarify missing or ambiguous information of patients. As shown in the following formula, utilize the diagnostic difference knowledge graph enhanced by the large language model and specially designed prompts to stimulate the inference ability of the large language model.
[0072] Different from most retrieval enhanced generation model systems that focus on answering short factual questions, this system is tailored for complex tasks in clinical scenarios. These prompts are specially designed to optimize the inference ability of the large language model, especially in differentiating diseases with similar symptom manifestations. The system conducts comprehensive inference by using the retrieved documents and the diagnostic difference knowledge graph extracted from the knowledge graph.
[0073] Use the electronic health record database as a document repository to retrieve the most relevant documents corresponding to the patient's symptom manifestations . Then, conduct a similarity search on the database to identify the most relevant 𝑘 records. After obtaining all the inputs, we designed a special prompt to guide the large language model to perform inference through the diagnostic difference knowledge graph to generate answers to help doctors differentiate similar diseases and actively generate follow-up questions.
[0074] For the emotional support expert model of the medical large model, in the medical escort scenario, the emotional needs of patients are complex and diverse. How to provide precise and personalized emotional support is crucial. Therefore, this module constructs a medical escort emotional support large model, aiming to provide better emotional care for patients.
[0075] In the emotional support scenario of medical escort, let the context composed of the patient's consultation content and conversation history be . Among them, represents the personal characteristic information of the patient, such as age, gender, condition, etc., and is used to depict the basic situation of the patient; represents the patient and the medical escort system The dialogue history between them reflects the communication process. The goal is to generate responses that meet the emotional needs of the patient and provide appropriate emotional support to the patient.
[0076] 1. Architecture Design The architecture of this emotional support large model mainly includes a soft prompt group, a dense retriever, and a frozen large language model (LLM).
[0077] Soft prompt group: Use to represent the soft prompt group, which consists of randomly initialized soft prompts. Each soft prompt has virtual tokens, is the hidden dimension of the LLM, is the prompt length. During the training process, the soft prompts will be fine-tuned while the LLM remains frozen.
[0078] Dense retriever: Responsible for selecting appropriate soft prompts from the soft prompt group. It calculates the similarity score between the context embedding and each soft prompt , sorts the soft prompts according to the score, and finds the soft prompt that best fits the current context.
[0079] LLM: Adopts a decoder-only causal language model with frozen weights, initialized by a pre-trained model. The selected soft prompt is merged with the context and then input into the LLM to guide it to generate an emotional support response.
[0080] 2. Calculate the similarity between the soft prompt and the context To reduce the computational overhead, the dense retriever uses two linear layers and to calculate the similarity score. The specific calculation process is as follows: First, obtain the context embedding through the word embedding layer of the LLM , and respectively pass through the linear layers and to get and . Then, take the average of and dimension by dimension to get and . Then, calculate the original similarity score Then, it is processed through the Softplus activation function to obtain the normalized similarity score , ensuring that the score is within the interval to improve the numerical stability during training.
[0081] 3. Learning Prompt Selection 3.1. Soft Prompt Loss Given the context and its corresponding ideal response , calculate the negative log-likelihood loss for each soft prompt. First, obtain the prediction result through , and then calculate the loss using , where represents the concatenation operation, is the forward propagation operation of the LLM, is the negative log-likelihood loss function. This will generate K loss values , which are used to measure the prediction ability of each soft prompt.
[0082] 3.2. Prompt Selection Loss Due to the lack of explicit annotation in the dialogue setting, it is quite challenging to update the retriever to select the best soft prompt. Use the loss of the soft prompt in the LLM to guide the selection. Align the performance evaluation of the LLM with the similarity score of the retriever through the KL divergence between the negative language model loss and the similarity score. Specifically, let be and the similarity scores of each soft prompt in the soft prompt group. The prompt selection loss is defined as: , , where is the Softmax function, is the temperature hyperparameter, is the KL divergence. This loss ensures that the selection of the dense retriever is consistent with the performance of the LLM, effectively reflecting the ability of the soft prompt to generate relevant responses.
[0083] 3.3. Context-Prompt Contrastive Learning To avoid the retriever always selecting a single soft prompt and promote prompt diversity, introduce the context-prompt contrastive loss. This loss adjusts the similarity score according to the text similarity of different contexts, and the formula is: , where is the distance function (such as BLEU), is the threshold, is the context and the cosine similarity score vector of the soft prompt in the soft prompt group, is the cosine similarity. This comparison strategy enhances the cosine similarity for similar context pairs and weakens it for dissimilar ones, ensuring the consistency between the retriever and the LLM evaluation while enhancing the diversity and uniqueness of the dialogue context and improving the model's adaptability.
[0084] 3.4. Prompt Fusion Learning To optimize the effectiveness of the soft prompts, a prompt fusion learning loss is adopted. This loss averages the prediction probabilities of all soft prompts in the soft prompt group, and the formula is: , ; In this way, by integrating the advantages of different soft prompts, smoothing the variance and bias of individual prompts, improving the accuracy and reliability of the overall prediction, and enhancing the model's ability to generate appropriate responses.
[0085] 3.5. Overall Objective Function The training of the emotional support large model relies on the collaborative effect of the above loss functions. The soft prompt loss ensures the accuracy of the LLM, the prompt selection loss makes the retriever consistent with the LLM output, the context-prompt contrast loss promotes the diversity of prompt selection, and the prompt fusion learning loss improves the overall performance of all soft prompts. The overall objective function is: where , , are hyperparameters that control the relative contributions of each loss component. During training, minimize to balance the fidelity of the LLM, the accuracy of the retriever, and the diversity of prompt selection, and build an adaptive emotional support dialogue generation system.
[0086] 3.6. Inference During the inference stage, the dense retriever selects the most appropriate soft prompt from the soft prompt group according to the given context, and inputs it together with the context into the LLM for decoding to generate the final response. The specific process is as follows: , , where is the selected soft prompt, is the response generated by the LLM.
[0087] Through the above methodology based on selective prompt adjustment, the medical escort emotional support large model can more effectively understand the emotional needs of patients, generate more targeted and personalized emotional support responses, improve the quality of medical escort services, and provide warm and caring emotional support for patients during the medical treatment process.
[0088] The operation processes of the consultation diversion dialogue system are as follows: 1. Receiving consultation content The system monitors the consultation information input by users in real time and inputs it in the form of natural language. These contents cover various consultation questions such as disease symptom descriptions, treatment-related questions, and concerns about the condition.
[0089] The input consultation content enters the demand judgment and classification module. First, the content will be vectorized and transformed into a format suitable for model analysis. Then, it will be mapped to the latent space of the Gaussian distribution through an encoder to obtain latent variables. Next, a graph structure is constructed based on the latent variables, and the structural entropy is calculated. By maximizing the structural entropy, the probability distribution of the latent variables is constrained, so that the data of information-type and emotion-type demands can be better separated in the latent space. After that, a Softmax classifier is used to process the newly input data. According to the probability distribution output by the classifier, it is judged whether the consultation content belongs to medical information consultation or emotional support consultation.
[0090] 2. Processing medical information consultation If it is determined to be a medical information consultation, the system calls the medical information consultation module. In this module, first, according to the patient's symptom description, the sentence is segmented, the symptoms are decomposed into detailed features, and they are matched with the clinical features in the diagnosis knowledge graph. By calculating the semantic similarity score, the nodes with a similarity higher than the threshold are screened out.
[0091] Then, using the method of upward traversal, based on the shortest path distance aggregation voting, the closest disease subcategory is determined, and then the related disease nodes and feature nodes are retrieved to form a diagnosis difference knowledge graph. At the same time, the system will use the large language model to reason according to the diagnosis difference knowledge graph and the information input by the patient to generate a diagnosis result, a personalized treatment plan, and medication suggestions. If the initial input information is incomplete, the system will also start an active diagnosis questioning mechanism to ask the patient targeted questions to obtain more information to improve the diagnosis.
[0092] 3. Processing emotional support consultation If it is determined to be an emotional support consultation, the system calls the emotional support module. The dense retriever in this module will select the most appropriate soft prompt from the soft prompt group according to the context composed of the patient's consultation content and the conversation history. The specific process is to first calculate the similarity score between the context embedding and each soft prompt, and then select the soft prompt with the highest score after sorting. Then, the selected soft prompt is merged with the context and input into the frozen large language model for decoding to generate a reply that meets the patient's emotional needs and give the patient emotional support.
[0093] The consultation diversion dialogue system based on the small expert model constructed by the present invention has brought innovative changes to the field of medical consultation. In terms of data processing, through multi-source collection and fine preprocessing, the quality and diversity of data are guaranteed, laying a solid foundation for subsequent analysis. The demand judgment and classification module uses the probability coding technology guided by structural entropy to accurately analyze the patient's consultation content, efficiently distinguish information and emotional needs, and realize the intelligent diversion of consultations.
[0094] The medical information consultation module improves the reasoning ability of the retrieval-augmented generation (RAG) model by constructing a diagnostic knowledge graph, and combines an active diagnosis and questioning mechanism to provide accurate and personalized medical advice for patients, effectively assisting doctors in making diagnosis decisions. The emotional support module uses a unique architecture design and co-training with multiple loss functions to deeply understand the patient's emotional needs and generate warm and targeted responses, significantly improving the patient's medical experience.
[0095] From the perspective of actual application effects, the system effectively solves many problems in traditional medical consultations. It greatly shortens the patient's waiting time and improves the consultation efficiency; relying on an authoritative knowledge base and advanced technology, it guarantees the accuracy and reliability of medical information; and the personalized emotional support based on the patient's historical data enhances the patient's treatment compliance.
[0096] In summary, the consultation diversion dialogue system of the present invention provides an effective solution for the optimization and upgrading of medical consultation services with its advantages in technological innovation, function realization, and application effects, and has important application value in the medical field.
[0097] What is disclosed above is only the preferred embodiment of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A consultation diversion dialogue system based on a small expert model, characterized in that, include: Demand judgment module, medical information consultation expert model and emotional support expert model; The user's consultation content is obtained through the demand judgment module, the consultation content is analyzed through the probability coding technology guided by structural entropy, the type of the consultation content is determined, the consultation content is diverted to the corresponding small expert model according to the type of the consultation content, and the response content generated by the small expert model for the consultation content is returned to the user; If the type of the consultation content is medical information, the consultation content is diverted to the medical information expert model, and the relevant information is obtained from the hierarchical diagnostic knowledge graph through the retrieval enhancement generation technology to generate a medical information response; If the type of the consultation content is emotional support, the consultation content is diverted to the emotional support expert model to generate an emotional support response according to the user's historical data context.
2. The consultation diversion dialogue system according to claim 1, characterized in that, The analysis of the consultation content by the probability coding technology guided by structural entropy includes: Convert the consultation content into a vector representation , and use an encoder to encode the vector representation to obtain a latent variable . Calculate the structural entropy of the latent variable by constructing a graph structure, and determine the type of the consultation content according to the structural entropy; Calculating the latent variable by constructing a graph structure The structural entropy includes: Based on the latent variable embedding to construct a graph ; Using information type and emotion type as demand type labels for the figure perform hierarchical division to obtain a three-layer coding tree 1; Calculate the structural entropy of the middle layer nodes of the three-layer coding tree, and constrain the probability distribution of the latent variables according to the maximization of the structural entropy ; The determining the type of the consultation content according to the structural entropy includes: Perform a linear transformation on the probability distribution of the latent variable to obtain the corresponding score vector . Based on the Softmax classification function, transform the score vector and calculate the information class demand probability and the sentiment class demand probability ; If the probability of the information type of demand is greater than the probability of the emotional type of demand , determine that the type of the consultation content is medical information , if the probability of the emotional type of demand is greater than the probability of the information type of demand , determine that the type of the consultation content is emotional support .
3. The consultation diversion dialogue system according to claim 2, characterized in that: The requirement judgment module uses an encoder architecture for probability encoding, and its overall loss function is: , where is the basic loss of probability encoding, is the regularization loss, is a hyperparameter used to control the weight of the structural entropy regularization loss . Regularization loss is calculated through the adjacency matrix and the assignment matrix with the formula as follows: , where is the full matrix with the shape , is the number of leaf nodes, is the number of intermediate nodes; the adjacency matrix is used to represent the relationship between the latent variables and is expressed as: , is the sigmoid activation function, which is used to ensure that the elements of the adjacency matrix are positive values.
4. The consultation diversion dialogue system according to claim 1, characterized in that: The hierarchical diagnosis knowledge graph includes four layers, and the multi-hop path from the top layer to the bottom layer is expressed as: , is the set of all disease names extracted from the electronic health record (EHR) database , represents the set of subcategories of is the set of categories aggregated from the set of subcategories is the set of disease manifestation characteristics, including and two subtypes represents the disease-specific characteristics enhanced by the large language model , represents the characteristics decomposed from the symptom manifestations extracted from the electronic health record (EHR) database ; represents the hierarchical or subordinate relationship is the feature relationship between the disease and the disease manifestation characteristics The relevant information is obtained from the hierarchical diagnostic knowledge graph through the retrieval-augmented generation technology to generate a medical information response, including: Semantically segment the description of the patient's symptom manifestations in the consultation content and decompose the patient's symptom manifestations into patient characteristics ; Match the patient characteristics with the nodes of the layer in the hierarchical diagnosis knowledge graph to obtain a set of clinical characteristic nodes , where the set of clinical characteristic nodes is a set of feature nodes in the layer with a matching degree greater than the threshold; For each node in the set of clinical characteristic nodes , the closest disease sub-category node in the layer is determined by upward traversal as the target sub-category node ; Using the target subcategory node as the parent node, traverse downwards to layer, and retrieve all nodes adjacent to the target subcategory node in the layer to form a node set . Furthermore, retrieve the nodes in the layer that are connected to the nodes in the node set to form a node set . Finally, combine the node set , the node set and the feature relationship to form a set of diagnostic difference knowledge graphs ; Using large language models , according to the patient's symptom manifestations , the electronic health record EHR database and the most relevant documents corresponding to the patient's symptom manifestations in the database , the diagnostic difference knowledge graph and its prompts to generate the medical response information A, that is .
5. The consultation diversion dialogue system according to claim 4 is characterized in that: Said matching the patient characteristics with the nodes of the layers in the hierarchical diagnosis knowledge graph comprises: For each patient feature, calculate the patient feature and the node semantic similarity score between them, and its calculation formula is as follows: where is the similarity model, is the embedding model applied to the patient feature and the node before calculating the similarity; Use an indicator function Take the semantic similarity score Greater than the threshold Of the nodes As the set of clinical feature nodes ; Among them, , 。 6. The consultation diversion dialogue system according to claim 4, characterized in that: The method of determining the disease subcategory node closest to it in the layer as the target subcategory node includes: For a node , by traversing upward in the hierarchical diagnosis knowledge graph to calculate the shortest path from the node to each disease subcategory , and selecting the disease subcategory node closest to the current node ; By an indicator function Determine the node with the highest number of votes Determine as the target subcategory node , where the indicator function Is defined as: , the target subcategory node Is expressed as: .
7. The consultation shunt dialogue system according to claim 4 or 5, characterized in that The consultation diversion dialogue system introduces an active diagnosis questioning mechanism, and also includes: For the set of clinical feature nodes, calculate the nodes in the hierarchical diagnosis knowledge graph the reciprocal of the degree centrality as the node discrimination score; In the active diagnosis questioning mechanism, the features corresponding to the node with the highest discrimination score in node are used to actively guide follow-up questions for a clear diagnosis.
8. The consultation diversion dialogue system according to claim 1, characterized in that The emotional support expert model includes a soft prompt group , a dense retriever, and a large language model with frozen parameters; The soft prompt group consists of randomly initialized soft prompts, each soft prompt having virtual tokens, which is the hidden dimension of the LLM, and is the prompt length; The dense retriever is used to calculate the context embeddings corresponding to the historical data context and each soft prompt to obtain a set of similarity scores , and select a suitable target soft prompt from the soft prompt group ; The parameter-frozen large language model is used to guide the input historical data context and target soft prompts to generate the emotional support response.
9. The consultation diversion dialogue system according to claim 8, characterized in that, The set of similarity scores is calculated as follows: For each soft prompt in the soft prompt group First, through the embedding layer Obtain the context embedding from the historical data context And respectively pass through two linear layers and To get the similarity scores and And for the similarity scores and Take the average by dimension and Calculate the original similarity score according to the following formula : ; Then process through the Softplus activation function to obtain the normalized similarity score That is , i where \(i\) is an integer greater than or equal to 1 and less than or equal to \(K\); The similarity scores corresponding to each soft prompt in the soft prompt group are aggregated to form the set of similarity scores for the soft prompt group .
10. The consultation diversion dialogue system according to claim 8, characterized in that, The overall objective function used during the training of the emotional support expert model is as follows: , where , , is a hyperparameter that controls the relative contribution of each loss component; is the negative log-likelihood loss, is the context-cue contrastive loss, is the cue selection loss, is the cue fusion learning loss.
Citation Information
Patent Citations
Online health consultation method and system based on intelligent AI and storage medium
CN119181517A
Retrieval enhancement method and system based on structure entropy hierarchical knowledge tree
CN119381009A
Lightweight large model-based electric power knowledge system construction and intelligent question and answer method
CN119597864A
Large model-based medical health consultation pushing method and system, and storage medium
CN119647579A
Cited By
Medical knowledge intelligent consultation system based on knowledge graph
CN120523920A
Human setting following method and system for role playing agent in decoding stage
CN121478960A