Mental health coach voice model optimization method based on large language model driving
By adopting a mental health coach voice model optimization method driven by a large language model in the intelligent voice assistant, integrating health knowledge graphs and improved models, the problems of insufficient professional depth, lack of personalization and stiff emotional interaction in the existing technology are solved, and the professionalization, personalization and emotionalization of mental health consultation are realized, and the user experience is improved.
Patent Information
- Application Number
- CN202411898453.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent voice assistants have problems such as insufficient professional depth, lack of personalization and stiff emotional interaction in mental health management.
The mental health coach voice model optimization method driven by a large language model is adopted, and the professionalization, personalization and emotionalization of mental health consultation is achieved by integrating health knowledge graphs, improving large language models, and strengthening personalization and emotional interaction.
It improves the comprehensive ability of voice assistants in mental health management, provides professional, accurate and personalized mental health advice, enhances the natural fluency of emotional interactions, and improves user experience.
Smart Images

Figure CN120032809A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of medical equipment and mental health management, and in particular to a mental health coach voice model optimization method driven by a large language model. Background Art
[0002] With the continuous advancement of artificial intelligence technology, intelligent voice assistants are increasingly used in the field of health management. Although certain results have been achieved, they still face many challenges. Traditional rule-based or template-based methods are unable to cope with complex and changing health consultation scenarios. Although existing voice assistants based on natural language processing (NLP) technology and dialogue generation models have made some progress in question-answering fluency and basic knowledge coverage, they have obvious limitations in the following aspects:
[0003] Insufficient professional depth: Most existing systems rely on general language models and lack deep learning and professional knowledge integration in the field of mental health. As a result, when providing specific health advice, it is often too general and unable to provide accurate and professional answers.
[0004] Lack of personalization: There are huge individual differences in users’ mental health, living habits, and psychological states, but existing voice assistants often adopt a “one-size-fits-all” approach to answering and lack customized mental health advice based on the user’s individual situation.
[0005] Stiff emotional interaction: In health consultation, the user's emotional state has an important impact on the effectiveness of the consultation. However, existing systems often ignore the importance of emotional communication, and the responses lack warmth, making it difficult to build user trust.
[0006] Therefore, there is an urgent need for a new technical route that integrates mental health knowledge and personalized features to enhance the comprehensive capabilities of voice assistants in mental health management. Summary of the invention
[0007] In order to address the deficiencies of the prior art, the disclosed embodiments provide a method for optimizing the voice model of a mental health coach driven by a large language model. The method aims to achieve professionalization, personalization, and emotionalization of mental health counseling by integrating health knowledge graphs, improving large language models, and strengthening personalization and emotional interaction, so as to address the problems of insufficient professional depth, lack of personalization, and stiff emotional interaction in traditional methods.
[0008] The present disclosure provides a method for optimizing a mental health coach voice model based on a large language model, including the following steps:
[0009] Collect health knowledge data and health consultation dialogue data from multiple channels and construct a multi-source heterogeneous data set;
[0010] Preprocessing the data in the multi-source heterogeneous data set, wherein the preprocessing includes data cleaning, data labeling, and data format unification;
[0011] Based on the preprocessed data, build a health knowledge graph and process the conversation context;
[0012] Setting an improved large language model; wherein the improved large language model includes an embedding layer, a multi-layer recursive structure, a residual connection and a fully connected layer, which are arranged in sequence according to the order of the data stream and are connected to each other in a specific manner;
[0013] Creating a knowledge question-answering task, embedding the health knowledge graph into the improved large language model to perform health knowledge fusion training;
[0014] Construct a dialogue scenario, embed the dialogue context and knowledge into the improved large model for dialogue training;
[0015] Determine the evaluation indicators, and analyze the deficiencies and problems of the model based on the feedback results of the evaluation indicators, and optimize them.
[0016] In a possible implementation, the embedding layer, the multi-layer recursive structure, the residual connection and the fully connected layer in the improved large language model are connected to each other in the following specific manner:
[0017] The embedding layer is used to convert discrete categorical data into a low-dimensional continuous vector representation, i.e., an embedding representation, wherein the categorical data includes words and category labels;
[0018] The multi-layer recursive structure is used to perform a recursive calculation operation on the embedding representation output by the embedding layer, which is composed of a plurality of identical or similar sub-layers stacked together, each layer performs a recursive calculation based on the output of the previous layer, and each layer includes the following three components: self-attention processing, Chebyshev feature extraction, and normalization operation;
[0019] The residual connection is used to add the input of each layer of the multi-layer recursive structure to the output of the layer to form a residual connection;
[0020] The fully connected layer is used to perform a linear transformation on the output of the last layer of the multi-layer recursive structure and output a prediction result.
[0021] In a possible implementation, the weight of each layer of the network of the improved large language model is represented by a Chebyshev polynomial coefficient matrix, which is dynamically adjusted through training, and the core parameters of the network include embedding dimension, number of attention heads, number of Chebyshev feature extractions and Chebyshev polynomial order.
[0022] In a possible implementation, the method includes:
[0023] Based on Chebyshev polynomial T degree (x) Expand or represent the input x:
[0024] T degree (x) = cos(degree arccos(x))
[0025] T 0 (x) = 1
[0026] T 1 (x) = x
[0027] T degree (x) = 2x·T degree-1 (x)-T degree-2 (x)
[0028] Among them, the input x is the health knowledge graph or conversation context; the Chebyshev polynomial T degree (x) is a degree polynomial, T degree-1 (x) is a degree-1 Chebyshev polynomial, T degree-2 (x) is a degree-2 Chebyshev polynomial; arccos(x) is the inverse cosine function, and degree is an index value used to index the order of the corresponding Chebyshev polynomial, thereby calculating the polynomial value at each order;
[0029] The input x is projected onto the coefficients of each Chebyshev polynomial order cheby coeffs Perform forward propagation calculation on:
[0030]
[0031] Among them, c degree is the coefficient, T degree (x) is the expansion term of Chebyshev polynomial, Degree is the order of Chebyshev polynomial; y m It represents the mth eigenvalue of the final output. The specific dimension of the output, i.e. the mth element of the output, is obtained by weighted summing each input feature on Chebyshev polynomials of different orders. m is the index of the feature in the output space, which represents the mth output feature currently calculated.
[0032] In a possible implementation, the health knowledge integration training process includes:
[0033] The discrete data of entities and relationships in the health knowledge graph are converted into low-dimensional continuous vector representations through the embedding layer, wherein the health knowledge graph is represented by a triple head entity, a relationship, and a tail entity, and the vector representation contains semantic information of the entities and relationships in the graph;
[0034] The vector representation output by the embedding layer is fed into the multi-layer recursive structure for deep feature extraction and sequence modeling; wherein, in each layer, the self-attention processing mechanism can capture the complex dependencies between entities and relationships in the graph; the Chebyshev feature extraction further extracts feature information from the graph by performing nonlinear transformation on the input vector through Chebyshev polynomials; and the normalization operation ensures the stability of the model during the training process;
[0035] The feature representation processed by the multi-layer recursive structure is sent to the fully connected layer, and the fully connected layer maps the feature representation to the output space according to the specific task and outputs the prediction result;
[0036] Furthermore, a joint loss function is used to optimize the model, so that the improved large language model can accurately associate and apply relevant knowledge when generating answers.
[0037] In a possible implementation, the joint loss function used in the health knowledge fusion training includes the loss of the entity classification task, the loss of the relationship classification task and a mixed loss function to enhance the entity recognition and relationship inference capabilities.
[0038] In a possible implementation, the dialogue training process includes:
[0039] The embedding layer converts the discrete data of words and phrases in the text into a low-dimensional continuous vector representation, wherein the vector representation contains semantic information in the conversation context;
[0040] The vector representation output by the embedding layer is fed into the multi-layer recursive structure for deep feature extraction and sequence modeling; wherein, in each layer, the self-attention processing mechanism can capture the semantic dependencies in the conversation context; the Chebyshev feature extraction performs nonlinear transformation on the input vector through the Chebyshev polynomial to further extract key information in the conversation context; and the normalization operation ensures the stability of the model during the training process;
[0041] The conversation context feature representation processed by the multi-layer recursive structure is sent to the fully connected layer, and the fully connected layer maps the feature representation to the output space according to the specific task and outputs the prediction result;
[0042] Furthermore, a joint loss function is used to optimize the model to improve the fluency and emotional suitability of the conversation.
[0043] In a possible implementation, the dialogue training process further includes:
[0044] Embed the dialogue context into a high-dimensional vector space. Each word in the context is processed through the embedding layer and the self-attention mechanism. After the context vector is fused with the knowledge embedding, it is input into the decoder. The decoder predicts the target response word by word based on the context and the historical part of the target response.
[0045] In a possible implementation, the joint loss function adopted in the dialogue training includes a dialogue generation loss, an emotional appropriateness constraint, and a fluency constraint to improve the fluency and emotional appropriateness of the dialogue.
[0046] In a possible implementation, determining the evaluation metrics, and based on the feedback results of the evaluation metrics, analyzing the deficiencies and problems existing in the model, and performing optimization, including:
[0047] Based on the feedback results of the evaluation metrics, analyze the deficiencies and problems existing in the model, and accordingly increase the training volume of relevant data, adjust the design of the training task, or optimize the parameters of the language generation part of the model.
[0048] Compared with the prior art, the embodiments of the present disclosure have the following beneficial effects:
[0049] 1. Improve the professionalism of health knowledge: By integrating the health knowledge graph with the improved large language model feature extraction technology, the present disclosure can more deeply understand and apply health knowledge, and provide professional and accurate mental health advice for users.
[0050] 2. Enhance dialogue personalization: The present disclosure fully considers the personalized characteristics of users in the dialogue generation process. By optimizing the model parameters and the design of the training task, the dialogue content is closer to the actual needs of users, improving the pertinence and practicality of the dialogue.
[0051] 3. Improve emotional appropriateness: By introducing an emotion classifier and an emotional appropriateness constraint, the present disclosure can better grasp the emotional factors in the dialogue generation process, making the dialogue more natural and fluent, and enhancing the emotional connection between users and the voice assistant.
[0052] 4. Improve comprehensive performance: After evaluation and optimization, the mental health coach voice model of the present disclosure has achieved significant improvements in aspects such as the accuracy of mental health knowledge, the relevance of answers, the effectiveness of suggestions, and the fluency and emotional appropriateness of the dialogue, providing users with a more high-quality and comprehensive health management service.
[0053] In summary, the optimization method of the mental health coach voice model driven by the large language model proposed by the present disclosure solves the problems of insufficient professional depth, lack of personalization, and rigid emotional interaction in traditional methods, and opens up a new path for the application of intelligent voice assistants in the field of mental health management. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Figure 1 A flowchart of a method for optimizing a mental health coaching voice model based on a large language model driven by an embodiment of the present disclosure is provided;
[0056] Figure 2 A schematic diagram of a mental health coach voice model optimization method based on a large language model driven by an embodiment of the present disclosure;
[0057] Figure 3 A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0058] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.
[0059] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor represent the necessary logical order between them. It should also be understood that in the embodiments of the present disclosure, "multiple" can refer to two or more, and "at least one" can refer to one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, in the absence of explicit limitation or contrary revelation given in the context, it can generally be understood as one or more. In addition, the term "and / or" in the present disclosure is only a description of the association relationship of the associated objects, indicating that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present disclosure generally indicates that the associated objects before and after are an "or" relationship. It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and the same or similar parts can refer to each other. For the sake of brevity, they will not be repeated one by one.
[0060] At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present disclosure and its application or use. The techniques, methods and devices known to ordinary technicians in the relevant fields may not be discussed in detail, but where appropriate, the techniques, methods and devices should be considered part of the specification. It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0061] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0062] This paper proposes a method for optimizing the voice model of a mental health coach driven by a large language model, aiming to address the shortcomings of existing technologies in the application of health knowledge, personalization of conversations, and emotional suitability. This paper combines the health knowledge graph with the improved large language model feature extraction technology to design a set of technical routes for health knowledge fusion training and dialogue generation optimization, so as to achieve the professionalization of voice assistants in health consultation and the improvement of conversation fluency. At the same time, the practicality of the voice assistant is enhanced by optimizing the knowledge embedding, context vector, and multi-round dialogue generation of the model.
[0063] Figure 1 A flowchart of a method for optimizing a mental health coaching voice model based on a large language model driven by an embodiment of the present disclosure is provided. Figure 2 The schematic diagram of a method for optimizing a mental health coaching voice model based on a large language model driven by an embodiment of the present disclosure is provided. Figure 1 and Figure 2 As shown, the method 100 comprises the following steps:
[0064] S110: Collect health knowledge data and health consultation dialogue data from multiple channels and construct a multi-source heterogeneous data set;
[0065] In step S110, first, various mental health data are collected from various channels such as professional mental health books, medical literature, authoritative health websites (World Health Organization and national health data), and consultation record databases of health institutions to build a rich and comprehensive multi-source heterogeneous data set. The data types include the following two:
[0066] (1) Health knowledge data: For example, collect mental health-related knowledge from professional books, journal articles, research reports, etc. in the field of mental health, such as anxiety, depression, stress management, emotion regulation, sleep disorders, interpersonal relationships, and other mental health topics; visit the official websites of authoritative health organizations such as the World Health Organization, the American Institute of Mental Health, and the China Mental Health Network to obtain mental health-related guidelines, manuals, research reports, etc. These websites usually provide mental health knowledge and suggestions that have been reviewed by experts and are highly authoritative and reliable. The professional knowledge text obtained through the above channels will serve as the basis for the voice assistant to answer users' health-related questions.
[0067] (2) Health consultation conversation data: Under the premise of ensuring compliance with relevant laws and regulations and privacy protection, conversation records and data in real health consultation scenarios are obtained from the consultation record database of health institutions after authorization and permission. These data include, for example, patient symptom descriptions, diagnosis results, treatment plans, as well as user questions during the consultation process, and responses from health experts or customer service staff. These data are used to enable the model to learn how to interact effectively in conversation scenarios.
[0068] S120: preprocessing the data in the multi-source heterogeneous data set, wherein the preprocessing includes data cleaning, data labeling, and data format unification;
[0069] S121: Data Cleansing
[0070] Remove noise information in the collected data, such as repeated content, format errors (garbled characters, incomplete sentences, etc.), irrelevant tags or special characters, etc., to ensure the quality and standardization of the data.
[0071] For data obtained from sources such as web pages, it is also necessary to remove irrelevant parts such as web page tags and advertising content, and only retain text content related to health.
[0072] S122: Annotated Data
[0073] For health knowledge data, the key entities involved (psychological symptoms, mental illness, emotional state, coping strategies, influencing factors, mental health resources, etc.) and their relationships (symptoms and diseases, emotions and symptoms, diseases and coping strategies, influencing factors and symptoms / diseases, resources and suggestions, etc.) are annotated to facilitate the model to learn the semantic structure of health knowledge.
[0074] Specifically, psychological symptoms: such as anxiety, depression, stress, insomnia, fear, anger, sadness, etc. Mental illness: such as anxiety disorder, depression, bipolar disorder, schizophrenia, obsessive-compulsive disorder, post-traumatic stress disorder (PTSD), etc. Emotional state: such as happiness, calmness, irritability, tension, frustration, excitement, etc. Coping strategies: such as cognitive behavioral therapy, relaxation training, mindfulness meditation, emotion regulation skills, social support, professional counseling, etc. Influencing factors: such as life changes, genetic factors, environmental factors, social support, personal character, coping style, etc.
[0075] Specifically, Symptoms & Diseases: Certain symptoms may indicate a mental illness, for example, persistent sadness and loss of interest may be symptoms of depression. Moods & Symptoms: Specific emotional states may trigger or exacerbate certain symptoms, for example, long-term stress may lead to anxiety. Diseases & Coping Strategies: Certain mental illnesses may require specific coping strategies, for example, people with depression may need cognitive behavioral therapy and medication. Contributing Factors & Symptoms / Diseases: Certain factors may increase the risk of developing a mental illness or exacerbate symptoms, for example, genetic factors and environmental stressors may work together to cause depression. Resources & Advice: Mental health resources may provide advice and support for specific symptoms or diseases, for example, people with anxiety disorders may benefit from mindfulness meditation and relaxation training.
[0076] When annotating data, you need to assign unique identifiers to these key entities and clearly annotate the relationships between them. For example, you can use triples (head entity-relationship-tail entity) to represent these relationships, such as (anxiety disorder-symptoms-nervousness), (cognitive behavioral therapy-targeting-anxiety disorder), etc.
[0077] For health consultation conversation data, labeling the conversation intent (asking about the condition, seeking advice, making an appointment for service, etc.), the conversation roles (user, consultant, etc.), and key information points helps the model understand the context and purpose of the conversation. For example, (1) Conversation intent: The user says, "I have been suffering from insomnia recently, what should I do?" - Intent: Seeking solutions to insomnia. (2) Conversation roles: User: "I have been suffering from insomnia recently, what should I do?" - Role: User; Voice assistant: "You can try some relaxation techniques, such as meditation." - Role: Mental health coach. (3) Key information points: The user mentioned "I have been suffering from insomnia recently", which is the key symptom information; the voice assistant suggests "try relaxation techniques", which is the key suggestion information.
[0078] The above-mentioned labeled data will be used to train the mental health knowledge graph and improved large language model, so that the model can understand and apply this mental health knowledge to provide users with accurate and professional advice and support.
[0079] S123: Unified data format
[0080] Unify data from different sources and in different formats into a format suitable for model input (JSON / CSV), convert text into a unified encoding method (UTF-8), and divide it according to fixed text paragraphs or sentence structures to ensure that the data can be processed correctly when entering the model training phase.
[0081] S130: Based on the preprocessed data, construct a health knowledge graph and process the conversation context;
[0082] S131: Building a health knowledge graph
[0083] (1) Determine the knowledge graph metamodel: This mainly includes the definition and types of ontology and ontology relationships in the knowledge graph. For example, in the health knowledge graph, ontology includes psychological symptoms, mental illness, emotional state, coping strategies, influencing factors, mental health resources, etc., while ontology relationships include symptoms and diseases, emotions and symptoms, diseases and coping strategies, influencing factors and symptoms / diseases, resources and suggestions, etc.
[0084] (2) Knowledge graph construction:
[0085] Initialization phase: Unlabeled corpus data is obtained from various sources such as relevant guidelines, expert consensus, books, literature, medical websites, etc., and researchers manually annotate the entities, relationships between entities, entity attributes, etc. to obtain the corpus for initial model training.
[0086] Ontology construction of knowledge graph: This is a process of continuous iteration of training models and new corpus. First, based on the ontology model of the initial model training phase, a new round of model ontology automatic annotation is performed on knowledge materials from various sources. At the same time, the newly identified ontology is imported into the tool for manual review and modification to update the ontology recognition model training set. Then the ontology recognition model is retrained with the updated training set, and the iterative process continues until no new entity pairs are recognized, that is, the knowledge graph ontology construction is completed.
[0087] Ontological relationship construction of knowledge graph: determine whether the relationship between entities is established.
[0088] Knowledge storage and query: Use graph databases such as Neo4j for knowledge storage and query. Neo4j graph database supports query reasoning of knowledge graphs and can efficiently store and retrieve entities, attributes, and relationships in knowledge graphs.
[0089] S132: Processing Dialogue Context
[0090] When processing the conversation context, the system needs to be able to understand and remember the previous conversation content in order to provide accurate and coherent responses in subsequent conversations, involving the following steps:
[0091] (1) Obtaining context information: The system needs to save historical conversation records or use a memory module to store conversation context. In this way, every time a new conversation occurs, the system can obtain the previous conversation content.
[0092] (2) Inferring contextual intent: By analyzing the context and keywords of the user’s question, the system needs to be able to infer the user’s intent. For example, if the user mentions the symptoms of a disease, the system needs to understand that the user may be asking about treatments or preventive measures for the disease.
[0093] (3) Generate responses using contextual information: When generating responses, the system needs to comprehensively consider the current user input and the previous conversation context. This way, the system can provide accurate, coherent responses that are consistent with the user’s intent.
[0094] In addition, in order to improve the efficiency and accuracy of dialogue processing, the system can also adopt some advanced technologies, such as dialogue strategy optimization based on reinforcement learning, memory network enhancement, etc. These technologies can help the system better understand contextual information and maintain the consistency of roles, intentions, and knowledge during the dialogue process.
[0095] S140: Setting an improved large language model; wherein the improved large language model includes an embedding layer, a multi-layer recursive structure, a residual connection, and a fully connected layer, which are arranged in sequence according to the order of the data stream and are connected to each other in a specific manner;
[0096] Specifically, the embedding layer, multi-layer recursive structure, residual connection and fully connected layer in the improved large language model are connected to each other in the following specific way:
[0097] Embedding layer, which is used to convert discrete categorical data (words, category labels, etc.) into a low-dimensional continuous vector representation x embed , that is, the embedding representation x embed ;
[0098] A multi-layer recursive structure is used to perform recursive calculations on the embedding representation output by the embedding layer. It is composed of multiple layers of identical or similar sub-layers. Each layer performs recursive calculations based on the output of the previous layer. Each layer contains the following three components: self-attention processing, Chebyshev feature extraction, and normalization operations.
[0099] Among them, the input is processed in parallel through a multi-head attention mechanism (8 or 16 heads) to capture the long-distance dependencies in the sequence. coeffs Extract features from the input to enhance the model's expressiveness. Normalize the input to prevent the model from overfitting or gradient disappearance / explosion during training.
[0100] Residual connection, which is used to add the input of each layer of the multi-layer recursive structure to the output of the layer to form a residual connection;
[0101] The fully connected layer is used to perform linear transformation on the output of the last layer of the multi-layer recursive structure and output the prediction result.
[0102] The input sequence passes through the embedding layer, which converts discrete categorical data (words, category labels, etc.) into a low-dimensional continuous vector representation, that is, the embedded representation x embed , the embedding representation x embed Through the stacked multi-layer structure, each layer is a recursive calculation based on the output of the previous layer; each layer consists of self-attention processing, Chebyshev feature extraction and normalization operations; the layers are connected through residual connections. The last layer is a fully connected layer (or classification head), which is used to output prediction results (such as classification labels, regression values, etc.).
[0103] The difference between this model and the traditional large language model is that the weight of each layer of the network in this model is composed of the Chebyshev polynomial coefficient matrix cheby coeffs Indicates that it is dynamically adjusted through training, that is, based on the Chebyshev polynomial T degree (x) is used to replace the basis function of the traditional large language model to expand or represent the input x:
[0104] T degree (x) = cos(degree arccos(x))
[0105] T 0 (x) = 1
[0106] T 1 (x) = x
[0107] T degree (x) = 2x·T degree-1 (x)-T degree-2 (x)
[0108] Among them, the input x is the health knowledge graph or conversation context; the Chebyshev polynomial T degree (x) is a degree polynomial, T degree-1 (x) is a degree-1 Chebyshev polynomial, T degree-2 (x) is a degree-2 Chebyshev polynomial; arccos(x) is the inverse cosine function, and degree is an index value used to index the order of the corresponding Chebyshev polynomial, thereby calculating the polynomial value at each order;
[0109] The input x is projected onto the coefficients of each Chebyshev polynomial order cheby coeffsPerform forward propagation calculation on:
[0110]
[0111] Among them, c degree is the coefficient, T degree (x) is the Chebyshev polynomial expansion term, and Degree is the Chebyshev polynomial order. m The mth eigenvalue of the final output is obtained by weighting each input feature over Chebyshev polynomials of different orders to obtain a specific dimension of the output (i.e., the mth element of the output). m is the index of the feature in the output space, representing the mth output feature currently calculated. In neural networks, the output can have multiple dimensions (a classification task may have multiple class probabilities), and each dimension of the output value corresponds to a specific output.
[0112] Among them, the core parameters of the network are set as follows:
[0113] Embedding dimension: 256 or 512.
[0114] Number of attention heads: Set multiple attention heads to improve parallel processing capabilities (8 or 16).
[0115] Number of Chebyshev feature extractions: determines the depth of feature learning (2 to 4).
[0116] Chebyshev polynomial order (Degree): The expansion used for Chebyshev feature extraction (5 to 10).
[0117] S150: Creating a knowledge question-answering task, embedding the health knowledge graph into the improved large language model for health knowledge fusion training; including:
[0118] (1) Create a knowledge question-answering task
[0119] Convert the knowledge points in the health knowledge text into questions and corresponding correct answers, and let the improved large language model (see step S140) learn how to retrieve and output accurate health knowledge based on the given questions; or design knowledge reasoning tasks, given some health-related prerequisites, and require the improved large language model to infer reasonable conclusions, so as to strengthen the improved large language model's understanding and application capabilities of entities and relationships in health knowledge.
[0120] (2) Health knowledge integration training
[0121] During the training process, we use the entity and relationship information in the annotated health knowledge to guide the improved large language model to focus on the key semantic structures in the health knowledge text, use the health knowledge graph to guide the training, and enhance the entity recognition and relationship inference capabilities through the loss function, so that the improved large language model can accurately associate and apply relevant knowledge when generating answers. For example, when a user asks about the treatment of a certain disease, the improved large language model can give reasonable suggestions based on the learned relationship between the disease and the treatment.
[0122] The health knowledge graph is represented by a triple (h, r, t), which can be represented as follows:
[0123] (h: anxiety disorder, r: symptoms, t: nervousness)
[0124] (h: depression, r: cause, t: genetic factors)
[0125] (h: stress, r: coping strategies, t: time management)
[0126] (h: meditation, r: effect, t: anxiety relief)
[0127] (h: low mood, r: possible direction, t: depression)
[0128] Specifically, the health knowledge integration training process includes:
[0129] The discrete data of entities and relationships in the health knowledge graph are converted into low-dimensional continuous vector representations through the embedding layer, where the vector representation contains the semantic information of entities and relationships in the graph;
[0130] The vector representation output by the embedding layer is fed into the multi-layer recursive structure for deep feature extraction and sequence modeling. In each layer, the self-attention processing mechanism can capture the complex dependencies between entities and relationships in the graph. The Chebyshev feature extraction further extracts feature information from the graph by performing nonlinear transformation on the input vector through Chebyshev polynomials. The normalization operation ensures the stability of the model during training.
[0131] The feature representation processed by the multi-layer recursive structure is sent to the fully connected layer, and the fully connected layer maps the feature representation to the output space according to the specific task and outputs the prediction result;
[0132] Furthermore, a joint loss function is used to optimize the model, so that the improved large language model can accurately associate and apply relevant knowledge when generating answers.
[0133] Among them, the joint loss function adopted in health knowledge fusion training includes the loss of entity classification task, the loss of relationship classification task and the mixed loss function to enhance the ability of entity recognition and relationship inference.
[0134] Furthermore, the joint loss function of network training is:
[0135]
[0136] Among them, β 1 , β 2 , β 3 is the weight coefficient, which is used to balance the losses of each task; The loss for the entity classification task; loss for relation classification tasks; is a mixed loss function.
[0137] Given a health question Q containing an entity e, the loss of the entity classification task is Defined as:
[0138]
[0139] y i Represents the label of the entity category. Each y i is a specific entity category label, usually a discrete classification label. For example, for mental health issues, entities may be psychological symptoms, mental illness, emotional state, coping strategies, influencing factors, mental health resources, etc. i represents the index of the entity category label. Assuming there are N entity categories, i is an index from 1 to N, indicating the number of entity categories. For example, i=1 corresponds to the entity category "psychological symptoms", i=2 corresponds to the entity category "mental illness", and so on.
[0140] P(y i ): The probability of the model predicting the entity category.
[0141] P(y i |Q, e): Given a health problem Q and an entity e (such as a psychological symptom, a mental illness, etc.), the model predicts that the entity belongs to category y i probability.
[0142] Predict the relationship r between the head entity h and the tail entity t, the loss of the relationship classification task Defined as:
[0143]
[0144] z j Represents the relationship category label, each z jis a specific relationship category label, indicating the type of relationship between the head entity h and the tail entity t. For example, in the field of mental health, possible relationships include "leading to", "treatment", "accompanying", etc. j represents the index of the relationship category label. Assuming there are M relationship categories, j is an index from 1 to M, indicating the number of the relationship category. For example, j = 1 may represent a "leading to" relationship, j = 2 may represent a "treatment" relationship, and so on.
[0145] P(z j ): The probability of the model predicting the relationship category.
[0146] P(z j | h, t, K): Given the head entity h, the tail entity t, and the knowledge base K, the model predicts the relationship between the two entities as category z j probability.
[0147] Assume that given a health question Q and an answer A, the goal is to generate a highly relevant answer through knowledge enhancement, and the mixed loss function is for:
[0148]
[0149] Text Generation Loss
[0150]
[0151] k represents a position index in the sequence, which refers to the position of the currently generated word when generating the sequence (the kth word).
[0152] Assume that answer A contains K words, and K in the formula represents the number of words in the generated answer.
[0153] P(A k ): Generate the kth word A in the model k probability.
[0154] A <k : represents all generated words before the kth word.
[0155] The knowledge graph is embedded and input into the improved large language model as context information (see step S140).
[0156] Knowledge graph loss.
[0157] α: Balance coefficient, which controls the weight of knowledge loss and text generation loss.
[0158] In order to let the model learn the health knowledge graph, the knowledge graph loss of the knowledge embedding model is used
[0159]
[0160] in, is the set of true triples (head entity, relation, tail entity). h, r, t: are the embedding vectors of the head entity, relation, and tail entity, respectively. h′ and t′ are the false entities generated by negative sampling (i.e., false triples that are not in the knowledge graph). d(·) is the distance function, usually the L2 norm or L1 norm, for example, d(h+r, t) = ||h+rt|| 2 γ is the marginal loss, controlling the separation distance between positive and negative triplets. [·] + : ReLU activation function, used to ensure non-negative.
[0161] The embedding vector h, r, t is optimized to minimize the distance of true triplets while maximizing the distance of false triplets.
[0162] S160: Constructing a dialogue scenario, embedding the dialogue context and knowledge into the improved large model, and performing dialogue training, including:
[0163] (1) Constructing a dialogue scenario
[0164] Construct simulated mental health consultation dialogue scenarios, invite professional counselors or mental health experts to participate in simulated dialogues, and record the content of the dialogues. These simulated dialogues can cover common mental health problems, counseling skills, and coping strategies, etc., and are used to train and optimize the mental health coach voice model. Using the collected health consultation dialogue data and artificially constructed simulated dialogue scenarios, let the model learn how to make appropriate responses in different dialogue situations, train the model to understand user intentions, generate responses that match the context, and let the model master the dialogue logic in health consultation scenarios. For example, when simulating a user asking for anxiety advice, the model can give relevant responses such as doing moderate exercise, such as walking, yoga, or running; seeking social support, and sharing your feelings with friends or family.
[0165] (2) Dialogue training
[0166] The dialogue training process includes:
[0167] The embedding layer converts the discrete data of words and phrases in the text into low-dimensional continuous vector representations, where the vector representation contains the semantic information in the conversation context;
[0168] The vector representation output by the embedding layer is fed into a multi-layer recursive structure for deep feature extraction and sequence modeling. In each layer, the self-attention processing mechanism can capture the semantic dependencies in the conversation context. The Chebyshev feature extraction uses Chebyshev polynomials to perform nonlinear transformations on the input vectors to further extract key information in the conversation context. The normalization operation ensures the stability of the model during training.
[0169] The conversation context feature representation processed by the multi-layer recursive structure is sent to the fully connected layer, and the fully connected layer maps the feature representation to the output space according to the specific task and outputs the prediction result;
[0170] Furthermore, a joint loss function is used to optimize the model to improve the fluency and emotional suitability of the conversation.
[0171] Furthermore, the dialogue training process also includes:
[0172] The improved large language model (see step S140) embeds the conversation context into a high-dimensional vector space. 1 , c 2 , ..., c m Each word c in i Processed by embedding layer and self-attention mechanism. Context vector C and knowledge embedding After fusion, it is input into the decoder. The decoder generates part R according to the context and the history of the target response. <i , predicting the target response word by word. Among them, the dialogue training adopts a joint loss function, including dialogue generation loss, emotional suitability constraint and fluency constraint, to improve the fluency and emotional suitability of the dialogue.
[0173] Specifically, the joint loss function is:
[0174]
[0175] λ 1 ,λ 2 is a hyperparameter used to balance the weights of various losses. Generate the loss for the dialogue. is the sentiment suitability constraint. PPL is the fluency constraint, which allows the dialogue to be generated smoothly. n is the number of words in the target response, which is used to calculate the generation fluency of the entire sequence.
[0176] The core of dialogue training is the sequence generation task. Given the dialogue context C and target response R, the dialogue generation loss Optimize the model by maximizing the probability of the target response:
[0177]
[0178] Where n is the number of words in the target response, which is used to calculate the fluency of the entire sequence (same as n in PPL above). C is the conversation context, including the previous conversation turn. R = {R 1 , R 2 , ..., R n} is the target response, which contains n words. <k Indicates the generation of the kth word R k All the words before, k represents a position index in the sequence, which refers to the position of the currently generated word when generating the sequence (the kth word). is a knowledge graph embedding used to enhance contextual information. k ): Generate the kth word R k The conditional probability is calculated by the improved large language model (see step S140).
[0179] The sentiment category s of the generated response is predicted by the sentiment classifier, and the sentiment suitability constraint is calculated with the target sentiment label s* (that is, the real sentiment category, that is, the label or actual value in the dataset):
[0180]
[0181] Where S is all possible emotion categories. For example, if the emotion classification task is binary (such as positive and negative), then S = {positive, negative}; if it is multi-classification (such as happy, sad, angry, etc.), then S contains all these categories. p(s) is the probability that the model predicts a certain emotion category s. This is the result of the model output layer after softmax (or similar function) processing, ensuring that the sum of the predicted probabilities of all categories is 1. p(s * ) is the model prediction for the true sentiment category s * probability.
[0182] S170: Determine the evaluation indicators, and analyze the deficiencies and problems of the model based on the feedback results of the evaluation indicators, and optimize them.
[0183] The specific step S170 includes:
[0184] (1) Determine the evaluation indicators:
[0185] Accuracy of mental health knowledge: By constructing a mental health knowledge test set, the degree of consistency between the mental health-related answers output by the model and authoritative mental health knowledge is examined. For example, whether the answers to the symptoms and treatment methods of mental illness are accurate.
[0186] Answer relevance: After given a user's mental health consultation question, evaluate the degree of relevance between the model's response and the question, that is, whether the response is actually targeted at the mental health topic asked by the user, avoiding irrelevant answers.
[0187] Effectiveness of recommendations: For the mental health recommendations given by the model (maintaining a regular schedule, exercise recommendations, etc.), professional mental health personnel or reference to authoritative standards will judge whether they are actually feasible and effective, and whether they can truly play a positive role in the user's mental health.
[0188] Conversational fluency and emotional appropriateness: This test examines whether the response sentences generated by the model during the conversation are fluent and natural, and whether they can reflect appropriate emotional attitudes (concern, friendliness, etc.) according to the conversation context, thereby improving the user's conversation experience.
[0189] (2) Model optimization:
[0190] Based on the feedback results of the above evaluation indicators (accuracy of mental health knowledge, relevance of answers, effectiveness of suggestions, and fluency and emotional appropriateness of conversations), the deficiencies and problems of the model were analyzed, and the following model optimization measures were implemented:
[0191] a. If the health-related answers output by the model are less consistent with authoritative health knowledge, that is, the accuracy of health knowledge is poor, then increase the amount of training related health knowledge data: collect data from more authoritative health books, medical literature, authoritative health websites and other channels to ensure that the model can learn more comprehensive and accurate mental health knowledge; adjust the design of training tasks: design more refined training tasks, such as knowledge question-and-answer tasks targeting mental illness or health issues, to strengthen the model's understanding and application of knowledge in the field of mental health.
[0192] b. If the content of the model's response is not closely related to the user's mental health counseling questions, that is, the answer is not relevant, then strengthen context understanding: By optimizing the model's context vector representation and attention mechanism, improve the model's ability to understand the conversation context, and ensure that the model can accurately capture the user's intentions and concerns; increase training of relevant conversation data: Collect more conversation data related to the user's mental health counseling questions, so that the model can learn how to give relevant responses to the user's questions during the training process.
[0193] c. If the health advice given by the model lacks practical feasibility and effectiveness, a psychological expert evaluation mechanism will be introduced: during the training process, professional mental health personnel will be invited to evaluate and provide feedback on the advice generated by the model, and the output logic and parameter settings of the model will be adjusted according to the feedback results; refer to authoritative mental health standards: when generating mental health advice, refer to authoritative mental health standards and guidelines to ensure the scientific nature and effectiveness of the advice; increase personalized features: combine the user's personal information and mental health status to generate personalized mental health advice for different users, and improve the pertinence and effectiveness of the advice.
[0194] d. If the response sentences generated by the model during the conversation are not smooth and natural, or fail to reflect the appropriate emotional attitude, that is, the conversation fluency or emotional suitability is insufficient, then optimize the language generation model: improve the fluency and naturalness of the model-generated responses by adjusting the parameter settings of the language generation part, introducing more advanced language generation algorithms, etc.; increase the training of high-quality conversation data: collect more high-quality, natural and fluent conversation data, so that the model can learn how to generate responses that conform to human communication habits during the training process; introduce emotion recognition and generation modules: add emotion recognition and generation modules to the model, perform emotion analysis on the user's input, and then generate responses with appropriate emotional attitudes based on the emotion analysis results.
[0195] This step S170 can effectively optimize the model based on the feedback results of each evaluation indicator, and improve the performance of the mental health coach voice model in terms of health knowledge accuracy, answer relevance, suggestion effectiveness, and conversation fluency and emotional suitability.
[0196] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0197] Figure 3 A schematic block diagram of an electronic device 300 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0198] The electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a ROM 302 or a computer program loaded from a storage unit 308 into a RAM 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An I / O interface 305 is also connected to the bus 304.
[0199] A number of components in the electronic device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0200] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above, such as the mental health coach voice model optimization method based on a large language model drive. For example, in some embodiments, the mental health coach voice model optimization method based on a large language model drive may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the mental health coach voice model optimization method based on a large language model drive described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to execute the described method in any other appropriate manner (for example, by means of firmware).
[0201] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0202] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0203] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0204] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0205] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0206] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0207] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0208] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for optimizing a mental health coach's speech model based on a large language model, characterized in that: The following steps are involved: Collect health knowledge data and health consultation dialogue data from multiple channels and construct a multi-source heterogeneous data set; Preprocessing the data in the multi-source heterogeneous data set, wherein the preprocessing includes data cleaning, data labeling, and data format unification; Based on the preprocessed data, build a health knowledge graph and process the conversation context; Setting an improved large language model; wherein the improved large language model includes an embedding layer, a multi-layer recursive structure, a residual connection and a fully connected layer, which are arranged in sequence according to the order of the data stream and are connected to each other in a specific manner; Creating a knowledge question-answering task, embedding the health knowledge graph into the improved large language model to perform health knowledge fusion training; Construct a dialogue scenario, embed the dialogue context and knowledge into the improved large model for dialogue training; Determine the evaluation indicators, and analyze the deficiencies and problems of the model based on the feedback results of the evaluation indicators, and optimize them.
2. The method according to claim 1, characterized in that in, The embedding layer, multi-layer recursive structure, residual connection and fully connected layer in the improved large language model are connected to each other in the following specific way: The embedding layer is used to convert discrete categorical data into a low-dimensional continuous vector representation, i.e., an embedding representation, wherein the categorical data includes words and category labels; The multi-layer recursive structure is used to perform a recursive calculation operation on the embedding representation output by the embedding layer, which is composed of a plurality of identical or similar sub-layers stacked together, each layer performs a recursive calculation based on the output of the previous layer, and each layer includes the following three components: self-attention processing, Chebyshev feature extraction, and normalization operation; The residual connection is used to add the input of each layer of the multi-layer recursive structure to the output of the layer to form a residual connection; The fully connected layer is used to perform a linear transformation on the output of the last layer of the multi-layer recursive structure and output a prediction result.
3. The method according to claim 1 or 2, characterized in that: in, The weight of each layer of the network of the improved large language model is represented by a Chebyshev polynomial coefficient matrix, which is dynamically adjusted through training, and the core parameters of the network include embedding dimension, number of attention heads, number of Chebyshev feature extractions and Chebyshev polynomial order.
4. The method according to claim 3, characterized in that Among them, include: Based on Chebyshev polynomial T degree (x) Expand or represent the input x: T degree (x)=cos(degree·arccos(x)) T0(x)=1 T1(x)=x T degree (x)=2x·T degree-1 (x)-T degree-2 (x) Among them, the input x is the health knowledge graph or conversation context; the Chebyshev polynomial T degree (x) is a degree polynomial, T degree-1 (x) is a degree-1 Chebyshev polynomial, T degree-2 (x) is a degree-2 Chebyshev polynomial; arccos(x) is the inverse cosine function, and degree is an index value used to index the order of the corresponding Chebyshev polynomial, thereby calculating the polynomial value at each order; The input x is projected onto the coefficients of each Chebyshev polynomial order cheby coeffs Perform forward propagation calculation on: Among them, c degree is the coefficient, T degree (x) is the expansion term of Chebyshev polynomial, Degree is the order of Chebyshev polynomial; y m It represents the mth eigenvalue of the final output. The specific dimension of the output, i.e. the mth element of the output, is obtained by weighted summing each input feature on Chebyshev polynomials of different orders. m is the index of the feature in the output space, which represents the mth output feature currently calculated.
5. The method according to claim 4, characterized in that in, The health knowledge integration training process includes: The discrete data of entities and relationships in the health knowledge graph are converted into low-dimensional continuous vector representations through the embedding layer, wherein the health knowledge graph is represented by a triple head entity, a relationship, and a tail entity, and the vector representation contains semantic information of the entities and relationships in the graph; The vector representation output by the embedding layer is fed into the multi-layer recursive structure for deep feature extraction and sequence modeling; wherein, in each layer, the self-attention processing mechanism can capture the complex dependencies between entities and relationships in the graph; the Chebyshev feature extraction further extracts feature information from the graph by performing nonlinear transformation on the input vector through Chebyshev polynomials; and the normalization operation ensures the stability of the model during the training process; The feature representation processed by the multi-layer recursive structure is sent to the fully connected layer, and the fully connected layer maps the feature representation to the output space according to the specific task and outputs the prediction result; Furthermore, a joint loss function is used to optimize the model, so that the improved large language model can accurately associate and apply relevant knowledge when generating answers.
6. The method according to claim 5, characterized in that in, The joint loss function adopted in the health knowledge fusion training includes the loss of entity classification task, the loss of relationship classification task and a mixed loss function to enhance the entity recognition and relationship inference capabilities.
7. The method according to claim 4, characterized in that in, The dialogue training process includes: The embedding layer converts the discrete data of words and phrases in the text into a low-dimensional continuous vector representation, wherein the vector representation contains semantic information in the conversation context; The vector representation output by the embedding layer is fed into the multi-layer recursive structure for deep feature extraction and sequence modeling; wherein, in each layer, the self-attention processing mechanism can capture the semantic dependencies in the conversation context; the Chebyshev feature extraction performs nonlinear transformation on the input vector through the Chebyshev polynomial to further extract key information in the conversation context; and the normalization operation ensures the stability of the model during the training process; The conversation context feature representation processed by the multi-layer recursive structure is sent to the fully connected layer, and the fully connected layer maps the feature representation to the output space according to the specific task and outputs the prediction result; Furthermore, a joint loss function is used to optimize the model to improve the fluency and emotional suitability of the conversation.
8. The method according to claim 7, characterized in that in, The dialogue training process also includes: The conversation context is embedded into a high-dimensional vector space, each word in the context is processed by the embedding layer and the self-attention mechanism, and the context vector is fused with the knowledge embedding and input into the decoder. The decoder predicts the target response word by word based on the context and the historical generation part of the target response.
9. The method according to claim 7 or 8, characterized in that: The joint loss function used in the dialogue training includes dialogue generation loss, emotional suitability constraint and fluency constraint to improve the fluency and emotional suitability of the dialogue.
10. The method according to claim 1, characterized in that in, Determining the evaluation indicators, and analyzing the deficiencies and problems of the model based on the feedback results of the evaluation indicators, and optimizing them, including: Based on the feedback from the evaluation indicators, analyze the deficiencies and problems of the model, and accordingly increase the amount of training data, adjust the design of the training tasks, or optimize the parameters of the language generation part of the model.