A mental health support dialogue generation method based on multi-modal input
By combining a multimodal input-based mental health support dialogue generation method with visual language models and large language models, the problem of insufficient multimodal information fusion in existing technologies is solved. This enables a deep understanding of users' emotional and behavioral states, generates more accurate and targeted supportive responses, and improves the effectiveness of mental health support.
Patent Information
- Application Number
- CN202411966795.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing mental health dialogue systems struggle to fully identify and present users' emotional and behavioral states, and their multimodal information fusion is insufficient, resulting in inadequate refinement and professional processing in the field of mental health.
We employ a multimodal input-based method for generating mental health support dialogues. By analyzing images through a visual language model and combining dialogue context and counseling strategies, we use a large language model (LLM) to generate more accurate supportive responses. We also introduce zero-shot and few-shot learning techniques and leverage common-sense knowledge graphs to enhance model adaptability.
It improves the depth and accuracy of the dialogue system's understanding of users' emotional and behavioral states, generates more targeted supportive responses, and enhances the responsiveness of mental health support.
Smart Images

Figure CN119785980B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal human-computer interaction and natural language processing, and in particular to a mental health support dialogue generation method based on multi-modal input. BACKGROUND
[0002] Currently, artificial intelligence technology is widely used in emotion state recognition and prediction of mental disorder development trajectory to achieve personalized medical support. Further, mental health agents based on artificial intelligence can simulate the dialogue between a mental therapist and a patient, effectively helping to reduce the user's stress, anxiety and depression symptoms. The application of such technology helps to improve the accessibility of mental health services in areas where professional personnel are relatively scarce, so that more people can benefit from high-quality mental health support and services.
[0003] The research of mental health dialogue systems has gone through a process from manual rule construction to the integration of advanced models and technologies. Early methods mainly relied on artificially developed rules to generate replies with empathy features, which were effective in specific situations but lacked adaptability. Subsequently, scholars combined rules and algorithms to analyze stress-related texts on social media to assist in generating supportive replies. This method improved the naturalness and relevance of the dialogue to some extent, but was still limited by the size and single features of the data. After that, generative dialogue models based on neural networks began to develop, and deep learning technology helped chat robots understand and respond to social support needs. However, these models mostly rely on large-scale and high-quality pre-training data, which is still insufficient for the mental health field, which requires fine and professional processing. Although existing research has attempted to introduce large language models (LLM) and affective computing, there are still problems in the design of data customization and professional motivation in the context of mental health.
[0004] On the other hand, most current mental health dialogue systems are mainly based on text input, making it difficult to fully identify and present the emotional and behavioral states of users. Emotional expression often has multi-modal characteristics, such as facial expressions and body movements, and relying solely on text will result in a lack of depth and accuracy in understanding the emotional state of the user. Although some attempts have been made to incorporate multi-modal information into mental health support, such research is still relatively rare.
[0005] How to effectively integrate multi-modal information (such as facial expressions and movements in images) with dialogue context and external knowledge to improve the depth and accuracy of the dialogue system's understanding of the user's emotional and behavioral state, thereby meeting the fine and professional data processing and intervention needs of the mental health field, is still a problem that needs to be researched and improved. SUMMARY
[0006] The present application overcomes the deficiencies in the prior art, and provides a mental health support dialogue generation method based on multi-modal input, so as to significantly improve the emotional recognition and response quality, thereby comprehensively understanding and responding to the emotional state of the user, and enhancing the response ability of the generated reply of the mental health support.
[0007] To achieve the above-mentioned application purposes, the present application adopts the following technical solutions:
[0008] The mental health support dialogue generation method based on multi-modal input has the following characteristics:
[0009] Step 1, obtaining a multi-modal data set , including a training set and a test set ; any training sample in the training set is denoted as , and any test sample in the test set is denoted as ; wherein and are training text utterance sequences containing n sentences of dialogue and test text utterance sequences containing m sentences of dialogue respectively, and are visual image data of , are visual image data of
[0010] Step 2, performing semantic processing on and , and calculating the similarity between the two to obtain a high-quality example set ;
[0011] Step 3, selecting a multi-modal dialogue language model as a visual language model, and constructing a prompt text template for generating facial expressions and actions , so as to generate facial expression and action description information of using formula (6):
[0012] (6)
[0013] Step 4, generating comprehensive common knowledge text of according to the relationship type;
[0014] Step 5, generating counseling strategy text of using a large language model LLM;
[0015] Step 6, combining 、 、 、 as well as Integrate into prompt text And input into the large language model LLM to generate Final response content The method for generating a mental health support dialogue based on multimodal input according to the present invention is also characterized in that step 2 is performed as follows:
[0016] Step 2.1, All text utterances except the last sentence in the dialogue are concatenated to form a test context sequence , and use sentence-level semantic encoder Will Mapping to a high-dimensional semantic vector space to obtain a test representation vector ;
[0017] Step 2.2, The m sentences in the dialogue are spliced to form a training context sequence , and use sentence-level semantic encoder Will Mapping to a high-dimensional semantic vector space to obtain a training representation vector ;
[0018] Step 2.3: Calculate using formula (5) and Similarity :
[0019] (5)
[0020] In formula (5), is the semantic similarity function represented by cosine similarity; represents the vector inner product operation, Represents the Euclidean norm of the vector; Step 2.4, according to the pre-set few sample parameters ,in, is a positive integer, from Select the top sample with the highest semantic similarity to the test sample training samples and form a high-quality example set .
[0021] Furthermore, step 4 is performed as follows:
[0022] Step 4.1: Use the fine-tuned common sense reasoning model right and relationship type The processing is performed, and a common sense inference text about is generated ; wherein, ∈ , represents a set of relationship types;
[0023] Step 4.2, fusing the common sense inference texts of all relationship types in R by using formula (8) to generate a comprehensive common sense knowledge text :
[0024] (8)
[0025] In formula (8), represents a text string splicing operation.
[0026] Further, step 5 is performed as follows:
[0027] Step 5.1, a large language model LLM infers the psychological emotional stage of by using formula (9) :
[0028] (9)
[0029] In formula (9), is a prompt text for psychological emotional stage recognition;
[0030] Step 5.2, constructing a counseling strategy guide prompt text , and inputting it together with and into the large language model LLM, so as to obtain the counseling strategy text of . .
[0031] The electronic device comprises a memory and a processor, and the feature is that the memory is used to store a program supporting the processor to execute the dialogue generation method, and the processor is configured to execute the program stored in the memory.
[0032] The computer readable storage medium stores a computer program, and the feature is that the computer program is executed by the processor to perform the steps of the dialogue generation method.
[0033] Compared with the prior art, the beneficial effects of the present application are as follows:
[0034] 1、The present application designs a new multi-modal framework that integrates visual information to enhance the LLM model's generation of mental health support responses. Compared with the current mainstream mental support methods, the method of the present application can more comprehensively understand the emotional state of the user and capture more rich emotional expressions, thereby generating more accurate and targeted supportive replies.
[0035] 2、The present application uses zero-shot and few-shot learning techniques to select the most relevant samples by calculating semantic similarity, thereby improving the adaptability of the LLM model in the field of psychotherapy without large-scale training.
[0036] 3、The method of the present application uses corresponding strategies to guide the generation of model replies, which can help the model generate more appropriate replies at different dialogue stages, improving the coherence and effectiveness of the dialogue.
[0037] 4、The present application constructs special data processing and LLM prompt techniques tailored for mental health environments, enhancing the responsiveness and sensitivity of the LLM system, which is crucial for effective mental health support. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 is a flowchart of the method of the present application;
[0039] Figure 2 is a template structure diagram of the prompt used in the present application. DETAILED DESCRIPTION
[0040] In this embodiment, a mental health support dialogue generation method based on multi-modal input is used to enhance the ability of LLM in generating responses for mental health support. This method integrates visual language models to analyze images and capture facial expressions, actions and emotions. Then, it combines this visual data with dialogue context and counseling strategies to make more subtle and supportive replies, such as Figure 1 As shown in the specific steps are as follows:
[0041] Step 1, obtain a multi-modal data set , including: a training set and a test set ; any training sample in the training set is denoted as , and any test sample in the test set is denoted as ; wherein the , of the test sample is the number of Chinese text utterances, , of the training sample is the number of Chinese text utterances, With respect to the visual image data, the visual image data.
[0042] Step 2, semantic processing is performed on and , and the similarity between the two is calculated to obtain a high-quality example set .
[0043] Step 2.1, all text utterances in except the last dialogue are concatenated into a string, and a test context sequence is formed using formula (1) :
[0044] (1) using a sentence-level semantic encoder to map to a high-dimensional semantic vector space, so as to obtain a test representation vector using formula (2) : (2)
[0045] Step 2.2, concatenate the m-turn dialogue in , and form a training context sequence using formula (3) : (3) using a sentence-level semantic encoder to map to a high-dimensional semantic vector space, so as to obtain a training representation vector using formula (4) : (4)
[0046] Step 2.3, the similarity between and is calculated using formula (5) :
[0047] (5)
[0048] In formula (5), is a semantic similarity function represented by cosine similarity; denotes vector inner product operation, denotes the Euclidean norm of the vector. Step 2.4, according to the pre-set few-sample parameter , wherein is a positive integer, the top training samples with the highest semantic similarity to the test sample are selected from , and a high-quality example set is formed.
[0049] Step 3, select a multi-modal dialogue language model with excellent image recognition capability (e.g. GPT-4) as a visual language model, and construct a prompt text template for generating facial expressions and actions , where is a predefined text prompt used to guide the model to extract the visual description of the image. The text prompt used here is: Please describe the expressions and actions of the characters in the image. Then use formula (6) to generate facial expression and action description information :
[0050] (6)
[0051] The generated as a visual context element in the further dialogue process is the final comprehensive embodiment of the visual content , together with the dialogue context and example set, to provide richer visual and semantic support for subsequent tasks.
[0052] Step 4, according to the relationship type, generate comprehensive common sense knowledge text ;
[0053] Step 4.1, traditional dialogue generation methods only rely on historical conversation data, making it difficult to fully understand the user's situation and emotions. The introduction of a common sense knowledge graph in mental health support dialogue can enhance the understanding of implicit information and enhance cognitive empathy, overcoming the limitations of reliance on dialogue history. In order to generate context-related common sense reasoning, the COMET model based on the GPT architecture and fine-tuned using the ATOMIC dataset is introduced to introduce common sense knowledge; the COMET model can generate common sense for four relationship types, including: representing the conditions required before speaking; representing the impact produced after speaking; representing the speaker's intention; representing the speaker's emotional or attitude reaction; using the fine-tuned common sense reasoning model processes and relationship types , thereby generating common sense inference text about using formula (7) :
[0054] (7)
[0055] In formula (7), ∈ , represents the set of relationship types.
[0056] Step 4.2, fuse the common sense inference text of all relationship types in R using formula (8) to generate comprehensive common sense knowledge text :
[0057] (8)
[0058] In formula (8), represents the text concatenation operation; It can provide more complete context common sense support for subsequent dialogue generation and understanding.
[0059] Step 5, generate counseling strategy text using large language model LLM ;
[0060] Step 5.1, introduce an emotional support dialogue framework to help the large language model identify key strategies consistent with the psychological treatment stage in the psychological counseling scenario, thereby promoting emotional exploration, deepening understanding, and behavior change for the dialogue, and ensuring its safety and autonomy; The large language model LLM uses formula (9) to infer the psychological emotional stage :
[0061] (9)
[0062] In formula (9), is the prompt text for psychological emotional stage recognition, and the prompt text here is specifically: "Based on the context, which of the three stages (Exploration: Help visitors identify problems; Insight: Help visitors reach a new depth of self-understanding; Action: Help visitors decide on actions to solve problems) does the above dialogue belong to? According to the stage, explain why you make such a judgment."
[0063] Step 5.2, after identifying the dialogue stage, build counseling strategy guidance prompt text and input it into the large language model LLM together with and , so as to use formula (10) to select counseling strategy text matching the identified stage from the pre-defined strategy set: (10)
[0064] In formula (10), represents the strategy text determined by LLM, The corresponding prompt text is: "According to different stages, tell me what strategy should be used in the next response. (One of the following seven stages: asking questions, restating or rephrasing questions, reflecting emotions, self-disclosure, affirmation and reassurance, providing suggestions and information.) (Exploration stage corresponds to: asking questions, restating or rephrasing questions, reflecting emotions and self-disclosure. Insight stage corresponds to: reflecting emotions, self-disclosure, affirmation and reassurance. Action stage corresponds to: self-disclosure, affirmation and reassurance, providing suggestions and information.) Use [] to enclose the selected strategy. Here's an example: Based on the context, this conversation is in the Action stage. The next response should use the (self-disclosure) strategy; in the above steps, the LLM will select one of the seven strategies provided based on the current stage of the conversation and mark the selected strategy in the form of "[]".
[0065] Step 6, integrate , , , and into the prompt text : (11) In formula (11), represents the mental health support conversation generation task prompt, and the specific prompt text is: "This is a psychological counseling task: In this task, the visitor (the first person) will share their feelings, experiences and challenges with the psychological counselor. The psychological counselor (the second person) needs to provide mental health and emotional support. The work of the psychological counselor is not to provide solutions or suggestions, but to help the visitor explore their feelings through listening, empathy and open-ended questions. The psychological counselor should promote the visitor's self-understanding and growth, help them identify and deal with inner conflicts or distress. In addition, the psychological counselor should maintain professional boundaries, keep confidential, and guide the visitor to seek more professional medical or mental health services when necessary. In this specific conversation, you will play the role of counselor, and based on the information provided by the visitor, give your next round of response." It is composed of multiple components, and the order will affect the generation result of the model; use a specially designed prompt template in the multi-modal psychological support scenario to assist the LLM in better meeting the counseling task requirements in subsequent responses. The content and structure of the prompt template are as follows Figure 2The task definition provides a comprehensive description of the tasks and functions required for the LLM. Samples are selected based on the zero-shot or few-shot setting to enhance the model's understanding of the task. Expression description analyzes visual images to capture the searcher's facial expressions and movements, enriching the text-based input. External knowledge supplements the dialogue context with additional relevant knowledge. Strategy guidance involves selecting strategic content based on the context to guide the model's response. The final dialogue context includes the dialogue history, excluding the last utterance of the counselor. The main goal of this arrangement is to enable the LLM to simulate the role of the counselor and effectively generate responses for the subsequent turns. The input large language model LLM, thereby generating the final reply content for using equation (12):
[0066] (12)
[0067] In this embodiment, an electronic device includes a memory for storing a program supporting the processor to execute the above method, and a processor configured to execute the program stored in the memory.
[0068] In this embodiment, a computer-readable storage medium has a computer program stored thereon, and the computer program, when executed by a processor, performs the steps of the above method.
Claims
1. A mental health support conversation generation method based on multi-modal input, characterized in that, is performed as follows: step 1, obtaining a multi-modal data set , comprising: a training set and a test set ; any training sample in the training set is recorded as , and any test sample in the test set is recorded as ; wherein, and are training text utterance sequences containing n dialogues and test text utterance sequences containing m dialogues, respectively, and are visual image data of , and is visual image data. Step 2, semantic processing on and and computing similarity between them to get high-quality example set ; Step 2.1, concatenate all text utterances except the last dialogue turn to form a test context sequence and use a sentence-level semantic encoder to map to a high-dimensional semantic vector space, resulting in a test representation vector ; Step 2.2, concatenate the dialogues in m sentences in to form training context sequences and use a sentence-level semantic encoder to map to a high-dimensional semantic vector space, resulting in training representation vectors ; Step 2.3, calculating with formula (5) Similarity to : (5) In formula (5), a semantic similarity function expressed by cosine similarity; denotes a vector inner product operation, denotes a Euclidean norm of a vector; Step 2.4, according to the pre-set few-sample parameters wherein, is a positive integer, from select the top training samples with the highest semantic similarity with the test sample from the training set, and constitute a high-quality example set ; Step 3, select a multi-modal dialogue language model as a visual language model, and construct a prompt text template for generating facial expressions and actions , so as to generate facial expression and action description information : (6) Step 4. Generating a comprehensive common sense text based on the relationship type ; Step 4.1: Use the fine-tuned common sense reasoning model right and relationship type Process and generate Common sense inference text ;in, ∈ , Represents a set of relationship types; Step 4.2, fuse the texts of all the relation types in R using the common sense inference text of formula (8) to generate the comprehensive common sense knowledge text : (8) In formula (8), represents a text string concatenation operation; Step 5, generating a consultation strategy text using a large language model LLM ; Step 5.1, the large language model LLM infers using equation (9) the psychological emotional phase : (9) In formula (9), is a prompt text for psychological emotional phase recognition; Step 5.2: Construct consultation strategy guidance text , and with and Input the large language model LLM together to get Consulting strategy text ; Step 6, combine , , , and into a prompt text and input into a large language model (LLM) to generate the final reply content for .
2. An electronic device comprising a memory and a processor, characterized in that The memory is configured to store a program supporting the processor to execute the dialogue generation method of claim 1.
3. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, performs the steps of the dialogue generation method of claim 1.
Citation Information
Patent Citations
Session type artificial intelligence driven personality simulation system based on context awareness and operation method
CN117874185A
Intelligent customer service question-answering system and method
CN118861223A