Psychological intervention chat agent system based on knowledge enhancement large language model
By integrating domain knowledge and emotional interaction mechanism into the large language model and combining three-dimensional display technology, a psychological intervention chat agent system is designed, which solves the limitations of artificial intelligence chat robots in the field of mental health in the existing technology, and achieves efficient and personalized mental health support.
Patent Information
- Application Number
- CN202510158770.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-16
AI Technical Summary
The existing AI chatbots have limitations in the integration of expertise and emotional communication in the field of mental health, resulting in their unsatisfactory results in providing personalized and professional psychological interventions.
By integrating domain knowledge into the large language model, a comprehensive mental health knowledge base is built, and combining emotional interaction mechanisms and stereoscopic display technology, a psychological intervention chat agent system based on knowledge-enhanced large language model is designed. The system can analyze user emotions in real time, dynamically select intervention strategies, and provide personalized and professional mental health support.
It significantly improves the naturalness and realism of emotional interactions, enhances the user's sense of immersion and participation, ensures the personalization and efficiency of psychological intervention, and significantly improves the effectiveness of automated consultation and personalized intervention capabilities in the field of mental health.
Smart Images

Figure CN120015243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a psychological intervention chat agent system based on a knowledge-enhanced large language model. The present invention provides intelligent psychological intervention and consulting services by integrating domain knowledge into the large language model, aiming to improve the automated consulting effect and personalized intervention capabilities in the field of mental health. Background Art
[0002] Psychological problems, especially common mental illnesses such as depression, have become a major challenge facing global health. Although there is a huge demand for patients with mental illness, the lack of existing treatment resources, the economic burden and misunderstanding of psychotherapy have led to a large number of patients failing to receive effective treatment. At the same time, the psychotherapy process is usually long and it is difficult for patients to see the treatment effect quickly, which further reduces their participation in treatment.
[0003] Currently, artificial intelligence is gradually being applied to the field of mental health. It can process large amounts of data in a short period of time and assist therapists in making decisions in real time, thereby improving treatment efficiency. In particular, chatbots based on large language models have attempted to replace traditional psychotherapy to a certain extent. However, existing artificial intelligence chatbots still have limitations in fully integrating professional knowledge in the field of mental health and achieving effective emotional communication, which makes them less effective in providing personalized and professional psychological interventions.
[0004] Therefore, there is an urgent need for an artificial intelligence psychological intervention chat agent system that can integrate professional knowledge and enhance emotional interaction capabilities to better meet the needs of intelligent and personalized intervention in the field of mental health. Summary of the invention
[0005] The present invention provides a psychological intervention chat agent system based on a knowledge-enhanced large language model, which aims to overcome the deficiencies of existing artificial intelligence chat robots in domain knowledge application and emotional interaction. By combining knowledge enhancement with emotional interaction mechanisms, the present invention provides a more personalized, professional and intelligent solution for psychological intervention, which can effectively improve the intervention effect in the field of mental health, especially for the non-strict psychological intervention needs of sub-healthy groups.
[0006] To achieve the above technical objectives, the present invention provides the following technical solutions: a psychological intervention chat agent system based on a knowledge-enhanced large language model, comprising the following steps:
[0007] S1. Optimize the large language model to enable it to handle complex mental health conversations, so that it can accurately understand the user's emotions and needs, generate personalized intervention content, and improve the intelligence level of psychological intervention;
[0008] S2. Build a comprehensive mental health knowledge base, conduct real-time analysis based on user conversation content, automatically identify potential psychological problems, and provide users with accurate and effective intervention strategies through professional information in the knowledge base;
[0009] S3. Design and implement an emotional interaction mechanism, through the comprehensive use of virtual images, dynamic facial expressions and emotional voice, so that the system can present natural emotional responses in the conversation and enhance the user's emotional resonance and interactive experience;
[0010] S4. Introduce advanced stereoscopic display technology to enhance the visual effects and realism of virtual images, thereby enhancing the user's sense of immersion and participation, and improving the realism and attractiveness of the interactive experience;
[0011] S5. The system dynamically selects and applies the most appropriate intervention strategy based on the user’s psychological needs and input through intelligent process optimization, ensuring personalized and efficient mental health support;
[0012] Preferably, the S1 specifically includes the following steps:
[0013] S11 uses ChatGLM2-6B as the basic large language model. It is based on the Transformer architecture and has powerful text generation and understanding capabilities. It is particularly suitable for processing long texts and capturing long-term dependencies, and can effectively support multi-round conversations.
[0014] S12 performs LoRA fine-tuning on the ChatGLM2-6B large language model. By introducing a low-dimensional bypass structure based on the original model, the model's performance in the field of mental health can be optimized, which can avoid large-scale parameter adjustments and improve efficiency.
[0015] By fine-tuning the large language model and optimizing its application in the field of mental health, it can efficiently generate personalized intervention content in specific fields, thereby improving the model's ability to handle complex mental health conversations.
[0016] Preferably, S2 specifically includes the following steps:
[0017] S21 builds a mental health knowledge base based on DSM-5 standards, including "disorder description", "client characteristics", "therapist characteristics", "intervention strategies", "prognosis" and "assessment methods";
[0018] This information provides professional support for the system when conducting mental health conversations, helping the system identify potential psychological problems of users and provide precise intervention strategies;
[0019] S22 analyzes the user's conversation content in real time and uses the TextRank algorithm to extract the text summary of the speech;
[0020] After S23 obtains the summary, it extracts keywords through TF-IDF and performs text similarity analysis with the “obstacle description” in the knowledge base;
[0021] After completing the initial matching at S24, the same method is used, with the “characteristics of the preferred therapist” and “assessment” of the identified disorder as the precondition, to continue asking questions to continuously obtain the client’s speech, which is then used to match the “client characteristics” of the identified disorder;
[0022] When the barrier match exceeds the threshold, the “intervention strategy” is used as the new preset to complete the rest of the conversation;
[0023] S25 Given that the traditional input structure of large language models is usually based on a prompt mechanism, the input structure of the model is optimized;
[0024] By effectively embedding structured psychological knowledge into the model's input process, the system can infer the required professional psychological knowledge based on the context of the current conversation and accurately apply this knowledge when generating responses, thereby providing intervention plans that are highly matched to the user's psychological state.
[0025] Preferably, S3 specifically includes the following steps:
[0026] S31 uses cartoon-style avatar design to optimize the emotional connection between the virtual agent and the user;
[0027] Cartoon-style design will reduce users’ expectations of the authenticity of virtual agents’ appearance, making them pay more attention to their interactions and emotional feedback with virtual agents, thus improving the naturalness and credibility of emotional interactions.
[0028] S32 uses Live2D technology to dynamically control the virtual image, giving it a near-3D visual effect. It can also adjust its expression and movements in real time according to the user's input during the interaction, further enhancing the realism and affinity of the interaction.
[0029] S33 performs sentiment annotation on the user's conversation text;
[0030] By adopting the sentiment annotation model, the system extracts the sentiment information in the conversation text and converts it into numerical representations of multiple sentiment dimensions;
[0031] These emotion dimension values are then mapped into facial expression activity units (AUs) to drive the facial expression changes of the virtual agent;
[0032] S34 uses the VITS2 speech synthesis model, which enables the speech output to be adjusted in real time according to the conversation content and emotional state;
[0033] Adjust the pitch, speed, and tone of speech through emotional annotation to ensure that the speech expression matches the user's emotional changes;
[0034] S35 uses Whisper voice recognition technology to recognize user voice input in real time to ensure smooth voice interaction;
[0035] Based on the facial expression generation, S36 uses forced alignment technology to synchronize the speech and facial expressions at the corresponding moments to ensure that the lip shape of the speech is consistent with the facial expression of the virtual agent;
[0036] Ensure the visual and auditory consistency of the virtual agent’s emotional expression, thereby improving the realism of the interaction and the accuracy of the emotional expression;
[0037] Through the synchronous adjustment of facial expressions and voice, the system can provide a more humane form of interaction, further enhance the user's sense of participation and the fluency of interaction, and provide more effective support for mental health intervention.
[0038] Preferably, the S4 specifically includes the following steps:
[0039] The S41 uses pseudo-holographic display technology to display the chat agent avatar;
[0040] Through this display method, the virtual image not only has a higher visual realism, but also can bring a three-dimensional visual effect during the interaction process, improving the visual effect limitations of traditional video consultation;
[0041] S42 designs and builds a stereo display, with an internal integrated Rockchip RK3588 processor as the main control, which enables efficient image processing and rendering;
[0042] The S43 video signal is transmitted to the semi-transparent screen via the HDMI interface for display, thus generating a pseudo-holographic effect;
[0043] A backlight system is added to the display design, and a semi-white transparent diffusion film is used as the background. This design makes the displayed virtual characters clearer and has a higher visual layering.
[0044] S44 optimizes the speaker structure so that the sound can be more accurately focused on the interaction area between the user and the virtual agent, thereby ensuring the perfect coordination of sound and visual effects and improving the overall interactive experience.
[0045] Preferably, the S5 specifically includes the following steps:
[0046] S51 integrates virtual images, emotional interaction, stereoscopic display technology and intelligent emotion analysis mechanism, and can dynamically select and apply the most appropriate intervention strategy according to the user's psychological needs and emotional changes;
[0047] The S52 system combines users' real-time input and feedback to comprehensively analyze users' emotional states, psychological needs, and specific performance during the interaction process, thereby providing personalized, flexible, and efficient mental health support.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] The present invention significantly improves the naturalness and realism of emotional interaction through multimodal interaction of virtual images, dynamic facial expressions and emotional voice. Through stereoscopic display technology and synchronous adjustment of facial expressions, the virtual agent can present emotional responses that match the user's emotional state in real time, thereby enhancing the user's sense of immersion and participation.
[0050] The present invention analyzes the user's emotional state in real time, intelligently matches and adjusts the mental health support plan, can ensure the personalization of intervention, and can flexibly respond to the user's emotional changes, thereby optimizing the intervention content in real time, significantly improving the effect of psychological intervention, and ensuring efficient and accurate mental health support. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the system structure of the present invention;
[0052] Figure 2 Optimizing the structure diagram for large language models;
[0053] Figure 3 This is a diagram of the process of psychological intervention dialogue for the large language model;
[0054] Figure 4 This is a hardware structure diagram of the psychological intervention chat agent system based on the knowledge-enhanced large language model;
[0055] Figure 5 Schematic diagram of the virtual image of the psychological intervention chat agent based on the knowledge-augmented large language model. DETAILED DESCRIPTION
[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] The present invention proposes a psychological intervention chat agent system based on a knowledge-enhanced large language model, comprising the steps of:
[0058] S1. Optimize the large language model to enable it to handle complex mental health conversations, so that it can accurately understand the user's emotions and needs, generate personalized intervention content, and improve the intelligence level of psychological intervention;
[0059] S2. Build a comprehensive mental health knowledge base, conduct real-time analysis based on user conversation content, automatically identify potential psychological problems, and provide users with accurate and effective intervention strategies through professional information in the knowledge base;
[0060] S3. Design and implement an emotional interaction mechanism, through the comprehensive use of virtual images, dynamic facial expressions and emotional voice, so that the system can present natural emotional responses in the conversation and enhance the user's emotional resonance and interactive experience;
[0061] S4. Introduce advanced stereoscopic display technology to enhance the visual effects and realism of virtual images, thereby enhancing the user's sense of immersion and participation, and improving the realism and attractiveness of the interactive experience;
[0062] S5. Through intelligent process optimization, the system dynamically selects and applies the most appropriate intervention strategy based on the user's psychological needs and input, ensuring personalized and efficient mental health support.
[0063] The S1 optimization of the large language model specifically includes the following steps:
[0064] S11 uses ChatGLM2-6B as the basic large language model. It is based on the Transformer architecture and has powerful text generation and understanding capabilities. It is particularly suitable for processing long texts and capturing long-term dependencies, and can effectively support multi-round conversations.
[0065] The model performs well in handling complex mental health conversations, being able to understand the user’s emotions and needs and generate appropriate conversation content;
[0066] S12 performs LoRA fine-tuning on the ChatGLM2-6B large language model. By introducing a low-dimensional bypass structure based on the original model, the model's performance in the field of mental health can be optimized, which can avoid large-scale parameter adjustments and improve efficiency.
[0067] The additional dimension reduction and dimension increase matrices for model training enable the model to generate more precise and personalized intervention content based on the specific needs of the mental health field;
[0068] By fine-tuning the large language model and optimizing its application in the field of mental health, it can efficiently generate personalized intervention content in specific fields, thereby improving the model's ability to handle complex mental health conversations.
[0069] The step S2 of building a mental health knowledge base and analyzing user conversations in real time specifically includes the following steps:
[0070] S21 builds a mental health knowledge base based on DSM-5 standards, including "disorder description", "client characteristics", "therapist characteristics", "intervention strategies", "prognosis" and "assessment methods";
[0071] This information provides professional support for the system when conducting mental health conversations, helping the system identify potential psychological problems of users and provide precise intervention strategies;
[0072] S22 analyzes the user's conversation content in real time and uses the TextRank algorithm to extract the text summary of the speech;
[0073] After S23 obtains the abstract, keywords are extracted through TF-IDF;
[0074] These keywords will be encoded and vectorized for word frequency, and then analyzed for text similarity with the "disorder description" in the knowledge base, and the cosine similarity will be calculated, that is, the cosine value of the angle between two vectors in a vector space is used as a measure of the difference between two individuals. The closer the cosine value is to 1 and the closer the angle is to 0, the more likely the visitor's speech is to match the disorder.
[0075] After completing the initial matching at S24, the same method is used, with the “characteristics of the preferred therapist” and “assessment” of the identified disorder as the precondition, to continue asking questions to continuously obtain the client’s speech, which is then used to match the “client characteristics” of the identified disorder;
[0076] When the barrier match exceeds the threshold, the “intervention strategy” is used as the new preset to complete the rest of the conversation;
[0077] S25 Given that the traditional input structure of large language models is usually based on a prompt mechanism, the input structure of the model is optimized;
[0078] By effectively embedding structured psychological knowledge into the model's input process, the system can infer the required professional psychological knowledge based on the context of the current conversation and accurately apply this knowledge when generating responses, thereby providing intervention plans that are highly matched to the user's psychological state.
[0079] The S3 design and implementation of the emotional interaction mechanism specifically includes the following steps:
[0080] S31 uses cartoon-style avatar design to optimize the emotional connection between the virtual agent and the user;
[0081] Cartoon-style design will reduce users’ expectations of the authenticity of virtual agents’ appearance, making them pay more attention to their interactions and emotional feedback with virtual agents, thus improving the naturalness and credibility of emotional interactions.
[0082] S32 uses Live2D technology to dynamically control the virtual image, giving it a near-3D visual effect. It can also adjust its expression and movements in real time according to the user's input during the interaction, further enhancing the realism and affinity of the interaction.
[0083] S33 performs sentiment annotation on the user's conversation text;
[0084] By adopting the sentiment annotation model, the system extracts the sentiment information in the conversation text and converts it into numerical representations of multiple sentiment dimensions;
[0085] These emotion dimension values are then mapped into facial expression activity units (AUs) to drive the facial expression changes of the virtual agent;
[0086] S34 uses the VITS2 speech synthesis model, which enables the speech output to be adjusted in real time according to the conversation content and emotional state;
[0087] Adjust the pitch, speed, and tone of speech through emotional annotation to ensure that the speech expression matches the user's emotional changes;
[0088] S35 uses Whisper voice recognition technology to recognize user voice input in real time to ensure smooth voice interaction;
[0089] Based on the facial expression generation, S36 uses forced alignment technology to synchronize the speech and facial expressions at the corresponding moments to ensure that the lip shape of the speech is consistent with the facial expression of the virtual agent;
[0090] Ensure the visual and auditory consistency of the virtual agent’s emotional expression, thereby improving the realism of the interaction and the accuracy of the emotional expression;
[0091] Through the synchronous adjustment of facial expressions and voice, the system can provide a more humane form of interaction, further enhance the user's sense of participation and the fluency of interaction, and provide more effective support for mental health intervention.
[0092] The S4 stereoscopic display design specifically includes the following steps:
[0093] The S41 uses pseudo-holographic display technology to display the chat agent avatar;
[0094] Through this display method, the virtual image not only has a higher visual realism, but also can bring a three-dimensional visual effect during the interaction process, improving the visual effect limitations of traditional video consultation;
[0095] S42 designed and built the stereoscopic display, using black resin 3D printing material as the shell to ensure the display's exterior structure is sturdy and modern;
[0096] The internal integrated Rockchip RK3588 processor is used as the main control, through which efficient image processing and rendering are achieved;
[0097] The S43 video signal is transmitted to the semi-transparent screen via the HDMI interface for display, thus generating a pseudo-holographic effect;
[0098] A backlight system is added to the display design, and a semi-white transparent diffusion film is used as the background. This design makes the displayed virtual characters clearer and has a higher visual layering.
[0099] S44 optimizes the speaker structure so that the sound can be more accurately focused on the interaction area between the user and the virtual agent, thereby ensuring the perfect coordination of sound and visual effects and improving the overall interactive experience.
[0100] The S5 system process optimization specifically includes the following steps:
[0101] S51 integrates virtual images, emotional interaction, stereoscopic display technology and intelligent emotion analysis mechanism, and can dynamically select and apply the most appropriate intervention strategy according to the user's psychological needs and emotional changes;
[0102] The S52 system combines users' real-time input and feedback to comprehensively analyze users' emotional states, psychological needs, and specific performance during the interaction process, thereby providing personalized, flexible, and efficient mental health support.
[0103] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A psychological intervention chat agent system based on knowledge-enhanced large language model, characterized by , including the following steps: S1. Optimize the large language model to enable it to handle complex mental health conversations, so that it can accurately understand the user's emotions and needs, generate personalized intervention content, and improve the intelligence level of psychological intervention; S2. Build a comprehensive mental health knowledge base, conduct real-time analysis based on user conversation content, automatically identify potential psychological problems, and provide users with accurate and effective intervention strategies through professional information in the knowledge base; S3. Design and implement an emotional interaction mechanism, through the comprehensive use of virtual images, dynamic facial expressions and emotional voice, so that the system can present natural emotional responses in the conversation and enhance the user's emotional resonance and interactive experience; S4. Introduce advanced stereoscopic display technology to enhance the visual effects and realism of virtual images, thereby enhancing the user's sense of immersion and participation, and improving the realism and attractiveness of the interactive experience; S5. Through intelligent process optimization, the system dynamically selects and applies the most appropriate intervention strategy based on the user's psychological needs and input, ensuring personalized and efficient mental health support.
2. The optimized large language model according to claim 1, characterized in that: (2-1) ChatGLM2-6B is selected as the basic large language model. It is based on the Transformer architecture and has powerful text generation and understanding capabilities. It is particularly suitable for processing long texts and capturing long-term dependencies, and can effectively support multi-round conversations. The model performs well in handling complex mental health conversations, being able to understand the user’s emotions and needs and generate appropriate conversation content; (2-2) LoRA fine-tuning of the ChatGLM2-6B large language model was performed. By introducing a low-dimensional bypass structure based on the original model, the model's performance in the field of mental health was optimized, which can avoid large-scale parameter adjustments and improve efficiency; The additional dimension reduction and dimension increase matrices for model training enable the model to generate more precise and personalized intervention content based on the specific needs of the mental health field; By fine-tuning the large language model and optimizing its application in the field of mental health, it can efficiently generate personalized intervention content in specific fields, thereby improving the model's ability to handle complex mental health conversations.
3. The method of constructing a mental health knowledge base and analyzing user dialogue in real time according to claim 1, characterized in that: (3-1) Construct a mental health knowledge base based on DSM-5 criteria, including "disorder description", "client characteristics", "therapist characteristics", "intervention strategies", "prognosis" and "assessment methods"; This information provides professional support for the system when conducting mental health conversations, helping the system identify potential psychological problems of users and provide precise intervention strategies; (3-2) Analyze the user's conversation content in real time and use the TextRank algorithm to extract the text summary of the speech; (3-3) After obtaining the abstract, extract keywords through TF-IDF; These keywords will be encoded and vectorized for word frequency, and then analyzed for text similarity with the "disorder description" in the knowledge base. The cosine similarity is calculated, that is, the cosine value of the angle between two vectors in a vector space is used as a measure of the difference between two individuals. The closer the cosine value is to 1 and the closer the angle is to 0, the more likely the visitor's speech is to match the disorder. (3-4) After completing the initial match, continue to ask questions using the same method, using the "characteristics of the preferred therapist" and "assessment" of the identified disorder as the default, to continuously obtain the client's speech, which is then matched with the "client characteristics" of the identified disorder; When the barrier matching degree exceeds the threshold, the "intervention strategy" is used as the new preset to complete the remaining conversation; (3-5) Given that the traditional input structure of large language models is usually based on a prompt mechanism, the input structure of the model is optimized; By effectively embedding structured psychological knowledge into the model's input process, the system can infer the required professional psychological knowledge based on the context of the current conversation and accurately apply this knowledge when generating responses, thereby providing intervention plans that are highly matched to the user's psychological state.
4. The design and implementation of the emotional interaction mechanism according to claim 1, characterized in that: (4-1) Use cartoon-style avatar design to optimize the emotional connection between the virtual agent and the user; Cartoon-style design will reduce users’ expectations of the authenticity of virtual agents’ appearance, making them pay more attention to their interactions and emotional feedback with virtual agents, thus improving the naturalness and credibility of emotional interactions. (4-2) Live2D technology is used to dynamically control the virtual image, giving it a visual effect close to 3D. It can also adjust its expression and movements in real time according to the user's input during the interaction, further enhancing the realism and affinity of the interaction. (4-3) Emotional annotation of user’s conversation text; By adopting the sentiment annotation model, the system extracts the sentiment information in the conversation text and converts it into numerical representations of multiple sentiment dimensions; These emotion dimension values are then mapped into facial expression activity units (AUs) to drive the facial expression changes of the virtual agent; (4-4) Using the VITS2 speech synthesis model, the speech output can be adjusted in real time according to the conversation content and emotional state; Adjust the pitch, speed, and tone of speech through emotional annotation to ensure that the speech expression matches the user's emotional changes; (4-5) Use Whisper voice recognition technology to recognize user voice input in real time to ensure smooth voice interaction; (4-6) Based on the facial expression generation, the forced alignment technology is used to synchronize the speech and facial expression at the corresponding moment to ensure that the lip shape of the speech is consistent with the facial expression of the virtual agent; Ensure the visual and auditory consistency of the virtual agent’s emotional expression, thereby improving the realism of the interaction and the accuracy of the emotional expression; Through the synchronous adjustment of facial expressions and voice, the system can provide a more humane form of interaction, further enhance the user's sense of participation and the fluency of interaction, and provide more effective support for mental health intervention.
5. The three-dimensional display design according to claim 1, characterized in that: (5-1) Using pseudo-holographic display technology to display the chat agent virtual image; Through this display method, the virtual image not only has a higher visual realism, but also can bring a three-dimensional visual effect during the interaction process, improving the visual effect limitations of traditional video consultation; (5-2) Design and construct a three-dimensional display, using black resin 3D printing material as the shell to ensure that the external structure of the display is sturdy and modern; The internal integrated Rockchip RK3588 processor is used as the main control, through which efficient image processing and rendering are achieved; (5-3) The video signal is transmitted to the semi-transparent screen via the HDMI interface for display, thereby generating a pseudo-holographic effect; A backlight system is added to the display design, and a semi-white transparent diffusion film is used as the background. This design makes the displayed virtual characters clearer and has a higher visual layering. (5-4) Optimize the speaker structure so that the sound can be more accurately focused on the interaction area between the user and the virtual agent, thereby ensuring the perfect coordination of sound and visual effects and improving the overall interactive experience.
6. The system process optimization according to claim 1, characterized in that: (6-1) Integrate virtual images, emotional interaction, stereoscopic display technology and intelligent emotion analysis mechanism to dynamically select and apply the most appropriate intervention strategy based on the user's psychological needs and emotional changes; (6-2) The system combines the user's real-time input and feedback to comprehensively analyze the user's emotional state, psychological needs, and specific performance during the interaction process, thereby providing personalized, flexible, and efficient mental health support.