Artificial intelligence psychotherapy equipment
AI-powered psychotherapy devices that integrate speech recognition and multimodal emotion recognition technologies address the problems of scarce resources and insufficient accuracy in emotion assessment in traditional psychotherapy. They enable personalized emotion soothing and cognitive correction, provide continuous conversational psychotherapy effects, and help users alleviate emotions and adjust their mindset.
Patent Information
- Application Number
- CN202511733736.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional psychotherapy relies on professional psychological counselors, which suffers from resource scarcity, geographical limitations, high consultation costs, and difficulties in protecting user privacy. Furthermore, traditional human-computer dialogue models have limitations in emotion perception and intervention adaptability, and rely on a single text modality for emotion recognition, resulting in insufficient accuracy in emotion judgment.
The device employs artificial intelligence-based psychotherapy equipment, integrating a speech recognition module, a semantic understanding and dialogue generation module, and a TTS speech synthesis module. By combining speech and text information and utilizing emotion recognition and dialogue strategy models, it achieves multimodal emotion recognition and personalized dialogue strategy generation, simulating the guidance style of a psychotherapist to provide emotional soothing, cognitive restructuring, and behavioral suggestions.
It achieves accurate judgment of emotional state and personalized response, provides continuous, gentle and psychologically therapeutic dialogue output, helps users relieve emotions, adjust mindset, and improve self-regulation ability, and has a good user immersion experience and intelligent companionship capabilities.
Smart Images

Figure CN121583244A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an artificial intelligence psychological treatment device. BACKGROUND
[0002] With the increase of social pressure and the growing demand for mental health, traditional psychological treatment relies on professional psychological consultants, which has problems such as resource scarcity, geographical restrictions, high consultation costs, and difficulty in protecting user privacy. The dialogue-based psychological support system based on artificial intelligence has become a research hotspot.
[0003] The traditional man-machine dialogue model has obvious limitations in emotion perception and intervention adaptability, relies on single text mode for emotion recognition, and ignores the acoustic emotion features in the voice, resulting in insufficient emotion judgment accuracy. SUMMARY
[0004] Therefore, it is necessary to provide an artificial intelligence psychological treatment device to achieve the purpose of relieving user emotions through man-machine dialogue.
[0005] To achieve the above purpose, the present application provides an artificial intelligence psychological treatment device, comprising: a speech recognition module, a semantic understanding and dialogue generation module, and a TTS speech synthesis module; The speech recognition module is configured to receive voice information of a user and convert the voice information into text information; The semantic understanding and dialogue generation module is configured to identify the voice information and the text information, determine the emotional state of the user, and generate text reply information for relieving the user's emotions based on the emotional state; The TTS speech synthesis module is configured to convert the text reply information into voice reply information and output the voice reply information.
[0006] In a possible implementation manner, the semantic understanding and dialogue generation module comprises an emotion recognition model; The emotion recognition model is configured to identify the emotional state of the user based on the voice information, the text information, and historical dialogue information.
[0007] In a possible implementation manner, the semantic understanding and dialogue generation module further comprises a dialogue strategy model; The dialogue strategy model comprises an input encoding layer, a strategy decision layer, and an output generation layer; The input encoding layer is configured to fuse the voice information, the text information, and historical dialogue information to obtain a fusion vector; The strategy decision layer is configured to select a corresponding conversation response based on the fusion vector and the emotional state. The output generation layer is configured to generate the text reply information based on the conversation response.
[0008] In a possible implementation, the input encoding layer includes a speech branch, a text branch, a context encoder, and a dynamic fusion module. The speech branch is configured to extract acoustic features from the speech information and generate an acoustic emotion vector. The text branch is configured to extract semantic features from the text information and generate a semantic emotion vector. The context encoder is configured to model an emotion sequence of the historical dialogue information and generate an emotion trend vector. The dynamic fusion module is configured to fuse the acoustic emotion vector, the semantic emotion vector, and the emotion trend vector to obtain a fusion vector.
[0009] In a possible implementation, the policy decision layer includes a policy selection network and a policy-action mapping module. The policy selection network is configured to predict a probability distribution of a dialogue policy corresponding to the fusion vector based on the fusion vector and the emotion state. The policy-action mapping module is configured to retrieve a corresponding conversation response from a preset template library based on the probability distribution.
[0010] In a possible implementation, the dialogue policy model is obtained by the following manner: Obtain a psychological counseling dialogue dataset of different users, and label an emotion category label and a dialogue policy label corresponding to each dialogue; Train an LMM model by using the dialogue dataset, the emotion category label, and the dialogue policy label to obtain the dialogue policy model.
[0011] In a possible implementation, the dialogue policy label includes: calming anxiety, guiding emotion awareness, suggesting behavior activation, relieving sadness, or channeling anger.
[0012] In a possible implementation, the dialogue policy model is constructed based on a counselor experience, a CBT, and an MBCT intervention framework.
[0013] In a possible implementation, further including a voiceprint recognition module. The voiceprint recognition module is configured to confirm identity information of a user based on the speech information.
[0014] In a possible implementation, the voiceprint recognition module, the speech recognition module, the semantic understanding and dialogue generation module and the TTS speech synthesis module are integrated on a master module. The master module comprises a Cortex-A7 embedded AI processing chip and an NPU coprocessor.
[0015] The artificial intelligence psychological treatment device provided by the present application has the advantages that the voice information of the user is received by the voice recognition module, and the voice information is converted into text information, so that the semantic understanding and dialogue generation module can recognize the voice information and the text information, accurately judge the emotional state of the user, and further generate text reply information according to the emotional state of the user, wherein the text reply information covers different levels of support strategies such as emotional pacification, cognitive reconstruction and behavior suggestion, and can be used to relieve the user's emotion, the TTS speech synthesis module converts the text reply information into natural speech reply information with emotional expression, realizes continuous, gentle and psychological healing effect dialogue output, and achieves the purpose of assisting the user to relieve emotion and adjust mentality. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 One of the embodiment structure schematic diagrams of the artificial intelligence psychological treatment device provided by the present application; Figure 2 One of the embodiment structure schematic diagrams of the artificial intelligence psychological treatment device provided by the present application; Figure 3 One of the embodiment structure schematic diagrams of the artificial intelligence psychological treatment device provided by the present application; DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] In the description of the embodiments of the application, unless otherwise specified, the meaning of "a plurality of" is two or more. The association relationship of the associated objects described by "and / or" indicates that there can be three relationships, for example: A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone.
[0020] The descriptions of "first", "second", and the like in the embodiments of the application are only for descriptive purposes and cannot be understood as indicating or implying relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the technical features defined with "first", "second" can explicitly or implicitly include at least one of the features.
[0021] In this document, the reference to "embodiments" means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0022] The application provides an artificial intelligence psychological treatment device, which is described below.
[0023] Figure 1 An embodiment structure schematic diagram of the artificial intelligence psychological treatment device provided by the application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the artificial intelligence psychological treatment device 100 comprises: a speech recognition module 101, a semantic understanding and dialogue generation module 102, and a TTS speech synthesis module 103. The speech recognition module 101 is configured to receive voice information of a user and convert the voice information into text information. The semantic understanding and dialogue generation module 102 is configured to identify the voice information and the text information, determine an emotional state of the user, and generate text reply information for relieving the emotional state of the user based on the emotional state. The TTS speech synthesis module 103 is configured to convert the text reply information into voice reply information and output the voice reply information.
[0024] The artificial intelligence psychological treatment device is started by voice wake-up, user voice input is collected by a microphone, the speech recognition module is configured to receive voice information of a user and convert the voice information into text information. After identification and identity confirmation by the speech recognition module (ESP-SR) and the voiceprint recognition module (SenseVoice), the information is sent to the semantic understanding and dialogue generation module based on a large language model (such as Qwen or DeepSeek).
[0025] The semantic understanding and dialogue generation module is used for recognizing voice information and text information, determining the emotional state of the user, and generating text reply information for relieving the emotion of the user based on the emotional state.
[0026] The voice information input by the user is first subjected to semantic analysis, and emotional features, psychological intentions and potential distress information are extracted therefrom. Based on the context of the historical dialogue of multiple rounds, the text reply information with psychological intervention features is generated by combining the built-in psychological health knowledge base and the emotion recognition model. The text reply information covers different levels of support strategies such as emotional pacification, cognitive reconstruction and behavior suggestion, and can be used to relieve the emotion of the user.
[0027] The dialogue system is composed of an input analysis module, a dialogue management module, a psychological intervention generation module and a voice output module, and can realize continuous tracking and dynamic response to the emotional state of the user.
[0028] The TTS voice synthesis module (Volcano Engine or CosyVoice) is used for converting the text reply information into voice reply information and outputting the voice reply information.
[0029] Finally, the text reply information is converted into natural voice reply information with emotional expression by the TTS voice synthesis module, realizing continuous, gentle and psychological healing effect dialogue output. The final output of the device not only includes emotional language response, but also generates operable behavior suggestions and prompts the user to seek professional help when necessary.
[0030] In summary, the artificial intelligence psychological treatment device provided by the embodiment of the present application receives the voice information of the user through the voice recognition module, and converts the voice information into text information, so that the semantic understanding and dialogue generation module can recognize the voice information and the text information, accurately determine the emotional state of the user, and further generate text reply information according to the emotional state of the user. The text reply information covers different levels of support strategies such as emotional pacification, cognitive reconstruction and behavior suggestion, and can be used to relieve the emotion of the user. The TTS voice synthesis module converts the text reply information into natural voice reply information with emotional expression, realizes continuous, gentle and psychological healing effect dialogue output, and realizes the purpose of assisting the user to relieve emotion and adjust mentality.
[0031] In some embodiments of the present application, the semantic understanding and dialogue generation module comprises an emotion recognition model; The emotion recognition model is used for recognizing the emotional state of the user based on the voice information, the text information and the historical dialogue information.
[0032] Figure 2 An embodiment structure schematic view two of the artificial intelligence psychological treatment device provided by the present application is as follows:Figure 2 As shown, the semantic understanding and dialogue generation module includes an emotion recognition model and a dialogue strategy model.
[0033] The judgment of the user emotional state is implemented through a multi-modal emotion recognition model, including the following dimensions: Voice feature analysis extracts the pitch, energy, speech rate, intonation, and other acoustic features of the user's voice, and combines a deep convolutional neural network (CNN) or a Transformer structure for emotion classification, such as identifying "sadness, anxiety, anger, and helplessness" emotional states.
[0034] Text semantic emotion recognition converts the user's voice into text and sends it to a large language model for emotion label recognition, combining keyword analysis and emotion tendency reasoning to determine the current psychological state.
[0035] Contextual emotion trajectory tracking combines the user's historical interaction emotion records to identify and continuously assess the emotional fluctuation trend across rounds of dialogue.
[0036] In some embodiments of the present application, the semantic understanding and dialogue generation module further includes a dialogue strategy model; The dialogue strategy model includes an input encoding layer, a strategy decision layer, and an output generation layer; The input encoding layer is used to fuse the voice information, text information, and historical dialogue information to obtain a fusion vector; The strategy decision layer is used to select a corresponding conversation response based on the fusion vector and the emotional state; The output generation layer is used to generate the text reply information based on the conversation response.
[0037] In some embodiments of the present application, the input encoding layer includes a voice branch, a text branch, a context encoder, and a dynamic fusion module; The voice branch is used to extract acoustic features from the voice information and generate an acoustic emotion vector; The text branch is used to extract semantic features from the text information and generate a semantic emotion vector; The context encoder is used to model the emotional sequence of the historical dialogue information to generate an emotional trend vector; The dynamic fusion module is used to fuse the acoustic emotion vector, the semantic emotion vector, and the emotional trend vector to obtain a fusion vector.
[0038] In some embodiments of the present application, the strategy decision layer includes a strategy selection network and a strategy-action mapping module: a policy selection network configured to predict a probability distribution of a dialogue policy corresponding to the fusion vector based on the fusion vector and the emotion state; a policy-action mapping module configured to retrieve a corresponding dialogue response from a preset template library based on the probability distribution.
[0039] The dialogue policy model structure mainly includes an input encoding layer, a policy decision layer, and an output generation layer.
[0040] I. Input encoding layer: (1) Speech branch: Local features of the spectrogram are extracted using a convolutional neural network, and time series patterns are captured through a bidirectional GRU or Conformer to generate twenty-dimensional acoustic emotion vectors containing fundamental frequency, energy, speech rate, pitch variation rate, MFCC, etc.
[0041] (2) Text branch: The system uses BERT or RoBERTa as a text encoder and designs a "triple fusion enhancement mechanism" as follows: First, cross-layer fusion (combining bottom, middle, and high layer outputs) is used to generate multi-layer information vectors rich in emotional and entity semantics; Then, a keyword and entity enhancement mechanism is introduced to automatically identify dimensions such as "anxiety" and "support" and give them additional weight to enhance emotion recognition and entity expression capabilities; In addition, a "context alignment loss" is added to guide the generation of sentence vectors to maintain consistency in emotions and semantics in multi-round continuous dialogue, thereby achieving more natural and coherent text emotion representation.
[0042] The final generated sentence vector (i.e., semantic emotion vector) not only contains keywords, emotional tendencies, and entity information but also has more rich multi-granularity semantic expression and dialogue continuity.
[0043] (3) Context encoder: Use multi-layer Transformer to aggregate and model the emotion sequence of the last ten rounds of dialogue, extract the trend of cross-round emotion fluctuations, and obtain an emotion trend vector.
[0044] (4) Dynamic fusion module: This module goes beyond the traditional BiLSTM + Attention fusion by improving fusion capabilities and structural innovation in the following ways: First, use a multi-level fusion strategy: use a fusion gate mechanism to fuse speech, text, and context features at the feature level; embed a cross-modal Transformer at the model level to explicitly model the dependency and collaboration between modalities through Cross-Modal Multi-Head Attention.
[0045] Secondly, the fused representation is further fused through a learnable graph structure fusion layer (Graph Fusion Layer) to use a graph neural network (GNN) to construct nodes and edges for "utterances at different times", so as to enhance the global structure modeling capability.
[0046] Finally, in order to prevent overfitting and improve expression stability, a hierarchical pooling strategy (such as Attention-based pooling + Layer-wise gating) is used instead of a simple global average pooling to obtain a final fusion vector, which provides a more fine and robust fusion representation for the final emotion classification.
[0047] II. Strategy decision layer: (1) Strategy selection network: the present application proposes a novel architecture of strategy selection network, which relies on the fused multi-modal representation (i.e. fusion vector) to predict the strategy probability distribution through the Transformer decoder, at the same time, generates the emotion prior vector by analyzing the user's short-term fluctuations and long-term trends in the dialogue, and injects it into the main decoder through the guide branch to dynamically adjust the strategy output; the network also integrates a modal preference guide mechanism to adjust the weighting bias of different modal information such as text, speech and vision in the multi-head attention according to the current situation, so as to realize efficient focusing on key emotional signals; in the training stage, the present application further adopts a strategy-emotion joint reinforcement learning optimization mechanism to use the user's emotional improvement degree as a reward signal to feedback the strategy selection, so that the network has the ability to continuously guide the positive development of user's emotion, and significantly improves the emotional sensitivity, consistency and long-term adaptability of strategy prediction.
[0048] (2) Strategy-action mapping module: not only retrieves the corresponding dialogue response from the template library according to the strategy label, but also fuses the emotion prior vector and modal preference parameters generated by the strategy selection network to guide the conditional generation model to dynamically adjust in the content expression and language style level, so that the output is more consistent with the user's current context in terms of emotional expression, modal adaptability and personalization.
[0049] At the same time, this module supports the linkage optimization between the generated result and the user's emotional feedback - the generated content not only meets the strategy constraints, but also has the adaptive ability of emotional guidance, and can incorporate the user's feedback results into the closed-loop mechanism for optimizing the integrated performance of subsequent strategy selection and generation mapping.
[0050] (3) Safety control module: adopt multi-stage risk gating and adaptive error correction mechanism, before generation, the system performs multi-task risk discrimination based on the emotion prior and modal preference parameters output by the policy network and dynamically sets the safety threshold; during generation, the "logit mask" is constructed through safety rules to directly affect the model decoding path, and the model self-evaluation-self-correction process is enabled to rewrite potential high-risk content instead of simply rejecting; after generation, an independent safety scorer scores the harmfulness of the text and voice content and compares the multi-modal consistency, once the threshold is exceeded, automatic interception, manual takeover or risk referral are triggered, and traceable logs are recorded to feedback the policy-emotion joint reinforcement learning, so as to gradually optimize the system risk identification ability while ensuring real-time safety, and ensure the high-risk recall rate.
[0051] III. Output generation layer: (1) Controllable text generation: introduce policy constraints during generation to ensure that the output content matches the policy label, and the style, tone and safety meet the preset requirements.
[0052] (2) Multi-modal synchronization: when voice output is required, call the TTS system and adjust the tone, speed and tone according to the policy to make the voice performance consistent with the text emotion, TTS output delay ≤ 200 milliseconds, emotion matching degree ≥ 95%.
[0053] (3) Response feedback: after completing a round of generation, update the user input and emotional state to the dialogue history to provide context for the next round of policy decision-making, and realize closed-loop optimization.
[0054] In some embodiments of the present application, the dialogue policy model is obtained by training in the following manner: Obtain a psychological counseling dialogue data set of different users, and label the emotion category label and dialogue policy label corresponding to each dialogue; Use the dialogue data set, the emotion category label and the dialogue policy label to train the LMM model to obtain the dialogue policy model.
[0055] In a possible implementation manner, the dialogue policy label includes: calm anxiety, guide emotional awareness, suggest behavior activation, relieve sadness or channel anger.
[0056] The training process of the dialogue policy model is as follows: I. Data preparation: (1) Multi-source data integration: Select and construct high-quality dialogue datasets covering multi-modal features, including public English datasets (DailyDialog, EmoryNLP, MELD, IEMOCAP) and company-built and purchased Chinese customer service, psychological support, medical consultation and other multi-turn dialogue data. The total size is about 500,000 rounds, covering more than 30 business scenarios and 20 types of emotion labels. All data are desensitized and retain text and voice information.
[0057] (2) Emotion-strategy pairing annotation: Double-blind two-round manual annotation with Cohen's κ coefficient not less than 0.85, annotating emotion category labels and corresponding strategy labels for each user input, and recording evidence sentences, emotion intensity classification and fluctuation direction. Dialogue strategy labels can include: calming anxiety, guiding emotion awareness, suggesting behavior activation, relieving sadness or soothing anger.
[0058] (3) Simulation data generation: Based on Markov state transition model or Transformer user simulator, generate long dialogue samples covering more than 80% of common emotion transition paths, each round length 5-15 rounds, total generation not less than 200,000 rounds, rare emotion transition ratio not less than 10%.
[0059] (4) Data augmentation: Add background noise with signal-to-noise ratio of 10-30 decibels to the voice sample, adjust the speech speed by ±20%, perform time stretching and spectrum shielding, perform synonym replacement (Top-5 similarity ≥0.8) on the text sample, Chinese-English translation (BLEU ≥0.75) and homonym disturbance (≤2% character modification), and balance the class distribution through class reweighting and minority class oversampling.
[0060] II. Algorithm training: (1) Supervised pre-training: Use BERT base or RoBERTa wwm ext as a text encoder, CNN+BiGRU or Conformer to extract acoustic features, and model the historical ten rounds of dialogue context through multi-layer Transformer.
[0061] On this basis, in order to significantly improve the performance and explanation ability of multi-modal emotion recognition in continuous dialogue, the present application introduces cross-modal contrastive learning (Cross-Modal Contrastive Learning), which aligns acoustic and text emotion representations through InfoNCE loss, and strengthens modality consistency and fusion robustness.
[0062] Further, an EmotionPropagation Graph is constructed with each round of dialogue as a node and modal and round dependence as an edge. Structured dialogue information is extracted based on graph neural networks, and the representation discrimination ability is improved by combining graph contrastive learning, which is fundamentally different from the traditional BiLSTM+Attention architecture.
[0063] Meanwhile, discrete emotion labels are mapped to continuous Valence-Arousal (VA) emotion space, and a VA distance-based contrastive loss is introduced to improve the semantic interpretability of emotion classification. Trend perception contrastive loss (Trend Contrastive Loss) is used to enhance the model's ability to perceive consistency between the current round and historical emotion trends.
[0064] The above mechanisms significantly improve the depth of multi-modal fusion, dialogue continuity understanding, and emotion semantic interpretation. Even with joint optimization, the strategy classification (Top1 accuracy ≥ 85%, Macro F1 ≥ 0.82) and emotion classification (Macro F1 ≥ 0.88) are improved.
[0065] (2) Strategy generation: When the strategy prediction confidence is ≥ 0.85, directly retrieve the template library, otherwise input the strategy label as a condition into the generative language model (such as T5-large) to generate text. The generated results need to meet BLEU ≥ 0.8 and emotion consistency ≥ 0.9, and through a safety dictionary, the sensitive word coverage rate is ≥ 99%, and the illegal interception rate is ≥ 98%.
[0066] Three, reinforcement learning optimization: (1) Environment construction: not only use user simulator for initial strategy training, but also real-time access to real dialogue scenarios to obtain emotion recognition feedback. The fusion of simulator and real feedback significantly improves the training generalization and reality adaptability.
[0067] (2) State definition: Based on the original twenty-dimensional emotion feature vector and seven hundred sixty-eight-dimensional context vector, an emotion trend vector is added to capture recent emotion evolution dynamics and enhance the model's perception of emotion continuity.
[0068] (3) Action space design: retain fifteen sub-strategies, and refine them into different intervention intensity levels (such as "mild comfort" and "deep guidance"), making strategy selection more accurate and flexible.
[0069] (4) Reward mechanism: adopt a multi-dimensional evaluation system, including short-term emotion improvement rate (average improvement ≥ 15%), emotion stability reward (trend consistency), user trust index (predicted score), task completion rate (≥ 90%), user satisfaction (≥ 4.5 / 5), and negative punishment for illegal responses (≤ 0.05%), achieving a balance between short-term effectiveness and long-term trust.
[0070] (5) Strategy update and training stability: PPO (clip=0.2, learning rate=1e-5, gamma=0.99) algorithm is adopted; on this basis, a "trend consistency loss" (encouraging consistent strategies in similar emotional states) is introduced, and the offline trajectory and online sampling experience are fused at a ratio of 1:1 to improve training efficiency and strategy stability.
[0071] IV. Deployment and optimization: (1) Model compression: Through knowledge distillation, quantization and structural pruning, the full multi-modal model is compressed into a lightweight version, with a parameter size reduction of more than 60% and a performance decrease of no more than 2%.
[0072] (2) Inference acceleration: Combined with INT8 quantization and no more than 30% structural pruning, the inference delay is reduced to 40% of the original.
[0073] (3) PC deployment: In a high-performance PC (such as Intel i9 / 32G RAM / RTX 3080) environment, the single-round response delay is controlled within 100 milliseconds, the continuous running stability is ≥99.9%, the violation content recall rate is ≥98%, and the false negative rate is ≤2%.
[0074] The artificial intelligence psychological treatment equipment provided in the embodiments of the present application simulates the question and answer logic and dialogue process of a psychologist through voice recognition, semantic understanding and emotion analysis technology, assists users in relieving emotions, adjusting mentality and improving self-regulation ability, and can be used as an auxiliary tool for psychological treatment.
[0075] In some embodiments of the present application, the dialogue strategy model is constructed based on the experience of psychological consultants, CBT and MBCT intervention framework.
[0076] Exemplarily, the dialogue strategy model is obtained in the following manner: (1) Expert knowledge rule construction: based on the experience of psychological consultants and mainstream intervention frameworks such as CBT (cognitive behavioral therapy) and MBCT (mindfulness cognitive therapy), the response logic and dialogue path tree in typical user emotional states are designed.
[0077] (2) Data-driven learning optimization: real psychological counseling dialogue data (desensitized) are used to train large language models, and through reinforcement learning (such as PPO) or human feedback fine-tuning (RLHF), diversified response strategies suitable for different emotional states, personality preferences, risk levels are constructed.
[0078] (3) Multi-round interaction strategy integration: The system introduces a memory module and user profile to dynamically adjust the dialogue rhythm and intervention method, so as to achieve more humane long-term guidance.
[0079] In some embodiments of the present invention, a voiceprint recognition module is also included; The voiceprint recognition module is used to confirm the user's identity information based on the voice information.
[0080] In some embodiments of the present invention, the voiceprint recognition module, the speech recognition module, the semantic understanding and dialogue generation module, and the TTS speech synthesis module are integrated on the main control module; The main control module includes a Cortex-A7 embedded AI processing chip and an NPU coprocessor.
[0081] The artificial intelligence psychotherapy device provided by this invention is an active electronic medical auxiliary device that uses a Cortex-A7+NPU chip and integrates core modules such as speech recognition, emotion recognition, language understanding and language generation.
[0082] The artificial intelligence psychotherapy device provided by this invention mainly consists of a speech recognition module (ESP-SR), a voiceprint recognition module (SenseVoice), a speech synthesis module (large-scale TTS), a semantic understanding and dialogue generation module (LLM, such as Qwen / DeepSeek), a main control chip (Cortex-A7+NPU), a 3D speaker, a microphone, a battery management module, an OLED or LCD display screen, a voice wake-up and interactive control module, a Wi-Fi / 4G communication module, a power adapter, and a casing.
[0083] The device is activated via voice activation. A microphone captures the user's voice input, which is then processed by a voice recognition module and a voiceprint recognition module for identification and identity verification. The input is then sent to a psychological dialogue system built on an LLM (Large Language Model, such as Qwen / DeepSeek) platform for semantic analysis and generation of psychologically interventional responses. Finally, a large-model TTS (Text-to-Speech) module (such as Volcano Engine or CosyVoice) generates the voice output, enabling continuous conversational psychological therapy intervention. The device supports emotion recognition, long-term dialogue memory, and personalized voiceprint customization, providing a superior user immersion experience and intelligent companionship capabilities.
[0084] The device works by using an AI intervention mechanism that combines CBT (Cognitive Behavioral Therapy) with emotion recognition. Through non-contact human-computer interaction, it simulates the guidance of a psychological counselor to achieve emotional relief and cognitive correction, and helps alleviate psychological symptoms such as depression, anxiety, obsessive-compulsive disorder, and autism.
[0085] The outer casing is made of high-strength, flame-retardant ABS engineering plastic, compliant with RoHS standards. Internal structural components are injection molded, and electronic components are assembled using industrial-grade welding integration methods. Key materials include PCB boards, silicone microphone diaphragms, electromagnetic shielding materials, and high-fidelity speaker chips. Overall, it complies with environmental and electromagnetic compatibility standards for electronic medical auxiliary products.
[0086] Figure 3 A schematic diagram of an embodiment of human-computer interaction provided by the present invention is shown below. Figure 3 As shown, place the device within 1 meter of the user and it will enter standby mode after being powered on. The user can activate the device by voice, such as "Hello, Heartfelt Words". The device will then begin a continuous dialogue with the user (including emotion recognition and psycholinguistic guidance). It is recommended to use the device for 15-30 minutes each time, 1-3 times a day. The user can interrupt or exit the conversation by pressing a button or using voice. All dialogue information is processed locally by default and is not uploaded to the cloud.
[0087] For example, a user inputs via voice: "I've been feeling very depressed lately, I can't sleep well at night, and work is making me feel suffocated." The system first identifies that the user is currently in a state of "anxiety + fatigue" and detects typical psychological characteristics of "stress overload".
[0088] Subsequently, the large language model generated a response based on the psychological intervention framework: "It sounds like you're experiencing a lot of stress, which is understandable. Try doing 5 minutes of deep breathing exercises before bed tonight to help you gradually unwind from the tension." This text content was converted into speech output by the TTS module, with a gentle tone and soothing effect. The system can then proactively ask, "Would you like to do a breathing exercise with me?" initiating an interactive healing process. The entire process demonstrates the psychological dialogue system's ability to "understand, comfort, and guide" people in real-world scenarios. The device supports emotion recognition, long-term dialogue memory, and personalized voiceprint customization, providing a good user immersion experience and intelligent companionship capabilities.
[0089] The device works by combining CBT (Cognitive Behavioral Therapy) with an AI intervention mechanism that integrates emotion recognition. It dynamically assesses the user's current emotional state (e.g., anxiety, depression, anger) through real-time emotion recognition and tracks the user's automatic thoughts, emotional reactions, and behavioral expressions within the context of the conversation. The large language model, based on the CBT "emotion-cognition-behavior" chain, simulates the typical guidance methods of a psychological counselor, intervening in the user's psychological process in stages to achieve effective intelligent intervention.
[0090] In practice, the system first identifies the user's emotions using voice or text input. For example, "I always feel like I'm not good enough, I can't do anything well" would be labeled as "self-doubt + depression." Then, the system responds gently, like an AI-simulated counselor: "You say you can't do anything well, do you often have similar thoughts? Let's see if this feeling is entirely true." This type of language embodies the "cognitive reappraisal" technique in CBT, guiding the user to identify and challenge irrational automatic thoughts. Next, the system further provides alternative thought construction (e.g., "You've already completed many important tasks at work, would you try to acknowledge your efforts?") and encourages behavioral activation (e.g., "Would you like to try making a small plan, such as completing a small task that you can control today?"). The AI intervention mechanism also supports "emotion tracking + feedback loop" functionality. The system continuously monitors the user's emotional response to the intervention content. If the user's emotions do not improve or fluctuate more, the system adjusts the intervention strategy, using more empathetic language or lower-level guidance methods, simulating the natural rhythm and technical structure of a human counselor's "listen-ask-guide-summarize" dialogue. This mechanism not only enhances the credibility of AI's interpersonal interaction, but also gives the psychological dialogue system a preliminary therapeutic structure, which helps to provide users with effective emotional support and cognitive guidance in non-crisis situations.
[0091] Furthermore, it should be noted that the artificial intelligence-based psychotherapy device provided by this invention does not have medical diagnostic or drug treatment functions, and is only used as an auxiliary psychological rehabilitation dialogue terminal; the system adopts edge AI deployment, runs locally by default, does not rely on the cloud, and ensures that data does not leave the device; the device can be updated and its functions iterated via OTA to ensure continuous optimization of the model algorithm; it has undergone multiple rounds of joint training and optimization by psychology experts, and has safety thresholds and buffered corpora to prevent the induction of induced negative emotions.
[0092] The artificial intelligence psychotherapy device provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An artificial intelligence-based psychotherapy device, characterized in that, include: The module includes a speech recognition module, a semantic understanding and dialogue generation module, and a TTS speech synthesis module. The speech recognition module is used to receive the user's voice information and convert the voice information into text information; The semantic understanding and dialogue generation module is used to identify the voice information and the text information, determine the user's emotional state, and generate text response information to alleviate the user's emotions based on the emotional state. The TTS speech synthesis module is used to convert the text reply information into speech reply information and output the speech reply information.
2. The artificial intelligence psychotherapy device according to claim 1, characterized in that, The semantic understanding and dialogue generation module includes an emotion recognition model; The emotion recognition model is used to identify the user's emotional state based on the voice information, the text information, and historical dialogue information.
3. The artificial intelligence psychotherapy device according to claim 2, characterized in that, The semantic understanding and dialogue generation module also includes a dialogue strategy model; The dialogue strategy model includes an input encoding layer, a strategy decision layer, and an output generation layer. The input coding layer is used to fuse the speech information, the text information, and the historical dialogue information to obtain a fusion vector; The strategy decision layer is used to select the corresponding conversation response based on the fusion vector and the emotional state; The output generation layer is used to generate the text reply information based on the session response.
4. The artificial intelligence psychotherapy device according to claim 3, characterized in that, The input coding layer includes a speech branch, a text branch, a context encoder, and a dynamic fusion module; The speech branch is used to extract acoustic features from the speech information and generate an acoustic emotion vector. The text branch is used to extract semantic features from the text information and generate a semantic sentiment vector; The context encoder is used to model the emotion sequence of the historical dialogue information and generate an emotion trend vector. The dynamic fusion module is used to fuse the acoustic emotion vector, the semantic emotion vector, and the emotion trend vector to obtain a fusion vector.
5. The artificial intelligence psychotherapy device according to claim 3, characterized in that, The policy decision layer includes a policy selection network and a policy-action mapping module: A strategy selection network is used to predict the probability distribution of the dialogue strategy corresponding to the fusion vector based on the fusion vector and the emotion state; The strategy-action mapping module is used to retrieve the corresponding session response from a preset template library based on the probability distribution.
6. The artificial intelligence psychotherapy device according to claim 3, characterized in that, The dialogue strategy model is trained in the following way: Obtain a dataset of psychological counseling dialogues from different users, and label each dialogue with an emotion category label and a dialogue strategy label. The LMM model is trained using the dialogue dataset, the emotion category labels, and the dialogue policy labels to obtain the dialogue policy model.
7. The artificial intelligence psychotherapy device according to claim 6, characterized in that, The dialogue strategy tags include: Soothe anxiety, guide emotional awareness, suggest behavioral activation, alleviate sadness, or channel anger.
8. The artificial intelligence psychotherapy device according to claim 3, characterized in that, The dialogue strategy model is constructed based on the experience of psychotherapists, CBT, and MBCT intervention frameworks.
9. The artificial intelligence psychotherapy device according to claim 1, characterized in that, It also includes a voiceprint recognition module; The voiceprint recognition module is used to confirm the user's identity information based on the voice information.
10. The artificial intelligence psychotherapy device according to claim 9, characterized in that, The voiceprint recognition module, the speech recognition module, the semantic understanding and dialogue generation module, and the TTS speech synthesis module are integrated on the main control module; The main control module includes a Cortex-A7 embedded AI processing chip and an NPU coprocessor.