Multi-mode psychological counseling system based on artificial intelligence
By combining facial expressions, voice and text data for comprehensive emotion recognition through a multimodal psychological counseling system, the problem of insufficient accuracy of facial expression recognition in traditional psychological counseling is solved, and more accurate and comprehensive emotion analysis and personalized counseling strategies are achieved.
Patent Information
- Application Number
- CN202510687772.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies lack accuracy and comprehensiveness in facial expression recognition in psychological counseling. When using facial expressions, voice emotion recognition or text sentiment analysis alone, there are interpretation biases and noise interference, making it difficult to achieve accurate emotion recognition.
A multimodal psychological counseling system is used to obtain facial expression data through video acquisition equipment and voice data through audio acquisition equipment. Natural language processing technology is combined to analyze text data. Deep learning and Transformer architecture are used to build expression recognition and voice emotion recognition models. Emotion recognition is performed through a multimodal fusion model, and personalized psychological counseling strategies are generated. Real-time adjustments are made through an interactive feedback module.
It improves the accuracy and comprehensiveness of emotion recognition, generates personalized psychological counseling strategies, and enhances the effectiveness of psychological counseling.
Smart Images

Figure CN120690385A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence and mental health services, and in particular to a multimodal psychological counseling system based on artificial intelligence. Background Art
[0002] With increasing social pressure, people are increasingly paying attention to their mental health. Traditional psychological counseling methods rely primarily on face-to-face interactions with professional counselors, which are subject to challenges such as limited resources, time and location constraints, and high costs. The development of artificial intelligence (AI) technology has provided new solutions for psychological counseling. Currently, technologies such as facial expression recognition, speech emotion recognition, and text sentiment analysis have made some progress in their respective fields. However, when used independently, recognition accuracy and comprehensiveness are insufficient due to the complexity and diversity of human emotional expression. For example, facial expressions can be interpreted erratically due to individual differences and cultural backgrounds; speech emotion recognition can be affected by environmental noise; and text sentiment analysis has difficulty processing complex semantics such as metaphors and irony.
[0003] In the field of facial expression recognition, while existing technologies can identify basic expressions, the accuracy of facial positioning and expression recognition in complex scenarios still needs to be improved. In psychological counseling scenarios, subtle changes in a user's expression are crucial for accurately judging their emotional state. Therefore, more advanced facial recognition technology is needed to support the accuracy of multimodal emotion recognition. Multimodal fusion technology can integrate multiple information sources, complementing and verifying each other, improving the accuracy and comprehensiveness of emotion recognition, and providing the possibility of building more intelligent and effective psychological counseling systems. Summary of the Invention
[0004] In view of this, the present application provides a multimodal psychological counseling system based on artificial intelligence to improve the accuracy and comprehensiveness of emotion recognition.
[0005] The present application provides a multimodal psychological counseling system based on artificial intelligence, which includes a multimodal data acquisition module, a multimodal emotion recognition module, a psychological counseling strategy generation module and an interactive feedback module; The multimodal data acquisition module is used to collect facial expressions of a target user through a video acquisition device to obtain image data of the target user, and to locate, crop and preprocess the image data to obtain facial image data that meets the input requirements of the expression recognition model; and also collecting the voice of the target user through an audio collection device to obtain voice data of the target user and text data corresponding to the voice data; The multimodal emotion recognition module is used to determine the facial expression recognition state of the target user through the facial image data, determine the voice emotion recognition state of the target user through the voice data, determine the text emotion analysis state of the target user through the text data, and fuse the facial expression recognition state, voice emotion recognition state and text emotion analysis state to obtain the comprehensive emotion state of the target user; The psychological counseling strategy generation module is used to match and generate personalized psychological counseling strategies from a preset knowledge base based on the user's comprehensive emotional state and the target user's personal information; The interactive feedback module is used to communicate with the target user according to the psychological counseling strategy, collect feedback information from the target user in real time during the communication process, and dynamically adjust the psychological counseling strategy according to the feedback information.
[0006] Optionally, the multimodal emotion recognition module includes a facial expression recognition submodule, a voice emotion recognition submodule, a text emotion analysis submodule and a multimodal fusion submodule; The facial expression recognition submodule is used to use a deep learning algorithm to build an expression recognition algorithm model, and to determine the facial key points of the target user in real time, and to determine the facial expression recognition status of the target user according to the coordinate changes of the facial key points through the expression recognition algorithm model; The speech emotion recognition submodule is used to extract and classify the speech data using a speech emotion recognition model to obtain acoustic features and prosodic features of the speech data, and then determine the speech emotion recognition state of the target user based on the acoustic features and prosodic features using the speech emotion recognition model; The text sentiment analysis submodule is used to construct a text sentiment analysis model based on the Transformer architecture through natural language processing technology, and analyze the text data through the text sentiment analysis model to obtain the text sentiment analysis status of the target user; The multimodal fusion submodule is used to calculate the weight values of the facial expression recognition submodule, the speech emotion recognition submodule, and the text emotion analysis submodule, and fuse the facial expression recognition state, the speech emotion recognition state, and the text emotion analysis state according to the weight values to obtain the comprehensive emotional state of the target user.
[0007] Optionally, the use of a deep learning algorithm to construct an expression recognition algorithm model includes: Pre-collecting a historical facial expression image dataset under different user conditions, and expanding the historical facial expression image dataset using data augmentation technology, wherein the user conditions include race, age, gender, cultural background, lighting conditions, angle, and posture; Build an initial expression recognition algorithm model and perform feature extraction and classification on the expanded facial expression image dataset by sequentially adding convolutional layers, pooling layers, and fully connected layers to obtain historical feature data for training. The initial expression recognition algorithm model is trained using the historical feature data to obtain a trained expression recognition algorithm model.
[0008] Optionally, the prosodic features include pitch, duration, and volume changes.
[0009] Optionally, the positioning, cropping, and preprocessing of the image data includes: Construct feature map pyramid layers P2 to P6, and determine classification networks C2 to C5 corresponding to the feature map pyramid layers P2 to P5 based on a preset data set, determine P2 to P5 through the output of the ResNet residual stage of C2 to C5, and determine P6 by performing a preset step-size convolution calculation at C5; Replace the convolution in each feature map pyramid layer with a deformable convolutional network; Anchor points of preset proportions are set on each feature map pyramid level from P2 to P6 so that the anchor points cover the image data, and then the image data is positioned, cropped, and preprocessed using the anchor points.
[0010] In the embodiment provided in this application, the system first collects the user's facial image data, voice data, and text data through a multimodal data acquisition module. Each of these data is then processed separately through a multimodal emotion recognition module, and the processed data is then integrated to obtain the user's comprehensive emotional state. The psychological counseling strategy generation module then determines the corresponding psychological counseling policy based on this comprehensive emotional state and the user's personal information. Finally, the interactive feedback module communicates with the user based on this policy and optimizes the policy based on the feedback received during the communication. This enables the user's emotional state to be judged through multiple information sources, including facial expressions, voice emotions, and voice text. This improves the accuracy and comprehensiveness of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 A system module diagram provided for an embodiment of the present application; Figure 2 A diagram of the device structure provided in an embodiment of the present application; Figure 3A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0013] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0014] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0015] This application provides a multimodal psychological counseling system based on artificial intelligence to improve the accuracy and comprehensiveness of user emotion recognition.
[0016] The following specific embodiments are used to describe the technical solution of the present application in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0017] like Figure 1 The figure shows a module diagram of a multimodal psychological counseling system based on artificial intelligence provided by this application. The system can be used in a closed psychological counseling room so that users can receive psychological counseling in a private environment. The functions and effects of each module are described below.
[0018] Module 1, Multimodal Data Acquisition Module. This module is used to capture facial expressions of a target user using a video acquisition device to obtain image data of the target user, and to locate, crop, and preprocess the image data to obtain facial image data that meets the input requirements of the expression recognition model. Furthermore, the target user's voice is collected through an audio collection device to obtain the target user's voice data and text data corresponding to the voice data.
[0019] In this embodiment, facial image data can be captured using a high-definition camera installed in the counseling room, employing image acquisition technology to capture the user's facial expressions in real time. The camera is equipped with autofocus and adaptive light adjustment features to ensure clear image data under varying lighting conditions. A pre-set algorithm is then used to locate the face, identifying the coordinates of the face within the image—that is, obtaining the coordinates represented by the facial bounding box. Based on the localization results, the facial region is cropped to remove irrelevant background interference and preserve the complete facial image. The cropped image is then normalized, adjusting parameters such as image size, brightness, and contrast to meet the input requirements of the subsequent expression recognition algorithm model.
[0020] For voice data, a high-sensitivity microphone array can be equipped to accurately capture user voice under different environmental noise conditions, while also capturing the audio characteristics of the voice, such as volume, pitch, speaking speed, etc. For text data, text extraction can be performed based on the voice data.
[0021] In another embodiment, positioning, cropping and preprocessing the image data includes: Construct feature map pyramid layers P2 to P6, and determine classification networks C2 to C5 corresponding to the feature map pyramid layers P2 to P5 based on a preset data set, determine P2 to P5 through the output of the ResNet residual stage of C2 to C5, and determine P6 by performing a preset step-size convolution calculation at C5; Replace the convolution in each feature map pyramid layer with a deformable convolutional network; Anchor points of preset proportions are set on each feature map pyramid level from P2 to P6 so that the anchor points cover the image data, and then the image data is positioned, cropped, and preprocessed using the anchor points.
[0022] In this embodiment, most cutting-edge face detection methods focus on single-stage designs. This design densely samples the location and scale of faces on a feature map pyramid, demonstrating superior performance and faster detection speed compared to two-stage approaches. Based on this background, we optimized and improved the single-stage face detection framework.
[0023] First, a self-supervised grid feature pyramid network (FPN) is constructed. The process is as follows: Get the ResNet-152 classification network weights pre-trained on the ImageNet-11k dataset from a public pre-trained model library (such as PyTorch or TensorFlow's model repository), load it into the local development environment, and use it as the basic network for feature extraction.
[0024] Then, using relevant functions and interfaces from deep learning frameworks such as PyTorch, the P2-P5 feature pyramid layers are calculated through a top-down path and lateral connections. Specifically, starting with the outputs of the C2-C5 stages of ResNet-152, ShenduFace uses upsampling, convolution, and other operations for fusion calculations to achieve cross-layer feature fusion and enhance feature expressiveness. For example, for upsampling, the nn.Upsample function can be used to set the appropriate upsampling factor and interpolation method, such as bilinear interpolation. For lateral convolution operations, the nn.Conv2d function is used to define appropriate convolution kernel parameters, such as kernel size, stride, and padding. Finally, based on the output of the C5 layer, the nn.Conv2d function is used to define a 3x3 convolution layer with a stride of 2. This convolution operation is performed on the feature map of the C5 layer to obtain the feature map of the P6 layer.
[0025] Then, build the context module. The process is as follows: For each feature pyramid layer (P2-P6), define and add a context module in the code. Context modules can be implemented using custom classes that inherit from the base module class in deep learning frameworks (such as nn.Module in PyTorch). In the class initialization function, define relevant member variables, such as the deformable convolution layer.
[0026] Third, determine the loss function of ShenduFace, which consists of four parts: 1. Face classification loss :in The model predicts The probability of belonging to a face, is the true label, with a value of 1 representing a positive anchor and a value of 0 representing a negative anchor. The classification loss uses softmax loss, which is used to process the binary classification task of face / non-face.
[0027] 2. Face frame regression loss : Represents the predicted box coordinates associated with the positive anchor, is the real box coordinate. According to the Fast R-CNN method, the center coordinate, width and height of the regression box are normalized, and the loss is calculated as follows: Defined as: 3. Face key point regression loss : are the predicted coordinates of the five facial key points, The regression of the five facial key points also uses the target normalization strategy based on the anchor center.
[0028] 4. Dense regression loss Based on the specific Dense Regression task definition and requirements, write the corresponding loss calculation function. For example, if you are regressing the 3D position and correspondence of facial pixels, you can define a custom loss function based on the difference between the predicted value and the true value. Assuming the predicted 3D position tensor is pred_3d_pixel and the true 3D position tensor is gt_3d_pixel, you can use the mean squared error (MSE) loss to calculate the loss.
[0029] The above four loss items are divided into two categories according to the set weights ( =0.25 , =0.1 , = 0.01) and perform weighted summation to obtain the final multi-task loss function: Next, anchors are set. Anchors are virtual boxes pre-placed by the model on the image, each representing a hypothetical "possible target region" (such as a face, object, etc.). ShenduFace uses a specific anchor ratio at feature map pyramid levels from P2 to P6. P2 can capture small faces by placing small anchors, but this increases computational effort and the risk of false positives. ShenduFace sets the scale step to 231 and the aspect ratio to 1:1. The input image size is 640*640, and the anchors cover the range of 16x16 to 406x406 in the feature map pyramid level. Anchors are tiled across five downsampled feature maps (4, 8, 16, 32, and 64). Each point in the original feature map is predicted to have three anchors, for a total of (160*160 + 80*80 + 40*40 + 400 + 100) * 3 = 102,300 anchors, 75% of which come from P2. However, in this embodiment, downsampling rates of 8, 16, and 32 can be used, namely the feature maps of the three layers P3, P4, and P5, and only two anchor points are placed at each point. For a 640*640 input image, when the downsampling output is 32, 16, and 8, the output format of each point is [(1,4,20,20), (1,8,20,20), (1,20,20,20)], [(1,4,40,40), (1,8,40,40), (1,20,40,40)], [(1,4,80,80), (1,8,80,80), (1,20,80,80)], where 4, 8, and 20 correspond to the number of categories of two anchor points per point (2 anchor points * 2 categories), (2 anchor points * box information), and (2 anchor points * 5 key point information (one point x, y)).
[0030] By reducing the number and layers of anchor points, the computational effort is significantly reduced, while relying on the P3 layer to retain the ability to detect small objects. The final facial localization algorithm model, ShenduFace, uses the aforementioned anchor points to locate, crop, and preprocess the captured image data.
[0031] Module 2, Multimodal Emotion Recognition Module. This module is used to determine the target user's facial expression recognition status based on the facial image data, determine the target user's voice emotion recognition status based on the voice data, and determine the target user's text emotion analysis status based on the text data. The module then fuses the facial expression recognition status, voice emotion recognition status, and text emotion analysis status to obtain the target user's comprehensive emotion state.
[0032] In this embodiment, facial expression recognition states, such as happiness, sadness, anger, surprise, fear, disgust, and neutrality, can be determined by an expression recognition algorithm model ShenduFacial, which can be obtained by training historical data.
[0033] For speech emotion recognition, a speech emotion recognition model can be used to extract and classify features from the collected speech signal. This model extracts acoustic and prosodic features of the speech, such as pitch, duration, and volume variations. The model is then trained to identify different speech emotion states, such as positive, negative, neutral, anxious, and depressed.
[0034] For the sentiment analysis status of the text, the overall sentiment analysis status of the text, such as positive, negative or neutral, can be determined by identifying the sentiment words, sentiment tendencies and semantic relationships in the text.
[0035] After determining the facial expression recognition status, voice emotion recognition status, and text sentiment analysis status, they can be fused through fusion algorithms such as tensor fusion networks to ultimately derive the user's comprehensive emotional state.
[0036] In another embodiment, the multimodal emotion recognition module includes a facial expression recognition submodule, a speech emotion recognition submodule, a text emotion analysis submodule and a multimodal fusion submodule; The facial expression recognition submodule is used to construct an expression recognition algorithm model using a deep learning algorithm, and determine the facial key points of the target user in real time, and determine the facial expression recognition status of the target user according to the coordinate changes of the facial key points through the expression recognition algorithm model; The speech emotion recognition submodule is used to extract and classify the speech data using a speech emotion recognition model to obtain acoustic features and prosodic features of the speech data, and then determine the speech emotion recognition state of the target user based on the acoustic features and prosodic features using the speech emotion recognition model; The text sentiment analysis submodule is used to construct a text sentiment analysis model based on the Transformer architecture through natural language processing technology, and analyze the text data through the text sentiment analysis model to obtain the text sentiment analysis status of the target user; The multimodal fusion submodule is used to calculate the weight values of the facial expression recognition submodule, the speech emotion recognition submodule, and the text emotion analysis submodule, and fuse the facial expression recognition state, the speech emotion recognition state, and the text emotion analysis state according to the weight values to obtain the comprehensive emotional state of the target user.
[0037] This embodiment analyzes the data obtained by the multimodal data acquisition module through three different models, namely, the expression recognition algorithm model analyzes the facial image data to obtain the facial expression recognition status; the speech emotion recognition model analyzes the speech data to obtain the speech emotion recognition status; the text sentiment analysis model analyzes the text data to obtain the text sentiment analysis status, and the various analysis states are fused through the multimodal fusion model.
[0038] The construction process of each model is described below.
[0039] 1. Expression recognition algorithm model.
[0040] We collected a large dataset of facial expression images from diverse ethnicities, ages, genders, cultural backgrounds, lighting conditions, angles, and postures, and meticulously annotated them. The annotations included key features of basic expression categories (happy, sad, angry, surprised, fearful, disgusted, and neutral) as well as micro-expressions. We built the model using a convolutional neural network (CNN) combined with the Transformer architecture, and conducted distributed training on a GPU cluster. Pre-training on large-scale public expression datasets (such as FER2013 and RAF-DB) leverages the local feature extraction capabilities of CNNs and the global modeling capabilities of Transformers to learn universal expression feature representations. During pre-training, multiple convolutional layers are used to extract low-level features such as edges and textures using kernels of varying sizes and numbers. Pooling layers are used to reduce data dimensionality and enhance the model's robustness to expression variations. Within the Transformer layer, a multi-head attention mechanism focuses on facial regions closely related to emotional expression (such as the eyes and mouth). In the training of the multimodal emotion recognition model, the training content of the facial expression recognition model was further refined. For the facial expression recognition model, a large dataset of facial expression images from diverse ethnicities, ages, genders, cultural backgrounds, lighting conditions, angles, and postures was collected and meticulously annotated. The model was constructed using a convolutional neural network (CNN) similar to [specific open source algorithm name] combined with the Transformer architecture (if a combination is used, the integration method will be detailed; otherwise, the CNN portion of the existing algorithm will be emphasized). The model was trained on a GPU cluster. The current model structure, as shown in the code, consists of multiple convolutional layers, pooling layers, and fully connected layers. Pre-training was first performed on large-scale public expression datasets (such as FER2013, which is an example of actual datasets used) to learn general expression feature representations. Then, fine-tuning was performed on a dataset specific to psychological counseling scenarios, adjusting model parameters to meet the expression recognition requirements of psychological counseling scenarios. During training, data augmentation techniques (such as rotation, scaling, flipping, and noise addition) were used to expand the dataset and improve the model's generalization capabilities. Continuously optimize the model structure and training parameters, such as adjusting the convolution kernel size, step size, number of neurons in the fully connected layer, and learning rate, to improve facial expression recognition accuracy. In practice, expression recognition is performed by calling the predict method in the EmotionClient class. This method first converts the input image to grayscale, resizes it to (48, 48), then performs dimensionality expansion, and finally inputs it into the model for prediction, returning the predicted probabilities for seven different expressions.
[0041] 2. Speech emotion recognition model.
[0042] We collected a rich speech dataset covering different emotional states, such as positive, negative, neutral, anxious, and depressed, as well as speech samples under different speech speeds, intonations, accents, and ambient noise conditions, and annotated them with corresponding emotion labels. We trained these models using recurrent neural networks (RNNs) and their variants (such as LSTMs and GRUs), optimizing them based on the data2vec self-supervised framework and the emotion2vec model. In terms of feature extraction, acoustic features of speech, such as Mel-Frequency Cepstral Coefficients (MFCCs), Linear Predictive Coding (LPC), and spectral entropy, as well as prosodic features such as pitch, duration, volume variation, and fundamental frequency, are extracted as model inputs. During training, sentence-level and frame-level losses are introduced. The sentence-level loss is calculated using the mean squared error (MSE). By temporally pooling the speech embeddings of the teacher network T and the student network S, the mean squared error (MSE) between the two is calculated to learn overall global emotion. The frame-level loss follows the Masked Language Model (MLM) approach, calculating only the loss of the masked portion. By calculating the mean squared error between the outputs of the teacher network T and the student network S on the masked frames, the model is encouraged to learn contextual emotional information. Using an online distillation strategy, the student network S updates its parameters via backpropagation, while the teacher network T updates its parameters via exponential moving average (EMA). The total loss L of the student network S is composed of a frame-level loss and a sentence-level loss, balanced by an adjustable weight alpha. Through extensive experiments, the model structure (such as the number of network layers and neurons), training parameters (such as the learning rate, batch size, and number of training epochs), and alpha value were adjusted to improve speech emotion recognition performance.
[0043] 3. Text sentiment analysis model.
[0044] Build a large-scale text corpus covering a variety of fields and sentiment types, including psychological counseling-related texts, social media texts, literary works, and popular psychology articles. Pre-train on Transformer architectures (such as BERT and GPT) and fine-tune for the psychological counseling field. During the pre-training phase, the model leverages large-scale, general-purpose text data to learn the basic semantics and grammatical knowledge of the language. During fine-tuning, appropriate training objectives and loss functions, such as the cross-entropy loss function, are designed to guide the model in learning the emotional expression characteristics of psychological counseling. Input text undergoes lexical analysis (word segmentation and part-of-speech tagging), syntactic analysis (phrase structure analysis, dependency analysis), and semantic analysis (emotional vocabulary identification, emotional tendency judgment, and semantic relationship extraction). The model identifies emotional vocabulary, emotional tendencies, and semantic relationships within the text, determines the overall emotional polarity of the text (positive, negative, or neutral), and further analyzes the emotional intensity and emotional category. By continuously optimizing model parameters (such as attention mechanism parameters, learning rate, and number of training rounds) and training strategies (such as data augmentation and model fusion), the model's accuracy in sentiment analysis of psychologically relevant text is improved.
[0045] 4. Multimodal fusion model.
[0046] We collected multimodal emotion data samples and used the results of facial expression recognition, speech emotion recognition, and text sentiment analysis as input, with the corresponding real-world comprehensive emotional states as labels, to train a multimodal fusion model. We used a tensor fusion network algorithm, taking into account the reliability and complementarity of data from different modalities, and calculated the weights of the results from each modality to achieve multimodal emotion fusion.
[0047] During the training process, cross-validation and other methods were used to evaluate model performance. First, the multimodal data samples were divided into training, validation, and test sets. The model was trained on the training set, and the parameters of the tensor fusion network algorithm (such as weight coefficients and fusion methods) were adjusted. Model performance was evaluated on the validation set, and model parameters were optimized based on the evaluation results. Finally, a final test was conducted on the test set to ensure the model's generalization ability. By continuously adjusting the fusion algorithm and parameters, the accuracy and reliability of the fusion results were improved, achieving a precise assessment of the user's comprehensive emotional state.
[0048] Module 3, psychological counseling strategy generation module: This module is used to match and generate personalized psychological counseling strategies from a preset knowledge base based on the user's comprehensive emotional state and the target user's personal information.
[0049] In this embodiment, the knowledge base construction process is as follows: Organize knowledge on psychological theories, psychotherapy methods, common psychological problems, and coping strategies. Store this knowledge in a structured knowledge base, managed using a database management system (e.g., MySQL, MongoDB). Regarding knowledge base management, a knowledge base management interface has been developed, providing a visual interface that allows psychology experts to add, edit, and delete knowledge. Knowledge is categorized and tagged, for example, by psychological problem type (anxiety, depression, stress, etc.), treatment type, and applicable population, to facilitate knowledge retrieval. Furthermore, a knowledge review mechanism has been established, requiring newly added or modified knowledge to undergo expert review before formal entry into the database, ensuring its accuracy and validity.
[0050] Then, based on the multimodal emotion recognition results and the user's personal information (age, gender, psychological profile, past consultation records, etc.), a matching algorithm is designed to filter and generate personalized psychological counseling strategies from the knowledge base. A method combining rule-based and machine learning is used: Rule-based matching pre-sets a series of rules. For example, when the user's emotion recognition result is anxiety and the age is under 30 years old, priority is given to screening anxiety coping methods suitable for young people from the knowledge base (such as relaxation training techniques, time management methods, positive self-suggestion statements, etc.); when the user's emotion recognition result is depression, priority is given to recommending strategies related to cognitive behavioral therapy. In terms of machine learning, algorithms such as decision trees and neural networks are used to further optimize and rank matching results. By learning from a large amount of historical user data and corresponding effective psychological counseling strategies, a mapping model is established between user emotional states, personal information, and psychological counseling strategies. For example, for young users identified as anxious, the model combines and prioritizes strategies selected from the knowledge base based on factors such as the user's specific anxiety level and personality traits, generating personalized psychological counseling strategies that best suit the user. This can include recommending specific relaxation training audio, guiding users through specific steps for cognitive restructuring, and recommending relevant psychological self-help books.
[0051] Module four, the interactive feedback module is used to communicate with the target user according to the psychological counseling strategy, collect feedback information from the target user in real time during the communication process, and dynamically adjust the psychological counseling strategy according to the feedback information.
[0052] In this embodiment, computer graphics and animation technology can be used to design a friendly and attractive virtual psychological companion character, including the character's appearance (facial features, hairstyle, skin color, etc.), clothing (different styles of clothing), expressions (rich facial expression variations), and movements (natural body movements). The virtual character model is created using 3D modeling software, and animation tools are used to design the character's various expressions and movement sequences. Deep learning-based speech synthesis technology then enables the system to generate natural and fluent voice responses based on psychological counseling strategies. A character emotion expression model is established to control the emotional expression of the virtual character based on the user's emotional state and psychological counseling strategies. For example, when a user is feeling down, the virtual character communicates with the user by slowing down their speech, adopting a gentle tone, and displaying caring and comforting expressions and gestures, providing emotional support and guidance. When the user is emotionally agitated, the virtual character responds in a calm and stable manner to help the user calm down. For feedback information, a variety of feedback entrances can be set up in the interactive interface, providing multiple feedback methods such as text input, voice feedback, rating selection (such as 1-5 point satisfaction rating), emoticon selection, etc., to facilitate users to input feedback information such as the degree of understanding of psychological counseling content, satisfaction, and whether there are new emotional changes. Finally, the collected feedback is analyzed and categorized in real time. For text and voice feedback, natural language processing technology is used to perform semantic understanding and sentiment analysis to extract key information. Ratings and emoji feedback are quantified. Based on the feedback analysis results, psychological counseling strategies are dynamically adjusted. For example, if a user reports that they don't understand a particular type of psychological advice, the system will reselect a more accessible explanation or recommend other relevant resources. If the user's emotions change, the system will regenerate a corresponding psychological counseling strategy based on the new emotional state to improve counseling effectiveness. So far, completed Figure 1 Functional description of each module in .
[0053] In an embodiment of the present application, the system first collects the user's facial image data, voice data, and text data through a multimodal data acquisition module. Each of these data is then processed separately through a multimodal emotion recognition module, and the processed data is then integrated to obtain the user's comprehensive emotional state. The psychological counseling strategy generation module then determines the corresponding psychological counseling policy based on this comprehensive emotional state and the user's personal information. Finally, the interactive feedback module communicates with the user based on this policy and optimizes the policy based on the feedback received during the communication. This enables the user's emotional state to be judged using multiple information sources, including facial expressions, voice emotions, and voice text. This improves the accuracy and comprehensiveness of emotion recognition.
[0054] like Figure 2 As shown, the present application also provides a multimodal psychological counseling device based on artificial intelligence, the device comprising: The multimodal data acquisition unit 201 is configured to acquire facial expressions of a target user through a video acquisition device to obtain image data of the target user, and to locate, crop, and preprocess the image data to obtain facial image data that meets the input requirements of an expression recognition model; and also collecting the voice of the target user through an audio collection device to obtain voice data of the target user and text data corresponding to the voice data; a multimodal emotion recognition unit 202 for determining a facial expression recognition state of the target user based on the facial image data, determining a speech emotion recognition state of the target user based on the speech data, determining a text emotion analysis state of the target user based on the text data, and fusing the facial expression recognition state, speech emotion recognition state, and text emotion analysis state to obtain a comprehensive emotion state of the target user; The psychological counseling strategy generating unit 203 is used to match and generate personalized psychological counseling strategies from a preset knowledge base based on the user's comprehensive emotional state and the target user's personal information; The interactive feedback unit 204 is used to communicate with the target user according to the psychological counseling strategy, collect feedback information from the target user in real time during the communication process, and dynamically adjust the psychological counseling strategy according to the feedback information.
[0055] In another embodiment, the multimodal emotion recognition unit includes a facial expression recognition submodule 202A, a speech emotion recognition submodule 202B, a text emotion analysis submodule 202C, and a multimodal fusion submodule 202D; The facial expression recognition submodule is used to use a deep learning algorithm to build an expression recognition algorithm model, and to determine the facial key points of the target user in real time, and to determine the facial expression recognition status of the target user according to the coordinate changes of the facial key points through the expression recognition algorithm model; The speech emotion recognition submodule is used to extract and classify the speech data using a speech emotion recognition model to obtain acoustic features and prosodic features of the speech data, and then determine the speech emotion recognition state of the target user based on the acoustic features and prosodic features using the speech emotion recognition model; The text sentiment analysis submodule is used to construct a text sentiment analysis model based on the Transformer architecture through natural language processing technology, and analyze the text data through the text sentiment analysis model to obtain the text sentiment analysis status of the target user; The multimodal fusion submodule is used to calculate the weight values of the facial expression recognition submodule, the speech emotion recognition submodule, and the text emotion analysis submodule, and fuse the facial expression recognition state, the speech emotion recognition state, and the text emotion analysis state according to the weight values to obtain the comprehensive emotional state of the target user.
[0056] In another embodiment, the multimodal emotion recognition unit uses a deep learning algorithm to construct an expression recognition algorithm model, including: Pre-collecting a historical facial expression image dataset under different user conditions, and expanding the historical facial expression image dataset using data augmentation technology, wherein the user conditions include race, age, gender, cultural background, lighting conditions, angle, and posture; Build an initial expression recognition algorithm model and perform feature extraction and classification on the expanded facial expression image dataset by sequentially adding convolutional layers, pooling layers, and fully connected layers to obtain historical feature data for training. The initial expression recognition algorithm model is trained using the historical feature data to obtain a trained expression recognition algorithm model.
[0057] In another embodiment, the prosodic features in the multimodal emotion recognition unit include pitch, duration, and volume changes.
[0058] In another embodiment, the positioning, cropping, and preprocessing of the image data in the multimodal data acquisition unit includes: Construct feature map pyramid layers P2 to P6, and determine classification networks C2 to C5 corresponding to the feature map pyramid layers P2 to P5 based on a preset data set, determine P2 to P5 through the output of the ResNet residual stage of C2 to C5, and determine P6 by performing a preset step-size convolution calculation at C5; Replace the convolution in each feature map pyramid layer with a deformable convolutional network; Anchor points of preset proportions are set on each feature map pyramid level from P2 to P6 so that the anchor points cover the image data, and then the image data is positioned, cropped, and preprocessed using the anchor points.
[0059] The above embodiments of the present invention provide a multimodal psychological counseling system based on artificial intelligence, and based on this system, provide a multimodal psychological counseling device based on artificial intelligence. Through the above system and device, the accuracy and comprehensiveness of emotion recognition can be improved.
[0060] This embodiment also discloses a computer device, such as Figure 3 As shown, the computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement any of the above-mentioned methods on the artificial intelligence-based multimodal psychological counseling system.
[0061] In addition, in the above-mentioned example implementation of the multimodal psychological counseling device based on artificial intelligence, the logical division of each program module is only an example. In actual application, the above-mentioned functions can be assigned to different program modules as needed, for example, for the configuration requirements of the corresponding hardware or the convenience of software implementation. That is, the internal structure of the multimodal psychological counseling device based on artificial intelligence is divided into different program modules to complete all or part of the functions described above.
[0062] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A multimodal psychological counseling system based on artificial intelligence, characterized by: The system includes a multimodal data acquisition module, a multimodal emotion recognition module, a psychological counseling strategy generation module and an interactive feedback module; The multimodal data acquisition module is used to collect facial expressions of a target user through a video acquisition device to obtain image data of the target user, and to locate, crop and preprocess the image data to obtain facial image data that meets the input requirements of the expression recognition model; and also collecting the voice of the target user through an audio collection device to obtain voice data of the target user and text data corresponding to the voice data; The multimodal emotion recognition module is used to determine the facial expression recognition state of the target user through the facial image data, determine the voice emotion recognition state of the target user through the voice data, determine the text emotion analysis state of the target user through the text data, and fuse the facial expression recognition state, voice emotion recognition state and text emotion analysis state to obtain the comprehensive emotion state of the target user; The psychological counseling strategy generation module is used to match and generate personalized psychological counseling strategies from a preset knowledge base based on the user's comprehensive emotional state and the target user's personal information; The interactive feedback module is used to communicate with the target user according to the psychological counseling strategy, collect feedback information from the target user in real time during the communication process, and dynamically adjust the psychological counseling strategy according to the feedback information.
2. The system according to claim 1, wherein: The multimodal emotion recognition module includes a facial expression recognition submodule, a voice emotion recognition submodule, a text emotion analysis submodule and a multimodal fusion submodule; The facial expression recognition submodule is used to use a deep learning algorithm to build an expression recognition algorithm model, and to determine the facial key points of the target user in real time, and to determine the facial expression recognition status of the target user according to the coordinate changes of the facial key points through the expression recognition algorithm model; The speech emotion recognition submodule is used to extract and classify the speech data using a speech emotion recognition model to obtain acoustic features and prosodic features of the speech data, and then determine the speech emotion recognition state of the target user based on the acoustic features and prosodic features using the speech emotion recognition model; The text sentiment analysis submodule is used to construct a text sentiment analysis model based on the Transformer architecture through natural language processing technology, and analyze the text data through the text sentiment analysis model to obtain the text sentiment analysis status of the target user; The multimodal fusion submodule is used to calculate the weight values of the facial expression recognition submodule, the speech emotion recognition submodule, and the text emotion analysis submodule, and fuse the facial expression recognition state, the speech emotion recognition state, and the text emotion analysis state according to the weight values to obtain the comprehensive emotional state of the target user.
3. The system according to claim 2, characterized in that The use of deep learning algorithms to construct an expression recognition algorithm model includes: Pre-collecting a historical facial expression image dataset under different user conditions, and expanding the historical facial expression image dataset using data augmentation technology, wherein the user conditions include race, age, gender, cultural background, lighting conditions, angle, and posture; Build an initial expression recognition algorithm model and perform feature extraction and classification on the expanded facial expression image dataset by sequentially adding convolutional layers, pooling layers, and fully connected layers to obtain historical feature data for training. The initial expression recognition algorithm model is trained using the historical feature data to obtain a trained expression recognition algorithm model.
4. The system according to claim 2, wherein: The rhythmic features include pitch, duration, and volume changes.
5. The system according to claim 1, wherein: Positioning, cropping and preprocessing the image data include: Construct feature map pyramid layers P2 to P6, and determine classification networks C2 to C5 corresponding to the feature map pyramid layers P2 to P5 based on a preset data set, determine P2 to P5 through the output of the ResNet residual stage of C2 to C5, and determine P6 by performing a preset step-size convolution calculation at C5; Replace the convolution in each feature map pyramid layer with a deformable convolutional network; Anchor points of preset proportions are set on each feature map pyramid level from P2 to P6 so that the anchor points cover the image data, and then the image data is positioned, cropped, and preprocessed using the anchor points.
Citation Information
Cited By
AI question and answer expert model construction method and system based on psychological counseling and medium
CN121171262A
Psychological counseling-based ai question and answer expert model construction method and system, and medium
CN121171262B