Emotion regulation instruction recommendation method, device, equipment and medium

By generating dialogue semantics, micro-expressions, and speech representation vectors from psychological counseling data, and performing comprehensive recognition and recommendation, the problem of low accuracy of emotion regulation instructions in existing technologies has been solved, achieving more accurate emotion regulation.

CN116776235BActive Publication Date: 2026-01-13PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310739433.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-20
Publication Date
2026-01-13
Estimated Expiration
2043-06-20

AI Technical Summary

Technical Problem

Existing technologies, when using artificial intelligence for psychological counseling data analysis, neglect the characteristic correlation information of psychosemantics, micro-expressions, and voice tone, resulting in low accuracy of emotion regulation instructions.

Method used

Pre-trained dialogue semantic representation models, facial expression representation models, and speech representation models are used to generate dialogue semantic, micro-expression, and speech representation vectors. An abnormal emotion comprehensive recognition model is used to generate comprehensive representation vectors, and an emotion regulation instruction recommendation model is used to recommend emotion regulation instructions.

Benefits of technology

It improved the accuracy of recommended emotion regulation instructions and enhanced the effectiveness of psychological counseling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116776235B_ABST
    Figure CN116776235B_ABST
Patent Text Reader

Abstract

The application relates to the fields of artificial intelligence and medical health technology, and discloses a mood regulation instruction recommendation method, device, equipment and medium, wherein the method comprises the following steps: adopting a pre-trained dialogue semantic representation model to generate a dialogue semantic representation vector of psychological counseling data; adopting a pre-trained facial expression representation model to generate a micro-expression representation vector of the psychological counseling data; adopting a pre-trained voice representation model to generate a voice representation vector of the psychological counseling data; adopting a pre-trained abnormal mood comprehensive recognition model to generate a comprehensive representation vector of the dialogue semantic representation vector, the micro-expression representation vector and the voice representation vector; and adopting a pre-trained mood regulation instruction recommendation model to recommend a mood regulation instruction for the comprehensive representation vector and a personal information feature vector corresponding to a counseling object, so as to obtain a recommendation result. Therefore, the accuracy of the recommendation result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and medical health technology, and in particular to a method, apparatus, device and medium for recommending emotion regulation instructions. Background Technology

[0002] With the development of Artificial Intelligence (AI), AI technology can provide users with increasing convenience. For example, in the healthcare field, AI is used to analyze psychological counseling data, providing users with more convenient services. Currently, when using AI to analyze psychological counseling data, abnormal emotions are predicted based on the features of psychosemantic characteristics, micro-expressions, and voice intonation, and the predicted abnormal emotions are used as the final output. This ignores the influence of the emotional correlation information of psychosemantic features, micro-expressions, and voice intonation on emotion regulation instructions, resulting in low accuracy of emotion regulation instructions. Summary of the Invention

[0003] Based on this, it is necessary to address the technical problem that existing technologies, when using artificial intelligence to analyze psychological counseling data, neglect the influence of emotional correlation information such as psychosemantic features, micro-expression features, and voice tone features on emotion regulation instructions, resulting in low accuracy of emotion regulation instructions. Therefore, a method, device, equipment, and medium for recommending emotion regulation instructions are proposed.

[0004] Firstly, a method for recommending emotion regulation instructions is provided, the method comprising:

[0005] Obtain the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee conducting psychological counseling with the counselor;

[0006] A pre-trained dialogue semantic representation model is used to generate dialogue semantic representation vectors from the psychological counseling data;

[0007] A pre-trained facial expression representation model is used to generate micro-expression representation vectors from the psychological counseling data;

[0008] A pre-trained speech representation model is used to generate speech representation vectors from the psychological counseling data;

[0009] A pre-trained abnormal emotion comprehensive recognition model is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector;

[0010] A pre-trained emotion regulation instruction recommendation model is used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, and the recommendation results are obtained.

[0011] Secondly, a device for recommending emotion regulation instructions is provided, the device comprising:

[0012] The data acquisition module is used to acquire the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee conducting psychological counseling with the counselor.

[0013] The first processing module is used to generate dialogue semantic representation vectors from the psychological counseling data using a pre-trained dialogue semantic representation model.

[0014] The second processing module is used to generate micro-expression vectors from the psychological counseling data using a pre-trained facial expression representation model.

[0015] The third processing module is used to generate speech representation vectors from the psychological counseling data using a pre-trained speech representation model.

[0016] The fourth processing module is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector using a pre-trained abnormal emotion comprehensive recognition model;

[0017] The recommendation module is used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object using a pre-trained emotion regulation instruction recommendation model, and obtain the recommendation result.

[0018] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the recommended method for the aforementioned emotion regulation instructions.

[0019] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the recommended method for the above-described emotion regulation instructions.

[0020] The emotion regulation instruction recommendation method of this application employs a pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector. Then, a pre-trained emotion regulation instruction recommendation model is used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, resulting in a recommendation result. By using the pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector, the method extracts emotional association information from the features of psychosemantic meaning, micro-expression, and speech tone. This allows the recommendation result to be determined based on the emotional association information of the features of psychosemantic meaning, micro-expression, and speech tone, thereby improving the accuracy of the recommendation result and enhancing the effectiveness of psychological counseling. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] in:

[0023] Figure 1 This is a diagram illustrating the application environment of a method for recommending emotion regulation instructions in one embodiment.

[0024] Figure 2 This is a flowchart of a method for recommending emotion regulation instructions in one embodiment;

[0025] Figure 3 This is a structural block diagram of a device for recommending emotion regulation instructions in one embodiment;

[0026] Figure 4 This is a structural block diagram of a computer device in one embodiment;

[0027] Figure 5 This is a structural block diagram of a computer device in one embodiment. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] The method for recommending emotion regulation instructions provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, client 110 communicates with server 120 via a network. Server 120 can obtain psychological counseling data corresponding to the current time window through client 110. This psychological counseling data is obtained from psychological counseling between the client and the counselor. Server 120 uses a pre-trained dialogue semantic representation model to generate dialogue semantic representation vectors from the psychological counseling data; a pre-trained facial expression representation model to generate micro-expression representation vectors from the psychological counseling data; a pre-trained speech representation model to generate speech representation vectors from the psychological counseling data; a pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vectors, micro-expression representation vectors, and speech representation vectors; and a pre-trained emotion regulation instruction recommendation model to recommend emotion regulation instructions based on the comprehensive representation vectors and the personal information feature vectors corresponding to the counselor, obtaining recommendation results. The recommendation results are then fed back to client 110; client 110 displays the recommendation results via interface or voice. By employing a pre-trained abnormal emotion comprehensive recognition model, the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector are used to generate a comprehensive representation vector that extracts emotional association information from the features of psychosemantics, micro-expressions, and speech intonation. This allows the recommendation results to be determined based on the emotional association information of the features of psychosemantics, micro-expressions, and speech intonation, thereby improving the accuracy of the recommendation results and the effectiveness of psychological counseling.

[0030] Optionally, client 110 acquires psychological counseling data corresponding to the current time window, wherein the psychological counseling data is data obtained from psychological counseling between the counselee and the counselor; client 110 uses a pre-trained dialogue semantic representation model to generate dialogue semantic representation vectors from the psychological counseling data, uses a pre-trained facial expression representation model to generate micro-expression representation vectors from the psychological counseling data, uses a pre-trained speech representation model to generate speech representation vectors from the psychological counseling data, and uses a pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector. A pre-trained emotion regulation instruction recommendation model is used to recommend emotion regulation instructions to the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, resulting in a recommendation result. The recommendation result is then displayed on an interface or via voice. By employing a pre-trained abnormal emotion comprehensive recognition model, the comprehensive representation vector generated from the dialogue semantic representation vector, the micro-expression representation vector, and the voice representation vector extracts emotion association information from the features of psychosemantic features, micro-expression features, and voice intonation features. This allows the recommendation result to be determined based on the emotion association information of psychosemantic features, micro-expression features, and voice intonation features, thereby improving the accuracy of the recommendation result and the effectiveness of psychological counseling.

[0031] The client 110 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server 120 can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0032] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a method for recommending emotion regulation instructions provided in an embodiment of the present invention includes the following steps:

[0033] S1: Obtain the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee from conducting psychological counseling with the counselor;

[0034] The person receiving counseling is the recipient of psychological counseling services, such as a consulting doctor. The person receiving counseling services refers to a mentally healthy individual or someone with poor mental health but no mental disorder.

[0035] During the psychological counseling process between clients, counseling data is acquired according to time windows. The current time window is the real-time time window during the psychological counseling process.

[0036] Specifically, it can obtain psychological counseling data corresponding to the current time window input by the user, psychological counseling data corresponding to the current time window from storage space, psychological counseling data corresponding to the current time window sent by third-party applications, psychological counseling data corresponding to the current time window sent by the client, and psychological counseling data corresponding to the current time window sent by counseling assistance devices.

[0037] Understandably, by obtaining the psychological counseling data corresponding to the current time window at its end, it is possible to recommend emotion regulation instructions based on the psychological counseling data. This helps to assist the counselee in timely intervention of the counselor's abnormal emotions and improves the effectiveness of psychological counseling.

[0038] S2: Using a pre-trained dialogue semantic representation model, dialogue semantic representation vectors are generated from the psychological counseling data;

[0039] Specifically, a pre-trained dialogue semantic representation model is used to generate dialogue semantic representation vectors for the dialogue text corresponding to the audio data segments in the psychological counseling data.

[0040] The dialogue semantic representation model is a model obtained by training a method based on the classification and prediction of abnormal emotions.

[0041] S3: Using a pre-trained facial expression representation model, micro-expression representation vectors are generated from the psychological counseling data;

[0042] Specifically, a pre-trained facial expression representation model is used to generate micro-expression representation vectors from the facial video segments in the psychological counseling data.

[0043] The facial expression representation model is a training method based on the classification and prediction of abnormal emotions.

[0044] S4: Using a pre-trained speech representation model, speech representation vectors are generated from the psychological counseling data;

[0045] Specifically, a pre-trained speech representation model is used to generate speech representation vectors for audio data segments in the psychological counseling data.

[0046] The speech representation model is a training method based on the classification and prediction of abnormal emotions.

[0047] S5: A pre-trained abnormal emotion comprehensive recognition model is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector;

[0048] Specifically, a pre-trained abnormal emotion comprehensive recognition model is used to generate a comprehensive representation vector based on psychosemantic features, micro-expression features, and speech tone features from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector. This extracts emotion-related information by considering the psychosemantic features, micro-expression features, and speech tone features.

[0049] The abnormal emotion comprehensive recognition model is a model obtained based on the training method of abnormal emotion classification and prediction.

[0050] S6: Using a pre-trained emotion regulation instruction recommendation model, emotion regulation instructions are recommended based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, and the recommendation results are obtained.

[0051] Specifically, the comprehensive representation vector and the personal information feature vector corresponding to the client are input into an emotion regulation instruction recommendation model to recommend emotion regulation instructions. The recommendation result is determined based on the data output by the emotion regulation instruction recommendation model. The client then determines an emotion regulation plan based on the recommendation result.

[0052] Optionally, the data output by the emotion regulation instruction recommendation model can be used as the recommendation result.

[0053] Emotion regulation instructions are directives for regulating emotions. Emotion regulation is the process by which an individual, based on monitoring and evaluating their emotions in response to internal and external environmental demands, employs specific behavioral strategies to influence and control those emotions. It is a functional response by which an individual maintains adaptation to both internal and external environments.

[0054] The emotion regulation instruction recommendation model is a classification model.

[0055] This embodiment employs a pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector. A pre-trained emotion regulation instruction recommendation model is then used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, resulting in a recommendation result. By using the pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector, the emotional association information of psychosemantic features, micro-expression features, and speech tone features is extracted. This allows the recommendation result to be determined based on the emotional association information of psychosemantic features, micro-expression features, and speech tone features, improving the accuracy of the recommendation result and enhancing the effectiveness of psychological counseling.

[0056] In one embodiment, the step of generating dialogue semantic representation vectors from the psychological counseling data using a pre-trained dialogue semantic representation model includes:

[0057] S21: Process the dialogue text corresponding to the counselor based on the audio data segment in the psychological counseling data to obtain the single-window dialogue text;

[0058] Optionally, a pre-trained speech-to-text model is used to convert the audio data segments in the psychological counseling data into speech and text. The dialogue text corresponding to the counselor is extracted from the converted text and used as the single-window dialogue text.

[0059] The pre-trained speech-to-text model is a model trained based on ASR (Automatic Speech Recognition) technology.

[0060] It is understandable that a model trained on a neural network can be used to separate the speech data corresponding to the counselor from the audio data segments in the psychological counseling data, and then a pre-trained speech-to-text model can be used to convert the separated speech data into text, and the converted text can be used as the single-window dialogue text.

[0061] S22: Input the single-window dialogue text into the dialogue semantic representation model to classify and predict abnormal emotions;

[0062] Specifically, the single-window dialogue text is input into the dialogue semantic representation model for abnormal emotion classification and prediction, wherein the dialogue semantic representation model performs semantic understanding and representation of the single-window dialogue text.

[0063] S23: Obtain the vector output by the flag bit of the dialogue semantic representation model, and use it as the dialogue semantic representation vector;

[0064] The flag bit of the dialogue semantic representation model is referred to as CLS-token.

[0065] The training steps of the dialogue semantic representation model include:

[0066] Obtain a first model and a first training sample set. The first model includes a BERT model, a first fully connected layer, and a first activation layer. The output of the flag bit of the BERT model is used as the input of the first fully connected layer, and the output of the first fully connected layer is used as the input of the first activation layer.

[0067] The first model is trained using the first training sample set to classify and predict abnormal emotions, and the BERT model in the first model after training is completed is used as the dialogue semantic representation model.

[0068] The first model is a multi-class classification model.

[0069] The BERT model stands for Bidirectional Encoder Representations from Transformers.

[0070] The first fully connected layer is used. The first activation layer is an activation layer using the sigmoid activation function, or an activation layer using the softmax activation function. The first fully connected layer and the first activation layer perform classification and prediction of various abnormal emotions based on the output of the BERT model's flag bits.

[0071] The sigmoid activation function is a common sigmoid function in biology, also known as an sigmoid growth curve. In information science, due to its monotonically increasing properties and the monotonically increasing properties of its inverse function, the sigmoid function is often used as the activation function of neural networks, mapping variables to the range of 0-1.

[0072] The Softmax activation function is often used as an activation function in the output layer of a neural network. It maps the values ​​of the output layer to the 0-1 interval, constructing a probability distribution of the neuron output. It is used in multi-class classification problems. The larger the value mapped by the Softmax activation function, the greater the probability of the true class.

[0073] The first training sample set includes multiple first training samples. The first training samples include: dialogue text samples and first abnormal emotion labeling data. The first abnormal emotion labeling data is the accurate labeling result of the abnormal emotion corresponding to the dialogue text samples.

[0074] The specific steps for training the first model to classify and predict abnormal emotions using the first training sample set are not detailed here.

[0075] The first model that has completed training refers to the first model whose loss value converges to a first preset value or the first model's training iterations reach a first preset threshold.

[0076] In this embodiment, the first training sample set is used to train the first model to classify and predict abnormal emotions. The BERT model in the first model after training is completed is used as the dialogue semantic representation model, so that the vector output by the dialogue semantic representation model is a representation vector based on the classification and prediction of abnormal emotions.

[0077] In one embodiment, the step of generating micro-expression vectors from the psychological counseling data using a pre-trained facial expression representation model includes:

[0078] S31: Extract key image frames from the face video segments in the psychological counseling data to obtain a key image frame set;

[0079] Key image frames are those with significant micro-expression changes.

[0080] Specifically, the image frames corresponding to the facial video segments in the psychological counseling data are input into the key image frame recognition model to classify and predict micro-expression changes, thereby obtaining a prediction vector. The key image frame recognition model is a model trained based on a 3D-CNN network. The image frame corresponding to each vector element in the prediction vector whose value is greater than a preset probability threshold is taken as the key image frame.

[0081] The key image frame recognition model is a multi-class classification model. It assigns a probability to each image frame as a key image frame. 3D-CNN networks, or 3D convolutional neural networks, are primarily used to address the correlation information between images, adding a new dimension of information.

[0082] Each element in the prediction vector corresponds to an image frame number. The value of each element in the prediction vector is the probability that the image frame corresponding to the image frame number is a key image frame. The higher the probability, the greater the likelihood that the image frame is a key image frame.

[0083] 3D-CNN networks are 3D convolutional neural networks, primarily designed to address the correlation between images by adding a new dimension of information.

[0084] S32: Input the key image frame set into the facial expression representation model to classify and predict abnormal emotions;

[0085] Specifically, the key image frame set is input into the facial expression representation model for abnormal emotion classification and prediction, wherein the facial expression representation model performs vector representation based on abnormal emotion classification and prediction.

[0086] S33: Obtain the vector output by the fully connected layer of the facial expression representation model, and use it as the micro-expression representation vector;

[0087] The training steps of the facial expression representation model include:

[0088] Obtain a second model and a second training sample set. The second model includes a 3D-CNN model, a second fully connected layer, and a second activation layer connected in sequence.

[0089] The second model is trained using the second training sample set to classify and predict abnormal emotions, and the 3D-CNN model and the second fully connected layer in the trained second model are used as the facial expression representation model.

[0090] The second model is a multi-class classification model.

[0091] The 3D-CNN model is a three-dimensional convolutional neural network.

[0092] The second fully connected layer is used. The second activation layer is an activation layer using the sigmoid activation function, or an activation layer using the softmax activation function. The second fully connected layer and the second activation layer perform classification and prediction of various abnormal emotions based on the output of the 3D-CNN model.

[0093] The second training sample set includes multiple second training samples. These include: face video data samples and second abnormal emotion labeling data, where the second abnormal emotion labeling data is the accurate labeling result of the abnormal emotion corresponding to the face video data samples.

[0094] The specific steps for training the second model to classify and predict abnormal emotions using the second training sample set are not detailed here.

[0095] The second model that has completed training refers to the loss value of the second model converging to a second preset value or the number of training iterations of the second model reaching a second preset threshold.

[0096] In this embodiment, the second training sample set is used to train the second model to classify and predict abnormal emotions. The 3D-CNN model and the second fully connected layer in the trained second model are used as the facial expression representation model, so that the vector output by the facial expression representation model is a representation vector based on the classification and prediction of abnormal emotions.

[0097] In one embodiment, the step of generating speech representation vectors from the psychological counseling data using a pre-trained speech representation model includes:

[0098] S41: Input the audio data segment of the psychological counseling data into the speech representation model to classify and predict abnormal emotions;

[0099] Specifically, the audio data segments of the psychological counseling data are input into the speech representation model for the classification and prediction of abnormal emotions, wherein the speech representation model performs vector representation based on the classification and prediction of abnormal emotions.

[0100] S42: Obtain the vector output by the flag bit of the speech representation model, and use it as the speech representation vector;

[0101] Specifically, the flag bit of the speech representation model is referred to as CLS-token.

[0102] The training steps of the speech representation model include:

[0103] Obtain a third model and a third training sample set. The third model includes a HuberT model, a third fully connected layer, and a third activation layer. The vector output by the flag bit of the HuberT model is used as the input of the third fully connected layer, and the output of the third fully connected layer is used as the input of the third activation layer.

[0104] The third training sample set is used to train the third model for classifying and predicting abnormal emotions, and the HuBERT model in the trained third model is used as the speech representation model.

[0105] The third model is a multi-class classification model.

[0106] The HuBERT model is Self-Supervised Speech Representation Learning by MaskedPrediction of Hidden Units.

[0107] The third fully connected layer is used. The third activation layer is an activation layer using the sigmoid activation function, or an activation layer using the softmax activation function. The third fully connected layer and the third activation layer perform classification and prediction of various abnormal emotions based on the output of the HuberT model.

[0108] The third training sample set includes multiple third training samples. The third training samples include: audio data samples and third abnormal emotion labeling data. The third abnormal emotion labeling data is the accurate labeling result of the abnormal emotion corresponding to the audio data samples.

[0109] The specific steps for training the third model to classify and predict abnormal emotions using the third training sample set are not elaborated here.

[0110] The third model that has completed training refers to the third model whose loss value converges to a third preset value or whose training iterations reach a third preset threshold.

[0111] In this embodiment, the third training sample set is used to train the third model for the classification and prediction of abnormal emotions. The HuBERT model in the third model after training is used as the speech representation model, so that the vector output by the speech representation model is a representation vector based on the classification and prediction of abnormal emotions.

[0112] In one embodiment, the step of generating a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector using a pre-trained abnormal emotion comprehensive recognition model includes:

[0113] S51: Perform average pooling operation on the dialogue semantic representation vector, the micro-expression representation vector and the speech representation vector respectively to obtain the dialogue semantic feature vector, the micro-expression feature vector and the speech feature vector, wherein the dialogue semantic feature vector, the micro-expression feature vector and the speech feature vector are all the same size;

[0114] Specifically, the dialogue semantic representation vector is subjected to average pooling to obtain the dialogue semantic feature vector; the micro-expression representation vector is subjected to average pooling to obtain the micro-expression feature vector; and the speech representation vector is subjected to average pooling to obtain the speech feature vector.

[0115] The fact that the dialogue semantic feature vector, the micro-expression feature vector, and the voice feature vector are all the same in size means that the vector widths of the dialogue semantic feature vector, the micro-expression feature vector, and the voice feature vector are the same, the vector heights of the dialogue semantic feature vector, the micro-expression feature vector, and the voice feature vector are the same, and the vector dimensions of the dialogue semantic feature vector, the micro-expression feature vector, and the voice feature vector are the same.

[0116] S52: Using preset splicing symbols, the dialogue semantic feature vector, the micro-expression feature vector, and the speech feature vector are spliced ​​together to obtain a first spliced ​​vector;

[0117] Optionally, the preset concatenation symbol is ";". For example, the first concatenation vector z = (u; v; w), where ; is the preset concatenation symbol, u is the dialogue semantic feature vector, v is the micro-expression feature vector, and w is the speech feature vector.

[0118] S53: Input the first concatenated vector into the abnormal emotion comprehensive recognition model to classify and predict abnormal emotions;

[0119] Specifically, the first concatenated vector is input into the abnormal emotion comprehensive recognition model for abnormal emotion classification and prediction, wherein the abnormal emotion comprehensive recognition model performs vector representation based on abnormal emotion classification and prediction.

[0120] S54: Obtain the vector output by the abnormal emotion comprehensive recognition model, and use it as the comprehensive representation vector;

[0121] The training steps of the abnormal emotion comprehensive recognition model include:

[0122] Obtain a fourth model and a fourth training sample set, wherein the fourth model comprises a multi-layer feedforward neural network and a fourth activation layer connected in sequence;

[0123] The fourth training sample set is used to train the fourth model for the classification and prediction of abnormal emotions, and the multi-layer feedforward neural network in the fourth model after training is completed is used as the comprehensive abnormal emotion recognition model.

[0124] The fourth model is a multi-class classification model.

[0125] Multilayer feedforward neural network, abbreviated as FFN, stands for Feedforward Neural Network.

[0126] The fourth activation layer is either an activation layer using the sigmoid activation function or an activation layer using the softmax activation function. The fourth activation layer performs classification and prediction of various abnormal emotions based on the output of the multi-layer feedforward neural network of the fourth model.

[0127] The fourth training sample set includes multiple fourth training samples. These fourth training samples include: feature vector samples and fourth abnormal emotion labeling data. The fourth abnormal emotion labeling data is the accurate labeling result of the abnormal emotion corresponding to the feature vector samples. The feature vector samples are feature vectors extracted based on facial video data samples and audio data samples corresponding to the same object.

[0128] The specific steps for training the fourth model to classify and predict abnormal emotions using the fourth training sample set are not detailed here.

[0129] The fourth model that has completed training refers to the fourth model whose loss value converges to a fourth preset value or whose training iterations reach a fourth preset threshold.

[0130] In this embodiment, the fourth training sample set is used to train the fourth model for the classification and prediction of abnormal emotions. The multilayer feedforward neural network in the fourth model after training is used as the comprehensive abnormal emotion recognition model. Thus, the vector output by the multilayer feedforward neural network is a representation vector based on the classification and prediction of abnormal emotions.

[0131] In one embodiment, the step of using a pre-trained emotion regulation instruction recommendation model to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, and obtaining the recommendation result, includes:

[0132] S61: Concatenate the comprehensive representation vector and the personal information feature vector corresponding to the consultation object to obtain a second concatenated vector;

[0133] Specifically, a second concatenated vector is obtained by concatenating the comprehensive representation vector and the personal information feature vector corresponding to the consulting object using a preset concatenation symbol. For example, the second concatenated vector is h = (x; y), where ; is a preset concatenation symbol, x is the comprehensive representation vector, and y is the personal information feature vector corresponding to the consulting object.

[0134] S62: Input the second concatenated vector into the emotion regulation instruction recommendation model to recommend emotion regulation instructions and obtain a recommendation vector;

[0135] Each vector element in the recommendation vector corresponds to an emotion regulation instruction, and the value of each vector element is the probability value of recommending the emotion regulation instruction corresponding to that vector element.

[0136] S63: Select the vector elements with the largest value and a quantity of K from the recommended vectors, and use them as a set of vector elements, where K is an integer greater than 0;

[0137] Specifically, the vector elements with the largest value and a quantity of K are selected from the recommended vectors, and all selected vector elements are used as a vector element set.

[0138] S64: Use each emotion regulation instruction corresponding to the vector element set as the recommendation result;

[0139] The emotion regulation instruction recommendation model includes a multi-layer feedforward neural network and a fifth activation layer connected in sequence.

[0140] Specifically, the emotion regulation instruction corresponding to each vector element in the vector element set is taken as the recommendation instruction, and all the recommendation instructions are taken as the recommendation result.

[0141] The fifth activation layer performs classification and prediction of various abnormal emotions based on the output of the multilayer feedforward neural network of the emotion regulation instruction recommendation model.

[0142] This embodiment implements emotion regulation instruction recommendation based on psychosemantic features, micro-expression features, and voice tone features, improving the accuracy of recommendation results and enhancing the effectiveness of psychological counseling. By combining the comprehensive representation vector with the personal information feature vector corresponding to the counseling subject, the recommendation of emotion regulation instructions is performed, taking into account both the data and personal information aspects of the psychological counseling process, thus improving the accuracy of the recommendation results.

[0143] In one embodiment, before the step of using a pre-trained emotion regulation instruction recommendation model to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, and obtaining the recommendation result, the method further includes:

[0144] S71: Obtain the structured data of the personal information of the person being consulted;

[0145] Specifically, it can obtain structured personal information data corresponding to the consultation object input by the user, obtain structured personal information data corresponding to the consultation object from storage space, obtain structured personal information data corresponding to the consultation object sent by a third-party application, obtain structured personal information data corresponding to the consultation object sent by the client, or obtain structured personal information data corresponding to the consultation object sent by the consultation assistance device.

[0146] The structured data of the personal information corresponding to the consultation recipient is a structured representation of the personal information of the consultation recipient. Personal information includes, but is not limited to: age, gender, occupation, education level, and city of residence.

[0147] S72: Input the structured data of the personal information into a pre-trained autoencoder network model to represent the high-dimensional sparse data into a low-dimensional compact vector, thereby obtaining the feature vector of the personal information corresponding to the consultation object.

[0148] An auto-encoder network model is a model where the output is equal to the input. It is a typical unsupervised learning model. The input data undergoes a series of feature transformations within the network model, but the output remains the same as the input.

[0149] Low-dimensional compact vectors, also known as low-dimensional dense vectors, can be viewed as features that are carefully selected data with high quality. Low dimensionality corresponds to a small number of variables, and density corresponds to the data having almost no missing values ​​and outliers.

[0150] The structured data of the personal information mentioned is high-dimensional sparse data.

[0151] This embodiment inputs the structured personal information data into a pre-trained autoencoder network model to represent high-dimensional sparse data as low-dimensional compact vectors, reducing the vector dimension and improving the quality of the data input into the emotion regulation instruction recommendation model, thus providing a foundation for improving the recommendation performance of the emotion regulation instruction recommendation model.

[0152] Please see Figure 3 As shown, in one embodiment, a device for recommending emotion regulation instructions is provided, the device comprising:

[0153] The data acquisition module 801 is used to acquire the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee in conducting psychological counseling with the counselor.

[0154] The first processing module 802 is used to generate dialogue semantic representation vectors from the psychological counseling data using a pre-trained dialogue semantic representation model.

[0155] The second processing module 803 is used to generate micro-expression vectors from the psychological counseling data using a pre-trained facial expression representation model.

[0156] The third processing module 804 is used to generate speech representation vectors from the psychological counseling data using a pre-trained speech representation model.

[0157] The fourth processing module 805 is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector using a pre-trained abnormal emotion comprehensive recognition model;

[0158] The recommendation module 806 is used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object using a pre-trained emotion regulation instruction recommendation model, and obtain the recommendation result.

[0159] This embodiment employs a pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector. A pre-trained emotion regulation instruction recommendation model is then used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, resulting in a recommendation result. By using the pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector, the emotional association information of psychosemantic features, micro-expression features, and speech tone features is extracted. This allows the recommendation result to be determined based on the emotional association information of psychosemantic features, micro-expression features, and speech tone features, improving the accuracy of the recommendation result and enhancing the effectiveness of psychological counseling.

[0160] In one embodiment, the step of the first processing module 802 generating dialogue semantic representation vectors from the psychological counseling data using a pre-trained dialogue semantic representation model includes:

[0161] Based on the audio data segments in the psychological counseling data, the corresponding dialogue text of the counseling object is processed to obtain the single-window dialogue text;

[0162] The single-window dialogue text is input into the dialogue semantic representation model for abnormal emotion classification and prediction;

[0163] Obtain the vector output by the flag bit of the dialogue semantic representation model, and use it as the dialogue semantic representation vector;

[0164] The training steps of the dialogue semantic representation model include:

[0165] Obtain a first model and a first training sample set. The first model includes a BERT model, a first fully connected layer, and a first activation layer. The output of the flag bit of the BERT model is used as the input of the first fully connected layer, and the output of the first fully connected layer is used as the input of the first activation layer.

[0166] The first model is trained using the first training sample set to classify and predict abnormal emotions, and the BERT model in the first model after training is completed is used as the dialogue semantic representation model.

[0167] In one embodiment, the step of generating micro-expression vectors from the psychological counseling data using a pre-trained facial expression representation model in the second processing module 803 includes:

[0168] Key image frames are extracted from the facial video segments in the psychological counseling data to obtain a set of key image frames.

[0169] The key image frame set is input into the facial expression representation model for the classification and prediction of abnormal emotions;

[0170] Obtain the vector output by the fully connected layer of the facial expression representation model, and use it as the micro-expression representation vector;

[0171] The training steps of the facial expression representation model include:

[0172] Obtain a second model and a second training sample set. The second model includes a 3D-CNN model, a second fully connected layer, and a second activation layer connected in sequence.

[0173] The second model is trained using the second training sample set to classify and predict abnormal emotions, and the 3D-CNN model and the second fully connected layer in the trained second model are used as the facial expression representation model.

[0174] In one embodiment, the step of the third processing module 804 generating speech representation vectors from the psychological counseling data using a pre-trained speech representation model includes:

[0175] The audio data segments of the psychological counseling data are input into the speech representation model for the classification and prediction of abnormal emotions;

[0176] Obtain the vector output by the flag bits of the speech representation model, and use it as the speech representation vector;

[0177] The training steps of the speech representation model include:

[0178] Obtain a third model and a third training sample set. The third model includes a HuberT model, a third fully connected layer, and a third activation layer. The vector output by the flag bit of the HuberT model is used as the input of the third fully connected layer, and the output of the third fully connected layer is used as the input of the third activation layer.

[0179] The third training sample set is used to train the third model for classifying and predicting abnormal emotions, and the HuBERT model in the trained third model is used as the speech representation model.

[0180] In one embodiment, the step of the fourth processing module 805 generating a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector using a pre-trained abnormal emotion comprehensive recognition model includes:

[0181] Average pooling is performed on the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector respectively to obtain the dialogue semantic feature vector, the micro-expression feature vector, and the speech feature vector, wherein the dialogue semantic feature vector, the micro-expression feature vector, and the speech feature vector are all of the same size;

[0182] Using preset splicing symbols, the dialogue semantic feature vector, the micro-expression feature vector, and the speech feature vector are spliced ​​together to obtain a first spliced ​​vector;

[0183] The first concatenated vector is input into the abnormal emotion comprehensive recognition model to classify and predict abnormal emotions;

[0184] Obtain the vector output by the abnormal emotion comprehensive recognition model, and use it as the comprehensive representation vector;

[0185] The training steps of the abnormal emotion comprehensive recognition model include:

[0186] Obtain a fourth model and a fourth training sample set, wherein the fourth model comprises a multi-layer feedforward neural network and a fourth activation layer connected in sequence;

[0187] The fourth training sample set is used to train the fourth model for the classification and prediction of abnormal emotions, and the multi-layer feedforward neural network in the fourth model after training is completed is used as the comprehensive abnormal emotion recognition model.

[0188] In one embodiment, the step of the recommendation module 806 using a pre-trained emotion regulation instruction recommendation model to perform emotion regulation instruction recommendation on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, and obtaining the recommendation result, includes:

[0189] The comprehensive representation vector and the personal information feature vector corresponding to the consultation object are concatenated to obtain the second concatenated vector;

[0190] The second concatenated vector is input into the emotion regulation instruction recommendation model to recommend emotion regulation instructions, and a recommendation vector is obtained.

[0191] Select the vector elements with the largest value and a quantity of K from the recommended vectors, and use them as the vector element set, where K is an integer greater than 0;

[0192] The emotion regulation instructions corresponding to the vector element set are used as the recommendation results;

[0193] The emotion regulation instruction recommendation model includes a multi-layer feedforward neural network and a fifth activation layer connected in sequence.

[0194] In one embodiment, the apparatus further includes:

[0195] The personal information feature vector generation module is used to obtain the structured personal information data corresponding to the consultation object, input the structured personal information data into a pre-trained autoencoder network model to represent high-dimensional sparse data into low-dimensional compact vectors, and obtain the personal information feature vector corresponding to the consultation object.

[0196] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements a recommended method for emotion regulation instructions, a server-side function, or steps.

[0197] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements a recommended method for emotion regulation instructions, providing client-side functions or steps.

[0198] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the following steps:

[0199] Obtain the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee conducting psychological counseling with the counselor;

[0200] A pre-trained dialogue semantic representation model is used to generate dialogue semantic representation vectors from the psychological counseling data;

[0201] A pre-trained facial expression representation model is used to generate micro-expression representation vectors from the psychological counseling data;

[0202] A pre-trained speech representation model is used to generate speech representation vectors from the psychological counseling data;

[0203] A pre-trained abnormal emotion comprehensive recognition model is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector;

[0204] A pre-trained emotion regulation instruction recommendation model is used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, and the recommendation results are obtained.

[0205] This embodiment employs a pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector. A pre-trained emotion regulation instruction recommendation model is then used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, resulting in a recommendation result. By using the pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector, the emotional association information of psychosemantic features, micro-expression features, and speech tone features is extracted. This allows the recommendation result to be determined based on the emotional association information of psychosemantic features, micro-expression features, and speech tone features, improving the accuracy of the recommendation result and enhancing the effectiveness of psychological counseling.

[0206] In one embodiment, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, performs the following steps:

[0207] Obtain the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee conducting psychological counseling with the counselor;

[0208] A pre-trained dialogue semantic representation model is used to generate dialogue semantic representation vectors from the psychological counseling data;

[0209] A pre-trained facial expression representation model is used to generate micro-expression representation vectors from the psychological counseling data;

[0210] A pre-trained speech representation model is used to generate speech representation vectors from the psychological counseling data;

[0211] A pre-trained abnormal emotion comprehensive recognition model is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector;

[0212] A pre-trained emotion regulation instruction recommendation model is used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, and the recommendation results are obtained.

[0213] This embodiment employs a pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector. A pre-trained emotion regulation instruction recommendation model is then used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, resulting in a recommendation result. By using the pre-trained abnormal emotion comprehensive recognition model to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector, the emotional association information of psychosemantic features, micro-expression features, and speech tone features is extracted. This allows the recommendation result to be determined based on the emotional association information of psychosemantic features, micro-expression features, and speech tone features, improving the accuracy of the recommendation result and enhancing the effectiveness of psychological counseling.

[0214] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0215] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0216] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0217] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for recommending emotion regulation instructions, the method comprising: Obtain the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee conducting psychological counseling with the counselor; A pre-trained dialogue semantic representation model is used to generate dialogue semantic representation vectors from the psychological counseling data; A pre-trained facial expression representation model is used to generate micro-expression representation vectors from the psychological counseling data; A pre-trained speech representation model is used to generate speech representation vectors from the psychological counseling data; A pre-trained abnormal emotion comprehensive recognition model is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector; A pre-trained emotion regulation instruction recommendation model is used to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, and the recommendation results are obtained. The step of generating micro-expression vectors from the psychological counseling data using a pre-trained facial expression representation model includes: Key image frames are extracted from the facial video segments in the psychological counseling data to obtain a set of key image frames. The key image frames are image frames with large micro-expression changes. The key image frame set is input into the facial expression representation model for the classification and prediction of abnormal emotions; Obtain the vector output by the fully connected layer of the facial expression representation model, and use it as the micro-expression representation vector; The training steps of the facial expression representation model include: Obtain a second model and a second training sample set. The second model includes a 3D-CNN model, a second fully connected layer, and a second activation layer connected in sequence. The second training sample set is used to train the second model to classify and predict abnormal emotions, and the 3D-CNN model and the second fully connected layer in the trained second model are used as the facial expression representation model. Extracting key image frames from the facial video segments in the psychological counseling data to obtain a key image frame set includes: inputting each image frame corresponding to the facial video segment in the psychological counseling data into a key image frame recognition model to classify and predict micro-expression changes, and obtaining a prediction vector; and taking the image frame corresponding to each vector element in the prediction vector whose value is greater than a preset probability threshold as a key image frame.

2. The method for recommending emotion regulation instructions according to claim 1, characterized in that, The step of generating dialogue semantic representation vectors from the psychological counseling data using a pre-trained dialogue semantic representation model includes: Based on the audio data segments in the psychological counseling data, the corresponding dialogue text of the counseling object is processed to obtain the single-window dialogue text; The single-window dialogue text is input into the dialogue semantic representation model for abnormal emotion classification and prediction; Obtain the vector output by the flag bit of the dialogue semantic representation model, and use it as the dialogue semantic representation vector; The training steps of the dialogue semantic representation model include: Obtain a first model and a first training sample set. The first model includes a BERT model, a first fully connected layer, and a first activation layer. The output of the flag bit of the BERT model is used as the input of the first fully connected layer, and the output of the first fully connected layer is used as the input of the first activation layer. The first model is trained using the first training sample set to classify and predict abnormal emotions, and the BERT model in the first model after training is completed is used as the dialogue semantic representation model.

3. The method for recommending emotion regulation instructions according to claim 1, characterized in that, The step of generating speech representation vectors from the psychological counseling data using a pre-trained speech representation model includes: The audio data segments of the psychological counseling data are input into the speech representation model for the classification and prediction of abnormal emotions; Obtain the vector output by the flag bits of the speech representation model, and use it as the speech representation vector; The training steps of the speech representation model include: Obtain a third model and a third training sample set. The third model includes a HuberT model, a third fully connected layer, and a third activation layer. The vector output by the flag bit of the HuberT model is used as the input of the third fully connected layer, and the output of the third fully connected layer is used as the input of the third activation layer. The third training sample set is used to train the third model for classifying and predicting abnormal emotions, and the HuBERT model in the trained third model is used as the speech representation model.

4. The method for recommending emotion regulation instructions according to claim 1, characterized in that, The step of generating a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector using a pre-trained abnormal emotion comprehensive recognition model includes: Average pooling is performed on the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector respectively to obtain the dialogue semantic feature vector, the micro-expression feature vector, and the speech feature vector, wherein the dialogue semantic feature vector, the micro-expression feature vector, and the speech feature vector are all of the same size; Using preset splicing symbols, the dialogue semantic feature vector, the micro-expression feature vector, and the speech feature vector are spliced ​​together to obtain a first spliced ​​vector; The first concatenated vector is input into the abnormal emotion comprehensive recognition model to classify and predict abnormal emotions; Obtain the vector output by the abnormal emotion comprehensive recognition model, and use it as the comprehensive representation vector; The training steps of the abnormal emotion comprehensive recognition model include: Obtain a fourth model and a fourth training sample set, wherein the fourth model comprises a multi-layer feedforward neural network and a fourth activation layer connected in sequence; The fourth training sample set is used to train the fourth model for the classification and prediction of abnormal emotions, and the multi-layer feedforward neural network in the fourth model after training is completed is used as the comprehensive abnormal emotion recognition model.

5. The method for recommending emotion regulation instructions according to claim 1, characterized in that, The step of using a pre-trained emotion regulation instruction recommendation model to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, and obtaining the recommendation result, includes: The comprehensive representation vector and the personal information feature vector corresponding to the consultation object are concatenated to obtain the second concatenated vector; The second concatenated vector is input into the emotion regulation instruction recommendation model to recommend emotion regulation instructions, and a recommendation vector is obtained. Select the vector elements with the largest value and a quantity of K from the recommended vectors, and use them as the vector element set, where K is an integer greater than 0; The emotion regulation instructions corresponding to the vector element set are used as the recommendation results; The emotion regulation instruction recommendation model includes a multi-layer feedforward neural network and a fifth activation layer connected in sequence.

6. The method for recommending emotion regulation instructions according to claim 1, characterized in that, Before the step of using a pre-trained emotion regulation instruction recommendation model to recommend emotion regulation instructions based on the comprehensive representation vector and the personal information feature vector corresponding to the consultation subject, and obtaining the recommendation result, the method further includes: Obtain the structured data of the personal information of the person being consulted; The structured data of the personal information is input into a pre-trained autoencoder network model to represent the high-dimensional sparse data as a low-dimensional compact vector, thereby obtaining the feature vector of the personal information corresponding to the consultation object.

7. A device for recommending emotion regulation instructions, characterized in that, The device includes: The data acquisition module is used to acquire the psychological counseling data corresponding to the current time window, wherein the psychological counseling data is the data obtained by the counselee conducting psychological counseling with the counselor. The first processing module is used to generate dialogue semantic representation vectors from the psychological counseling data using a pre-trained dialogue semantic representation model. The second processing module is used to generate micro-expression vectors from the psychological counseling data using a pre-trained facial expression representation model. The third processing module is used to generate speech representation vectors from the psychological counseling data using a pre-trained speech representation model. The fourth processing module is used to generate a comprehensive representation vector from the dialogue semantic representation vector, the micro-expression representation vector, and the speech representation vector using a pre-trained abnormal emotion comprehensive recognition model; The recommendation module is used to use a pre-trained emotion regulation instruction recommendation model to recommend emotion regulation instructions to the comprehensive representation vector and the personal information feature vector corresponding to the consultation object, and obtain the recommendation result. The step of generating micro-expression vectors from the psychological counseling data using a pre-trained facial expression representation model includes: Key image frames are extracted from the facial video segments in the psychological counseling data to obtain a set of key image frames. The key image frames are image frames with large micro-expression changes. The key image frame set is input into the facial expression representation model for the classification and prediction of abnormal emotions; Obtain the vector output by the fully connected layer of the facial expression representation model, and use it as the micro-expression representation vector; The training steps of the facial expression representation model include: Obtain a second model and a second training sample set. The second model includes a 3D-CNN model, a second fully connected layer, and a second activation layer connected in sequence. The second training sample set is used to train the second model to classify and predict abnormal emotions, and the 3D-CNN model and the second fully connected layer in the trained second model are used as the facial expression representation model. Extracting key image frames from the facial video segments in the psychological counseling data to obtain a key image frame set includes: inputting each image frame corresponding to the facial video segment in the psychological counseling data into a key image frame recognition model to classify and predict micro-expression changes, and obtaining a prediction vector; and taking the image frame corresponding to each vector element in the prediction vector whose value is greater than a preset probability threshold as a key image frame.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the recommended method of the emotion regulation instructions as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the recommended method of emotion regulation instructions as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Demand recommendation method and device, electronic equipment and storage medium

    CN112489808A

  • Speech expression multi-mode emotion recognition method based on weighted feature fusion

    CN115691559A

  • Three-mode sentiment analysis method based on long and short time feature and decision fusion

    CN115758218A