Millimeter-wave radar-based multimodal emotion recognition method and device for psychological counseling

Through the millimeter-wave radar-based multimodal emotion recognition method for psychological counseling, combined with millimeter-wave radar data and voice data, the privacy issues caused by camera collection are solved and more accurate emotion recognition is achieved.

CN119606379BActive Publication Date: 2025-09-19SHENZHEN QINGWEN DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411810066.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-09-19
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

When existing technologies use emotion recognition technology based on visual recognition in psychological counseling, they rely on cameras to capture facial images or videos, which causes privacy concerns among visitors, affects the natural expression of emotions, and thus affects the accuracy of emotion recognition results.

Method used

A multimodal emotion recognition method for psychological counseling based on millimeter-wave radar is adopted. By acquiring millimeter-wave radar data and user voice data, heart rate variability and voice emotion features are extracted. The preset data fusion strategy is used for time series alignment and feature fusion, and the data is input into a pre-trained emotion recognition model for emotional state classification.

Benefits of technology

It improves the accuracy of emotion recognition, avoids privacy concerns caused by the use of cameras, and ensures the naturalness of emotional expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119606379B_ABST
    Figure CN119606379B_ABST
Patent Text Reader

Abstract

The present invention discloses a millimeter-wave radar-based multimodal emotion recognition method and device for psychological counseling. The method obtains a low-frequency-high-frequency ratio feature sequence and a speech emotion feature sequence of a target user to be identified, performs temporal alignment and feature fusion, and obtains a multimodal feature vector set. The multimodal feature vector set is then input into an emotion recognition model to obtain an emotion state classification result corresponding to each multimodal feature vector in the multimodal feature vector set, thereby forming an output sequence of emotion state classification results. This embodiment of the present invention collects the current millimeter-wave radar dataset and the current user speech dataset of the target user to be identified, performs feature extraction, temporal alignment, and feature fusion on each dataset to obtain a multimodal feature vector set, and then inputs the resulting multimodal feature vector set into the emotion recognition model to obtain a recognition result, thereby improving the accuracy of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent decision-making technology of artificial intelligence, and in particular to a method and device for multimodal emotion recognition in psychological counseling based on millimeter-wave radar. Background Art

[0002] In fields such as psychological counseling and emotion recognition, a variety of technologies and solutions have been developed for emotion recognition. These include visual recognition-based emotion recognition technologies, which typically rely on cameras to capture facial images or videos of users and then use computer vision algorithms (such as facial landmark detection and facial expression recognition) to analyze facial expression changes to determine emotional state. While camera-based vision technology can capture subtle changes in facial expressions, its application in psychological counseling has significant limitations. The use of cameras can raise privacy concerns for clients and hinder their ability to express emotions naturally, thus affecting the accuracy of emotion recognition results.

[0003] Currently, voice-based emotion recognition technology is also being used. This technology uses extracted voice features combined with machine learning or deep learning models to determine emotional states. Voice emotion recognition has been applied in many fields, but distinguishing the source of voices in psychological counseling settings presents challenges. During counseling, the voices of counselors and clients must be accurately distinguished; otherwise, the accuracy of emotion analysis may be affected. Summary of the Invention

[0004] The embodiments of the present invention provide a multimodal emotion recognition method and device for psychological counseling based on millimeter-wave radar, which aims to solve the problem that when using emotion recognition technology based on visual recognition in the field of psychological counseling in the existing technology, it often relies on cameras to capture facial images or videos of users and uses computer vision algorithms to judge the user's emotional state. However, the use of cameras will bring privacy concerns to visitors, affect their natural expression of emotions, and thus affect the accuracy of emotion recognition results.

[0005] In a first aspect, an embodiment of the present invention provides a multimodal emotion recognition method for psychological counseling based on millimeter wave radar, which includes:

[0006] In response to a multimodal emotion recognition instruction, obtaining to-be-recognized user information corresponding to the multimodal emotion recognition instruction and a to-be-recognized target user corresponding to the to-be-recognized user information;

[0007] Obtaining a current millimeter-wave radar data set corresponding to the target user to be identified; wherein the current millimeter-wave radar data set is composed of millimeter-wave radar data collected by the millimeter-wave radar module within the target collection time interval corresponding to the multimodal emotion recognition instruction after the millimeter-wave radar module is aimed at the target user to be identified;

[0008] Acquire a current user voice data set corresponding to the target user to be identified; wherein the current user voice data set is composed of user voice data collected by the sound pickup and processing module from the target user to be identified within a target collection time interval;

[0009] Based on a preset millimeter-wave radar feature extraction strategy, a heart rate variability data sequence corresponding to the current millimeter-wave radar data set and a low-frequency and high-frequency ratio feature sequence corresponding to the heart rate variability data sequence are acquired in time sequence;

[0010] Based on a preset voice feature extraction model, a voice emotion feature sequence corresponding to the current user voice data set is obtained in time sequence;

[0011] Based on a preset data fusion strategy, each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence is time-aligned and feature-fused with each voice emotion feature in the voice emotion feature sequence to obtain a multimodal feature vector set corresponding to the target user to be identified; wherein the multimodal feature vector set includes multiple multimodal feature vectors, and each multimodal feature vector is composed of a low-frequency and high-frequency ratio feature and a corresponding voice emotion feature that have completed time alignment;

[0012] The multimodal feature vector set is input into a pre-trained emotion recognition model to obtain an emotion state classification result corresponding to each multimodal feature vector in the multimodal feature vector set to form an emotion state classification result output sequence.

[0013] In a second aspect, an embodiment of the present invention further provides a multimodal emotion recognition device for psychological counseling based on millimeter-wave radar, which includes:

[0014] a target user information acquisition unit, configured to, in response to a multimodal emotion recognition instruction, acquire user information to be identified corresponding to the multimodal emotion recognition instruction and a target user to be identified corresponding to the user information to be identified;

[0015] a millimeter-wave radar data acquisition unit, configured to acquire a current millimeter-wave radar data set corresponding to the target user to be identified; wherein the current millimeter-wave radar data set is composed of millimeter-wave radar data collected by the millimeter-wave radar module within a target collection time interval corresponding to the multimodal emotion recognition instruction after the millimeter-wave radar module is aligned with the target user to be identified;

[0016] A sound data acquisition unit, configured to acquire a current user voice data set corresponding to the target user to be identified; wherein the current user voice data set is composed of user voice data collected by the sound pickup and processing module from the target user to be identified within a target collection time interval;

[0017] a low-frequency-high-frequency ratio feature acquisition unit, configured to acquire, in time sequence, a heart rate variability data sequence corresponding to the current millimeter-wave radar data set and a low-frequency-high-frequency ratio feature sequence corresponding to the heart rate variability data sequence based on a preset millimeter-wave radar feature extraction strategy;

[0018] A voice emotion feature acquisition unit, configured to acquire, based on a preset voice feature extraction model, a voice emotion feature sequence corresponding to the current user voice data set in a time sequence;

[0019] a multimodal feature fusion unit, configured to perform time-series alignment and feature fusion on each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence and each voice emotion feature in the voice emotion feature sequence based on a preset data fusion strategy, thereby obtaining a multimodal feature vector set corresponding to the target user to be identified; wherein the multimodal feature vector set includes a plurality of multimodal feature vectors, and each multimodal feature vector is composed of a low-frequency and high-frequency ratio feature and a corresponding voice emotion feature that have completed time-series alignment;

[0020] The classification result output unit is used to input the multimodal feature vector set into a pre-trained emotion recognition model to obtain the emotional state classification results corresponding to each multimodal feature vector in the multimodal feature vector set to form an emotional state classification result output sequence.

[0021] In a third aspect, an embodiment of the present invention further provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in the first aspect is implemented.

[0022] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method described in the first aspect can be implemented.

[0023] The embodiment of the present invention provides a multimodal emotion recognition method and device for psychological counseling based on millimeter-wave radar, the method comprising: in response to a multimodal emotion recognition instruction, obtaining user information to be identified corresponding to the multimodal emotion recognition instruction, and a target user to be identified corresponding to the user information to be identified; obtaining a current millimeter-wave radar data set corresponding to the target user to be identified; wherein the current millimeter-wave radar data set is composed of millimeter-wave radar data collected by the millimeter-wave radar module within the target collection time interval corresponding to the multimodal emotion recognition instruction after aiming at the target user to be identified; obtaining a current user voice data set corresponding to the target user to be identified; wherein the current user voice data set is composed of user voice data collected by the sound pickup and processing module within the target collection time interval for the target user to be identified; based on a preset millimeter-wave radar feature extraction strategy, obtaining the corresponding data set in time sequence according to the current millimeter-wave radar data set. A heart rate variability data sequence and a low-frequency and high-frequency ratio feature sequence corresponding to the heart rate variability data sequence are obtained; based on a preset speech feature extraction model, a speech emotion feature sequence corresponding to the current user speech dataset is obtained in time sequence; based on a preset data fusion strategy, each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence is time-aligned and feature-fused with each speech emotion feature in the speech emotion feature sequence to obtain a multimodal feature vector set corresponding to the target user to be identified; wherein the multimodal feature vector set includes multiple multimodal feature vectors, and each multimodal feature vector is composed of a low-frequency and high-frequency ratio feature that has completed time alignment and a corresponding speech emotion feature; the multimodal feature vector set is input into a pre-trained emotion recognition model to obtain an emotion state classification result corresponding to each multimodal feature vector in the multimodal feature vector set, thereby forming an emotion state classification result output sequence. The embodiment of the present invention can collect the current millimeter-wave radar dataset and the current user speech dataset of the target user to be identified, perform feature extraction, time alignment, and feature fusion on each of them to obtain a multimodal feature vector set, and then input it into the emotion recognition model to obtain a recognition result, thereby improving the accuracy of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 A schematic diagram of an application scenario of a multimodal emotion recognition method for psychological counseling based on millimeter-wave radar provided by an embodiment of the present invention;

[0026] Figure 2A flowchart of a multimodal emotion recognition method for psychological counseling based on millimeter-wave radar provided by an embodiment of the present invention;

[0027] Figure 3 A schematic diagram of a sub-process of a multimodal emotion recognition method for psychological counseling based on millimeter-wave radar provided in an embodiment of the present invention;

[0028] Figure 4 A schematic diagram of a sub-process of a multimodal emotion recognition method for psychological counseling based on millimeter-wave radar provided in an embodiment of the present invention;

[0029] Figure 5 A schematic diagram of a sub-process of a multimodal emotion recognition method for psychological counseling based on millimeter-wave radar provided in an embodiment of the present invention;

[0030] Figure 6 A schematic block diagram of a millimeter-wave radar-based multimodal emotion recognition device for psychological counseling provided by an embodiment of the present invention;

[0031] Figure 7 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0033] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0034] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0035] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0036] Please also refer to Figure 1 and Figure 2 ,in Figure 1 This is a schematic diagram of a scenario of a multimodal emotion recognition method for psychological counseling based on millimeter-wave radar according to an embodiment of the present invention. Figure 2 FIG. 1 is a flow chart of a multimodal emotion recognition method for psychological counseling based on millimeter wave radar provided by an embodiment of the present invention. Figure 1 As shown, the millimeter-wave radar-based multimodal emotion recognition method for psychological counseling provided in an embodiment of the present invention is applied to a millimeter-wave radar-based multimodal emotion recognition device for psychological counseling. The millimeter-wave radar-based multimodal emotion recognition device for psychological counseling includes a server 10, a millimeter-wave radar module 20, and a sound pickup and processing module 30. The millimeter-wave radar module 20 and the sound pickup and processing module 30 are both communicatively connected to the server 10. Moreover, during specific implementation, the millimeter-wave radar module 20 and the sound pickup and processing module 30 can be located in the same indoor location, such as a psychological counseling room, and the server 10 can be located in a computer room outside the psychological counseling room (of course, to ensure data transmission efficiency, the server 10 can also be located in the psychological counseling room).

[0037] like Figure 2 As shown, the method includes the following steps S110-S170.

[0038] S110 : In response to a multimodal emotion recognition instruction, obtain user information to be recognized corresponding to the multimodal emotion recognition instruction and a target user to be recognized corresponding to the user information to be recognized.

[0039] In this embodiment, when a user enters a psychological counseling room to communicate with a user in the role of a counselor, the user in the role of a counselor can first log in to the user interaction interface corresponding to the server using their user terminal (such as a desktop computer, tablet computer, smart phone, etc.). After obtaining the information of the user to be identified (such as the user information to be identified at least including the user's unique identity identification code, user age, user gender, user contact number, etc.) through manual entry or voice recognition, click the virtual start button on the user interaction interface to trigger the generation of a multimodal emotion recognition instruction, and send the multimodal emotion recognition instruction to the server. At the same time, the target user to be identified corresponding to the user information to be identified is sent to the server to notify the target user to be identified as the current object to be identified. During the time interval between detecting the multimodal emotion recognition instruction and detecting the stop recognition instruction, the server will control the millimeter-wave radar module to collect millimeter-wave radar data of the target user to be identified, and control the sound pickup and processing module to collect user voice data of the target user to be identified.

[0040] S120: Acquire a current millimeter-wave radar data set corresponding to the target user to be identified.

[0041] The current millimeter-wave radar data set is composed of millimeter-wave radar data collected by the millimeter-wave radar module within the target collection time interval corresponding to the multimodal emotion recognition instruction after the millimeter-wave radar module is aimed at the target user to be identified.

[0042] In this embodiment, when determining the target collection time interval corresponding to the multimodal emotion recognition instruction, in addition to obtaining the time when the multimodal emotion recognition instruction was generated, it is also necessary to obtain the time when the consultant user clicked the virtual "End" button on the user interface using their user terminal to generate the stop recognition instruction. The time when the multimodal emotion recognition instruction was generated is then used as the starting time point of the target collection time interval, and the time when the stop recognition instruction was generated is used as the ending time point of the target collection time interval.

[0043] The millimeter-wave radar module installed in the psychological counseling room is specifically aimed at the chest area of ​​the target user to be identified (in specific implementation, the millimeter-wave radar module can also be set as a clip-on card-type millimeter-wave radar module and worn on the chest of the user playing the role of the counselor. When the server is specifically a desktop computer, the millimeter-wave radar module can be connected to the server via a USB interface, and powered and data transmitted via USB), and the current millimeter-wave radar data set corresponding to the target user to be identified is formed after detecting the tiny movements of the chest of the target user to be identified (the chest fluctuations caused by the heartbeat).

[0044] S130: Acquire a current user voice data set corresponding to the target user to be identified.

[0045] The current user voice data set is composed of user voice data collected by the sound pickup and processing module from the target user to be identified within a target collection time interval.

[0046] In this embodiment, the target acquisition time interval corresponding to the current user voice data set is exactly the same as the target acquisition time interval corresponding to the current millimeter-wave radar data set. The sound pickup and processing module set in the psychological counseling room can collect the current user voice data corresponding to the target user to be identified in the target acquisition time interval and form the current user voice data set.

[0047] In one embodiment, before step S130, the method further includes:

[0048] Acquire a current communication process speech data set corresponding to the target user to be identified;

[0049] Based on the pre-collected voiceprint features of the user to be filtered, the voice signals in the current communication process voice data set and the voiceprint features of the user to be filtered are filtered to obtain the current user voice data set.

[0050] In this embodiment, in order to ensure that the current user voice data set only includes the voice data of the target user to be identified, the sound pickup and processing module (in specific implementation, the sound pickup and processing module can also use a microphone module) can first pre-collect the user voice data of the consultant role as the initial sample sound data, and then perform audio noise reduction, normalization processing, time domain feature extraction (such as extracting energy features, short-time average amplitude), and frequency domain feature extraction (such as extracting linear prediction cepstral coefficients, Mel-frequency cepstral coefficients, spectral centroid, etc.) on the initial sample sound data, and then form the voiceprint features of the user to be filtered based on the time domain features and frequency domain features corresponding to the initial sample sound data (in this case, the user to be filtered refers to the user of the consultant role).

[0051] After obtaining the current communication process speech data set including the voice data of the target user to be identified and the voice data of the user to be filtered, the speech signal in the current communication process speech data set and the voiceprint feature of the user to be filtered can be matched and filtered based on the dynamic time warping (DTW) algorithm or the Gaussian mixture model (GMM), so as to obtain the current user speech data set that only includes the voice data of the target user to be identified, thereby retaining only the voice of the target user to be identified for subsequent emotion analysis.

[0052] In one embodiment, after step S130, the method further includes:

[0053] Speech recognition is performed on the current user speech data set based on a pre-trained speech recognition model to obtain current user text data including text time sequence information.

[0054] In this embodiment, the server also includes an automatic speech recognition module (also referred to as an ASR module). After the sound pickup and processing module obtains the current user speech data set corresponding to the target user to be identified, the sound pickup and processing module can send the current user speech data set to the server, and the automatic speech recognition module in the server performs speech recognition. More specifically, the speech recognition model deployed in the automatic speech recognition module (such as a hidden Markov model, a Gaussian mixture model, a deep neural network, a convolutional neural network, etc.) performs speech recognition on the current user speech data set to obtain the current user text data including text timing information. It should be noted that after each speech data in the current user speech data set is identified and converted into text, the feature of the collection time corresponding to the speech data is also inherited by the text. This is also to facilitate the subsequent timing alignment of the emotion recognition results based on the time-series sequence.

[0055] In one embodiment, after step S130 and before step S140, the method further includes:

[0056] Get the preset sliding window period;

[0057] Dividing the current millimeter-wave radar data set based on the sliding window period to obtain a plurality of current millimeter-wave radar data subsets; wherein the plurality of current millimeter-wave radar data subsets constitute the current millimeter-wave radar data set;

[0058] The current user voice data set is divided based on the sliding window period to obtain multiple current user voice data subsets; wherein the multiple current user voice data subsets constitute the current user voice data.

[0059] In this embodiment, since the target acquisition time interval corresponding to the current user voice dataset is exactly the same as the target acquisition time interval corresponding to the current millimeter-wave radar dataset, the same sliding window period is used (for example, it is set to 5s, 10s, 15s, 20s, etc., of course, the specific real-time performance is not limited to the time intervals listed in the above examples and can also be flexibly set according to the actual needs of the user). After the current millimeter-wave radar dataset and the current user voice dataset are divided respectively, the total number of first subsets corresponding to the multiple current millimeter-wave radar data subsets is equal to the total number of second subsets corresponding to the multiple current user voice data subsets.

[0060] Afterwards, when performing millimeter-wave radar feature extraction on the current millimeter-wave radar data set, it can be decomposed into performing millimeter-wave radar feature extraction on each current millimeter-wave radar data subset, thereby forming a heart rate variability data sequence and a low-frequency and high-frequency ratio feature sequence. Similarly, when performing voice emotion recognition on the current user voice data set, it can be decomposed into performing voice emotion feature extraction on each current user voice data subset, thereby forming a voice emotion feature sequence. It can be seen that through the above method, the current millimeter-wave radar data set and the current user voice data set can be divided into equal sliding window lengths, which facilitates subsequent time series alignment. Among them, the current millimeter-wave radar data set and the current user voice data set, as well as various processed data obtained subsequently, can all be encrypted and stored in the server to improve data security.

[0061] S140. Based on a preset millimeter-wave radar feature extraction strategy, acquire, in time sequence, a heart rate variability data sequence corresponding to the current millimeter-wave radar data set, and a low-frequency to high-frequency ratio feature sequence corresponding to the heart rate variability data sequence.

[0062] In this embodiment, after the server obtains the current millimeter-wave radar dataset, it can fit a waveform graph according to the acquisition time sequence corresponding to each current millimeter-wave radar data in the current millimeter-wave radar dataset. From this waveform, the heart rate data corresponding to the target user to be identified can be extracted. Ultimately, the heart rate variability data sequence can be determined by combining the intervals between consecutive heartbeats (such as NN intervals, RR intervals, etc.). The heart rate variability data sequence includes multiple heart rate variability data (HRV data, the full name of which is Heart Rate Variability). Each heart rate variability data can be further decomposed into a low-frequency component (i.e., LF component, corresponding to the low-frequency range of 0.04-0.15Hz) and a high-frequency component (i.e., HF component, corresponding to the high-frequency range of 0.15-0.4Hz). By calculating the low-frequency component / high-frequency component ratio of each heart rate variability data, the corresponding low-frequency-high-frequency ratio feature can be obtained. After obtaining the heart rate variability data sequence corresponding to the current millimeter-wave radar dataset in a time-sequential manner, a low-frequency-high-frequency ratio feature sequence with the same time sequence can also be obtained.

[0063] In one embodiment, if Figure 3 As shown, step S140 includes:

[0064] S141. Acquire a current heart rate fitting dataset corresponding to the current millimeter-wave radar data according to a time sequence of acquiring each current millimeter-wave radar data subset in the current millimeter-wave radar data set;

[0065] S142, performing data cleaning on the current heart rate fitting dataset based on a preset data cleaning strategy to obtain a cleaned current heart rate fitting dataset;

[0066] S143, obtaining a current RR interval sequence corresponding to the cleaned current heart rate fitting data set, and obtaining the heart rate variability data sequence based on the current RR interval sequence and a preset power spectrum density estimation model;

[0067] S144. For each heart rate variability data in the heart rate variability data sequence, a ratio of low-frequency component to high-frequency cost is obtained as a low-frequency-high-frequency ratio feature, and the low-frequency-high-frequency ratio feature of each heart rate variability data in the heart rate variability data sequence is arranged in time series to form the low-frequency-high-frequency ratio feature sequence.

[0068] In this embodiment, after obtaining the current millimeter-wave radar data set, a waveform can be fitted according to the acquisition time sequence corresponding to each current millimeter-wave radar data subset in the current millimeter-wave radar data set, and the current heart rate fitting data set corresponding to the target user to be identified can be extracted from the waveform. Then, the current heart rate fitting data set is removed from the outliers using a data cleaning strategy, such as an outlier cleaning strategy, to obtain a cleaned current heart rate fitting data set. Thereafter, the time interval values ​​between two adjacent heartbeat signals are determined in sequence according to the time sequence in the cleaned current heart rate fitting data set, thereby obtaining the current RR interval sequence (wherein the RR interval refers to the time interval between two adjacent heartbeat signals). In addition, the current RR interval sequence can be further subjected to a fast Fourier transform in combination with a power spectral density estimation model to perform power spectral density estimation, thereby decomposing the variability of the current RR interval sequence into the power of different frequency components.

[0069] Finally, because each heart rate variability data has been decomposed into a low-frequency component (i.e., LF component, and the low-frequency band corresponding to the low-frequency component is 0.04-0.15Hz) and a high-frequency component (i.e., HF component, and the high-frequency band corresponding to the high-frequency component is 0.15-0.4Hz) based on the power spectrum density estimation model, the corresponding low-frequency and high-frequency ratio features can be obtained by calculating the low-frequency component / high-frequency cost ratio of each heart rate variability data. After obtaining the heart rate variability data sequence corresponding to the current millimeter-wave radar data set in time sequence, the low-frequency and high-frequency ratio feature sequence of the same time sequence can also be obtained. Among them, the total number of low-frequency and high-frequency ratio features included in the low-frequency and high-frequency ratio feature sequence is the same as the number of the first total subset. It can be seen that the above method, combined with the data cleaning strategy and the power spectrum density estimation model, can quickly extract the low-frequency and high-frequency ratio feature sequence corresponding to the current millimeter-wave radar data set.

[0070] S150 , based on a preset speech feature extraction model, obtaining a speech emotion feature sequence corresponding to the current user speech data set in time sequence.

[0071] In this embodiment, because the current user speech data set includes multiple current user speech data subsets arranged in chronological order, speech emotion recognition can be performed on each of the multiple current user speech data subsets corresponding to the current user speech data set. After obtaining speech emotion features corresponding to each current user speech data subset, the speech emotion feature sequence is formed. The total number of speech emotion features included in the speech emotion feature sequence is the same as the total number of the first subsets.

[0072] In one embodiment, if Figure 4 As shown, step S150 includes:

[0073] S151. For each current user voice data subset in the current user voice data set, input the current user voice data subset into the voice feature extraction model to perform voice emotion feature extraction, thereby obtaining a current voice emotion feature corresponding to the current user voice data subset; wherein the current voice emotion feature includes at least user pitch feature data, user voice intensity feature data, and user speech rate feature data;

[0074] S152: Arrange the current voice emotion features corresponding to each current user voice data subset in the current user voice data set in time sequence to form the voice emotion feature sequence.

[0075] In this embodiment, the speech feature extraction model preset in the server includes at least a speech noise reduction processing submodel, a speech framing submodel, a pitch feature extraction submodel, a sound intensity feature extraction submodel, and a speech rate feature extraction submodel. When the current user speech data subset is input into the speech feature extraction model, the current user speech data subset is obtained after the noise reduction processing (such as adaptive filtering noise reduction processing) of the extraction submodel and the framing processing (such as first determining the frame length and frame shift, and then performing a windowing operation to achieve framing processing of the speech signal) of the speech framing submodel are sequentially performed. Then, the pre-processed current user speech data subset is subjected to corresponding feature extraction based on the pitch feature extraction submodel (such as the harmonic peak detection model), the sound intensity feature extraction submodel, and the speech rate feature extraction submodel (such as the phoneme interval statistical model). The current speech emotion feature including the user pitch feature data, the user sound intensity feature data, and the user speech rate feature data can be obtained. It can be seen that based on the above method, the speech feature extraction model can quickly extract the speech emotion feature sequence corresponding to the current user speech data set.

[0076] S160. Based on a preset data fusion strategy, each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence is time-aligned and feature-fused with each voice emotion feature in the voice emotion feature sequence to obtain a multimodal feature vector set corresponding to the target user to be identified.

[0077] The multimodal feature vector set includes a plurality of multimodal feature vectors, and each multimodal feature vector is composed of a low-frequency and high-frequency ratio feature with time alignment and a corresponding speech emotion feature.

[0078] In this embodiment, after the low-frequency and high-frequency ratio feature sequence and the voice emotion feature sequence are obtained in the server, in order to make the obtained multimodal feature vector set also have time series characteristics, the low-frequency and high-frequency ratio features in the low-frequency and high-frequency ratio feature sequence and the voice emotion features in the voice emotion feature sequence can be time-aligned based on a preset data fusion strategy, and then feature fusion is performed to obtain a multimodal feature vector set corresponding to the target user to be identified.

[0079] In one embodiment, if Figure 5 As shown, step S160 includes:

[0080] S161, obtaining a first timestamp for each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence, and obtaining a second timestamp for each voice emotion feature in the voice emotion feature sequence;

[0081] S162, aligning the low-frequency and high-frequency ratio features with the voice emotion features having the same first timestamp and second timestamp and performing feature fusion to form a multimodal feature vector;

[0082] S163 , arranging the multiple multimodal feature vectors in chronological order to form the multimodal feature vector set.

[0083] In this embodiment, when temporally aligning each low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence with each voice emotion feature in the voice emotion feature sequence, it is necessary to obtain the first timestamp corresponding to each low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence, and to obtain the second timestamp corresponding to each voice emotion feature in the voice emotion feature sequence. For example, if the first low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence corresponds to the first timestamp T10, and the first voice emotion feature in the voice emotion feature sequence corresponds to the second timestamp T20, when the first timestamp T10 is equal to the second timestamp T20, the two features can be considered to be temporally aligned. The first low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence can also be fused with the first voice emotion feature in the voice emotion feature sequence to form a multimodal feature vector corresponding to the first timestamp T10 or the second timestamp T20. Subsequently, when temporally aligning other low-frequency-high-frequency ratio features in the low-frequency-high-frequency ratio feature sequence with the corresponding voice emotion features in the voice emotion feature sequence, the above process is also referred to. Since the low-frequency and high-frequency ratio feature includes the low-frequency and high-frequency ratio, and the voice emotion feature includes the user pitch feature data, the user sound intensity feature data and the user speech rate feature data, the obtained multimodal feature vector includes the low-frequency and high-frequency ratio, the user pitch feature data, the user sound intensity feature data and the user speech rate feature data. Moreover, the multiple multimodal feature vectors included in the multimodal feature vector set corresponding to the target user to be identified are also arranged in sequence according to the time sequence of the timestamps. It can be seen that through the above method, the low-frequency and high-frequency ratio feature sequence and the voice emotion feature sequence can be quickly aligned and fused to obtain multiple multimodal feature vectors arranged in time sequence.

[0084] S170: Input the multimodal feature vector set into a pre-trained emotion recognition model to obtain an emotion state classification result corresponding to each multimodal feature vector in the multimodal feature vector set, so as to form an emotion state classification result output sequence.

[0085] In this embodiment, the emotion recognition model adopted is a long short-term memory network integrated with a self-attention mechanism or a multimodal convolutional neural network integrated with a self-attention mechanism. The multiple multimodal feature vectors included in the multimodal feature vector set also have the temporal sorting characteristics of the time series. Each of the multimodal feature vectors in the multimodal feature vector set is input into the pre-trained emotion recognition model in sequence according to the temporal sorting order, and the emotion state classification results corresponding to each multimodal feature vector can be obtained. The above-mentioned multiple emotion state classification results are combined in the temporal order to obtain the emotion state classification result output sequence. Specifically, each of the multimodal feature vectors corresponds to an emotional state classification result corresponding to a 1*2 vector, and the two vector values ​​included therein are both probability values ​​ranging from 0 to 1. The first probability value can determine the emotional category based on the specific value interval to which it belongs, and the second probability value can determine the emotional intensity based on the specific value interval to which it belongs. For example, the emotional state classification result is [0.3, 0.8]. Based on the value interval to which the first probability value 0.3 in the emotional state classification result belongs, the emotional category is determined to be anxiety, and based on the value interval to which the second probability value 0.8 in the emotional state classification result belongs, the emotional intensity is determined to be high intensity. Because the emotional state classification results in the emotional state classification result output sequence are also arranged in chronological order, an emotional state diagram can also be constructed with the time axis as the horizontal axis and the probability values ​​corresponding to the emotional categories in the emotional state classification results as the vertical axis to characterize the emotional fluctuation state of the target user to be identified. Of course, after performing speech recognition on the current user's speech data set to obtain the current user's text data including text time series information, the current user's text data can also be integrated and displayed based on the timeline of the emotional state diagram. That is, each text record is accompanied by a timestamp and displayed together with the corresponding emotional state, making it easier for the user in the counselor role to view the communication text content and emotional state at different moments. When the emotion category corresponding to a time point in the emotional state diagram changes significantly, for example, from relaxation to anger, an abnormal fluctuation prompt information can be automatically generated and displayed on the corresponding display device of the server.

[0086] Among them, when the emotion recognition model adopts a long short-term memory network that integrates a self-attention mechanism, it can focus more on capturing the dynamic changes of the low-frequency and high-frequency ratio and speech emotion features; and when the emotion recognition model adopts a multimodal convolutional neural network that integrates a self-attention mechanism, it can perform more refined fusion at the feature level. Moreover, the self-attention mechanism used is to dynamically adjust the weight of each feature in the multimodal feature vector when each multimodal feature vector is input into the emotion recognition model to cope with the difference in the importance of features under different emotional states. For example, when the low-frequency and high-frequency ratio in the multimodal feature vector increases significantly, the weight of the low-frequency and high-frequency ratio feature will be automatically increased; when the change in speech features is more significant, the weight of the speech emotion feature will be increased.

[0087] Moreover, when the emotional state classification results corresponding to each multimodal feature vector are obtained based on the emotion recognition model in a chronological order, specifically, an emotional state classification sub-result can be output for the low-frequency and high-frequency ratio features in the multimodal feature vector, or an emotional state classification sub-result can be output for the speech emotion features in the multimodal feature vector, and the above two emotional state classification sub-results are weighted and summed to obtain a comprehensive emotional state classification result. However, it should be noted that when there is a significant difference between outputting an emotional state classification sub-result for the low-frequency and high-frequency ratio features in the multimodal feature vector and outputting an emotional state classification sub-result for the speech emotion features in the multimodal feature vector, performing the weighted summation can balance the emotional judgment results.

[0088] It can be seen that the embodiment of this method can collect the current millimeter-wave radar dataset and the current user voice dataset of the target user to be identified, perform feature extraction, time alignment and feature fusion respectively to obtain a multimodal feature vector set, and then input it into the emotion recognition model to obtain the recognition result, thereby improving the accuracy of emotion recognition.

[0089] Figure 6 This is a schematic block diagram of a multimodal emotion recognition device for psychological counseling based on millimeter wave radar provided by an embodiment of the present invention. Figure 6 As shown, corresponding to the above-mentioned psychological consultation multimodal emotion recognition method based on millimeter wave radar, the present invention also provides a psychological consultation multimodal emotion recognition device 100 based on millimeter wave radar. The psychological consultation multimodal emotion recognition device 100 based on millimeter wave radar includes a unit for executing the above-mentioned psychological consultation multimodal emotion recognition method based on millimeter wave radar. Figure 6The millimeter-wave radar-based multimodal emotion recognition device 100 for psychological counseling includes: a target user information acquisition unit 110, a millimeter-wave radar data acquisition unit 120, a sound data acquisition unit 130, a low-frequency and high-frequency ratio feature acquisition unit 140, a voice emotion feature acquisition unit 150, a multimodal feature fusion unit 160 and a classification result output unit 170.

[0090] The target user information acquiring unit 110 is configured to, in response to a multimodal emotion recognition instruction, acquire user information to be identified corresponding to the multimodal emotion recognition instruction and a target user to be identified corresponding to the user information to be identified.

[0091] In this embodiment, when a user enters a psychological counseling room to communicate with a user in the role of a counselor, the user in the role of a counselor can first log in to the user interaction interface corresponding to the server using their user terminal (such as a desktop computer, tablet computer, smart phone, etc.). After obtaining the information of the user to be identified (such as the user information to be identified at least including the user's unique identity identification code, user age, user gender, user contact number, etc.) through manual entry or voice recognition, click the virtual start button on the user interaction interface to trigger the generation of a multimodal emotion recognition instruction, and send the multimodal emotion recognition instruction to the server. At the same time, the target user to be identified corresponding to the user information to be identified is sent to the server to notify the target user to be identified as the current object to be identified. During the time interval between detecting the multimodal emotion recognition instruction and detecting the stop recognition instruction, the server will control the millimeter-wave radar module to collect millimeter-wave radar data of the target user to be identified, and control the sound pickup and processing module to collect user voice data of the target user to be identified.

[0092] The millimeter wave radar data acquisition unit 120 is configured to acquire a current millimeter wave radar data set corresponding to the target user to be identified.

[0093] The current millimeter-wave radar data set is composed of millimeter-wave radar data collected by the millimeter-wave radar module within the target collection time interval corresponding to the multimodal emotion recognition instruction after the millimeter-wave radar module is aimed at the target user to be identified.

[0094] In this embodiment, when determining the target collection time interval corresponding to the multimodal emotion recognition instruction, in addition to obtaining the time when the multimodal emotion recognition instruction was generated, it is also necessary to obtain the time when the consultant user clicked the virtual "End" button on the user interface using their user terminal to generate the stop recognition instruction. The time when the multimodal emotion recognition instruction was generated is then used as the starting time point of the target collection time interval, and the time when the stop recognition instruction was generated is used as the ending time point of the target collection time interval.

[0095] The millimeter-wave radar module installed in the psychological counseling room is specifically aimed at the chest area of ​​the target user to be identified (in specific implementation, the millimeter-wave radar module can also be set as a clip-on card-type millimeter-wave radar module and worn on the chest of the user playing the role of the counselor. When the server is specifically a desktop computer, the millimeter-wave radar module can be connected to the server via a USB interface, and powered and data transmitted via USB), and the current millimeter-wave radar data set corresponding to the target user to be identified is formed after detecting the tiny movements of the chest of the target user to be identified (the chest fluctuations caused by the heartbeat).

[0096] The voice data acquisition unit 130 is configured to acquire a current user voice data set corresponding to the target user to be identified.

[0097] The current user voice data set is composed of user voice data collected by the sound pickup and processing module from the target user to be identified within a target collection time interval.

[0098] In this embodiment, the target acquisition time interval corresponding to the current user voice data set is exactly the same as the target acquisition time interval corresponding to the current millimeter-wave radar data set. The sound pickup and processing module set in the psychological counseling room can collect the current user voice data corresponding to the target user to be identified in the target acquisition time interval and form the current user voice data set.

[0099] In one embodiment, the millimeter-wave radar-based multimodal emotion recognition device 100 for psychological counseling further includes:

[0100] A current communication process speech data set acquisition unit, configured to acquire a current communication process speech data set corresponding to the target user to be identified;

[0101] The voice signal filtering unit is used to filter the voice signals in the current communication process voice data set and the voice print features of the user to be filtered based on the pre-collected voice print features of the user to be filtered, so as to obtain the current user voice data set.

[0102] In this embodiment, in order to ensure that the current user voice data set only includes the voice data of the target user to be identified, the sound pickup and processing module (in specific implementation, the sound pickup and processing module can also use a microphone module) can first pre-collect the user voice data of the consultant role as the initial sample sound data, and then perform audio noise reduction, normalization processing, time domain feature extraction (such as extracting energy features, short-time average amplitude), and frequency domain feature extraction (such as extracting linear prediction cepstral coefficients, Mel-frequency cepstral coefficients, spectral centroid, etc.) on the initial sample sound data, and then form the voiceprint features of the user to be filtered based on the time domain features and frequency domain features corresponding to the initial sample sound data (in this case, the user to be filtered refers to the user of the consultant role).

[0103] After obtaining the current communication process speech data set including the voice data of the target user to be identified and the voice data of the user to be filtered, the speech signal in the current communication process speech data set and the voiceprint feature of the user to be filtered can be matched and filtered based on the dynamic time warping (DTW) algorithm or the Gaussian mixture model (GMM), so as to obtain the current user speech data set that only includes the voice data of the target user to be identified, thereby retaining only the voice of the target user to be identified for subsequent emotion analysis.

[0104] In one embodiment, the millimeter-wave radar-based multimodal emotion recognition device 100 for psychological counseling further includes:

[0105] The speech-to-text recognition unit is configured to perform speech recognition on the current user speech data set based on a pre-trained speech recognition model to obtain current user text data including text time sequence information.

[0106] In this embodiment, the server also includes an automatic speech recognition module (also referred to as an ASR module). After the sound pickup and processing module obtains the current user speech data set corresponding to the target user to be identified, the sound pickup and processing module can send the current user speech data set to the server, and the automatic speech recognition module in the server performs speech recognition. More specifically, the speech recognition model deployed in the automatic speech recognition module (such as a hidden Markov model, a Gaussian mixture model, a deep neural network, a convolutional neural network, etc.) performs speech recognition on the current user speech data set to obtain the current user text data including text timing information. It should be noted that after each speech data in the current user speech data set is identified and converted into text, the feature of the collection time corresponding to the speech data is also inherited by the text. This is also to facilitate the subsequent timing alignment of the emotion recognition results based on the time-series sequence.

[0107] In one embodiment, the millimeter-wave radar-based multimodal emotion recognition device 100 for psychological counseling further includes:

[0108] A sliding window period acquisition unit, configured to acquire a preset sliding window period;

[0109] a first dividing unit, configured to divide the current millimeter-wave radar data set based on the sliding window period to obtain a plurality of current millimeter-wave radar data subsets; wherein the plurality of current millimeter-wave radar data subsets constitute the current millimeter-wave radar data set;

[0110] The second dividing unit is configured to divide the current user voice data set based on the sliding window period to obtain a plurality of current user voice data subsets; wherein the plurality of current user voice data subsets constitute the current user voice data.

[0111] In this embodiment, since the target acquisition time interval corresponding to the current user voice dataset is exactly the same as the target acquisition time interval corresponding to the current millimeter-wave radar dataset, the same sliding window period is used (for example, it is set to 5s, 10s, 15s, 20s, etc., of course, the specific real-time performance is not limited to the time intervals listed in the above examples and can also be flexibly set according to the actual needs of the user). After the current millimeter-wave radar dataset and the current user voice dataset are divided respectively, the total number of first subsets corresponding to the multiple current millimeter-wave radar data subsets is equal to the total number of second subsets corresponding to the multiple current user voice data subsets.

[0112] Afterwards, when performing millimeter-wave radar feature extraction on the current millimeter-wave radar data set, it can be decomposed into performing millimeter-wave radar feature extraction on each current millimeter-wave radar data subset, thereby forming a heart rate variability data sequence and a low-frequency and high-frequency ratio feature sequence. Similarly, when performing voice emotion recognition on the current user voice data set, it can be decomposed into performing voice emotion feature extraction on each current user voice data subset, thereby forming a voice emotion feature sequence. It can be seen that through the above method, the current millimeter-wave radar data set and the current user voice data set can be divided into equal sliding window lengths, which facilitates subsequent time series alignment. Among them, the current millimeter-wave radar data set and the current user voice data set, as well as various processed data obtained subsequently, can all be encrypted and stored in the server to improve data security.

[0113] The low-frequency-high-frequency ratio feature acquisition unit 140 is used to acquire, in time sequence, a heart rate variability data sequence corresponding to the current millimeter-wave radar data set and a low-frequency-high-frequency ratio feature sequence corresponding to the heart rate variability data sequence based on a preset millimeter-wave radar feature extraction strategy.

[0114] In this embodiment, after the server obtains the current millimeter-wave radar dataset, it can fit a waveform graph according to the acquisition time sequence corresponding to each current millimeter-wave radar data in the current millimeter-wave radar dataset. From this waveform, the heart rate data corresponding to the target user to be identified can be extracted. Ultimately, the heart rate variability data sequence can be determined by combining the intervals between consecutive heartbeats (such as NN intervals, RR intervals, etc.). The heart rate variability data sequence includes multiple heart rate variability data (HRV data, the full name of which is Heart Rate Variability). Each heart rate variability data can be further decomposed into a low-frequency component (i.e., LF component, corresponding to the low-frequency range of 0.04-0.15Hz) and a high-frequency component (i.e., HF component, corresponding to the high-frequency range of 0.15-0.4Hz). By calculating the low-frequency component / high-frequency component ratio of each heart rate variability data, the corresponding low-frequency-high-frequency ratio feature can be obtained. After obtaining the heart rate variability data sequence corresponding to the current millimeter-wave radar dataset in a time-sequential manner, a low-frequency-high-frequency ratio feature sequence with the same time sequence can also be obtained.

[0115] In one embodiment, the low-frequency-high-frequency ratio feature acquisition unit 140 is specifically configured to:

[0116] Acquire a current heart rate fitting dataset corresponding to the current millimeter-wave radar data according to a time sequence of acquiring each current millimeter-wave radar data subset in the current millimeter-wave radar data set;

[0117] Performing data cleaning on the current heart rate fitting data set based on a preset data cleaning strategy to obtain a cleaned current heart rate fitting data set;

[0118] Acquire a current RR interval sequence corresponding to the cleaned current heart rate fitting data set, and acquire the heart rate variability data sequence based on the current RR interval sequence and a preset power spectrum density estimation model;

[0119] For each heart rate variability data in the heart rate variability data sequence, the ratio of low-frequency component / high-frequency cost is obtained as a low-frequency-high-frequency ratio feature, and the low-frequency-high-frequency ratio feature of each heart rate variability data in the heart rate variability data sequence is arranged in time series to form the low-frequency-high-frequency ratio feature sequence.

[0120] In this embodiment, after obtaining the current millimeter-wave radar data set, a waveform can be fitted according to the acquisition time sequence corresponding to each current millimeter-wave radar data subset in the current millimeter-wave radar data set, and the current heart rate fitting data set corresponding to the target user to be identified can be extracted from the waveform. Then, the current heart rate fitting data set is removed from the outliers using a data cleaning strategy, such as an outlier cleaning strategy, to obtain a cleaned current heart rate fitting data set. Thereafter, the time interval values ​​between two adjacent heartbeat signals are determined in sequence according to the time sequence in the cleaned current heart rate fitting data set, thereby obtaining the current RR interval sequence (wherein the RR interval refers to the time interval between two adjacent heartbeat signals). In addition, the current RR interval sequence can be further subjected to a fast Fourier transform in combination with a power spectral density estimation model to perform power spectral density estimation, thereby decomposing the variability of the current RR interval sequence into the power of different frequency components.

[0121] Finally, because each heart rate variability data has been decomposed into a low-frequency component (i.e., LF component, and the low-frequency band corresponding to the low-frequency component is 0.04-0.15Hz) and a high-frequency component (i.e., HF component, and the high-frequency band corresponding to the high-frequency component is 0.15-0.4Hz) based on the power spectrum density estimation model, the corresponding low-frequency and high-frequency ratio features can be obtained by calculating the low-frequency component / high-frequency cost ratio of each heart rate variability data. After obtaining the heart rate variability data sequence corresponding to the current millimeter-wave radar data set in time sequence, the low-frequency and high-frequency ratio feature sequence of the same time sequence can also be obtained. Among them, the total number of low-frequency and high-frequency ratio features included in the low-frequency and high-frequency ratio feature sequence is the same as the number of the first total subset. It can be seen that the above method, combined with the data cleaning strategy and the power spectrum density estimation model, can quickly extract the low-frequency and high-frequency ratio feature sequence corresponding to the current millimeter-wave radar data set.

[0122] The speech emotion feature acquisition unit 150 is configured to acquire, based on a preset speech feature extraction model, a speech emotion feature sequence corresponding to the current user speech data set in a time sequence.

[0123] In this embodiment, because the current user speech data set includes multiple current user speech data subsets arranged in chronological order, speech emotion recognition can be performed on each of the multiple current user speech data subsets corresponding to the current user speech data set. After obtaining speech emotion features corresponding to each current user speech data subset, the speech emotion feature sequence is formed. The total number of speech emotion features included in the speech emotion feature sequence is the same as the total number of the first subsets.

[0124] In one embodiment, the speech emotion feature acquisition unit 150 is specifically configured to:

[0125] For each current user voice data subset in the current user voice data set, input the current user voice data subset into the voice feature extraction model to perform voice emotion feature extraction to obtain a current voice emotion feature corresponding to the current user voice data subset; wherein the current voice emotion feature includes at least user pitch feature data, user voice intensity feature data, and user speech rate feature data;

[0126] The current voice emotion features corresponding to each current user voice data subset in the current user voice data set are arranged in time sequence to form the voice emotion feature sequence.

[0127] In this embodiment, the speech feature extraction model preset in the server includes at least a speech noise reduction processing submodel, a speech framing submodel, a pitch feature extraction submodel, a sound intensity feature extraction submodel, and a speech rate feature extraction submodel. When the current user speech data subset is input into the speech feature extraction model, the current user speech data subset is obtained after the noise reduction processing (such as adaptive filtering noise reduction processing) of the extraction submodel and the framing processing (such as first determining the frame length and frame shift, and then performing a windowing operation to achieve framing processing of the speech signal) of the speech framing submodel are sequentially performed. Then, the pre-processed current user speech data subset is subjected to corresponding feature extraction based on the pitch feature extraction submodel (such as the harmonic peak detection model), the sound intensity feature extraction submodel, and the speech rate feature extraction submodel (such as the phoneme interval statistical model). The current speech emotion feature including the user pitch feature data, the user sound intensity feature data, and the user speech rate feature data can be obtained. It can be seen that based on the above method, the speech feature extraction model can quickly extract the speech emotion feature sequence corresponding to the current user speech data set.

[0128] The multimodal feature fusion unit 160 is used to perform time alignment and feature fusion on each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence and each voice emotion feature in the voice emotion feature sequence based on a preset data fusion strategy to obtain a multimodal feature vector set corresponding to the target user to be identified.

[0129] The multimodal feature vector set includes a plurality of multimodal feature vectors, and each multimodal feature vector is composed of a low-frequency and high-frequency ratio feature with time alignment and a corresponding speech emotion feature.

[0130] In this embodiment, after the low-frequency and high-frequency ratio feature sequence and the voice emotion feature sequence are obtained in the server, in order to make the obtained multimodal feature vector set also have time series characteristics, the low-frequency and high-frequency ratio features in the low-frequency and high-frequency ratio feature sequence and the voice emotion features in the voice emotion feature sequence can be time-aligned based on a preset data fusion strategy, and then feature fusion is performed to obtain a multimodal feature vector set corresponding to the target user to be identified.

[0131] In one embodiment, the multimodal feature fusion unit 160 is specifically configured to:

[0132] Acquire a first timestamp for each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence, and acquire a second timestamp for each voice emotion feature in the voice emotion feature sequence;

[0133] Aligning the low-frequency and high-frequency ratio features with the same first and second timestamps with the speech emotion features and performing feature fusion to form a multimodal feature vector;

[0134] Arrange multiple multimodal feature vectors in chronological order to form the multimodal feature vector set.

[0135] In this embodiment, when temporally aligning each low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence with each voice emotion feature in the voice emotion feature sequence, it is necessary to obtain the first timestamp corresponding to each low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence, and to obtain the second timestamp corresponding to each voice emotion feature in the voice emotion feature sequence. For example, if the first low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence corresponds to the first timestamp T10, and the first voice emotion feature in the voice emotion feature sequence corresponds to the second timestamp T20, when the first timestamp T10 is equal to the second timestamp T20, the two features can be considered to be temporally aligned. The first low-frequency-high-frequency ratio feature in the low-frequency-high-frequency ratio feature sequence can also be fused with the first voice emotion feature in the voice emotion feature sequence to form a multimodal feature vector corresponding to the first timestamp T10 or the second timestamp T20. Subsequently, when temporally aligning other low-frequency-high-frequency ratio features in the low-frequency-high-frequency ratio feature sequence with the corresponding voice emotion features in the voice emotion feature sequence, the above process is also referred to. Since the low-frequency and high-frequency ratio feature includes the low-frequency and high-frequency ratio, and the voice emotion feature includes the user pitch feature data, the user sound intensity feature data and the user speech rate feature data, the obtained multimodal feature vector includes the low-frequency and high-frequency ratio, the user pitch feature data, the user sound intensity feature data and the user speech rate feature data. Moreover, the multiple multimodal feature vectors included in the multimodal feature vector set corresponding to the target user to be identified are also arranged in sequence according to the time sequence of the timestamps. It can be seen that through the above method, the low-frequency and high-frequency ratio feature sequence and the voice emotion feature sequence can be quickly aligned and fused to obtain multiple multimodal feature vectors arranged in time sequence.

[0136] The classification result output unit 170 is used to input the multimodal feature vector set into a pre-trained emotion recognition model to obtain the emotional state classification results corresponding to each multimodal feature vector in the multimodal feature vector set to form an emotional state classification result output sequence.

[0137] In this embodiment, the emotion recognition model adopted is a long short-term memory network integrated with a self-attention mechanism or a multimodal convolutional neural network integrated with a self-attention mechanism. The multiple multimodal feature vectors included in the multimodal feature vector set also have the temporal sorting characteristics of the time series. Each of the multimodal feature vectors in the multimodal feature vector set is input into the pre-trained emotion recognition model in sequence according to the temporal sorting order, and the emotion state classification results corresponding to each multimodal feature vector can be obtained. The above-mentioned multiple emotion state classification results are combined in the temporal order to obtain the emotion state classification result output sequence. Specifically, each of the multimodal feature vectors corresponds to an emotional state classification result corresponding to a 1*2 vector, and the two vector values ​​included therein are both probability values ​​ranging from 0 to 1. The first probability value can determine the emotional category based on the specific value interval to which it belongs, and the second probability value can determine the emotional intensity based on the specific value interval to which it belongs. For example, the emotional state classification result is [0.3, 0.8]. Based on the value interval to which the first probability value 0.3 in the emotional state classification result belongs, the emotional category is determined to be anxiety, and based on the value interval to which the second probability value 0.8 in the emotional state classification result belongs, the emotional intensity is determined to be high intensity. Because the emotional state classification results in the emotional state classification result output sequence are also arranged in chronological order, an emotional state diagram can also be constructed with the time axis as the horizontal axis and the probability values ​​corresponding to the emotional categories in the emotional state classification results as the vertical axis to characterize the emotional fluctuation state of the target user to be identified. Of course, after performing speech recognition on the current user's speech data set to obtain the current user's text data including text time series information, the current user's text data can also be integrated and displayed based on the timeline of the emotional state diagram. That is, each text record is accompanied by a timestamp and displayed together with the corresponding emotional state, making it easier for the user in the counselor role to view the communication text content and emotional state at different moments. When the emotion category corresponding to a time point in the emotional state diagram changes significantly, for example, from relaxation to anger, an abnormal fluctuation prompt information can be automatically generated and displayed on the corresponding display device of the server.

[0138] Among them, when the emotion recognition model adopts a long short-term memory network that integrates a self-attention mechanism, it can focus more on capturing the dynamic changes of the low-frequency and high-frequency ratio and speech emotion features; and when the emotion recognition model adopts a multimodal convolutional neural network that integrates a self-attention mechanism, it can perform more refined fusion at the feature level. Moreover, the self-attention mechanism used is to dynamically adjust the weight of each feature in the multimodal feature vector when each multimodal feature vector is input into the emotion recognition model to cope with the difference in the importance of features under different emotional states. For example, when the low-frequency and high-frequency ratio in the multimodal feature vector increases significantly, the weight of the low-frequency and high-frequency ratio feature will be automatically increased; when the change in speech features is more significant, the weight of the speech emotion feature will be increased.

[0139] Moreover, when the emotional state classification results corresponding to each multimodal feature vector are obtained based on the emotion recognition model in a chronological order, specifically, an emotional state classification sub-result can be output for the low-frequency and high-frequency ratio features in the multimodal feature vector, or an emotional state classification sub-result can be output for the speech emotion features in the multimodal feature vector, and the above two emotional state classification sub-results are weighted and summed to obtain a comprehensive emotional state classification result. However, it should be noted that when there is a significant difference between outputting an emotional state classification sub-result for the low-frequency and high-frequency ratio features in the multimodal feature vector and outputting an emotional state classification sub-result for the speech emotion features in the multimodal feature vector, performing the weighted summation can balance the emotional judgment results.

[0140] It can be seen that the embodiment of the device can collect the current millimeter-wave radar data set and the current user voice data set of the target user to be identified, perform feature extraction, time alignment and feature fusion respectively to obtain a multimodal feature vector set, and then input it into the emotion recognition model to obtain the recognition result, thereby improving the accuracy of emotion recognition.

[0141] The above-mentioned psychological counseling multimodal emotion recognition device based on millimeter wave radar can be implemented in the form of a computer program. The computer program can be used in Figure 7 Runs on the computer device shown.

[0142] See also Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device 400 integrates any of the millimeter-wave radar-based multimodal emotion recognition devices for psychological counseling provided by an embodiment of the present invention.

[0143] See Figure 7 The computer device 400 includes a processor 402 , a memory, and a network interface 405 connected via a system bus 401 , wherein the memory may include a storage medium 403 and an internal memory 404 .

[0144] The storage medium 403 can store an operating system 4031 and a computer program 4032. The computer program 4032 includes program instructions, which, when executed, can enable the processor 402 to execute a multimodal emotion recognition method for psychological counseling based on millimeter-wave radar.

[0145] The processor 402 is used to provide computing and control capabilities to support the operation of the entire computer device.

[0146] The internal memory 404 provides an environment for the operation of the computer program 4032 in the storage medium 403. When the computer program 4032 is executed by the processor 402, the processor 402 can execute the above-mentioned millimeter-wave radar-based multimodal emotion recognition method for psychological counseling.

[0147] The network interface 405 is used to communicate with other devices through the network. Figure 7 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0148] The processor 402 is configured to run a computer program 4032 stored in the memory to implement the multimodal emotion recognition method for psychological counseling based on millimeter-wave radar as described in any of the above embodiments.

[0149] It should be understood that in the embodiment of the present invention, the processor 402 may be a central processing unit (CPU), and the processor 402 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0150] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0151] Therefore, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to execute the multimodal emotion recognition method for psychological counseling based on millimeter-wave radar as described in any of the above embodiments.

[0152] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0153] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0154] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0155] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0156] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0157] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A multimodal emotion recognition method for psychological counseling based on millimeter wave radar, characterized in that: include: In response to a multimodal emotion recognition instruction, obtaining to-be-recognized user information corresponding to the multimodal emotion recognition instruction and a to-be-recognized target user corresponding to the to-be-recognized user information; Obtaining a current millimeter-wave radar data set corresponding to the target user to be identified; wherein the current millimeter-wave radar data set is composed of millimeter-wave radar data collected by the millimeter-wave radar module within the target collection time interval corresponding to the multimodal emotion recognition instruction after the millimeter-wave radar module is aimed at the target user to be identified; Acquire a current user voice data set corresponding to the target user to be identified; wherein the current user voice data set is composed of user voice data collected by the sound pickup and processing module from the target user to be identified within a target collection time interval; Based on a preset millimeter-wave radar feature extraction strategy, a heart rate variability data sequence corresponding to the current millimeter-wave radar data set and a low-frequency and high-frequency ratio feature sequence corresponding to the heart rate variability data sequence are acquired in time sequence; Based on a preset speech feature extraction model, a speech emotion feature sequence corresponding to the current user speech data set is obtained in time sequence; Based on a preset data fusion strategy, each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence is time-aligned and feature-fused with each voice emotion feature in the voice emotion feature sequence to obtain a multimodal feature vector set corresponding to the target user to be identified; wherein the multimodal feature vector set includes multiple multimodal feature vectors, and each multimodal feature vector is composed of a low-frequency and high-frequency ratio feature and a corresponding voice emotion feature that have completed time alignment; The multimodal feature vector set is input into a pre-trained emotion recognition model to obtain an emotion state classification result corresponding to each multimodal feature vector in the multimodal feature vector set to form an emotion state classification result output sequence.

2. The method according to claim 1, characterized in that Before the step of obtaining a current user voice data set corresponding to the target user to be identified, the method further includes: Acquire a current communication process speech data set corresponding to the target user to be identified; Based on the pre-collected voiceprint features of the user to be filtered, the voice signals in the current communication process voice data set and the voiceprint features of the user to be filtered are filtered to obtain the current user voice data set.

3. The method according to claim 1, characterized in that After the step of obtaining a current user voice data set corresponding to the target user to be identified, the method further includes: Speech recognition is performed on the current user speech data set based on a pre-trained speech recognition model to obtain current user text data including text time sequence information.

4. The method according to claim 1, wherein After the step of obtaining a current user voice dataset corresponding to the target user to be identified, and before the step of obtaining, in time sequence based on a preset millimeter-wave radar feature extraction strategy, a heart rate variability data sequence corresponding to the current millimeter-wave radar dataset and a low-frequency-high-frequency ratio feature sequence corresponding to the heart rate variability data sequence, the method further includes: Get the preset sliding window period; Dividing the current millimeter-wave radar data set based on the sliding window period to obtain a plurality of current millimeter-wave radar data subsets; wherein the plurality of current millimeter-wave radar data subsets constitute the current millimeter-wave radar data set; The current user voice data set is divided based on the sliding window period to obtain multiple current user voice data subsets; wherein the multiple current user voice data subsets constitute the current user voice data.

5. The method according to claim 1, wherein The method of obtaining a heart rate variability data sequence corresponding to the current millimeter wave radar data set and a low-frequency to high-frequency ratio feature sequence corresponding to the heart rate variability data sequence in a time sequence based on a preset millimeter wave radar feature extraction strategy includes: Acquire a current heart rate fitting dataset corresponding to the current millimeter-wave radar data according to a time sequence of acquiring each current millimeter-wave radar data subset in the current millimeter-wave radar data set; Performing data cleaning on the current heart rate fitting data set based on a preset data cleaning strategy to obtain a cleaned current heart rate fitting data set; Acquire a current RR interval sequence corresponding to the cleaned current heart rate fitting data set, and acquire the heart rate variability data sequence based on the current RR interval sequence and a preset power spectrum density estimation model; For each heart rate variability data in the heart rate variability data sequence, the ratio of low-frequency component / high-frequency cost is obtained as a low-frequency-high-frequency ratio feature, and the low-frequency-high-frequency ratio feature of each heart rate variability data in the heart rate variability data sequence is arranged in time series to form the low-frequency-high-frequency ratio feature sequence.

6. The method according to claim 1, characterized in that The method of obtaining a speech emotion feature sequence corresponding to the current user speech data set in time sequence based on a preset speech feature extraction model includes: For each current user voice data subset in the current user voice data set, input the current user voice data subset into the voice feature extraction model to perform voice emotion feature extraction to obtain a current voice emotion feature corresponding to the current user voice data subset; wherein the current voice emotion feature includes at least user pitch feature data, user voice intensity feature data, and user speech rate feature data; The current voice emotion features corresponding to each current user voice data subset in the current user voice data set are arranged in time sequence to form the voice emotion feature sequence.

7. The method according to claim 1, characterized in that The method of performing time sequence alignment and feature fusion on each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence and each voice emotion feature in the voice emotion feature sequence based on a preset data fusion strategy to obtain a multimodal feature vector set corresponding to the target user to be identified includes: Acquire a first timestamp for each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence, and acquire a second timestamp for each voice emotion feature in the voice emotion feature sequence; Aligning the low-frequency and high-frequency ratio features with the same first and second timestamps with the speech emotion features and performing feature fusion to form a multimodal feature vector; Arrange multiple multimodal feature vectors in chronological order to form the multimodal feature vector set.

8. A multimodal emotion recognition device for psychological counseling based on millimeter wave radar, characterized in that: include: a target user information acquisition unit, configured to, in response to a multimodal emotion recognition instruction, acquire user information to be identified corresponding to the multimodal emotion recognition instruction and a target user to be identified corresponding to the user information to be identified; a millimeter-wave radar data acquisition unit, configured to acquire a current millimeter-wave radar data set corresponding to the target user to be identified; wherein the current millimeter-wave radar data set is composed of millimeter-wave radar data collected by the millimeter-wave radar module within a target collection time interval corresponding to the multimodal emotion recognition instruction after the millimeter-wave radar module is aligned with the target user to be identified; A sound data acquisition unit, configured to acquire a current user voice data set corresponding to the target user to be identified; wherein the current user voice data set is composed of user voice data collected by the sound pickup and processing module from the target user to be identified within a target collection time interval; a low-frequency-high-frequency ratio feature acquisition unit, configured to acquire, in time sequence, a heart rate variability data sequence corresponding to the current millimeter-wave radar data set and a low-frequency-high-frequency ratio feature sequence corresponding to the heart rate variability data sequence based on a preset millimeter-wave radar feature extraction strategy; A voice emotion feature acquisition unit, configured to acquire, based on a preset voice feature extraction model, a voice emotion feature sequence corresponding to the current user voice data set in a time sequence; a multimodal feature fusion unit, configured to perform time-series alignment and feature fusion on each low-frequency and high-frequency ratio feature in the low-frequency and high-frequency ratio feature sequence and each voice emotion feature in the voice emotion feature sequence based on a preset data fusion strategy, thereby obtaining a multimodal feature vector set corresponding to the target user to be identified; wherein the multimodal feature vector set includes a plurality of multimodal feature vectors, and each multimodal feature vector is composed of a low-frequency and high-frequency ratio feature and a corresponding voice emotion feature that have completed time-series alignment; The classification result output unit is used to input the multimodal feature vector set into a pre-trained emotion recognition model to obtain the emotional state classification results corresponding to each multimodal feature vector in the multimodal feature vector set to form an emotional state classification result output sequence.

9. A computer device, characterized in that: The computer device includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the multimodal emotion recognition method for psychological counseling based on millimeter-wave radar as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the multimodal emotion recognition method for psychological counseling based on millimeter-wave radar as described in any one of claims 1 to 7 can be implemented.

Citation Information

Patent Citations

  • Multi-modal face emotion recognition method and device

    CN114399818A

  • Emotion recognition method and device and electronic equipment

    CN114767112A