Multimodal Mental State Detection Method and System Based on Spatiotemporal Attention Mechanism

By adopting a multimodal detection method based on the spatiotemporal attention mechanism in psychological state detection, the spatial and temporal information of facial expressions and speech is integrated, the problem of low detection accuracy in the prior art is solved, and more efficient psychological state detection is achieved.

CN117315738BActive Publication Date: 2025-06-27LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311083083.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2025-06-27
Estimated Expiration
2043-08-25

AI Technical Summary

Technical Problem

The existing psychological state detection methods ignore the information interaction between facial and speech data, and mine less semantic information in single-modal data, resulting in low detection accuracy.

Method used

A multimodal detection method based on the spatiotemporal attention mechanism is adopted. By obtaining facial expressions and speech data in the video stream data, facial expression features and speech features are extracted, and feature dimensions are unified, and inputted to the spatiotemporal attention converter to obtain spatiotemporal fusion features. Finally, the features are input to the psychological state classifier for detection.

Benefits of technology

By integrating multimodal spatiotemporal information of facial expressions and speech, the complementarity and collaboration effect of features between modals is improved, and the characteristics that distinguish different psychological states are extracted, which significantly improves the accuracy of psychological state detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315738B_ABST
    Figure CN117315738B_ABST
Patent Text Reader

Abstract

The present application provides a multi-modal mental state detection method and system based on a spatio-temporal attention mechanism. After obtaining facial expression data and speech data, the method can extract facial expression features and speech features from the facial expression data and speech data, and unify the feature dimensions. Then, the facial expression features and speech features with unified feature dimensions are input into a spatio-temporal attention transformer to obtain spatio-temporal fusion features. Finally, the spatio-temporal fusion features are input into a mental state classifier to obtain a classification result. The method uses two types of modal data, namely, human faces and speech in video stream data, extracts temporal features and spatial features in each single modality respectively, and fuses the spatio-temporal features of the two modalities, improving the complementary and collaborative effects of features between modalities. Furthermore, it can extract features that can distinguish different mental states, realize the detection of mental states using the social media data of users, and improve the accuracy of mental state detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and particularly to a multi-modal mental state detection method and system based on a spatio-temporal attention mechanism. Background Art

[0002] A mental state refers to the complete characteristics of mental activities within a certain period of time. According to the characteristics of mental states, various mental states can be distinguished, such as attention, fatigue, tension, relaxation, sadness, joy, etc. Different mental states can produce different manifestations. Therefore, the mental state of a user can be detected and classified based on the user's manifestations, and then the user's mental state can be detected. For example, users with depression have specific manifestations in terms of mental states, including emotional depression, slow thinking, reduced speech and motor activities, slowness, etc. Therefore, the mental state detection method can assist in detecting depression.

[0003] The mental state detection method can use facial expression data or voice data to extract relevant features, and then detect the mental state. However, this method ignores the information interaction between facial and voice data. It is also possible to use text, voice, and facial expressions as inputs to a mental state detection model for detection. However, this method mines less semantic information in single-modal data and cannot simultaneously focus on the semantic information and temporal information of single-modal data, reducing the detection accuracy. Summary of the Invention

[0004] This application provides a multi-modal mental state detection method and system based on a spatio-temporal attention mechanism to solve the problem of low mental state detection accuracy.

[0005] In a first aspect, this application provides a multi-modal mental state detection method based on a spatio-temporal attention mechanism, including:

[0006] Obtain video stream data, where the video stream data includes facial expression data and voice data;

[0007] Extract facial expression features from the facial expression data and extract voice features from the voice data;

[0008] Unify the feature dimensions of the facial expression features and the voice features;

[0009] Input the facial expression features and the speech features after unifying the feature dimensions into a spatio-temporal attention transformer to obtain the spatio-temporal fusion features output by the spatio-temporal attention transformer. The spatio-temporal fusion features include the spatio-temporal fusion features of facial expressions and the spatio-temporal fusion features of speech. The spatio-temporal attention transformer includes a spatial attention module, a temporal attention module, and a multi-modal fusion transformer. The spatial attention module is used to extract the unimodal spatial information of facial expressions and speech, the temporal attention module is used to extract the unimodal temporal information of facial expressions and speech, and the multi-modal fusion transformer is used to fuse the multi-modal spatio-temporal information of facial expressions and speech.

[0010] Input the spatio-temporal fusion features into a mental state classifier to obtain the classification result output by the mental state classifier.

[0011] Optionally, the method further includes: identifying the face region in the facial expression data;

[0012] Extract the facial feature points of the face region to obtain facial expression features;

[0013] Extract a preset number of speech features depicting speech in the speech data to obtain speech features.

[0014] Optionally, the method further includes: respectively detecting the data lengths of the facial expression features and the speech features;

[0015] If the data length is greater than a preset length threshold, then crop the facial expression features and the speech features according to the preset length threshold;

[0016] If the data length is less than a preset length threshold, then perform an interpolation operation on the facial expression features and the speech features according to the preset length threshold.

[0017] Optionally, the method further includes: performing data normalization processing on the facial expression features and the speech features;

[0018] Perform smoothing processing on the facial expression features and the speech features.

[0019] Optionally, the step of unifying the feature dimensions of the facial expression features and the speech features includes:

[0020] Input the facial expression features and the speech features into a linear mapping layer, and the linear mapping layer includes at least two layers of one-dimensional convolution;

[0021] Obtain the facial expression features and the speech features with the same matrix shape output by the linear mapping layer.

[0022] Optionally, the spatial attention module extracts the unimodal spatial information of facial expressions and speech according to the following formula:

[0023]

[0024]

[0025]

[0026] where Re() is a matrix shape transformation function, X Sm is the facial expression feature or speech feature after unifying the feature dimensions, tanh() is an activation function, Q is the query element in the self-attention mechanism, K is the key element in the self-attention mechanism, V is the value element in the self-attention mechanism, LN() is a data normalization layer, and X′ Sm is the spatial feature encoded by the spatial attention module.

[0027] Optionally, the temporal attention module extracts the unimodal temporal information of facial expressions and speech according to the following formula:

[0028] X″ Sm =LN(Att s (Re′(X′ Sm ))+Re′(X′ Sm ));

[0029] where X″ Sm is the temporal feature encoded by the temporal attention module.

[0030] Optionally, the multimodal fusion converter fuses the multimodal spatio-temporal information of facial expressions and speech according to the following formula:

[0031]

[0032]

[0033] where MultiHead() is a multi-head attention mechanism, Q a is the query element in the facial modality, K v is the key element in the speech modality, V v is the value element in the speech modality, X″ Sa is the spatio-temporal feature in the facial modality or speech modality obtained through the spatial attention module and the temporal attention module, FFN() is a two-layer fully connected layer with an activation function, is the spatio-temporal fusion feature of facial expressions or speech.

[0034] Optionally, the mental state classifier includes a multi-layer perceptron layer and a fully connected layer, and the multi-layer perceptron layer includes multiple layers of fully connected linear layers.

[0035] In a second aspect, the present application provides a multi-modal mental state detection system based on a spatio-temporal attention mechanism, including:

[0036] A data acquisition module for acquiring video stream data, where the video stream data includes facial expression data and voice data;

[0037] A feature extraction module for extracting facial expression features from the facial expression data and voice features from the voice data;

[0038] A preprocessing module for unifying the feature dimensions of the facial expression features and the voice features;

[0039] A multi-modal fusion module for inputting the facial expression features and the voice features with unified feature dimensions into a spatio-temporal attention transformer to obtain spatio-temporal fusion features output by the spatio-temporal attention transformer. The spatio-temporal fusion features include spatio-temporal fusion features of facial expressions and spatio-temporal fusion features of voice. The spatio-temporal attention transformer includes a spatial attention module, a temporal attention module, and a multi-modal fusion transformer. The spatial attention module is used to extract single-modal spatial information of facial expressions and voices, the temporal attention module is used to extract single-modal temporal information of facial expressions and voices, and the multi-modal fusion transformer is used to fuse multi-modal spatio-temporal information of facial expressions and voices.

[0040] A detection and classification module for inputting the spatio-temporal fusion features into a mental state classifier to obtain a classification result output by the mental state classifier.

[0041] As can be seen from the above technical solutions, the present application provides a multi-modal mental state detection method and system based on a spatio-temporal attention mechanism. After obtaining facial expression data and speech data, the method can extract facial expression features and speech features from the facial expression data and speech data, and unify the feature dimensions. Then, the facial expression features and speech features with unified feature dimensions are input into a spatio-temporal attention transformer to obtain spatio-temporal fusion features. Among them, the spatio-temporal fusion features include the spatio-temporal fusion features of facial expressions and the spatio-temporal fusion features of speech. The spatio-temporal attention transformer includes a spatial attention module, a temporal attention module, and a multi-modal fusion transformer. The spatial attention module is used to extract the single-modal spatial information of facial expressions and speech, the temporal attention module is used to extract the single-modal temporal information of facial expressions and speech, and the multi-modal fusion transformer is used to fuse the multi-modal spatio-temporal information of facial expressions and speech. Finally, the spatio-temporal fusion features are input into a mental state classifier to obtain a classification result. By using two-modal data of faces and voices from social media data, time features and spatial features are respectively extracted in a single modality, and the spatio-temporal features of the two modalities are fused to improve the complementary and collaborative effects of features between modalities. Furthermore, features that can distinguish different mental states can be extracted, enabling the use of users' social media data to detect mental states and improving the accuracy of mental state detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 It is a flowchart of the multi-modal mental state detection method based on the spatio-temporal attention mechanism provided by the embodiment of the present application;

[0044] Figure 2 It is a flowchart of extracting facial expression features and speech features provided by the embodiment of the present application;

[0045] Figure 3 It is a flowchart of mental state detection provided by the embodiment of the present application;

[0046] Figure 4 It is a structural diagram of the spatio-temporal attention transformer provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] Embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following embodiments do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0048] To improve the detection accuracy of mental states, an embodiment of the present application provides a multi-modal mental state detection method based on a spatio-temporal attention mechanism. The method applies the spatio-temporal attention mechanism to multi-modal mental state detection. By using two-modal data of face and voice from social media data, temporal features and spatial features are extracted separately in a single modality, and the spatio-temporal features of the two modalities are fused to improve the complementarity and cooperation of features between modalities. Furthermore, features capable of distinguishing different mental states can be extracted, enabling the detection of mental states using the user's social media data. It should be noted that the multi-modal mental state detection method based on the spatio-temporal attention mechanism provided in the embodiment of the present application can be applied to the detection of mental states based on facial expression data and voice data to improve the detection accuracy. For example, it can assist in detecting whether a user has a risk of depression.

[0049] As Figure 1 shown, Figure 1 is a schematic flowchart of the multi-modal mental state detection method based on the spatio-temporal attention mechanism in the embodiment of the present application, which specifically includes the following content:

[0050] S100: Obtain video stream data.

[0051] Among them, the video stream data is the video stream data of the user to be detected for mental state, including facial expression data and voice data. The source of the video stream data can be set according to actual needs. For example, for assisting in detecting whether a user has depression, since users with depression are more inclined to disclose their mental conditions on social media, social media data of the user can be obtained, and then by analyzing the facial expression data and voice data in the social media data, the mental state of the user can be detected.

[0052] It can be understood that the video stream data is stacked by semantic information in the time dimension. For facial expression data, temporal features more express the trend of emotional changes, while semantic features more carry the mutual information between facial feature points. And compared with data such as gait and physiological signals, facial expression data and voice data are easier to capture. Therefore, by analyzing the information interaction between facial expression data and voice data, and integrating the spatio-temporal information embedded in the data, and then detecting the mental state, the detection accuracy and efficiency can be improved.

[0053] S200: Extract facial expression features from facial expression data and extract speech features from speech data.

[0054] After obtaining the facial expression data and speech data, to prevent problems such as data overfitting and privacy leakage, facial expression features can be extracted from the facial expression data, and speech features can be extracted from the speech data.

[0055] As Figure 2 shown, for the facial expression data, face detection can be performed on the facial expression data to identify the face region in the facial expression data, and the facial feature points in the face region can be extracted to obtain facial expression features. For example, based on the CNN network, face detection is performed on the facial expression data, and 68 facial feature points are extracted as facial expression features.

[0056] In addition, since the facial expression data captured in the actual scenario may have the face angle in the facial expression data tilted due to reasons such as the camera angle and shooting distance. To avoid the influence of factors such as the tilted face angle on the stability of the later model, the facial feature points can be corrected, and the original facial feature points are subjected to coordinate transformation to align the facial feature points in different poses.

[0057] For the speech data, a preset number of speech features can be extracted from the speech data to obtain speech features. For example, the opensmile toolkit is used to extract 25 speech features from the speech data as the speech features of the speech data, including loudness, mel cepstrum parameters, frame energy, critical band spectrum, auditory spectrum, etc.

[0058] In addition, since there may be noise interference in the speech data, to improve the detection accuracy, the speech data can be filtered. For example, the speech data is filtered based on Kalman filtering to filter out the noise interference in the speech data.

[0059] In some embodiments, when training a model based on the spatio-temporal attention mechanism (spatio-temporal attention transformer and mental state classifier), the model can be trained based on the sample video stream data, where the sample video stream data includes sample video stream data marked with different mental state labels. For example, for assisting in detecting whether a user has depression, social media data including facial expression data and speech data of users with depression and normal users can be used as the sample video stream data.

[0060] Since the lengths of video stream data posted by different users in social media are different. Therefore, the lengths of video stream data of different lengths can be normalized to increase the amount of data and prevent overfitting during model training. Specifically, a length threshold L can be set. For data with a length greater than the length threshold L, the data is cropped to the data length L. For data with a length less than the length threshold L, interpolation operations can be performed, such as using polynomial interpolation fitting methods for interpolation operations. In this way, the number of samples in the training set, validation set, and test set will increase.

[0061] Therefore, after extracting the facial expression features and speech features, length normalization processing can be performed on them. The lengths of the facial expression features and speech features are detected respectively. If the data length is greater than the preset length threshold L, the facial expression features and speech features are cropped according to the preset length threshold L. If the data length is less than the preset length threshold L, interpolation operations are performed on the facial expression features and speech features according to the preset length threshold L.

[0062] Among them, for data with a short data duration, using an overly large length threshold L will result in information loss. And using an overly small length threshold L will result in the segmented data having less temporal information. Therefore, the length threshold L can be set to 60.

[0063] In some embodiments, in order to improve the convergence speed of the model, data normalization processing can be performed on the facial expression features and speech features, such as max-min normalization, and smoothing processing can be performed on the facial expression features and speech features.

[0064] S300: Unify the feature dimensions of the facial expression features and speech features.

[0065] Since the dimensions of the facial expression features and speech features do not match, in order to ensure the consistency of dimensions in feature fusion, the facial expression features and speech features can be preprocessed to unify the feature dimensions of the facial expression features and speech features.

[0066] In some embodiments, as Figure 3 shown, the facial expression features and speech features can be input into a linear mapping layer to obtain facial expression features and speech features with the same matrix shape as the output of the linear mapping layer. Among them, the linear mapping layer includes at least two layers of one-dimensional convolution (Conv1d). The linear mapping layer is a linear mapping function Re() implemented by using two layers of Conv1d to change the matrix shape before inputting the data into the model, that is, a matrix shape transformation function (reshape function).

[0067] The feature dimensions of the facial expression features and speech features can be unified according to the following formula:

[0068]

[0069] Among them, Re() is a matrix shape transformation function, which is a linear mapping layer implemented using at least two layers of Conv1d, and is used to make the facial expression features and speech features have the same matrix shape. represents facial expression features or speech features, where m ∈ {a, v}, a represents the facial modality, v represents the speech modality, and n is the dimension of the single-modal data features. X Sm is the data obtained after unifying the dimensions of the data of the two modalities, that is, the facial expression features or speech features after the feature dimensions are unified.

[0070] In order to enable the encoder of the model to utilize the temporal and spatial order relationships in the video stream data, additional learnable temporal and spatial position information can be expressed as:

[0071] X t = X t + PE N ; X n = X n + PE T ;

[0072] Among them, PE N represents the matrix of N semantic word segments (tokens) shared among all frames. X n represents the sequence composed of data points of the same type in all frames. All types of data points share the tokens matrix PE T at time T, X t represents the sequence composed of all data points at time T. That is to say, all data points in the same frame share the same temporal position encoding. The entire model jointly trains PE N and PE T .

[0073] S400: Input the facial expression features and speech features with unified feature dimensions into the spatio-temporal attention transformer to obtain the spatio-temporal fusion features output by the spatio-temporal attention transformer.

[0074] Among them, the spatio-temporal fusion features include the spatio-temporal fusion features of facial expressions and the spatio-temporal fusion features of speech. Such as Figure 3 , Figure 4As shown in the figure, the spatial-temporal attention Transformer (STAT) can capture the spatial-temporal correlations of multi-modal data simultaneously, including the Spatial Attention Block (SAB), the Temporal Attention Block (TAB), and the Multimodal Fusion Transformer (MTB). To deepen the model depth and extract more effective feature representations, multiple spatial-temporal attention transformers can be used, that is, multiple spatial-temporal attention transformers are stacked. The number of stacked spatial-temporal attention transformers can be set according to actual needs.

[0075] The spatial attention module models the facial expressions or speech states represented by each data point in each frame, and is used to extract the unimodal spatial information of facial expressions and speech. The temporal attention module captures the change patterns of facial expressions or speech, and is used to extract the unimodal temporal information of facial expressions and speech. The multimodal fusion transformer is used to fuse the multimodal spatial-temporal information of facial expressions and speech.

[0076] It can be understood that the model provided in this application based on the spatial-temporal attention mechanism is based on the core idea of the attention mechanism model (Transformer), that is, the use of the attention mechanism. However, it does not directly adopt the Transformer structure, but involves a new structure, namely STAT, to make it more suitable for the detection of multi-modal mental states.

[0077] In some embodiments, the spatial attention module is used to extract the spatial features of the data and perform self-attention mechanism calculations in the form of residuals. And since the tanh function is more flexible than the softmax function, the tanh function is used as the activation function. That is, the spatial attention module extracts the unimodal spatial information of facial expressions and speech according to the following formula:

[0078]

[0079]

[0080]

[0081] where tanh() is the activation function, and the self-attention value can be calculated. In the self-attention mechanism, Q is the query element in the self-attention mechanism (Query), K is the key element in the self-attention mechanism (Key), and V is the value element in the self-attention mechanism (Value), which can be obtained through linear transformation. Re() is the matrix shape transformation function, X SmThe facial expression features or speech features after unifying the feature dimensions, LN() is the data normalization layer, X′ Sm is the spatial feature encoded by the spatial attention module, that is, the unimodal spatial information.

[0082] The correlation between frames in the time dimension is related to the length of the time interval. To better construct complex and uncertain correlations, the time attention module provided in this application can capture long-distance and short-distance correlations in the time dimension of the input data. The unimodal time information is captured through the self-attention mechanism in the form of residuals. That is, the time attention module extracts the unimodal time information of facial expressions and speech according to the following formula:

[0083] X″ Sm =LN(Att s (Re′(X′ Sm ))+Re′(X′ Sm ));

[0084] Among them, X″ Sm is the time feature encoded by the time attention module, that is, the unimodal time information.

[0085] In this embodiment, the data of the two single modes are complementary. For example, when the user's mental state is happy, their facial expression will express positive emotions and their intonation will be more relaxed. Therefore, the multimodal fusion converter is a multi-head self-attention conversion structure that can fuse spatio-temporal information from different modalities and adaptively adjust the weights of various types of features. That is, the multimodal fusion converter fuses the multimodal spatio-temporal information of facial expressions and speech according to the following formula:

[0086]

[0087]

[0088] Among them, MultiHead() is the multi-head attention mechanism, Q a is the query element (Query) in the facial modality, K v is the key element (Key) in the speech modality, V v is the value element (Value) in the speech modality, X″ Sa is the spatio-temporal feature in the facial modality or speech modality obtained through the spatial attention module and the time attention module, FFN() is a two-layer fully connected layer with an activation function, is the spatio-temporal fusion feature of facial expressions or speech. Thus, the spatio-temporal fusion feature of facial expressions and the spatio-temporal fusion feature of speech can be obtained.

[0089] S500: Input the spatio-temporal fusion feature into the mental state classifier to obtain the classification result output by the mental state classifier.

[0090] After obtaining the spatio-temporal fusion features of facial expressions and the spatio-temporal fusion features of speech, input them into the mental state classifier for mental state detection, and then obtain the classification results output by the mental state classifier. The classification results are the mental states detected according to the video stream data. For example, for assisting in detecting whether a user has depression, the classification results include two detection results: having depression and not having depression.

[0091] Among them, the mental state classifier includes a Multi-Layer Perceptron (MLP) layer and a fully connected layer (softmax layer), and the multi-layer perceptron layer includes multiple fully connected linear layers.

[0092] Since a user has multiple segments of data (which can be seen in the above normalization description), multiple classification results will be generated. To avoid the instability of the classification results caused by data fluctuations, a mental state classifier with a voting mechanism can be set, and the voting mechanism is used to take the category with the most classifications as the final classification category.

[0093] In this embodiment, in order to improve the accuracy of mental state detection, a spatio-temporal attention multi-modal mental state detection method based on the Transformer model is proposed. A spatio-temporal attention converter is set to extract the spatio-temporal features of the data and effectively fuse the spatio-temporal information of multi-modal data. The mutual complementarity of spatial features and temporal features makes the model performance more stable. And using a classifier with a voting mechanism can better classify the mental state.

[0094] Based on the above multi-modal mental state detection method based on the spatio-temporal attention mechanism, the embodiment of the present application also provides a multi-modal mental state detection system based on the spatio-temporal attention mechanism, including:

[0095] A data acquisition module for acquiring video stream data.

[0096] Among them, the video stream data includes facial expression data and voice data.

[0097] A feature extraction module for extracting facial expression features from facial expression data and voice features from voice data.

[0098] A preprocessing module for unifying the feature dimensions of facial expression features and voice features.

[0099] A multi-modal fusion module for inputting the facial expression features and voice features with unified feature dimensions into the spatio-temporal attention converter to obtain the spatio-temporal fusion features output by the spatio-temporal attention converter.

[0100] Among them, the spatio-temporal fusion features include the spatio-temporal fusion features of facial expressions and the spatio-temporal fusion features of speech. The spatio-temporal attention transformer includes a spatial attention module, a temporal attention module, and a multimodal fusion transformer. The spatial attention module is used to extract the unimodal spatial information of facial expressions and speech. The temporal attention module is used to extract the unimodal temporal information of facial expressions and speech. The multimodal fusion transformer is used to fuse the multimodal spatio-temporal information of facial expressions and speech.

[0101] The detection and classification module is used to input the spatio-temporal fusion features into the mental state classifier to obtain the classification result output by the mental state classifier.

[0102] It can be seen from the above technical solutions that the present application provides a multimodal mental state detection method and system based on a spatio-temporal attention mechanism. After obtaining facial expression data and speech data, the method can extract facial expression features and speech features from the facial expression data and speech data and unify the feature dimensions. Then, the facial expression features and speech features with unified feature dimensions are input into the spatio-temporal attention transformer to obtain spatio-temporal fusion features. Among them, the spatio-temporal fusion features include the spatio-temporal fusion features of facial expressions and the spatio-temporal fusion features of speech. The spatio-temporal attention transformer includes a spatial attention module, a temporal attention module, and a multimodal fusion transformer. The spatial attention module is used to extract the unimodal spatial information of facial expressions and speech. The temporal attention module is used to extract the unimodal temporal information of facial expressions and speech. The multimodal fusion transformer is used to fuse the multimodal spatio-temporal information of facial expressions and speech. Finally, the spatio-temporal fusion features are input into the mental state classifier to obtain the classification result. By using two-modal data of face and speech from social media data, temporal features and spatial features are respectively extracted in a single modality, and the spatio-temporal features of the two modalities are fused to improve the complementarity and cooperation of features between modalities, and then features that can distinguish different mental states can be extracted, realizing the detection of mental states using the social media data of users.

[0103] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other implementation manner extended based on the solution of the present application without creative efforts belongs to the protection scope of the present application.

Claims

1. A multi-modal mental state detection method based on spatio-temporal attention mechanism, characterized in that, Including: Obtain video stream data, where the video stream data includes facial expression data and voice data; Extract facial expression features from the facial expression data and extract voice features from the voice data; Unify the feature dimensions of the facial expression features and the voice features; Input the facial expression features and the voice features with unified feature dimensions into a spatio-temporal attention transformer to obtain spatio-temporal fusion features output by the spatio-temporal attention transformer. The spatio-temporal fusion features include spatio-temporal fusion features of facial expressions and spatio-temporal fusion features of voices. The spatio-temporal attention transformer includes a spatial attention module, a temporal attention module, and a multimodal fusion transformer. The spatial attention module is used to extract unimodal spatial information of facial expressions and voices, the temporal attention module is used to extract unimodal temporal information of facial expressions and voices, and the multimodal fusion transformer is used to fuse multimodal spatio-temporal information of facial expressions and voices; Input the spatio-temporal fusion features into a mental state classifier to obtain a classification result output by the mental state classifier; The spatial attention module extracts unimodal spatial information of facial expressions and voices according to the following formula: where Re( ) is a matrix shape transformation function, and X Sm is the facial expression feature or speech feature after unifying the feature dimensions, tanh( ) is an activation function, Q is the query element in the self-attention mechanism, K is the key element in the self-attention mechanism, V is the value element in the self-attention mechanism, LN( ) is a data normalization layer, and X' Sm is the spatial feature encoded by the spatial attention module; The temporal attention module extracts unimodal temporal information of facial expressions and voices according to the following formula: X″ Sm = LN(Att s (Re′(X′ Sm )) + Re′(X′ Sm )); Among them, X″ Sm is the time feature after encoding by the time attention module; The multimodal fusion transformer fuses multimodal spatio-temporal information of facial expressions and voices according to the following formula: Among them, MultiHead( ) is the multi-head attention mechanism, and Q a is the query element in the facial modality, and K v is the key element in the speech modality, and V v is the value element in the speech modality, and X' S ' a is the spatio-temporal feature in the facial modality or speech modality obtained by the spatial attention module and the temporal attention module. FFN( ) is a two-layer fully connected layer with an activation function, which is the spatio-temporal fusion feature of facial expressions or speech.

2. The multi-modal mental state detection method based on spatio-temporal attention mechanism according to claim 1, wherein, Also including: Identify the face region in the facial expression data; Extract facial feature points of the face region to obtain facial expression features; Extract a preset number of voice features depicting voice features in the voice data to obtain voice features.

3. The multi-modal mental state detection method based on spatio-temporal attention mechanism according to claim 1, wherein Also including: Detect the data lengths of the facial expression features and the voice features respectively; If the data length is greater than a preset length threshold, crop the facial expression features and the voice features according to the preset length threshold; If the data length is less than the preset length threshold, perform an interpolation operation on the facial expression features and the voice features according to the preset length threshold.

4. The multi-modal mental state detection method based on spatio-temporal attention mechanism according to claim 1, wherein, Also including: Perform data normalization processing on the facial expression features and the voice features; Perform smoothing processing on the facial expression features and the voice features.

5. The multi-modal mental state detection method based on spatio-temporal attention mechanism according to claim 1, wherein The step of unifying the feature dimensions of the facial expression features and the voice features includes: Input the facial expression features and the voice features into a linear mapping layer, and the linear mapping layer includes at least two layers of one-dimensional convolution; Obtain the facial expression features and the voice features with the same matrix shape output by the linear mapping layer.

6. The multimodal mental state detection method based on spatio-temporal attention mechanism according to claim 1, wherein, The mental state classifier includes a multi-layer perceptron layer and a fully connected layer, and the multi-layer perceptron layer includes multiple fully connected linear layers.

7. A multi-modal mental state detection system based on a spatio-temporal attention mechanism, characterized in that, Including: A data acquisition module for obtaining video stream data, where the video stream data includes facial expression data and voice data; A feature extraction module for extracting facial expression features from the facial expression data and extracting voice features from the voice data; A preprocessing module for unifying the feature dimensions of the facial expression features and the voice features; The multi-modal fusion module is used to input the facial expression features and the speech features with unified feature dimensions into the spatio-temporal attention transformer to obtain the spatio-temporal fusion features output by the spatio-temporal attention transformer. The spatio-temporal fusion features include the spatio-temporal fusion features of facial expressions and the spatio-temporal fusion features of speech. The spatio-temporal attention transformer includes a spatial attention module, a temporal attention module, and a multi-modal fusion transformer. The spatial attention module is used to extract the unimodal spatial information of facial expressions and speech. The temporal attention module is used to extract the unimodal temporal information of facial expressions and speech. The multi-modal fusion transformer is used to fuse the multi-modal spatio-temporal information of facial expressions and speech; The detection and classification module is used to input the spatio-temporal fusion features into the mental state classifier to obtain the classification result output by the mental state classifier; The spatial attention module extracts the unimodal spatial information of facial expressions and speech according to the following formula: where Re( ) is a matrix shape transformation function, and X Sm is the facial expression feature or voice feature after the unification of the feature dimension, tanh( ) is an activation function, Q is the query element in the self-attention mechanism, K is the key element in the self-attention mechanism, V is the value element in the self-attention mechanism, LN( ) is a data normalization layer, and X' Sm is the spatial feature encoded by the spatial attention module; The temporal attention module extracts the unimodal temporal information of facial expressions and speech according to the following formula: X″ Sm = LN(Att s (Re′(X′ Sm )) + Re′(X′ Sm )); wherein, X″ Sm is the time feature after encoding by the time attention module; The multi-modal fusion transformer fuses the multi-modal spatio-temporal information of facial expressions and speech according to the following formula: Among them, MultiHead() is the multi-head attention mechanism, and Q a is the query element in the facial modality, and K v is the key element in the speech modality, and V v is the value element in the speech modality, X' S ' a is the spatio-temporal feature in the facial modality or speech modality obtained through the spatial attention module and the temporal attention module. FFN() is a two-layer fully connected layer with an activation function, which is the spatio-temporal fusion feature of facial expressions or speech.