Multimodal feature consistency mental health abnormality recognition method and system

By processing video data through micro-expression amplification and feature enhancement, and combining it with voice features, a multimodal feature consistency method for identifying mental health abnormalities is constructed. This method solves the problem of failing to effectively utilize micro-expressions and modal feature consistency in existing technologies, improves recognition accuracy and detection performance, and provides an objective reference for mental health diagnosis.

CN116230234BActive Publication Date: 2026-03-17HEBEI UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310265823.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2026-03-17
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

In existing technologies, mental health abnormality identification methods based on audio and video multimodal information fail to effectively utilize micro-expression features and do not consider the consistency issues between different modal features, resulting in insufficient recognition accuracy.

Method used

By processing video data through micro-expression amplification and feature enhancement, and combining it with speech features, a multimodal feature consistency method for identifying mental health abnormalities is constructed. Deep neural networks are used for feature extraction and fusion, and an audio-video multimodal consistency loss function is used for training.

Benefits of technology

It improves the accuracy and detection performance of identifying mental health abnormalities, provides more objective diagnostic evidence, enhances the model's performance, and is suitable for self-testing in mobile app and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116230234B_ABST
    Figure CN116230234B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for identifying mental health abnormalities based on multimodal feature consistency, comprising the following steps: acquiring raw data files containing both audio and video files from the same scene; preprocessing the video and audio data within the acquired raw data files; normalizing the micro-expression severity scores of each frame in a continuous frame sequence to obtain micro-expression keyframe vectors; and inputting the frequency features extracted from the audio data into an audio stream deep feature extraction network to obtain deep speech features F. A And audio feature prediction results y A A continuous sequence of frames is simultaneously fed into a video stream depth feature extraction network to obtain depth video features F. V And video feature prediction results y V ; after that, F A and F V Feature fusion is performed, and then the data is fed into a mental health classification network for final prediction. This invention effectively utilizes facial micro-expression features and voice features to identify mental health abnormalities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mental health abnormality recognition technology, specifically to a method and system for recognizing mental health abnormalities based on multimodal feature consistency of micro-expression amplification and voice features. This method uses micro-expression and voice features of audio and video as the main judgment criteria to recognize mental health abnormalities. Background Technology

[0002] With the continuous development of the economy and society, the mental health of the people has gradually received widespread attention from all sectors of society. Therefore, it is necessary to study a method for identifying mental health abnormalities, improve doctors' diagnostic efficiency, and provide doctors with a relatively objective reference basis for diagnosis.

[0003] Microexpressions, as involuntary movements of facial muscles, occur when people attempt to conceal their inner emotions; they cannot be faked or suppressed. Compared to consciously made expressions, microexpressions reveal people's true feelings and motivations more clearly, thus microexpression features are often used as an important objective indicator in automated depression detection. Similar to microexpression facial features, vocal features can also reflect an individual's emotional changes and abnormal psychological states, objectively and reliably reflecting the speaker's true psychological state. Therefore, these two indicators have important reference value for the analysis of mental illnesses. However, microexpressions are very subtle and short-lived, making them difficult to capture, thus often proving challenging when using video for mental health research.

[0004] In recent years, research in the field of artificial intelligence and machine learning has emerged on identifying abnormal mental states based on multimodal audio and video information, with automated depression detection being the most widespread. However, most existing studies utilizing multimodal audio and video information for mental health anomaly identification do not consider micro-expressions, a feature that significantly reflects psychological and emotional states. Furthermore, existing micro-expression-based research is mostly unimodal, failing to incorporate speech features for mental health analysis. Simultaneously, existing research often handles multimodal features in a simplistic manner, neglecting the consistency issues between different modalities within the same sample. For example, Chinese patent CN 112560811 B proposes an end-to-end automated depression detection method based on audio and video. This method, based on audio and video data, first preprocesses the raw data, which includes long-duration audio and video files, into segments. Then, the audio and video segments are input into audio and video feature extraction networks respectively, yielding deep audio and video features. A multi-head attention mechanism is used to calculate these deep speech and video features, resulting in attention-based audio and video features. These features are then aggregated into audio-video features via a feature aggregation module and finally fed into a decision network to predict an individual's depression level. While this method also uses audio and video data as input, it fails to consider the micro-expression features of the video data, thus failing to effectively capture micro-expression information and lacking consideration for the feature consistency issue between the audio and video modalities.

[0005] Based on the above reasons, this invention proposes a method and system for identifying mental health abnormalities based on the consistency of multimodal features of micro-expressions and speech features. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the technical problem this invention aims to solve is to provide a method and system for identifying mental health abnormalities based on multimodal feature consistency of micro-expression and voice features, which can better handle the task of identifying mental health abnormalities.

[0007] The technical solution adopted by the present invention to solve the aforementioned technical problem is as follows:

[0008] In a first aspect, the present invention provides a method for identifying mental health abnormalities based on multimodal feature consistency, the method comprising the following:

[0009] Obtain raw data files containing both audio and video files from the same scenario, including normal and abnormal samples, with disease type labels set for the abnormal samples;

[0010] The video data containing complete facial images in the obtained raw data file is subjected to micro-expression magnification processing. Then, the video frames are scored according to the magnitude of the facial micro-expression movements in each frame, and the micro-expression degree score of each frame is recorded. The facial images are then registered and aligned, thus completing the micro-expression magnification processing of the video data. Then, the magnified video data is sampled continuously at a set sampling interval to obtain multiple continuous frame sequences with the same dimensions, thus completing the preprocessing of the video data.

[0011] The micro-expression severity score of each frame in the continuous frame sequence is normalized to obtain the micro-expression keyframe vector;

[0012] The audio data in the obtained raw data file is subjected to noise reduction and noise removal processing to extract the frequency features of the audio data.

[0013] The frequency features extracted from the audio data are fed into an audio stream deep feature extraction network to extract audio deep features, resulting in deep speech features F. A Simultaneously, speech feature prediction is performed on the speech information in the audio stream deep feature extraction network to obtain the audio feature prediction result. ;

[0014] A deep feature extraction network for video streams is constructed, comprising a C3D ResNet50 with a depooling layer, a spatial attention module, a channel attention module, a temporal attention module, and a classification network connected in sequence. The input to the temporal attention module is the product of the output of the channel attention module and the micro-expression keyframe vectors. The output of the temporal attention module is averaged across the frame sequence to obtain the deep video features. ;

[0015] Multiple consecutive frame sequences with the same dimensions obtained from video data are simultaneously fed into a video stream depth feature extraction network to extract video depth features, resulting in depth video features. In a video stream deep feature extraction network, video feature prediction is performed on video information from a continuous frame sequence. The predicted video feature results are then obtained through a classification network. ;

[0016] The deep video features and deep speech features are fused using a tensor fusion network to obtain audio-video multimodal fusion features; these features are then fed into a mental health classification network for final prediction.

[0017] Calculate the audio-visual multimodal consistency loss according to formula (2). ,

[0018] (2)

[0019] Where n is the number of samples, and M is the number of label categories. Let represent the label of the i-th sample. This represents the model's predicted probability that the i-th sample belongs to class c. The penalty coefficient; function The penalty function increases the penalty strength of the loss function when the audio feature prediction results and video feature prediction results are inconsistent. The penalty function is defined as follows:

[0020] (3).

[0021] Furthermore, the aforementioned The set value is 1.0-1.6; preferably 1.55.

[0022] Furthermore, the audio stream deep feature extraction network includes a 2D ResNet18 network, a temporal attention mechanism, and a fully connected layer. The output of the pre-trained 2D ResNet18 network is connected to the temporal attention mechanism, and the output of the temporal attention mechanism is connected to a fully connected layer.

[0023] The temporal attention mechanism consists of one 1D convolutional layer, one fully connected layer, and one softmax function, which compresses features in both spatial and channel dimensions, extracts features in the temporal dimension, and obtains deep speech features.

[0024] Furthermore, the backbone network of the video stream deep feature extraction network is a pre-trained C3DResNet50. After feature extraction by the backbone network, the features are sequentially fed into the spatial attention module and the channel attention module for spatial and channel dimension information integration and weight allocation. Next, the feature vector Fsc obtained by the channel attention module is multiplied with the micro-expression key frame vector Fem. Then, it is fed into the temporal attention module for temporal dimension information integration at the frame sequence level.

[0025] Each residual block of C3D ResNet50 consists of three cascaded 3D convolutions, with the kernel size of the first 3D convolution being [missing value]. With a stride of 1, the kernel size of the second 3D convolution is... With a stride of 2, the kernel size of the third 3D convolution is... The stride is 1; each 3D convolution is followed by a 3D batch regularization and a ReLU activation function.

[0026] Furthermore, the micro-expression magnification processing uses the Euler motion magnification algorithm to magnify micro-expressions. The scoring process is as follows: the keyframe detection algorithm is used to find the frame with the greatest change in micro-expression in the video after micro-expression magnification, and then all frames of the video are scored according to the degree of change in micro-expression.

[0027] The process of obtaining a continuous frame sequence is to divide all frames of the entire video into approximately 10 segments, and then extract 30 consecutive frames from each segment to obtain 10 continuous frame sequences with 30 frames in each dimension.

[0028] Furthermore, the aforementioned method for identifying mental health abnormalities based on multimodal feature consistency can be integrated into a mobile app or application.

[0029] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the steps of the method.

[0030] Thirdly, the present invention provides a multimodal feature-consistent mental health abnormality identification system, comprising:

[0031] The human-computer interaction module is used to record changes in micro-expressions and micro-expression speech features of samples to obtain audio and video data;

[0032] The video data preprocessing module is used to extract micro-expression data from videos;

[0033] The micro-expression keyframe vector acquisition module is used to normalize the micro-expression severity score of each frame to obtain weights.

[0034] The audio data preprocessing module is used to obtain the frequency characteristics of the audio data;

[0035] The mental health anomaly identification model includes a video stream deep feature extraction network, an audio stream deep feature extraction network, a tensor fusion network, and a mental health classification network.

[0036] Video stream depth feature extraction networks are used to obtain depth video features. Video feature prediction results ;

[0037] Audio stream deep feature extraction network is used to obtain deep speech features F A and audio feature prediction results ;

[0038] Tensor fusion networks are used to fuse deep video features and deep speech features to obtain audio-video multimodal fusion features;

[0039] A mental health classification network is used to classify multimodal features fused from audio and video.

[0040] The mental health anomaly identification model uses multimodal consistency loss. Video feature prediction results and audio feature prediction results Apply training constraints.

[0041] Furthermore, the process of the human-computer interaction module is as follows: The computer plays the video, automatically displaying the next step. The subject's information is pre-entered using an ID card reader, and the patient's identity information is desensitized within the program. Then, clicking the "Start" button enters the interface for interaction with the subject. The human-computer interaction steps include three sections: "Watching Video," "Reading Text," and "Answering Questions." The "Watching Video" section involves the computer playing an emotionally stimulating video for the subject to watch. This step aims to elicit a response from the subject to external stimuli, thereby allowing for a clearer and more intuitive capture of the subject's facial micro-expression features. The "Reading Text" section... The first part involved the computer displaying a text message, which the participant was asked to read aloud. This step was designed to allow the participant to release vocal information, thus enabling the clear capture of the participant's vocal frequency and texture characteristics. The second part involved the computer displaying a question, which the participant was asked to think about and answer. This step captured both visual and vocal information. The entire human-computer interaction process was recorded by a headset microphone and a camera with a resolution of 1080*640 and a frame rate of 60fps. The "watching video" and "reading text" parts each lasted 30 seconds, and the "answering questions" part lasted 40 seconds, for a total video and audio duration of 100 seconds.

[0042] Compared with the prior art, the beneficial effects of the present invention are:

[0043] This invention effectively utilizes facial micro-expression features and voice features to identify mental health abnormalities. Both facial micro-expressions and voice features are important indicators that reflect the emotions and psychology of individuals. Studies have shown that multimodal mental health abnormality identification generally outperforms single-modal identification in terms of accuracy and detection performance. Therefore, this invention comprehensively considers the multimodal features of micro-expression and voice features, enabling the model to incorporate multimodal feature information and enhance its performance. Micro-expression features are an important reference indicator in this invention. The visual features of this invention emphasize and focus on micro-expression features, amplifying and enhancing the micro-expression features contained in video data. The results show that the model performance is significantly better when using video data with amplified micro-expressions than when using ordinary video data without micro-expression amplification.

[0044] 2. This paper presents a method that uses an end-to-end deep neural network to enable the model to automatically learn deep features that are helpful for identifying mental health abnormalities. This avoids the time-consuming and laborious manual feature labeling and extraction. The deep neural network automatically extracts features for identifying mental health abnormalities. The results show that the model performs well, and the extracted features can be used effectively for the task of identifying mental health abnormalities. They can be used as a basis for initial screening and provide relatively objective reference indicators for the clinical diagnosis of psychologists. They can also be integrated into mobile app or application for patients' self-examination and self-testing.

[0045] 3. This invention considers the consistency between multimodal feature information. That is, for the same sample, the prediction result of its speech features should be consistent with the prediction result of its micro-expression video features. It utilizes audio-video multimodal consistency to construct a loss function for the learning and training of the audio-video feature extraction network, fusion network, and mental health classification network. This more fully utilizes the feature information of each modality and increases the exchange of feature information between the two modalities. Compared with the traditional cross-entropy classification loss function, it considers the consistency of multimodal data prediction results, rather than simply performing feature fusion. In other words, for the same sample, whether based on video or audio, the ideal final prediction result should be consistent with its label. Therefore, if the video stream network predicts the sample as normal while the audio stream network predicts it as abnormal, the inconsistency between the two modal prediction results indicates that the network needs further training, the model performance needs further improvement, and the penalty of the loss function needs to be increased. This approach better integrates multimodal information and effectively improves the model's classification performance and capabilities. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is an overall flowchart of the multimodal mental health abnormality recognition method based on micro-expression and voice features provided in the embodiments of this application.

[0048] Figure 2 This is a schematic diagram of the video stream depth feature extraction network in this invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the present invention will be described in detail below with reference to the embodiments and accompanying drawings, but this is not intended to limit the scope of protection of this application.

[0050] This application utilizes micro-expression features, performs micro-expression amplification and feature enhancement, and combines audio to identify multimodal mental health abnormalities. The consistency of multimodal features is considered during the training process.

[0051] This invention provides a method for identifying mental health abnormalities based on multimodal feature consistency (hereinafter referred to as the method), which uses micro-expression and voice features as the basis for judgment, and includes the following steps:

[0052] Step 1: Raw audio and video data acquisition:

[0053] Design a human-computer interaction program to automatically collect data and obtain raw data files of two modalities, audio and video files, in the same scene;

[0054] Step 2: Audio and video data preprocessing:

[0055] For video data: First, the acquired video data containing complete facial images is cropped and reconstructed into a video containing only the head region to facilitate subsequent micro-expression magnification processing. Then, addressing the issue of subtle and difficult-to-capture micro-expressions, a micro-expression motion magnification method is used to magnify the head region video. Specifically, the Euler motion magnification algorithm is employed to amplify the micro-expressions. Next, a keyframe detection algorithm is used to find the frames with the greatest micro-expression changes in the magnified video. All frames are then scored based on the degree of micro-expression change, and the facial images are registered and aligned, thus completing the micro-expression magnification processing of the video data. Finally, each magnified video data point is sampled continuously at a certain sampling interval to obtain multiple consecutive frame sequences with the same dimensions. This completes the preprocessing of the video data.

[0056] For audio data: The acquired audio data is subjected to noise reduction and noise removal. First, the voice separation model for solving the "cocktail party" problem is used to separate the human voice from the noisy audio data. Then, the spectral characteristics of the audio are observed using Fast Fourier Transform, and noise reduction is performed using filters with different passbands and stopbands. After amplification, clearer audio data is obtained, thus completing the noise reduction and noise removal of the audio data. Then, the Mel-frequency cepstral coefficients of the noise-reduced audio data are calculated, and the frequency characteristics of the Mel-frequency cepstral coefficients of the audio data are extracted.

[0057] Step 3: Audio depth feature extraction:

[0058] The Mel-frequency cepstral coefficients obtained from the audio data in step two are fed into the audio stream deep feature extraction network for audio deep feature extraction, resulting in deep speech features; specifically:

[0059] The deep feature extraction network for audio streams consists of a pre-trained 2DResNet18 backbone incorporating a temporal attention mechanism. The Mel-frequency cepstral coefficients of the audio data are first processed by 2DResNet18 for deep feature extraction. The resulting feature vectors are then processed by the temporal attention mechanism to compress spatial and channel-level features, extracting temporal features. The final layer of the deep feature extraction network is a fully connected layer that performs single-modal psychological health anomaly identification, yielding audio feature prediction results. After the audio data is processed by 2DResNet18, and then further processed by the temporal attention mechanism, the temporal feature vector is obtained. This is what we call deep speech features. Audio feature prediction results obtained through a fully connected circuit The calculation formula can be expressed as follows:

[0060]

[0061] in These represent the parameters of the fully connected layer, which are learnable parameters;

[0062] Step 4: Video depth feature extraction:

[0063] The video data obtained from step two consists of multiple consecutive frame sequences with the same dimensions, which are simultaneously fed into the video stream depth feature extraction network to extract video depth features and obtain depth video features.

[0064] Video Stream Deep Feature Extraction Network: The backbone of the video stream deep feature extraction network is a pre-trained C3D ResNet50. After feature extraction by the backbone network, the resulting feature vectors are sequentially fed into the spatial attention module and the channel attention module for spatial and channel dimension information integration and weight allocation. Next, the obtained feature vectors are multiplied by the micro-expression keyframe vectors, which are obtained in step two using a keyframe detection algorithm. Specifically, the keyframe detection algorithm finds the frames with the greatest micro-expression changes, scores all frames in the video based on the degree of micro-expression change, and then normalizes the micro-expression score for each frame. The normalized result is used to construct the micro-expression keyframe vector, whose dimension is the same as the dimension of the continuous frame sequence. This vector is then multiplied by the output of the channel attention module as a weight. This approach addresses the problem that micro-expressions are short-lived and difficult to capture and detect in long videos, making them difficult for the model to recognize. Finally, the features containing micro-expression identification information are fed into the temporal attention module for temporal dimension information integration at the frame sequence level. The calculation formulas for the temporal, spatial, and channel attention mechanisms can be expressed as:

[0065]

[0066] In the formula, , where S represents spatial attention, C represents channel attention, and T represents temporal attention; This represents the feature vector fed into the attention module. This represents a one-dimensional convolution operation matrix. Represents a fully connected matrix. and These are learnable parameters;

[0067] Specifically, the continuous frame sequence is first processed by C3D ResNet50 for depth feature extraction, resulting in a video feature vector after C3DResNet50 processing. The network then sequentially enters the spatial attention module, channel attention module, and temporal attention module (the three attention modules are connected in series via a fully connected layer to form the entire network). This process compresses the features of the other two dimensions, extracting and integrating the spatial, channel, and temporal features respectively. First, it passes through the spatial attention module to obtain the spatial feature vector. Then, after passing through the channel attention module, spatial channel feature vectors are obtained. The calculation formula is:

[0068]

[0069] Among them, symbols This represents a vector product; then the processed spatial channel feature vector... Combine it with micro-expression keyframe vectors Multiplying them together yields a spatial channel feature vector that incorporates micro-expression identifier information. The calculation formula is:

[0070] (8)

[0071] Then it is fed into the temporal attention module to obtain the spatial channel temporal feature vector. The calculation formula is:

[0072]

[0073] The obtained spatial channel time feature vector The average value of the last dimension is used to obtain the deep video features. The final layer of the video stream deep feature extraction network is a fully connected layer, which performs single-modal mental health anomaly identification in the video to obtain video feature prediction results. The calculation formula can be expressed as follows:

[0074]

[0075] in The parameters representing a fully connected network are learnable parameters;

[0076] Step 5: Audio and video feature fusion:

[0077] The obtained deep video features and deep speech features are fused using a tensor fusion network to obtain audio-video multimodal fusion features. ;

[0078] Step Six: Category Prediction

[0079] Multimodal fusion features of audio and video The data is fed into a mental health classification network for final prediction, identifying mental health abnormalities. The mental health classification network consists of a three-layer fully connected network. After passing through this network, the final predicted audio and video feature prediction result y is obtained, predicting whether mental health is abnormal. The calculation formula can be expressed as follows:

[0080]

[0081] in The parameters representing the mental health classification network are learnable parameters;

[0082] Step 7: Loss Function Construction:

[0083] Considering the multimodal feature consistency problem, a multimodal consistency loss function for audio and video feature extraction and classification models is designed based on the traditional cross-entropy function. This is then applied to the video feature prediction results obtained in step three. and audio feature prediction results And the final predicted audio and video feature prediction result y.

[0084] The traditional formula for calculating cross-entropy loss is:

[0085]

[0086] Where n is the number of samples and M is the number of categories (in this embodiment, the number of categories is 2, namely normal and abnormal categories). Let represent the label of the i-th sample. This represents the predicted probability that the i-th sample belongs to category c.

[0087] Based on the cross-entropy loss, consider the audio feature prediction results. Video feature prediction results To address the consistency issue of prediction results, when audio feature prediction results are inconsistent with video feature prediction results, a larger penalty weight should be applied. Therefore, the audio-video multimodal consistency loss is calculated. The formula is as follows:

[0088]

[0089] Among them, the introduced The penalty coefficient is a hyperparameter of the coefficient, and the function is... The penalty function increases the penalty intensity of the loss function when the audio feature prediction results and video feature prediction results are inconsistent. It is defined as follows:

[0090]

[0091] As can be seen, when the categories are inconsistent, the value of the loss function will increase due to the existence of the penalty function and penalty coefficient.

[0092] Example 1

[0093] This embodiment of the multimodal feature consistency mental health anomaly identification method identifies mental health anomalies using a specific mental health dataset. The input audio and video files are self-collected and include the following:

[0094] Step 1: Raw audio and video data acquisition:

[0095] Design a human-computer interaction program to collect data. The program records micro-expression changes and micro-expression voice features of the samples. The program is played on a computer and automatically appears the next step. The subject's information is entered in advance using an ID card reader, and the patient's identity and other information are desensitized in the program. Then, clicking the "start" button will enter the interface for interacting with the subject. To better capture the micro-expression facial features and voice features of the subjects, this disclosure sets up a human-computer interaction process for data collection, including three stages: "watching a video" (for more detailed acquisition of micro-expressions), "reading text aloud" (for capturing the texture features of sound), and "answering questions" (for capturing the thought process). The "watching a video" stage involves the computer playing an emotionally evoked video for the subjects to watch. This stage aims to elicit a response from the subjects to external stimuli, thereby enabling clearer and more intuitive capture of their facial micro-expression features. The "reading text aloud" stage involves the computer displaying a text and instructing the subjects to read it aloud. This stage aims to elicit vocal information from the subjects, thereby enabling clear capture of their vocal frequency and texture features. The "answering questions" stage involves the computer displaying a question and instructing the subjects to think and answer it. This stage captures both visual and vocal information. The entire human-computer interaction process was recorded by a headset microphone and a camera with a resolution of 1080*640 and a frame rate of 60fps. The "watching video" and "reading text" segments each lasted 30 seconds, and the "answering questions" segment lasted 40 seconds, for a total video and audio duration of 100 seconds each. This yielded raw data files containing both audio and video files (100 seconds long, real-time corresponding, and of equal length). The audio files were in WAV format, and the video files were in AVI format. The dataset contained 1469 samples, including 1278 normal samples and 371 abnormal samples. Each sample was labeled. These 1469 samples constituted the dataset. The data in the dataset could be categorized into normal and abnormal categories, or further categorized by specific abnormal disease type plus normal category. Each sample's label was professionally assessed and annotated by a qualified psychologist.

[0096] The following example uses two labels. Normal and abnormal samples are divided in a 3:1 ratio. Three-quarters of the normal samples and three-quarters of the abnormal samples are mixed as the training set and saved to the `train.csv` file in the directory. The remaining one-quarter of the normal and abnormal samples are mixed as the test set and saved to the `test.csv` file in the directory. The `train.csv` and `test.csv` files are then used to generate a JSON file for the model to load and read data.

[0097] Step 2: Audio and video data preprocessing:

[0098] For video data: First, the acquired AVI videos are converted into frames, and all frames of each video are stored in a folder named after the sample. The samples are then placed in the training and test set folders. Next, the acquired video image data containing complete facial images is cropped to recreate videos containing only the head region, facilitating subsequent micro-expression magnification processing. For micro-expression magnification processing of the head region video, the Euler motion magnification algorithm is first used to magnify the micro-expressions, with a magnification factor of 5. Then, a keyframe detection algorithm is used for top-frame detection and localization. Micro-expression level detection and scoring are performed, and the calculation results are saved in .mat format and then regularized. Next, facial images are registered and aligned to generate a mean standard image, thus completing the micro-expression amplification processing of the video data. Each frame is then cropped to retain only facial features, with a cropped size of 300*300. Then, for each video sample after micro-expression amplification, all its frames are divided into 10 equal segments, and 30 consecutive frames are extracted from each segment, resulting in a continuous frame sequence of 30 frames in each of the 10 dimensions. This completes the preprocessing of the video data.

[0099] For audio data: The acquired audio data is subjected to noise reduction and noise removal. First, the voice separation model for solving the "cocktail party" problem is used to separate the human voice from the noisy audio data. The audio data is then subjected to Fast Fourier Transform and filtered using filters with different passbands and stopbands. After amplification, clearer audio data is obtained, thus completing the noise reduction and noise removal of the audio data. Then, the Mel-frequency cepstral coefficients of the noise-reduced audio data are calculated to extract the frequency features of the audio data.

[0100] Step 3: Audio depth feature extraction:

[0101] The Mel-frequency cepstral coefficients obtained from the audio data in step two are fed into an audio stream feature extraction network for audio depth feature extraction, resulting in deep speech features; specifically:

[0102] The deep feature extraction network for audio streams consists of a pre-trained 2D ResNet18 backbone incorporating a temporal attention mechanism. This mechanism comprises one 1D convolutional layer, one fully connected layer, and one softmax layer. Mel-spectral coefficients of the audio data are first processed by 2D ResNet18 for deep feature extraction. The extracted feature vectors then enter the temporal attention mechanism, which compresses the spatial and channel-level features to extract the temporal dimension. Finally, a fully connected layer is added to the deep feature extraction network to identify psychological health anomalies in a single audio modality, resulting in audio feature predictions. Similar to video streams, audio data is processed by 2D ResNet18, and then further processed by a temporal attention mechanism to obtain deep speech features. The audio feature prediction results obtained after full connection The calculation formula can be expressed as,

[0103]

[0104] in These represent the parameters of the fully connected layer, which are learnable parameters;

[0105] Step 4: Video depth feature extraction:

[0106] The video data obtained from step two consists of multiple consecutive frame sequences with the same dimensions, which are simultaneously fed into the video stream depth feature extraction network to extract video depth features and obtain depth video features.

[0107] Video Stream Deep Feature Extraction Network: The video frame sequence is processed into a dimension of [dimension value missing] before being fed into the video stream deep feature extraction network. The vectors are denoted by the following dimensions: number of consecutive frame sequences, batch size, number of RGB channels, consecutive frame sequence length, and image sampling size, respectively. The backbone network of the video stream deep feature extraction network is a C3D ResNet50 residual network pre-trained on the ImageNet and Kinetic datasets. Each residual block of C3D ResNet50 consists of three 3D convolutions, with the kernel size of the first 3D convolution being [missing value]. With a stride of 1, the kernel size of the second 3D convolution is... With a stride of 2, the kernel size of the third 3D convolution is... The stride is 1; each convolutional layer is followed by a 3D batch regularization and a ReLU activation function; in the network settings, considering the subtle and easily lost features of micro-expressions, all pooling layers are removed in C3D ResNet50, and the stride of the second convolutional kernel in each residual block is set to 2. Micro-expression features are not processed by pooling layers, and this setting reduces computation without losing details; the dimension of the video feature vector F after processing by C3D ResNet50 is... The first dimension is the product of the continuous frame sequence and the batch size, the second dimension is the spatial dimension after processing, and the third dimension is the channel dimension, which is also the length of the continuous frame sequence.

[0108] Then, the video feature vector F processed by C3D ResNet50 is successively fed into the spatial attention module and the channel attention module to integrate information and assign weights in the spatial and channel dimensions, thereby obtaining the spatial feature vector. and spatial channel feature vectors The spatial channel feature vectors obtained next Micro-expression keyframe vectors Multiplication, where the micro-expression keyframe vector is obtained from step two using the keyframe detection algorithm. After obtaining the micro-expression severity score for each frame, it is normalized here, with a dimension of [missing value]. Then used as weights with Multiplying them yields a spatial channel feature vector containing micro-expression identifiers. Next will The data is fed into a temporal attention module to integrate temporal information at the frame sequence level, resulting in deep video features. The spatial, channel, and temporal attention modules are all structured as one 1D convolutional layer, one fully connected layer, and one softmax layer. The calculation formula for the attention mechanism can be expressed as:

[0109] (5)

[0110] In the formula, , where S represents spatial attention, C represents channel attention, and T represents temporal attention; This represents the feature vector fed into the attention module. This represents a one-dimensional convolution operation matrix. Represents a fully connected matrix. and These are learnable parameters;

[0111] Specifically, the video feature vector F processed by C3D ResNet50 first passes through a spatial attention module to obtain a spatial feature vector. Then, after passing through the channel attention module, spatial channel feature vectors are obtained. The calculation formula is:

[0112]

[0113] Among them, symbols Represents vector product; the spatial attention module and channel attention module do not change the size and dimension of the vector, therefore the spatial feature vector and spatial channel feature vectors The size is still Then the spatial channel feature vector The product of the number of consecutive frame sequences and the batch size is split along the first dimension to obtain... Size is Then, here it is compared with the micro-expression keyframe vector. Multiply, The dimension is This yields a spatial channel feature vector that incorporates micro-expression identifier information. The calculation formula is:

[0114] (8)

[0115] get The dimension remains the same. Then get To change the size and swap the dimensions, specifically: calculate the average of the last dimension, and then swap the second and third dimensions of its feature vectors, resulting in a dimension of... eigenvectors Finally, it is fed into the time attention module, and the calculation formula is as follows:

[0116] (9)

[0117] Obtain the spatial channel temporal feature vector Size is After a simple integration operation, the average is calculated again in the third dimension, i.e., the frame sequence dimension, to obtain the final size. Deep video features ;

[0118] At the end of the deep feature extraction network for the video stream is a fully connected layer that performs single-modal mental health anomaly identification in the video, obtaining the video feature prediction results. The calculation formula can be expressed as follows:

[0119] (10)

[0120] in These represent the parameters of the fully connected layer, which are learnable parameters;

[0121] During training, the C3D ResNet50 network is pre-trained and fine-tuned, and its parameters are frozen and do not participate in gradient backpropagation. The parameters of spatial attention, channel attention, temporal attention, and the classification network are involved in gradient backpropagation.

[0122] Step 5: Audio and video feature fusion:

[0123] The obtained deep video features and deep speech features are fused using a tensor fusion network to obtain audio-video multimodal fusion features. Simple splicing fusion methods may lead to the loss of modal dynamic information when fusing multimodal features. In contrast, tensor fusion networks focus on the correlation of multimodal features in high-dimensional space compared to splicing-based methods. Although their computation is more complex, they fuse features with better correlation and tighter coupling, resulting in better fusion performance.

[0124] Step Six: Category Prediction

[0125] Multimodal fusion features of audio and video The data is fed into a mental health classification network for final prediction and identification of mental health abnormalities. The mental health classification network consists of a three-layer fully connected network. After passing through the mental health classification network, the final predicted audio and video feature prediction result y is obtained, which predicts whether mental health is abnormal and, if so, what type of abnormality it is. The calculation formula can be expressed as follows:

[0126] (11)

[0127] in The parameters representing the mental health classification network are learnable parameters;

[0128] Step 7: Loss Function Construction:

[0129] To address the multimodal feature consistency problem, a new audio-video multimodal consistency loss function is proposed, building upon the traditional cross-entropy function, to constrain the results across different modalities. This loss function is based on the video feature prediction results obtained in step three. and audio feature prediction results And the final predicted audio and video feature prediction result y. The traditional cross-entropy loss calculation formula is:

[0130] (1)

[0131] Where n is the number of samples and M is the number of categories. Let represent the label of the i-th sample. Let represent the model's predicted probability that the i-th sample belongs to class c. Based on the cross-entropy loss, the audio feature prediction results are considered. Video feature prediction results To address the consistency issue, when audio feature prediction results differ from video feature prediction results, a larger penalty weight should be applied. Therefore, the calculation of the audio-video multimodal consistency loss is derived. The formula is as follows:

[0132]

[0133] Among them, the introduced The penalty coefficient is a hyperparameter, and in this embodiment, The value is set to 1.55, function The penalty function increases the penalty intensity of the loss function when the audio prediction result and the video prediction result are inconsistent. It is defined as follows:

[0134] (3)

[0135] As can be seen, when the categories are inconsistent, the value of the loss function will increase due to the existence of the penalty function and penalty coefficient.

[0136] This invention effectively utilizes facial micro-expression features and voice features for the identification of mental health abnormalities. It amplifies and enhances the micro-expression features contained in video data, comprehensively considering two important physiological indicators of psychological well-being—micro-expression and voice features—allowing the model to comprehensively consider multimodal feature information and enhance its performance. In terms of classification evaluation metrics precision and recall, using raw video and audio data for mental health abnormality identification yields a precision of 0.85 and a recall of 0.61; using video and audio data with amplified micro-expressions and keyframe annotation, the precision is 0.91 and the recall is 0.67, representing a 5 percentage point improvement in precision and a 6 percentage point improvement in recall. Furthermore, it provides a method using an end-to-end deep neural network to allow the model to automatically learn deep features helpful for identifying mental health abnormalities, avoiding time-consuming and laborious manual feature labeling and extraction. Results show that the method performs well and can serve as a basis for initial screening, providing relatively objective reference indicators for clinical diagnosis by mental health professionals. It can also be integrated into mobile applets or applications for patient self-assessment.

[0137] This invention utilizes the audio-video multimodal feature consistency loss function for learning and training audio and video stream model networks, making fuller use of the feature information of each modality and increasing the exchange of feature information between the two modalities. Compared with the traditional cross-entropy classification loss function, it considers the consistency of multimodal data prediction results rather than simply performing feature fusion, thus better integrating multimodal information and effectively improving the model's classification performance and capabilities. Experimental results show that when the input data consists of video and audio data with magnified micro-expressions, the precision using the traditional cross-entropy loss function is 0.82, and the recall is 0.58; while the precision using the multimodal feature consistency loss function is 0.91, and the recall is 0.67, representing a 9 percentage point improvement in both precision and recall. Therefore, the multimodal feature consistency method based on micro-expression and speech features in this invention is effective.

[0138] Any aspects not covered in this invention are applicable to existing technologies.

Claims

1. A multimodal feature consistency mental health anomaly recognition method, comprising the following steps: obtaining original data files containing audio files and video files of two modalities under the same scene, containing normal samples and abnormal samples, and setting disease category labels for abnormal samples; performing micro-expression amplification processing on video data containing complete facial images in the obtained original data files, then scoring each frame of the video according to the micro-expression action amplitude size of each frame, recording the micro-expression degree score of each frame, and then aligning the face images, thereby completing the micro-expression amplification processing of the video data; then sampling the continuous frame sequences of the amplified video data according to the set sampling interval to obtain multiple continuous frame sequences with the same dimension, thereby completing the preprocessing of the video data; normalizing the micro-expression degree score of each frame in the continuous frame sequence to obtain a micro-expression key frame vector; performing noise removal and impurity removal processing on the audio data in the obtained original data files to extract the frequency features of the audio data; The frequency features extracted from the audio data are sent to an audio stream deep feature extraction network for audio deep feature extraction, to obtain deep speech features F A Meanwhile, speech feature result prediction is performed on the speech information in the audio stream deep feature extraction network, to obtain an audio feature prediction result ; The video stream deep feature extraction network comprises a C3D ResNet50 of a de-pooling layer connected in sequence, a spatial attention module, a channel attention module, a time attention module and a classification network; wherein the input of the time attention module is the multiplication result of the output of the channel attention module and the micro-expression key frame vector, and the output of the time attention module is averaged in the frame sequence dimension to obtain a deep video feature ​ The plurality of dimensionally same continuous frame sequences obtained from the video data are simultaneously fed into a video stream deep feature extraction network for video deep feature extraction, to obtain deep video features ; video feature prediction is performed on the video information of the continuous frame sequences in the video stream deep feature extraction network, and a video feature prediction result is obtained through a classification network ; performing feature fusion on the deep video features and deep speech features using a tensor fusion network to obtain audio-video multimodal fusion features; sending the audio-video multimodal fusion features into a mental health classification network for final prediction, and obtaining the final predicted audio-video feature prediction result y after passing through the mental health classification network; The audio-video multi-modal consistency loss is calculated according to formula (2) , (2) where n is the number of samples, M is the number of label categories, represents the label of the i-th sample, represents the model prediction probability that the i-th sample belongs to category c, is a penalty coefficient; the function is a penalty function that increases the penalty of the loss function when the audio feature prediction result and the video feature prediction result are inconsistent, and the definition of the penalty function is (3)。 2. The multi-modal feature consistency mental health abnormality recognition method of claim 1, wherein, The The set value is 1.0-1.

6.

3. The multi-modal feature consistency mental health abnormality recognition method of claim 1, wherein, The set value is 1.

55.

4. The multi-modal feature consistency mental health abnormality recognition method of claim 1, wherein, the audio stream deep feature extraction network comprises a 2D ResNet18 network, a time attention mechanism and a fully connected layer, the output of the pre-trained 2D ResNet18 network is connected to the time attention mechanism, and the output of the time attention mechanism is connected to a fully connected layer; the time attention mechanism comprises one layer of one-dimensional convolution, one layer of fully connected layer and one layer of softmax function, which compresses the features of two dimensions of space and channel, extracts the features of time dimension, and obtains deep speech features.

5. The multi-modal feature consistency mental health abnormality recognition method of claim 1, wherein, the backbone network of the video stream deep feature extraction network is a pre-trained C3D ResNet50; after feature extraction by the backbone network, the features are sequentially sent into a spatial attention module and a channel attention module for spatial dimension and channel dimension information integration and weight distribution; then the feature vector Fsc obtained by processing the channel attention module is multiplied by the micro-expression key frame vector Fem; and then sent into a time attention module for time dimension information integration at the frame sequence level. Each residual block of C3D ResNet50 is composed of three 3D convolutions in series, the first 3D convolution has a kernel size of with a stride of 1, the second 3D convolution has a kernel size of with a stride of 2, and the third 3D convolution has a kernel size of with a stride of 1; each 3D convolution is followed by a 3D batch normalization and a ReLU activation function.

6. The multi-modal feature consistency mental health abnormality recognition method of claim 1, wherein, the micro-expression amplification processing is to amplify the micro-expression using Euler motion amplification algorithm, and the scoring process is to find the frame with the largest micro-expression change in the amplified video using key frame detection algorithm, and then score all frames of the video according to the micro-expression change degree; the continuous frame sequence obtaining process is to divide all frames of the video into 10 segments on average, and then extract 30 continuous frames from each segment to obtain 10 continuous frame sequences with the dimension of 30 frames. 7.Integrating the multimodal feature consistency mental health anomaly recognition method of any one of claims 1-6 into a mobile phone applet or application.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program can implement the steps of the method of any one of claims 1-6 when executed by a processor.

9. A multi-modal feature consistency mental health abnormality recognition system, characterized in that, The system performs the identification method of any one of claims 1-6, comprising: A human-computer interaction module for recording sample micro-expression changes and micro-expression voice feature information to obtain audio and video data; A video data preprocessing module for extracting micro-expression data in the video; A micro-expression key frame vector acquisition module for normalizing the micro-expression degree score of each frame to obtain a weight An audio data preprocessing module for obtaining frequency characteristics of the audio data; A mental health anomaly identification model comprising a video stream deep feature extraction network, an audio stream deep feature extraction network, a tensor fusion network, and a mental health classification network, The video stream deep feature extraction network is used to obtain deep video features and video feature prediction results ; The audio stream deep feature extraction network is used to obtain deep speech features F A and the audio feature prediction result ; The tensor fusion network is used for the fusion of deep video features and deep voice features to obtain audio and video multi-modal fusion features; The mental health classification network is used for classifying the audio and video multi-modal fusion features; Psychological health abnormality recognition model uses multi-modal consistency loss To video feature prediction results And audio feature prediction results Training constraints.

Citation Information

Patent Citations

  • End-to-end automated audio-visual detection method for depression

    CN112560811B

  • Depression tendency evaluation system and method based on multi-modal characteristics

    CN114241599A

  • Multi-modal face emotion recognition method and device

    CN114399818A