Audio and video fusion multi-modal evaluation method for depression detection
The audio-visual fusion method using F3DA, AEF, AV-CAM, and AMFM enhances depression detection by integrating facial and audio features, addressing the limitations of current methods and improving accuracy and robustness.
Patent Information
- Application Number
- CN202510498415.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
AI Technical Summary
The existing depression detection methods have shortcomings in feature extraction and multimodal fusion, and it is difficult to fully capture the space-time dependence and emotional characteristics of facial expressions and speech data, resulting in limited detection accuracy.
The facial 3D attention module (F3DA) and audio emotion extraction module (AEF) are used to extract initial features, modal alignment is achieved through the audio-visual cross-modal attention module (AV-CAM), and the weight is dynamically adjusted using the adaptive multi-scoring fusion module (AMFM) to generate more accurate depression detection results.
It improves the accuracy and robustness of depression detection, can better capture emotional characteristics in facial micro-expressions and speech signals, and dynamically adjust weights through the adaptive multi-score fusion module to generate more reliable detection results.
Smart Images

Figure CN120319480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of intelligent technology and image detection technology, and specifically to an audio-visual fusion multi-modal evaluation method for depression detection. Background Technique
[0002] Current depression detection methods mainly rely on the analysis of facial expressions and voice data, but there are still many deficiencies in feature extraction and multi-modal fusion in the prior art. In terms of facial expression detection, early research mainly adopted manual feature extraction methods, including Local Phase Quantization (LPQ), Motion History Histogram (MBH), Local Binary Patterns (LBP), and Local Binary Patterns on Three Orthogonal Planes (LBP-TOP), etc. These methods can extract effective visual features to a certain extent, but there are limitations in capturing the complex spatio-temporal dynamic characteristics of depression patients. With the development of deep learning technology, the research has gradually turned to automatic feature extraction methods based on convolutional neural networks (CNNs), where two-dimensional convolutional neural networks (2D-CNNs) such as VGG and ResNet are widely used to extract the spatial features of facial expressions and combine long short-term memory networks (LSTMs) to model the temporal dynamic characteristics. In addition, three-dimensional convolutional neural networks (3D-CNNs) have also been applied to the spatio-temporal modeling of video data. For example, the C3D (Convolutional 3D) model can analyze the temporal changes of facial expressions. However, single-scale 2D or 3D convolutions often have difficulty in comprehensively capturing multi-scale spatio-temporal dependencies, thus limiting the performance of the model in depression detection.
[0003] In the depression detection of the voice modality, early methods also relied on manual feature extraction, such as Mel Frequency Cepstral Coefficients (MFCCs), spectral features, and prosodic features, etc. These features can capture the information related to emotional changes in the voice signal and be used as the input of a deep neural network (DNN) for modeling. However, these traditional features often have difficulty in capturing the complex emotional features in voice data, resulting in limited detection accuracy. In recent years, researchers have gradually introduced more advanced voice feature modeling methods, including Mel filter bank features, spectrogram features, and Self-Attention mechanism, which can better capture the emotional dynamic changes in voice data and show higher detection performance in the depression detection task.
[0004] With the introduction of multimodal data, multimodal fusion has become a key component of depression detection systems. Multimodal fusion techniques mainly include three strategies: feature-level fusion, model-level fusion, and decision-level fusion. Feature-level fusion (early fusion) concatenates features of different modalities into a unified feature vector at the input stage to form a unified representation for downstream task modeling. However, this method is prone to the "curse of dimensionality", significantly increasing the complexity of model training. Decision-level fusion (late fusion) generates the final detection result by weighting, voting, or summing the prediction results of independent models. However, this method ignores the potential interaction information between modalities, resulting in information loss. In contrast, model-level fusion can effectively capture the interaction information between modalities by jointly modeling the data distributions of different modalities during the training stage, offering higher flexibility and adaptability. However, existing model-level fusion methods often fail to explicitly model the characteristics of different modalities, leading to insufficient feature extraction and affecting the performance improvement of the model. Additionally, although some studies have introduced modality-specific feature modeling, most of them only use simple weighting operations for modality fusion, failing to fully consider the differential contributions of different modalities to depression detection, resulting in unstable and inaccurate detection results.
[0005] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention
[0006] The objective of the present invention is to provide an audio-visual fusion multimodal evaluation method for depression detection, aiming to improve the accuracy and robustness of depression detection by integrating audio-visual modality information to solve the technical problems presented in the background art.
[0007] To achieve the above objective, the present invention provides the following technical solution: An audio-visual fusion multimodal evaluation method for depression detection, which at least includes the following steps:
[0008] S1: Separate the input audio-visual data to obtain audio data and video data, extract initial features from the audio data and video data, and process the extracted initial features through a facial 3D attention module and an audio emotion extraction module. The facial 3D attention module is F3DA, and the audio emotion extraction module is AEF.
[0009] S2: Jointly input the features extracted in S1 into an audio-visual cross-modal attention module. The audio-visual cross-modal attention module is AV-CAM. AV-CAM achieves modality alignment through an audio-visual cross-attention mechanism, enabling the interaction and fusion of key information between the audio and video modalities, enhancing the complementarity between modalities, and outputting an audio-visual interaction prediction result.
[0010] S3: The scores generated from the output results of F3DA, AEF, and AV-CAM are passed into the Adaptive Multi-Score Fusion Module, namely AMFM. AMFM can dynamically adjust weights according to the contribution degrees of different modalities, thereby generating a more comprehensive and reliable depression detection result.
[0011] Furthermore, the separation of the audio and video is performed using the integrated software Kingshiper Vocal Remover, which can effectively separate the audio and video and generate separate audio and video files.
[0012] Furthermore, F3DA is designed to extract micro-expression and facial dynamic features, that is, to obtain image features, so as to capture fine-grained emotional cues in the video modality, and then output video prediction results;
[0013] F3DA is equipped with a facial 3D attention mechanism. The facial 3D attention mechanism combines the channel attention mechanism and the spatial attention mechanism to enhance the attention to local facial features related to depression. The application of F3DA includes at least the following steps:
[0014] First, a channel descriptor is obtained through global average pooling (GAP), and the channel attention weight is generated through the following formula:
[0015]
[0016] a c =σ(W2δ(W1z c )) (2)
[0017] X c =a c ⊙X (3)
[0018] where X is the input feature map; F gp is the global pooling operation; a c is the channel attention weight; X c is the channel feature; z c is the channel descriptor; T is the time step; H is the height; W is the width; σ is the Sigmoid activation function; W1 and W2 are weights; δ is the ReLU function; ⊙ is element-wise multiplication;
[0019] Then, the spatial attention mechanism performs pooling operations, namely average pooling and max pooling, on the channel-enhanced feature map, and generates the spatial attention weight through the following formula:
[0020]
[0021] a s =σ(F conv(X avg ; X max )) ∈ R T×H×W×1 (5)
[0022] Among them, X avg is average pooling; X max is max pooling; C represents the channel; F conv represents the convolution operation; a s is the spatial attention weight;
[0023] Finally, by applying channel and spatial attention through element-wise multiplication, the optimized feature map X out is obtained:
[0024] X out = a s ⊙ X c (6)
[0025] Generate the final depression prediction feature through global average pooling operation:
[0026] z out = F gp (X out ) (7)
[0027] S v = MLP(reshape(z out , B, D)) (8)
[0028] Among them, reshape means changing the tensor; B is the batch size; D is the feature dimension.
[0029] Furthermore, the AEF is used to model the speech features in the audio modality, that is, to obtain audio features to identify intonation, rhythm, and emotional changes, and then output the audio prediction result;
[0030] The AEF is based on the Transformer encoder and the multi-head self-attention mechanism. The AEF effectively improves the prediction ability of depression-related audio features by extracting high-level spatio-temporal features in the audio time series data. The application of the AEF at least includes the following steps:
[0031] For the input audio sequence X ∈ R N×L×D , where N represents the batch size, L represents the sequence length, and D represents the feature dimension of each time step;
[0032] After passing through the Transformer encoder, the encoded representation X encoded ∈ R N×L×D is obtained:
[0033] X encoded = Encoder(X) (9)
[0034] The Transformer encoder effectively captures the global dependencies of the sequence and generates an encoded representation containing temporal and feature dimension information;
[0035] Subsequently, the multi-head self-attention (MHA) mechanism is applied to weight the encoded features, generating the output X with attention weights atten ∈R N×L×D :
[0036] X attn = MHA(X encoded , X encoded , X encoded ) (10)
[0037] The multi-head self-attention mechanism can aggregate the information of the input sequence and emphasize the features of key time steps according to the attention weights, which helps to capture the emotional changes and temporal features;
[0038] Then, by summing over the time dimension L, the final feature representation X is obtained sum ∈R N×L×D , which compresses the temporal information and provides a more compact feature representation for the subsequent regression task:
[0039]
[0040] The pooling operation aggregates the time series information, ensuring that the model focuses on the key parts and extracts the core temporal features for depression detection;
[0041] Finally, the aggregated feature X sum is processed by the multi-layer perceptron (MLP) regression module to generate the final audio depression score:
[0042] S a = MLP(X sum ) (12)
[0043] By combining the characteristics of the Transformer encoder and the self-attention mechanism, the accuracy and robustness of depression detection can be effectively improved.
[0044] Furthermore, the AV-CAM is used to align the audio features and the image features and map them to a unified three-dimensional feature space D alignment , and an A-V feature alignment function is adopted to project the audio features onto thus converting the two into feature spaces of the same dimension;
[0045] The alignment process is achieved through convolutional mapping or linear mapping to ensure that the audio and image features are converted to a unified feature dimension;
[0046] After the feature alignment is completed, the audio features and image features become:
[0047]
[0048] The aligned audio query feature Q audio and the key-value pairs K image and V image of the image features are input into the audio-visual cross-attention mechanism as inputs, and this process is expressed as:
[0049]
[0050] The feature I av after the interaction is further input into the multi-layer perceptron (MLP) regression module to generate the final depression prediction score:
[0051] S av = MLP(I av ) (15)
[0052] where I av represents the feature representation generated after the interaction between the audio query feature Q audio and the key-value pairs K image and V image of the image features.
[0053] Furthermore, the AMFM aims to dynamically combine the prediction results from different modalities to ensure that the final prediction result is more accurate and reliable;
[0054] The AMFM integrates the scores of the prediction results from different sources of F3DA, AEF, and AV-CAM. The input scores of F3DA, AV-CAM, and AEF are respectively represented as S a , S av and S v ;
[0055] The scores are initially represented in the form of a vector of size n×1, where n is the number of samples in the batch. For the convenience of fusion, the modality-specific scores are concatenated along the feature dimension to form a matrix, and the concatenation process is expressed as:
[0056] S concat = [S v , S av , S a ∈ R n×3 (16)
[0057] Next, the Gating Network calculates a weight vector w ∈ R n×3, weights are assigned to each modality, and the weights are calculated based on the concatenated score matrix;
[0058] Each component of the weight vector represents the relative importance of the video, video-audio interaction, and audio modalities respectively;
[0059] The calculation of the weights is completed by a multi-layer perceptron (MLP) and combined with a Sigmoid activation function to ensure that the value of each weight is between 0 and 1. The calculation formula of the weight vector is:
[0060]
[0061] where, w v , w av and w a are the learnable weights for the video, video-audio interaction, and audio respectively. The calculated weight vector w is used to perform a weighted sum of the modality-specific prediction scores to obtain the final fused prediction result y. The specific formula is:
[0062] y = w v ·S v + w av ·S av + w a ·S a (18)
[0063] The final output y represents the result of dynamically weighted fusion of the scores of each modality, where the weights are adaptively adjusted through learning to reflect the relative importance of each modality in predicting the target variable, thereby significantly improving the prediction accuracy and the robustness of the model.
[0064] Compared with the prior art, the beneficial effects of the present invention are:
[0065] The present invention is a method for implementing evaluation by a multi-modal depression assessment framework, which can effectively extract deep features in facial expressions and speech data, and achieve efficient fusion of multi-modal information through an optimized cross-modal interaction mechanism. Moreover, the framework proposed by the present invention can not only capture the spatio-temporal dynamic changes of facial micro-expressions, but also extract the emotional features in speech signals, thereby improving the accuracy and robustness of depression detection. And through the adaptive multi-score fusion module (AMFM), the present invention can dynamically adjust the weights according to the importance of different modalities to achieve a more accurate assessment of the severity of depression, effectively solving the problems of insufficient modal-specific modeling and utilization of information complementarity in the existing methods. Brief Description of the Drawings
[0066] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0067] Figure 1 It is the overall flowchart of the present invention;
[0068] Figure 2 It is the schematic diagram of the overall framework of the present invention;
[0069] Figure 3 It is the schematic diagram of AMFM of the present invention. Specific embodiments
[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.
[0071] Based on the corresponding method, the present invention proposes a multi-modal depression assessment framework (MDAF). Refer to Figure 2 , during the dataset pairing process of the multi-modal depression assessment framework, 32 video frames and the entire audio segment are selected as inputs during the training phase, while in the inference phase, the complete video frames and audio are input into the network to achieve more comprehensive feature extraction and prediction. For the MDAF network architecture, the present invention adopts a facial 3D attention mechanism (Face3D Attention, F3DA) to extract facial micro-expressions and dynamic behavior features in the video sequence, thereby accurately capturing the temporal changes of the face. For the audio modality, the present invention designs an Audio Emotion Former (AEF) to extract key information such as intonation, rhythm, and emotional patterns in the speech, thereby improving the accuracy of emotion recognition. The present invention further designs an Audio-Visual Cross Attention Mechanism (AV-CAM) to promote information interaction and alignment between the two modalities, ensuring that their homogeneous information is fully integrated. Finally, through the Adaptive Multi-Score Fusion Module (AMFM), the present invention dynamically fuses the depression prediction scores from different modalities to generate a more comprehensive and reliable final evaluation result. The present invention uses RMSE as the loss function, emphasizing reducing large prediction errors to improve the overall accuracy of the model. At the same time, Adam is used as the optimizer. With its adaptive learning rate and stability, the model can be efficiently trained and quickly converge on complex datasets.
[0072] Please refer to Figure 1 , for the application of MDAF, a multi-modal evaluation method for audio-visual fusion for depression detection is proposed, which at least includes the following steps:
[0073] S1: Separate the input audio-visual data to obtain audio data and video data. Extract initial features from the audio data and video data, and process the extracted initial features through a facial 3D attention module and an audio emotion extraction module. The facial 3D attention module is F3DA, and the audio emotion extraction module is AEF;
[0074] S2: The features extracted in S1 are jointly input into an audio-visual cross-modal attention module. The audio-visual cross-modal attention module is AV-CAM. AV-CAM achieves modal alignment through an audio-visual cross-attention mechanism, enabling the interaction and fusion of key information between the audio and video modalities, enhancing the complementarity between modalities, and outputting an audio-visual interaction prediction result;
[0075] S3: The scores generated from the output results of F3DA, AEF, and AV-CAM are passed into an adaptive multi-score fusion module. The adaptive multi-score fusion module is AMFM. AMFM can dynamically adjust the weights according to the contribution degrees of different modalities, thereby generating a more comprehensive and reliable depression detection result. Please refer to Figure 3 .
[0076] The separation of audio-visual data uses the integrated software Kingshiper Vocal Remover, which can effectively separate audio-visual data and generate separate audio and video files; Kingshiper Vocal Remover is based on AI technology and uses deep learning models (such as U-Net or Spleeter) to perform short-time Fourier transform (STFT) processing on audio signals, thereby extracting time-frequency features and achieving precise separation of human voices and background music.
[0077] F3DA aims to extract micro-expression and facial dynamic features, that is, to obtain image features, so as to capture fine-grained emotional cues in the video modality, and then output a video prediction result. In order to overcome the insufficient capture of subtle changes by facial 3D convolutional neural networks (CNNs) in the existing technology;
[0078] F3DA is equipped with a facial 3D attention mechanism. The facial 3D attention mechanism combines a channel attention mechanism and a spatial attention mechanism to enhance the attention to local facial features related to depression. The application of F3DA at least includes the following steps:
[0079] First, obtain a channel descriptor through global average pooling (GAP) and generate channel attention weights through the following formula:
[0080]
[0081] a c = σ(W2δ(W1z c )) (2)
[0082] X c = a c ⊙X (3)
[0083] where X is the input feature map; F gp is the global pooling operation; a c is the channel attention weight; X c is the channel feature; z c is the channel descriptor; T is the time step; H is the height; W is the width; σ is the Sigmoid activation function; W1 and W2 are weights; δ is the ReLU function; ⊙ is element-wise multiplication;
[0084] Then, the spatial attention mechanism generates spatial attention weights through pooling operations on the channel-enhanced feature map, namely average pooling and max pooling, and by the following formula:
[0085]
[0086] a s = σ(F conv (X avg ; X max )) ∈ R T×H×W×1 (5)
[0087] where X avg is the average pooling; X max is the max pooling; C represents the channel; F conv represents the convolution operation; a s is the spatial attention weight;
[0088] Finally, the channel and spatial attention are applied through element-wise multiplication to obtain the optimized feature map X out :
[0089] X out = a s ⊙X c (6)
[0090] The final depression prediction feature is generated through the global average pooling operation:
[0091] z out = F gp (X out ) (7)
[0092] Sv = MLP(reshape(z out , B, D))(8)
[0093] where reshape means to change the tensor; B is the batch size; D is the feature dimension.
[0094] AEF is used to model speech features in the audio modality, that is, to obtain audio features to identify intonation, rhythm, and emotional changes, and then output audio prediction results; in order to overcome the deficiencies of traditional deep learning methods in capturing emotional changes and speech features in the audio depression prediction task;
[0095] AEF is based on the Transformer encoder and the multi-head self-attention mechanism. The Transformer encoder, with its powerful global modeling ability, overcomes the deficiencies of traditional methods in capturing long-range dependencies, especially in modeling emotional changes and speech features. In addition, the multi-head self-attention mechanism further enhances the model's attention and identification ability for key depression features by adaptively weighting the contributions of different time steps and feature dimensions. AEF effectively improves the prediction ability of depression-related audio features by extracting high-level spatio-temporal features from audio time series data. The application of AEF at least includes the following steps:
[0096] For the input audio sequence X ∈ R N×L×D , where N represents the batch size, L represents the sequence length, and D represents the feature dimension of each time step;
[0097] After passing through the Transformer encoder, the encoded representation X encoded ∈ R N×L×D is obtained:
[0098] X encoded = Encoder(X) (9)
[0099] The Transformer encoder effectively captures the global dependencies of the sequence and generates an encoded representation containing temporal and feature dimension information;
[0100] Subsequently, the multi-head self-attention (MHA) mechanism is applied to weight the encoded features to generate the output X atten ∈ R N×L×D after attention weighting:
[0101] X attn = MHA(X encoded , X encoded , X encoded ) (10)
[0102] The multi-head self-attention mechanism can aggregate the information of the input sequence and emphasize the features of key time steps according to the attention weights, which helps to capture the emotional changes and temporal features;
[0103] Then, by performing a summation operation on the time dimension L, the final feature representation X is obtained sum ∈R N×L×D , compressing the temporal information and providing a more compact feature representation for the subsequent regression task:
[0104]
[0105] The pooling operation aggregates the time series information, ensuring that the model focuses on the key parts and extracts the core temporal features for depression detection;
[0106] Finally, the aggregated feature X sum is processed by a multi-layer perceptron (MLP) regression module to generate the final audio depression score:
[0107] S a = MLP(X sum ) (12)
[0108] By combining the characteristics of the Transformer encoder and the self-attention mechanism, the accuracy and robustness of depression detection can be effectively improved.
[0109] AV-CAM is used to align the audio features and the image features and map them to a unified three-dimensional feature space D alignment , and an A-V feature alignment function is adopted to project the audio features into thus converting the two into feature spaces of the same dimension;
[0110] The alignment process is achieved through convolutional mapping or linear mapping to ensure that the audio and image features are converted to a unified feature dimension;
[0111] After the feature alignment is completed, the audio features and the image features become:
[0112]
[0113] The aligned audio query feature Q audio and the key-value pair K of the image features image and V image are input into the audio-visual cross-attention mechanism, and this process is expressed as:
[0114]
[0115] The interacted feature I avis further input into a multi-layer perceptron (MLP) regression module to generate the final depression prediction score:
[0116] S av = MLP(I av ) (15)
[0117] where I av represents the feature representation generated after the audio query feature Q audio interacts with the key-value pairs K image and V image of the image features.
[0118] AMFM aims to dynamically combine the prediction results from different modalities to ensure that the final prediction result is more accurate and reliable;
[0119] AMFM integrates the scores of the prediction results from different sources of F3DA, AEF, and AV-CAM. The input scores of F3DA, AV-CAM, and AEF are respectively denoted as S a , S av and S v ;
[0120] The scores are initially represented in the form of a vector of size n×1, where n is the number of samples in the batch. For ease of fusion, the modality-specific scores are concatenated along the feature dimension to form a matrix. The concatenation process is expressed as:
[0121] S concat = [S v , S av , S a ∈ R n×3 (16)
[0122] Next, the Gating Network calculates a weight vector w ∈ R n×3 , assigns weights to each modality, and this weight is calculated based on the concatenated score matrix;
[0123] Each component of the weight vector respectively represents the relative importance of the video, audio-visual interaction, and audio modalities;
[0124] The calculation of the weights is completed by a multi-layer perceptron (MLP) and combined with the Sigmoid activation function to ensure that the value of each weight is between 0 and 1. The calculation formula of the weight vector is:
[0125]
[0126] where w v , w av and w aThey are the learnable weights for video, audiovisual interaction, and audio respectively. The calculated weight vector w is used to perform a weighted sum of the modality-specific prediction scores to obtain the final fused prediction result y. The specific formula is:
[0127] Y = w v ·S v + w av ·S av + w a ·S a (18)
[0128] The final output y represents the result of dynamically weighted fusion of the scores of each modality. The weights are adaptively adjusted through learning to reflect the relative importance of each modality in predicting the target variable, thereby significantly improving the prediction accuracy and the robustness of the model.
[0129] In summary:
[0130] The present invention proposes a method for implementing an innovative multi-modal depression assessment framework. This framework can effectively integrate facial videos and speech signals, thereby significantly improving the accuracy and robustness of depression detection. By comprehensively considering the emotional information in the visual and auditory modalities, the comprehensiveness and reliability of the detection results are enhanced. In order to capture more detailed emotional cues, the present invention introduces a facial 3D attention mechanism, which can accurately extract facial micro-expressions and thus better reflect the emotional changes of depression patients. In addition, the present invention also develops an audiovisual cross-modal attention mechanism for feature alignment and dynamically adjusts the contributions of different modalities through an adaptive multi-score fusion module to ensure the best fusion effect, thereby further improving the overall performance of the model.
[0131] Based on the above content, the following verification is proposed:
[0132] To comprehensively verify the effectiveness of the method proposed in the present invention, as shown in Table 1 and Table 2, the present invention systematically evaluated the proposed method on the AVEC2013 and AVEC2014 datasets and achieved remarkable detection performance. In the AVEC2013 dataset, the method of the present invention obtained the best result in terms of MAE (Mean Absolute Error), reaching 7.30, indicating that the model can generally predict the severity of depression more accurately. At the same time, although the result of RMSE (Root Mean Square Error) was 9.53, slightly higher than that of MAE, the overall error level was still better than that of most existing methods. The task setting of the AVEC2013 dataset was a free task, and the performance of the participants was more natural and unrestricted, which posed a higher challenge for the model to capture subtle emotional changes and potential patterns. However, by introducing a multi-modal attention mechanism and an Adaptive Multi-Score Fusion Module (AMFM), the method of the present invention effectively improved the ability to capture depression-related features, thus still achieving excellent prediction results in the context of free tasks.
[0133] In the AVEC2014 dataset, the method of the present invention achieved the best performance in both MAE and RMSE evaluation metrics, which were 7.25 and 9.47 respectively. Compared with the AVEC2013 dataset, the task setting of the AVEC2014 dataset was more specific, including the "North Wind Task" and the "Free Form Task", and the participants needed to complete specific reading aloud or Q&A tasks. This task setting provided more restrictive conditions for the feature learning of the model. Therefore, the model of the present invention showed higher stability and accuracy on the AVEC2014 dataset. Especially in terms of RMSE, compared with the AVEC2013 dataset, the smaller error value indicated that the method of the present invention could capture the subtle changes of depression features more precisely. In addition, the improvement of the method of the present invention in multi-modal feature alignment and fusion strategies enabled the model to maintain good detection performance on data of different task types, thus demonstrating strong generalization ability and robustness.
[0134] Table 1 Comparison with previous advanced methods on the AVEC2013 dataset
[0135]
[0136] Table 2 Comparison with previous advanced methods on the AVEC2014 dataset
[0137]
[0138]
[0139] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. An audio-visual fusion multimodal evaluation method for depression detection, characterized in that: At least include the following steps: S1: Separate the input audio and video to obtain audio data and video data. Extract initial features from the audio data and video data, and process the extracted initial features through a facial 3D attention module and an audio emotion extraction module. The facial 3D attention module is the F3DA, and the audio emotion extraction module is the AEF; S2: The features extracted in S1 are jointly input into an audio-visual cross-modal attention module. The audio-visual cross-modal attention module is the AV-CAM. The AV-CAM achieves modality alignment through an audio-visual cross-attention mechanism, enabling the interaction and fusion of key information between the audio and video modalities, enhancing the complementarity between modalities, and outputting an audio-visual interaction prediction result; S3: The scores generated from the output results of the F3DA, AEF, and AV-CAM are passed into an adaptive multi-score fusion module. The adaptive multi-score fusion module is the AMFM. The AMFM can dynamically adjust the weights according to the contribution degrees of different modalities, thereby generating a more comprehensive and reliable depression detection result.
2. The multimodal evaluation method for depression detection by audio-visual fusion according to claim 1, wherein: The separation of the audio and video uses the integrated software Kingshiper Vocal Remover, which can effectively separate the audio and video and generate separate audio and video files.
3. The multimodal evaluation method for depression detection by audio-visual fusion according to claim 1, wherein: The F3DA aims to extract micro-expressions and facial dynamic features, that is, obtain image features, so as to capture fine-grained emotional cues in the video modality, and then output a video prediction result; The F3DA is equipped with a facial 3D attention mechanism. The facial 3D attention mechanism combines a channel attention mechanism and a spatial attention mechanism to enhance the attention to local facial features related to depression. The application of the F3DA at least includes the following steps: First, obtain a channel descriptor through global average pooling, and generate channel attention weights through the following formula: a c = σ(W2δ(W1z c )) (2) X c = a c ⊙X (3) Among them, X is the input feature map; F gp is the global pooling operation; a c is the channel attention weight; X c is the channel feature; a c is the channel descriptor; T is the time step; H is the height; W is the width; σ is the Sigmoid activation function; W1 and W2 are weights; δ is the ReLU function; ⊙ is the element-wise multiplication; Then, the spatial attention mechanism performs pooling operations on the feature map enhanced by the channels, that is, average pooling and max pooling, and generates spatial attention weights through the following formula: a s = σ(F conv (X avg ; X max )) ∈ R T×H×W×1 (5) Among them, X avg is average pooling; X max is max pooling; C represents channels; F conv represents the convolution operation; a s is the spatial attention weight; Finally, the optimized feature map X is obtained by applying channel and spatial attention through element-wise multiplication out : X out = a s ⊙X c (6) Generate the final depression prediction feature through global average pooling operation: z out = F gp (X out ) (7) S v = MLP(reshape(z out , B, D))(8) Among them, reshape means to change the tensor; B is the batch size; D is the feature dimension.
4. The audio-visual fusion multi-modal evaluation method for depression detection according to claim 3, characterized in that: The AEF is used to model the speech features in the audio modality, that is, obtain audio features, to identify intonation, rhythm, and emotional changes, and then output an audio prediction result; The AEF is based on a Transformer encoder and a multi-head self-attention mechanism. The AEF effectively improves the prediction ability of depression-related audio features by extracting high-level spatio-temporal features in audio time series data. The application of the AEF at least includes the following steps: For the input audio sequence X ∈ R N×L×D , where N represents the batch size, L represents the sequence length, and D represents the feature dimension at each time step; After passing through the Transformer encoder, the encoded representation X is obtained encoded ∈R N×L×D : X encoded = Encoder(X) (9) The Transformer encoder effectively captures the global dependencies of the sequence and generates an encoded representation containing temporal and feature dimension information; Subsequently, the multi-head self-attention mechanism is applied to weight the encoded features to generate the output X with attention weights atten ∈R N×L×D : X attn = MHA(X encoded , X encoded , X encoded ) (10) The multi-head self-attention mechanism can aggregate the information of the input sequence and emphasize the features of key time steps according to the attention weights, which helps to capture emotional changes and temporal features; Then, by performing a summation operation on the time dimension L, the final feature representation X is obtained sum ∈R N×L×D , which compresses the temporal information and provides a more compact feature representation for subsequent regression tasks: Aggregate the time series information through a pooling operation to ensure that the model focuses on the key parts and extracts the core temporal features for depression detection; Finally, the aggregated feature X sum is processed by a multi-layer perceptron regression module to generate the final audio depression score: S a = MLP(X sum ) (12) By combining the characteristics of the Transformer encoder and the self-attention mechanism, the accuracy and robustness of depression detection can be effectively improved.
5. The multimodal evaluation method for depression detection by audio-visual fusion according to claim 4, characterized in that: The AV-CAM is used to align audio features with image features and map them to a unified three-dimensional feature space D alignment , and an A-V feature alignment function is adopted to project the audio features into so as to convert the two into feature spaces of the same dimension; The alignment process is achieved through convolutional mapping or linear mapping to ensure that audio and image features are transformed into a unified feature dimension; After feature alignment is completed, the audio features and image features become: Aligned audio query feature Q audio Key-value pair K of image features image and V image Are passed as inputs into the audio-visual cross-attention mechanism, and this process is expressed as: Feature I after interaction av is further input into a multi-layer perceptron regression module to generate the final depression prediction score: S av = MLP(I av ) (15) Among them, I av represents the audio query feature Q audio and the key-value pair K image and V image of the image feature after interaction to generate the feature representation.
6. The multimodal assessment method for depression detection by audio-visual fusion according to claim 5, characterized in that: The AMFM aims to dynamically combine the prediction results from different modalities, thus ensuring that the final prediction result is more accurate and reliable; The AMFM integrates the scores of prediction results from different sources of F3DA, AEF, and AV-CAM. The input scores of F3DA, AV-CAM, and AEF are respectively represented as S a , S av and S v ; The scores are initially represented in the form of a vector of size n×1, where n is the number of samples in the batch. For ease of fusion, the modality-specific scores are concatenated along the feature dimension to form a matrix. The concatenation process is expressed as: S concat = [S v , S av , S a ∈ R n×3 (16) Next, the gating network computes a weight vector w ∈ R n×3 , which assigns weights to each modality, based on the concatenated score matrix; Each component of the weight vector \(w = [w v , w av , w a T represents the relative importance of video, video-audio interaction, and audio modalities respectively; The calculation of the weights is completed through a multi-layer perceptron and combined with the Sigmoid activation function to ensure that the value of each weight is between 0 and 1. The calculation formula for the weight vector is: w = σ(MLP(S concat )) = [w v , w av , w a T ∈R n×3 (17) where, w v , w av and w a are the learnable weights for video, video-audio interaction, and audio respectively. The calculated weight vector w is used to perform a weighted sum of the modality-specific prediction scores to obtain the final fused prediction result y. The specific formula is as follows: y = w v ·S v +w av ·S av +w a ·S a (18) The final output y represents the result of dynamically weighted fusion of the scores of each modality, where the weights are adaptively adjusted through learning to reflect the relative importance of each modality in predicting the target variable, thereby significantly improving the prediction accuracy and the robustness of the model.
Citation Information
Cited By
Pancreatitis severity assessment method based on layered reliable evidence fusion
CN121075666A
Audio and video multi-mode identification method based on artificial intelligence
CN121095997A
An audio and video multi-modal recognition method based on artificial intelligence
CN121095997B