Automatic depression detection method and system
By constructing a multi-scale spatiotemporal enhancement subnet and a multi-resolution temporal difference molecular network, combined with an adaptive fusion strategy, the problem of incomplete feature extraction in the existing depression detection methods is solved, and the effective fusion of fine-grained global spatiotemporal semantics and dynamic information is achieved, and the accuracy of depression detection is improved.
Patent Information
- Application Number
- CN202510350873.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing depression detection methods fail to effectively extract fine-grained global spatiotemporal semantics and dynamic information in multiple time ranges, and cannot effectively utilize the correlation between different semantics for adaptive fusion, resulting in insufficient feature extraction and inaccurate detection results.
A multi-scale spatiotemporal enhancement subnet was constructed to perform global spatiotemporal semantic feature extraction, combined with a multi-resolution temporal difference molecular network for dynamic differential semantic feature extraction, and the two were fused through an adaptive fusion strategy and input to the full connection layer to achieve depression detection.
Effectively extract occult facial behavior characteristics and dynamic information within multiple time ranges, achieving more accurate depression detection and improving the accuracy and comprehensiveness of the test results.
Smart Images

Figure CN120280144A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of facial expression recognition, and particularly relates to an automatic depression detection method and system. Background Art
[0002] With the acceleration of the modern life rhythm and the increase of psychological pressure, social mental health problems have become more and more prominent. Early identification, accurate diagnosis and timely treatment of depression patients are crucial. However, current depression detection methods, such as clinical interviews and structured questionnaires, etc., are difficult to meet the needs of large-scale screening. In addition, such methods are not only susceptible to the subjective biases of patients and deliberate concealment of symptoms, but also increase the risk of misdiagnosis due to differences in the understanding of diagnostic criteria, experience accumulation and judgment strategies of clinicians. Therefore, how to explore an automatic depression detection (ADD) method with the help of artificial intelligence technology has become an important research direction in the field of mental health.
[0003] Currently, physiological signals (electroencephalogram, brain imaging) and behavioral signals (eye movement, facial expression) can both be used as important clues for automatic depression detection. Given the unique advantages of facial expressions in emotional state representation, non-verbal communication, non-contact acquisition and obtaining, etc., depression detection based on facial expressions has received extensive attention from researchers. In recent years, researchers have proposed a variety of depression detection networks based on convolutional neural networks, recurrent neural networks and attention mechanisms, etc. Although these methods can capture macro facial expression changes, they ignore the hidden facial behavior characteristics such as low visual saliency and spatio-temporal sparsity that patients often show due to emotional depression or deliberate concealment, resulting in a lack of attention to fine-grained features. For example, a depression assessment method and device based on audio-visual multi-modal data fusion disclosed in Chinese Patent Publication No. CN118173267A only extracts the facial video and audio of the subject to obtain corresponding features, without capturing the hidden facial behavior and extracting fine-grained global spatio-temporal semantics. On the other hand, there are progressive changes (such as continuously low mood, loss of interest) in the faces of depression patients within a relatively long time range and rapid fluctuations (such as sudden strong emotional fluctuations, emotional instability) within a relatively short time range. The time dispersion of the facial behavior motion sequence further increases the difficulty of automatic depression detection. Traditional fixed-resolution modeling methods only focus on the motion information within a single time range and are difficult to take into account both the short-term changes and long-term trends of facial behavior. Therefore, it is necessary to capture the dynamic information within multiple time ranges. In addition, to enhance the discriminability of spatio-temporal semantics, it is necessary to fuse features from different sources. However, existing feature concatenation strategies are difficult to effectively utilize the complementary information between different features. How to effectively utilize the correlation between different semantics for adaptive fusion is an urgent problem to be solved. Summary of the Invention
[0004] The technical problem to be solved by the present invention is that existing depression detection methods do not extract fine-grained global spatio-temporal semantics and dynamic information within multiple time ranges, and cannot utilize the correlation between different semantics for adaptive fusion, resulting in incomplete feature extraction and inaccurate automatic detection results.
[0005] The present invention solves the above technical problems through the following technical means: An automatic depression detection method, comprising:
[0006] S1. Construct a multi-scale spatio-temporal enhancement subnet to extract global spatio-temporal semantic features;
[0007] S2. Construct a multi-resolution time difference subnet to extract dynamic difference semantic features;
[0008] S3. Fuse the global spatio-temporal semantic features and the dynamic difference semantic features to obtain adaptive fusion features, and input them into a fully connected layer to achieve depression detection.
[0009] Beneficial effects: The present invention uses a multi-scale spatio-temporal enhancement subnet to extract global spatio-temporal semantic features, and uses a multi-resolution time difference subnet to extract dynamic difference semantic features, thereby effectively extracting hidden facial behavior features and dynamic information within multiple time ranges, and fusing the global spatio-temporal semantic features and the dynamic difference semantic features to obtain adaptive fusion features, thereby utilizing the correlation between different semantics for adaptive fusion, with relatively comprehensive feature extraction and relatively accurate detection results.
[0010] Further, before S1, it also includes constructing a 3D inverted residual module, and the construction process of the 3D inverted residual module is as follows:
[0011] Perform a pointwise convolution operation on the input feature map and pass through a ReLU activation function to obtain a channel-expanded feature map;
[0012] Perform a depthwise convolution operation on the channel-expanded feature map and pass through another ReLU activation function to obtain a depthwise convolution feature map;
[0013] Perform a pointwise convolution operation on the depthwise convolution feature map and pass through a linear activation function to obtain a channel-compressed feature map;
[0014] Add the channel-compressed feature map to the input feature map to obtain an output feature map.
[0015] Furthermore, the construction process of the multi-scale spatio-temporal enhancement subnet is as follows:
[0016] S11. Evenly divide the input feature tensor X into l feature sub-blocks along the time dimension
[0017] S12. Input the feature sub - blocks into the 3D inverted residual module to perform layer - by - layer convolutional operations along the forward and reverse directions respectively to obtain spatio - temporal multi - scale features and i = 1, 2, … l, 3D - IR is the abbreviation of the 3D inverted residual module;
[0018] S13. Concatenate the spatio - temporal multi - scale features and along the time dimension, and perform a residual connection with the input feature tensor X to obtain multi - scale enhanced semantic features;
[0019] S14. Input the multi - scale enhanced semantic features into N layers of 3D - IR, and then perform global average pooling to obtain global spatio - temporal semantic features.
[0020] Furthermore, S12 includes:
[0021]
[0022] Among them, 3D - IR(·) represents a single - layer 3D inverted residual module.
[0023] Furthermore, the process of constructing the multi - resolution time - difference subnet is as follows:
[0024] S21. Divide the input feature tensor X into T frame feature maps and perform a first - order difference operation with a step size of m to obtain a first - order difference map with a step size of m, m ∈ [1, M], M represents the resolution, that is, the total number of different step sizes;
[0025] S22. Perform a second - order difference operation on each frame of the first - order difference map with a step size of m to obtain a second - order difference map with a step size of m;
[0026] S23. Concatenate the first - order difference map and the second - order difference map along the channel dimension, and use a single - layer point - wise convolution to fuse the channel features, thereby obtaining a difference map E with a step size of m m ;
[0027] S24. Input the difference map E m into N layers of 3D - IR, and then perform a global average pooling layer to obtain dynamic difference semantic features with a step size of m.
[0028] Furthermore, S21 includes:
[0029] Divide the input feature tensor X into T frame feature maps and perform a first - order difference operation:
[0030]
[0031] where \(m\in[1,M]\) is the difference step size; represents the first-order difference operation result of the \(t\)-th frame with step size \(m\); \(U\) m represents the first-order difference map with step size \(m\).
[0032] Furthermore, S22 includes:
[0033] Perform a second-order difference operation on the basis of the first-order difference map \(U\) m :
[0034]
[0035] where represents the second-order difference operation result of the \(t\)-th frame with step size \(m\); \(G\) m represents the second-order difference map with step size \(m\).
[0036] Furthermore, S3 includes:
[0037] S31. Concatenate the dynamic difference semantic features of each step size along the token dimension to obtain a mapping matrix \(F\), and multiply the mapping matrix \(F\) by three different first weight matrices respectively to obtain the query matrix, key matrix, and value matrix of each attention head of the multi-head self-attention mechanism;
[0038] S32. Perform scaled dot-product attention calculation on the query matrix, key matrix, and value matrix of the attention heads of the multi-head self-attention mechanism to obtain the calculation result of each attention head, and perform a multi-head concatenation operation on the calculation results of each attention head to obtain the self-attention difference feature;
[0039] S33. Preset three second weight matrices, multiply the global spatio-temporal semantic feature by a second weight matrix to construct the query matrix of each attention head of the multi-head cross-attention mechanism, and multiply the self-attention difference feature by the other two second weight matrices respectively to obtain the key matrix and value matrix of each attention head of the multi-head cross-attention mechanism;
[0040] S34. Perform scaled dot-product attention calculation on the query matrix, key matrix, and value matrix of the attention heads of the multi-head cross-attention mechanism to obtain the calculation result of each attention head, and perform a multi-head concatenation operation on the calculation results of each attention head to obtain the cross-attention difference feature;
[0041] S35. Concatenate the cross-attention difference feature and the global spatio-temporal semantic feature along the channel dimension and input them into a fully connected layer to predict the depression level.
[0042] The present invention also provides an automatic depression detection system, including:
[0043] A first feature extraction module, configured to construct a multi-scale spatio-temporal enhancement subnet and perform global spatio-temporal semantic feature extraction;
[0044] The second feature extraction module is used to construct a multi-resolution time-difference subnet for dynamic differential semantic feature extraction;
[0045] The result output module is used to fuse the global spatio-temporal semantic features and the dynamic differential semantic features to obtain an adaptive fusion feature, and input it into the fully connected layer to realize depression detection.
[0046] The advantages of the present invention are as follows:
[0047] (1) The present invention uses a multi-scale spatio-temporal enhancement subnet for global spatio-temporal semantic feature extraction, and a multi-resolution time-difference subnet for dynamic differential semantic feature extraction, so as to effectively extract hidden facial behavior features and dynamic information within multiple time ranges, and fuse the global spatio-temporal semantic features and the dynamic differential semantic features to obtain an adaptive fusion feature, thereby performing adaptive fusion using the correlation between different semantics. The feature extraction is relatively comprehensive and the automatic detection result is relatively accurate.
[0048] (2) Aiming at the concealment of facial behavior features and the dispersion of key motion sequences in depression detection, the present invention proposes a depression detection network combining multi-scale spatio-temporal enhancement and multi-resolution time difference. In addition, considering that it is difficult to train a network model with a large number of parameters due to the limited number of samples in the depression detection dataset, a plug-and-play 3D inverted residual module (3D-IR: 3D Inverted Residual) is further constructed based on depthwise separable convolution to replace the 3D-CNN with a large number of parameters, so as to reduce the model parameter scale while maintaining the feature representation ability. Compared with the baseline network, the proposed end-to-end network can significantly improve the accuracy of automatic depression detection through the multi-scale enhancement structure and the multi-resolution difference mechanism.
[0049] (3) To solve the problem of difficult spatio-temporal semantic extraction caused by the concealment of facial behavior features, a multi-scale enhancement structure (MES: Multi-scale Enhancement Structure) is constructed by using symmetric multi-scale convolution operations. By performing two-way convolution operations on the spatio-temporal context of multiple scales, the capture of fine-grained global spatio-temporal semantics is realized.
[0050] (4) To overcome the problem of difficult dynamic information encoding caused by the dispersion of key motion time series, a multi-resolution difference mechanism (MDM: Multi-resolution Difference Mechanism) is proposed based on time difference operations. Through hierarchical cascaded difference operations, the first-order difference is used to capture the instantaneous emotion change of the time series feature map at multiple time resolutions, and the second-order difference is used to model the emotion fluctuation trend, so as to encode the feature changes in different time windows and realize the extraction of multi-resolution dynamic differential semantics. Brief Description of the Drawings
[0051] Figure 1 Schematic diagram of the structure of the depression detection network (SETD-Net) in an automatic depression detection method disclosed in an embodiment of the present invention;
[0052] Figure 2 Schematic diagram of the detailed structures of 3D-CNN and 3D-IR in an automatic depression detection method disclosed in an embodiment of the present invention, where, Figure 2 (a) Schematic diagram of the detailed structure of 3D-CNN, Figure 2 (b) Schematic diagram of the detailed structure of 3D-IR;
[0053] Figure 3 Schematic diagram of the multi-scale spatio-temporal enhancement subnet in an automatic depression detection method disclosed in an embodiment of the present invention;
[0054] Figure 4 Schematic diagram of facial expressions in different time ranges in an automatic depression detection method disclosed in an embodiment of the present invention;
[0055] Figure 5 Schematic diagram of the multi-resolution time difference subnet in an automatic depression detection method disclosed in an embodiment of the present invention;
[0056] Figure 6 Schematic diagram of the performance comparison of different parameter configurations in MES in an automatic depression detection method disclosed in an embodiment of the present invention, Figure 6 (a) Schematic diagram of the performance comparison of different parameter configurations in MES on the AVEC2013 dataset, Figure 6 (b) Schematic diagram of the performance comparison of different parameter configurations in MES on the AVEC2014 dataset;
[0057] Figure 7 Schematic diagram of the performance comparison of different feature fusion methods in an automatic depression detection method disclosed in an embodiment of the present invention, Figure 7 (a) Schematic diagram of the performance comparison of different feature fusion methods on the AVEC2013 dataset, Figure 7 (b) Schematic diagram of the performance comparison of different feature fusion methods on the AVEC2014 dataset;
[0058] Figure 8 Schematic diagram of the influence of the balance factor on the network performance in an automatic depression detection method disclosed in an embodiment of the present invention, Figure 8 (a) Schematic diagram of the influence of the balance factor on the network performance on the AVEC2013 dataset, Figure 8 (b) Schematic diagram of the influence of the balance factor on the network performance on the AVEC2014 dataset;
[0059] Figure 9 Schematic diagram of the influence of the number of 3D-IR stacking layers on the network performance in an automatic depression detection method disclosed in an embodiment of the present invention Figure 9 (a) Schematic diagram of the influence of the number of 3D-IR stacking layers on the network performance on the AVEC2013 dataset Figure 9 (b) Schematic diagram of the influence of the number of 3D-IR stacking layers on the network performance on the AVEC2014 dataset
[0060] Figure 10 Schematic diagram of the visualization of the detection results of the SETD-Net network in an automatic depression detection method disclosed in an embodiment of the present invention Detailed implementation manners
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] Embodiment 1
[0063] As Figure 1 shown, Embodiment 1 of the present invention provides an automatic depression detection method, which uses a depression detection network (SETD-Net: Multi-scale Spatiotemporal Enhancement and Multi-resolution Temporal Difference Newtork) for depression detection. The depression detection network includes three main parts: a multi-scale spatiotemporal enhancement subnet, a multi-resolution temporal difference subnet, and an adaptive fusion subnet. In particular, in order to reduce network parameters to adapt to small-sample depression datasets, a 3D inverted residual module (3D-IR) is constructed to replace 3D-CNN and is embedded in the multi-scale spatiotemporal enhancement subnet and the multi-resolution temporal difference subnet. The method includes the following steps:
[0064] 1. Construct a 3D inverted residual module (3D-IR)
[0065] 3D-CNN shows unique advantages in capturing spatiotemporal features in the processing of time-series data such as video analysis. Its multi-dimensional convolutional kernel structure can synchronously extract dynamic information in the spatial dimension and the time dimension, but its high number of parameters leads to a significant increase in model complexity. Figure 2 Schematic diagrams of the detailed structures of 3D-CNN and 3D-IR Figure 2(a) is a schematic diagram of the detailed structure of 3D-CNN, Figure 2 (b) is a schematic diagram of the detailed structure of 3D-IR. Figure 2 Among them, PWConv and DWConv respectively represent pointwise convolution and depthwise separable convolution; β represents the channel expansion coefficient, and the value in the present invention is 4; s represents the stride of the convolution kernel, and the value in the present invention is 1 or 2; C represents the number of input channels of the module. The dashed skip connection means that the residual connection is used when s = 1. The 3D-CNN structure with a convolution kernel size of 3×3×3 (such as Figure 2 (a)) The parameter size P 3D-CNN and the computational cost F 3D-CNN can be expressed as:
[0066]
[0067] In Equation (1), T, H, and W respectively represent the time length, height, and width of the input feature map.
[0068] In the automatic depression detection task, due to the influence of factors such as the difficulty of obtaining medical data and privacy protection, the currently available datasets and scales are relatively small. As can be seen from Equation (1), the deep network model embedded with 3D-CNN is prone to being difficult to adapt to the small-sample depression dataset due to its large parameters. To ensure that the proposed network can adapt to the small-sample depression dataset and reduce the risk of overfitting, the present invention constructs a plug-and-play 3D inverted residual module (3D-IR) based on depthwise separable convolution, such as Figure 2 (b) shown.
[0069] First, in order to enhance the feature representation ability and fuse the information between different channels, pointwise convolution with a kernel size of 1×1×1 is used to integrate the cross-channel information and increase the number of channels:
[0070] Y pw1 = Relu(PWConv(X)) (2)
[0071] In Equation (2), represents the input feature tensor; PWConv(·) represents the pointwise convolution operation; Relu(·) represents the ReLU activation function; represents the channel-expanded feature map.
[0072] Secondly, considering that the computational complexity of full-channel convolution is relatively high, depthwise convolution with a kernel size of 3×3×3 is used to perform convolution operations independently on each channel:
[0073] Y dw = Relu(DWConv(Y pw1 )) (3)
[0074] In Equation (3), DWConv(·) represents the depth convolution operation; represents the depth convolution feature map, and its number of channels is the same as that of Y pw1 and remains consistent.
[0075] In addition, in order to further fuse channel information and reduce feature redundancy, a pointwise convolution with a kernel size of 1×1×1 is further used to fuse multi-channel features and reduce the number of channels:
[0076] Y pw2 = Line(PWConv(Y dw )) (4)
[0077] In Equation (4), Line(·) represents the linear activation function; represents the channel compression feature map. The specific parameters of the 3D-IR constructed in the present invention are shown in Table 1.
[0078] Table 1. Specific parameters of 3D-IR
[0079]
[0080] Finally, to achieve cross-layer reuse of shallow features and alleviate the gradient decay problem in deep networks, a residual connection is added between Y pw2 and X to retain the original information and improve the robustness of the model. The output Y of 3D-IR can be expressed as:
[0081] Y = Y pw2 + X (5)
[0082] In Equation (5), represents the output feature map of 3D-IR.
[0083] From the above steps, it can be seen that the number of parameters P 3D-IR and the computational cost F 3D-IR of the constructed 3D-IR module can be expressed as:
[0084]
[0085] Combining Equation (1) and Equation (6), when β = 4 and s = 1, the parameter ratio R1 and the computational cost ratio R2 of 3D-IR and 3D-CNN can be expressed as:
[0086]
[0087] It can be seen from Equation (7) that when the number of channels C ≥ 32, the values of R1 and R2 are between 0.296 and 0.421. Thus, it can be known that the number of parameters and the floating-point operation amount of 3D-IR are about 1 / 3 of those of 3D-CNN.
[0088] To extract rich spatio-temporal context information, N layers of 3D-IR are stacked in the network to obtain deep semantics:
[0089] Y = 3D-IR N (X) (8)
[0090] In Equation (8), 3D-IR N (·) represents N layers of 3D-IR; X and Y respectively represent the input and output of 3D-IR N (·).
[0091] II. Construct a multi-scale spatio-temporal enhancement subnet to achieve fine-grained spatio-temporal semantic extraction based on the multi-scale enhancement structure (MES)
[0092] Compared with healthy people, the facial behaviors of patients with depression are more concealed and have low visual saliency and spatio-temporal sparsity (such as reduced activation of the periorbital muscle group and attenuation of the dynamic characteristics of the corners of the mouth). These characteristics lead to a significant increase in the difficulty of extracting fine-grained spatio-temporal features. To extract the concealed facial behavior features and capture the fine-grained spatio-temporal context information, the present invention constructs a multi-scale enhancement structure (MES), as Figure 3 shown.
[0093] First, to model spatio-temporal features at multiple scales, the input feature tensor X with a frame length of T is equally divided into l (l = 4) feature sub-blocks along the time dimension
[0094]
[0095] In Equation (9), Group(·) represents the grouping operation; l represents the division scale.
[0096] Second, to enhance the representation ability of the global spatio-temporal context, the constructed 3D-IR module is used to perform layer-by-layer convolutional operations along the forward and reverse directions respectively, so as to capture spatio-temporal multi-scale features and
[0097]
[0098] In Equation (10), and respectively represent the spatio-temporal multi-scale features obtained by 3D-IR convolution from left to right and from right to left.
[0099] In addition, to capture multi-scale enhanced semantics, the spatio-temporal multi-scale features and are concatenated along the time dimension and connected with X by residual connection:
[0100]
[0101] In formula (11), represents multi-scale enhanced semantics; Concate t (·) represents a concatenation operation along the time dimension.
[0102] Finally, in order to capture the fine-grained global spatio-temporal semantics G, the multi-scale enhanced semantics P is input into the N-layer 3D-IR, and global average pooling is performed:
[0103] G = GAP(3D-IR N (P)) (12)
[0104] In formula (12), represents the global spatio-temporal semantics, where C′ represents the output channel number of 3D-IR N (·); GAP(·) represents the global average pooling operation.
[0105] It can be seen from the above construction process that the constructed MES structure features the 3D-IR module as the feature extraction unit, hierarchically processes the multi-scale feature maps through two-way layer-by-layer convolution operations, not only can focus on the hidden facial behaviors related to depression, but also realizes the extraction of fine-grained global spatio-temporal semantics.
[0106] III. Construct a multi-resolution time difference subnet to achieve dynamic difference semantic extraction based on the multi-resolution difference mechanism (MDM);
[0107] The global spatio-temporal semantics G captured by the multi-scale spatio-temporal enhancement subnet represents the hidden facial behavior information related to depression. However, depressive symptoms usually manifest in multiple time ranges, such as Figure 4 in, the short-term behavior features (step = 1) can express the patient's instantaneous emotional response, while the long-term behavior features (step = 4) can express the overall emotional migration.
[0108] In order to capture the patient's facial dynamic information in multiple time ranges, the present invention proposes a multi-resolution difference mechanism (MDM) to parallelly extract the difference maps of M resolutions, as Figure 5 shown.
[0109] The following will take the extraction process of the m-th difference map as an example for specific illustration. First, in order to capture the instantaneous intensity of the patient's emotional fluctuations, the input feature tensor X is divided into T-frame feature maps and perform a first-order difference operation on:
[0110]
[0111] In formula (13), m ∈ [1, M] is the difference step; Denote the first-order difference operation result of the t-th frame with step size m; Denote the first-order difference graph with step size m. The first-order difference graph U m Can reflect the instantaneous manifestation of the patient's emotional changes, but lacks attention to the long-term trend of the patient's emotional fluctuations. To characterize the evolution trend of the patient's emotional fluctuations, perform a second-order difference operation on the first-order difference graph U m Based on:
[0112]
[0113] In Equation (14), Denote the output of the t-th frame of the second-order difference operation with step size m; Denote the second-order difference graph with step size m.
[0114] Then, to enhance the characterization ability of the difference features, concatenate the first-order difference graph U m And the second-order difference graph G m Along the channel dimension, and use a single-layer pointwise convolution to fuse the channel features, thereby obtaining the difference graph with step size m
[0115] E m = PWConv(Concate c (U m , G m )) (m ∈ [1, M]) (15)
[0116] In Equation (15), Concate c (·) represents the concatenation operation along the channel dimension.
[0117] In addition, to capture complex spatial relationships and extract deep semantics, input the difference graph E m Into the N-layer 3D-IR and global average pooling layer to capture dynamic difference semantics:
[0118] F m = GAP(3D-IR N (E m )) (m ∈ [1, M]) (16)
[0119] In Equation (16), Denote the dynamic difference semantics with step size m.
[0120] Finally, to model the dynamic difference features of facial behaviors at multiple resolutions, parallelly compute the dynamic difference semantics of M resolutions:
[0121] {F 1 , …, F m , …, F M} (17)
[0122] As can be seen from the above construction process, the MDM mechanism extracts dynamic differential semantics with different time granularities in parallel through differential operations with multi-level time resolutions. At each level of time resolution, the MDM mechanism uses first-order differences and second-order differences respectively to capture the instantaneous manifestations and evolution trends of patients' emotional fluctuations. Therefore, it can not only model the motion information scattered at different time granularities, but also complete the extraction of multi-resolution dynamic differential semantics.
[0123] IV. Fuse the global spatio-temporal semantic features and dynamic differential semantic features to obtain adaptive fusion features, and input them into the fully connected layer for depression detection, realizing adaptive fusion based on the dual-stage attention strategy (DAS).
[0124] After extracting the global spatio-temporal semantics G and dynamic differential semantics Traditional strategies use feature concatenation methods for feature fusion. However, the feature concatenation method strategy is difficult to fully utilize different semantic information, thus affecting the feature representation ability. To give full play to the complementary advantages of global spatio-temporal semantics and dynamic differential semantics, the present invention designs a dual-stage attention strategy (DAS) based on the self-attention mechanism and cross-attention mechanism, as Figure 1 shown.
[0125] In the first stage, a multi-head self-attention mechanism is introduced to capture the complementary information between multi-resolution dynamic differential semantics First, construct a mapping matrix and calculate the query matrix, key matrix, and value matrix of the multi-head self-attention mechanism respectively based on the mapping matrix F: In Equation (18), Concate
[0126]
[0127] (·) represents the concatenation operation along the token dimension; f (·) denotes the concatenation operation along the token dimension; and represent the query matrix, key matrix, and value matrix of the i-th attention head based on F respectively, where i k = C′ / h, h is the number of attention heads; and represent trainable weight matrices.
[0128] Then, use scaled dot-product attention to perform weighted fusion on multi-resolution dynamic differential semantics and introduce a residual connection to retain the original dynamic information:
[0129]
[0130] In Equation (19), softmax(·) represents the normalized exponential function; represents a trainable weight matrix; represents self-attention differential features; Concate k (·) represents a multi-head concatenation operation.
[0131] The self-attention differential feature S can represent complementary information between multi-resolution dynamic differential semantics However, it lacks the correlation information with the global spatio-temporal semantics G. Therefore, in the second stage, a multi-head cross-attention mechanism is introduced to model the correlation between S and G. First, a query matrix of the multi-head cross-attention mechanism is constructed with the global spatio-temporal semantics G, and a key matrix and a value matrix of the multi-head cross-attention mechanism are constructed with the self-attention differential feature S:
[0132]
[0133] In Equation (20), represents the query matrix of the i-th attention head based on G; and
[0134] respectively represent the key matrix and the value matrix of the i-th attention head based on S; W i Qd 、W i Kd and
[0135] represent trainable weight matrices.
[0136] Then, the global spatio-temporal semantics G is used to guide the self-attention differential feature S:
[0137]
[0138] In Equation (21), is a trainable weight matrix; represents the cross-attention differential feature.
[0139] Finally, in order to retain the original global context information, the cross-attention differential feature D and the global spatio-temporal semantics G are concatenated along the channel dimension and input into a fully connected layer to predict the depression level
[0140]
[0141] In Equation (22), represents the final depression feature; FC(·) represents the fully connected layer; represents the predicted depression level.
[0142] To accelerate the convergence speed of the SETD-Net network, the present invention adopts a joint optimization strategy of single-branch independent optimization and double-branch collaborative optimization, and its overall loss function can be expressed as:
[0143] L = λL g +(1 - λ)L d +L f (23)
[0144] In Equation (23), λ ∈ [0, 1] is a balance factor; L g and L d respectively represent the loss functions of two single-branch independent optimizations based on the multi-scale spatio-temporal enhancement subnet and the multi-resolution temporal difference subnet; L f represents the loss function of double-branch collaborative optimization combining the two subnets. In addition, during the single-branch training process based on the multi-resolution temporal difference subnet, to further improve the optimization efficiency, the differential process with a step size from 1 to M is trained separately. The calculation methods of various loss functions can be expressed as:
[0145]
[0146] In Equation (24), and respectively represent the prediction results of the k-th training sample in the two single-branch independent trainings based on the multi-scale spatio-temporal enhancement subnet and the multi-resolution temporal difference subnet; represents the prediction result of the k-th training sample in the double-branch collaborative training; y k represents the label of the k-th training sample; K represents the total number of training samples.
[0147] V. Experimental Results and Analysis
[0148] 5.1 Dataset and Evaluation Metrics
[0149] To illustrate the depression detection performance of SETD-Net, the present invention conducts experiments on the AVEC2013 and AVEC2014 datasets. This dataset is released by the International Audio-Visual Emotion Challenge (AVEC) and is widely used in the field of automatic depression detection research. The AVEC2013 dataset comes from the Audio-Visual Depression Language Corpus (AViD-Corpus), containing audio-visual records of 82 German-speaking subjects aged between 18 and 63 years old (average age 31.5 ± 12.3 years). The data in the dataset is collected under 14 human-computer interaction tasks, including long vowel pronunciation, problem-solving, etc. AVEC2013 contains a total of 150 videos, which have been standardly divided into a training set, a validation set, and a test set, with each video having a frame rate of 30fps. The AVEC2014 dataset contains audio-visual records of two structured tasks (i.e., Freeform and Northwind). In Freeform, subjects are required to answer questions related to their personal lives (such as "Describe a sad childhood memory"); in Northwind, subjects are required to read a specified fable text. The dataset contains a total of 300 videos, which have been standardly divided into a training set, a validation set, and a test set, with each video having an average duration of about 2 minutes. Both datasets are annotated using the second edition of the Beck Depression Inventory (BDI-II), with a scoring range of 0 - 63 points, specifically divided into: no depressive symptoms (0 - 13 points), mild depression (14 - 19 points), moderate depression (20 - 28 points), and severe depression (29 - 63 points).
[0150] To evaluate the detection accuracy of SETD-Net, the present invention uses the Mean Absolute Error (MAE) and the Root Mean Square Error (RMSE) as evaluation metrics:
[0151]
[0152] Among them, and y q respectively represent the predicted result and the true label of the q-th test sample, and Q represents the total number of test samples.
[0153] 5.2 Experimental Details
[0154] Considering the slow-changing characteristics of the facial expressions of depression patients, to reduce temporal redundancy, all videos in AVEC2013 and AVEC2014 are downsampled, and the frame rate is adjusted to 6fps. Secondly, to achieve spatially consistent alignment of facial features, OpenFace is used to detect facial key points and perform geometric correction, and the facial region is cropped to a size of 224×224 pixels. Finally, to expand the scale of the dataset, a 16-frame sliding window sampling strategy (step size 8 frames) is adopted to construct a sequence of video segments with 50% overlap. In addition, in the differential operation of MDM, to make the output of the differential operation consistent with the input size, zero-padding method is used for padding; in the grouping operation of MES, when the data does not meet the equal division condition, zero-padding method is also used for processing. To illustrate the effectiveness of the proposed SETD-Net, ResNet18 is selected as the baseline network, and the same fully connected layer structure as in Equation (22) is used to achieve depression level prediction. The proposed SETD-Net network is implemented based on the PyTorch deep learning framework and is trained and tested on the NVIDIA A30 GPU platform. The main network parameters are shown in Table 2.
[0155] Table 2. SETD-Net Network Parameter Settings
[0156]
[0157] 5.3 Experimental Analysis
[0158] 5.3.1 Performance Analysis of SETD-Net
[0159] To analyze the depression detection performance of the SETD-Net network, the present invention conducts experiments on the AVEC2013 and AVEC2014 datasets for regression tasks and classification tasks respectively, and the results are shown in Table 3. In the regression task, the sample labels are represented by depression severity scores from 0 to 63; in the classification task, the sample labels are divided into four categories: no depressive symptoms, mild depression, moderate depression, and severe depression according to BDI-II. In addition, to evaluate the recognition performance of SETD-Net for depression categories, F1 and accuracy are added as evaluation indicators in the classification task.
[0160] The experimental results of the regression task show that the MAE of the SETD-Net network on the AVEC2013 dataset is 6.06 and the RMSE is 7.60; the MAE on the AVEC2014 dataset is 6.03 and the RMSE is 7.59. Compared with the baseline network, the MAE of the proposed method on the AVEC2013 and AVEC2014 datasets is reduced by 0.74 and 0.69 respectively, and the RMSE is reduced by 1.30 and 1.34 respectively.
[0161] The experimental results of the classification task show that the SETD-Net network significantly outperforms the baseline network in the depression detection performance on two datasets. Specifically, on the AVEC2013 dataset, compared with the baseline network, the average F1 score and accuracy of the proposed network are improved by 18.05% and 22.40% respectively. On the "no depressive symptoms" and "severe depression" categories where the baseline network has the best recognition effect, the F1 values of SETD-Net reach 67.22% and 68.07% respectively; on the "mild depression" and "moderate depression" categories that are more difficult to recognize by the baseline network, the F1 values of SETD-Net are improved by 18.25% and 14.84% respectively. On the AVEC2014 dataset, SETD-Net also has a significant improvement compared with the baseline network, and its average F1 score and accuracy are increased by 22.40% and 20.53% respectively.
[0162] The above experimental results show that the proposed SETD-Net network uses a multi-scale enhancement structure to focus on the hidden facial behavior features and encodes the motion information at different time granularities with the help of a multi-resolution difference mechanism. Therefore, the detection performance on the two datasets is effectively improved. In addition, the RMSE values on the two datasets are both lower than 7.60, which also shows that the proposed network has good generalization ability and prediction stability.
[0163] Table 3. Detection performance of SETD-NET
[0164]
[0165] 5.3.2 Ablation experiment
[0166] To analyze the influence of the modules constructed in the present invention on the detection performance of the SETD-Net network, the present invention conducts 5 groups of ablation experiments on the AVEC2013 and AVEC2014 datasets, as shown in Table 4. In Table 4, Strategy 1 and Strategy 2 respectively embed a multi-scale enhancement structure and a multi-resolution difference mechanism in the baseline network; Strategy 3 extracts global spatio-temporal semantics and dynamic difference semantics in parallel and uses a cascade strategy to adaptively fuse the two types of semantics.
[0167] Compared with the baseline network, Strategy 1 with the embedded MES structure reduces the MAE and RMSE on the AVEC2013 dataset by 0.40 and 0.39 respectively, and reduces the MAE and RMSE on the AVEC2014 dataset by 0.35 and 0.41 respectively. This indicates that the MES structure can capture fine-grained global spatio-temporal semantics through two-way layer-by-layer convolution operations, thereby paying attention to the hidden facial behavior information. Similarly, compared with the baseline network, Strategy 2 with the embedded MDM mechanism reduces the MAE and RMSE on AVEC2013 by 0.38 and 0.43 respectively, and reduces the MAE and RMSE on AVEC2014 by 0.26 and 0.53 respectively. This shows that the MDM mechanism can extract multi-resolution dynamic differential semantics through differential operations at multiple time resolutions, thereby paying attention to the motion information scattered at different time granularities.
[0168] Compared with Strategy 3, the SETD-Net network with the embedded DAS strategy reduces the MAE and RMSE on AVEC2013 by 0.24 and 0.76 respectively, and reduces the MAE and RMSE on AVEC2014 by 0.26 and 0.74 respectively. This shows that DAS can model the correlation between global spatio-temporal semantics and multi-resolution dynamic differential semantics, and achieve the adaptive fusion of the two.
[0169] Compared with the relevant strategies, the SETD-Net integrating MES, MDM and DAS reduces the MAE and RMSE on AVEC2013 to 6.06 and 7.60 respectively, and reduces the MAE and RMSE on AVEC2014 to 6.03 and 7.59 respectively. Compared with the baseline network, the MAE of the two datasets is reduced by 0.74 and 0.69 respectively, and the RMSE is reduced by 1.30 and 1.34 respectively. This also shows that SETD-Net can overcome the difficulty of dynamic feature extraction caused by the concealment of facial behavior features and the dispersion of key motion time series, thus improving the depression detection performance.
[0170] Table 4. Detection performance of SETD-Net under different strategies
[0171]
[0172] 5.3.3 Performance comparison of different configurations of MES
[0173] To illustrate the influence of the division scale l in the MES structure on the detection performance, the present invention conducts comparative experiments at different division scales, as Figure 6 shown Figure 6 (a) is a schematic diagram of the performance comparison of different parameter configurations in MES on the AVEC2013 dataset. Figure 6(b) Schematic diagram of performance comparison of different parameter configurations in MES on the AVEC2014 dataset. The experimental results show that as the division scale l increases, the detection error of SETD-Net gradually decreases. When the input features are equally divided into 4 blocks, the performance of the network reaches the best. Specifically, the MAE and RMSE on AVEC2013 are reduced to 6.06 and 7.60 respectively; the MAE and RMSE on AVEC2014 are reduced to 6.03 and 7.59 respectively. However, when the division scale l further increases, the detection error of SETD-Net shows an upward trend. This is mainly because: when the number of divisions is too large, the structural complexity increases, resulting in the network relying too much on local features, thus weakening the ability to model global spatio-temporal information. The division scale l in the multi-scale enhancement structure of the present invention is set to 4.
[0174] 5.3.4 Performance Comparison of Different Configurations of MDM
[0175] To evaluate the influence of the time resolution M and different orders in the MDM mechanism on the detection performance, the present invention conducted comparative experiments at different time resolutions and orders, and the results are shown in Table 5.
[0176] From the perspective of the number of different resolutions, as M increases, the detection performance of the SETD-Net network gradually improves. When the total number of resolutions M = 6, the performance of the network reaches the best. Specifically, the MAE and RMSE on the AVEC2013 dataset are 6.06 and 7.60 respectively; the MAE and RMSE on the AVEC2014 dataset are 6.03 and 7.59 respectively. However, when the total number of resolutions M further increases, the detection error of the SETD-Net network gradually increases. The main reason is that although increasing the time resolution can capture richer spatio-temporal dynamic features, the differential operation on too large a time window will lead to information redundancy and interference, thus having a negative impact on the depression detection performance. The total number of resolutions M in the multi-resolution difference mechanism of the present invention is set to 6.
[0177] From the perspective of different orders, the method including the second-order difference operation generally performs better than the method only using the first-order difference operation. This shows that: the second-order difference can better capture the acceleration of facial behavior changes and reflect the intensity and trend of the patient's emotional fluctuations, thus being conducive to improving the representation ability of the MDM mechanism for dynamic difference semantics.
[0178] Table 5. Performance Comparison of Different Parameter Configurations in MDM
[0179]
[0180] 5.3.5 Fusion Performance Analysis of DAS
[0181] To illustrate the feature fusion performance of the DAS strategy, the present invention conducted comparative experiments on DAS and several common feature fusion strategies, such as Figure 7 as shown Figure 7 (a) is a schematic diagram of the performance comparison of feature fusion methods on the AVEC2013 dataset, Figure 7 (b) is a schematic diagram of the performance comparison of feature fusion methods on the AVEC2014 dataset. Among them, the cascade strategy means directly splicing the global spatio-temporal semantics and the dynamic differential semantics in the channel dimension; the element-wise addition strategy means performing element-wise summation on the global spatio-temporal semantics and the dynamic differential semantics; the weighted summation strategy means performing element-wise weighted summation after manually assigning fixed weights to the global spatio-temporal semantics and the dynamic differential semantics.
[0182] Compared with the cascade strategy that is prone to feature redundancy, the DAS strategy uses the attention mechanism to adaptively allocate feature weights and suppresses the information redundancy caused by the expansion of the channel dimension. The MAE and RMSE on the AVEC2013 dataset are reduced by 0.24 and 0.76 respectively; the MAE and RMSE on the AVEC2014 dataset are reduced by 0.26 and 0.74 respectively.
[0183] Compared with the element-wise addition strategy that lacks feature discrimination ability, the DAS strategy can accurately identify the importance of different features through the attention mechanism and effectively reduces the risk of information loss. The MAE and RMSE on the AVEC2013 dataset are reduced by 0.34 and 0.80 respectively; the MAE and RMSE on the AVEC2014 dataset are reduced by 0.28 and 0.80 respectively.
[0184] Compared with the weighted summation strategy limited by the fixed weight allocation method, the DAS strategy dynamically adjusts the weights based on the attention mechanism and fully utilizes the complementary information between the global spatio-temporal semantics and the dynamic differential semantics. The MAE and RMSE on the AVEC2013 dataset are reduced by 0.25 and 0.78 respectively; the MAE and RMSE on the AVEC2014 dataset are reduced by 0.23 and 0.77 respectively.
[0185] 5.3.6 Influence of Hyperparameters on Network Performance
[0186] The present invention further discusses the influence of the balance factor λ in the optimization loss function and the stacking layer number N of 3D-IR in the network on the detection performance of the SETD-Net network. The experimental results are respectively as Figure 8 and Figure 9 shown.
[0187] Referring to Figure 8 , Figure 8 (a) is a schematic diagram of the influence of the balance factor on the network performance on the AVEC2013 dataset, Figure 8(b) Schematic diagram of the influence of the balance factor on the network performance on the AVEC2014 dataset. When λ = 0.0, the SETD-Net network jointly optimizes the single-branch based on the multi-scale spatio-temporal enhancement subnet and the double-branch of the collaborative multi-scale spatio-temporal enhancement subnet and the multi-resolution time difference subnet; when λ = 1.0, the SETD-Net network jointly optimizes the single-branch based on the multi-resolution time difference subnet and the double-branch of the collaborative multi-scale spatio-temporal enhancement subnet and the multi-resolution time difference subnet. Due to the lack of complementary information between the multi-scale spatio-temporal enhancement subnet and the multi-resolution time difference subnet, the detection effects of these two optimization strategies are not good. As λ increases, the loss function in Equation (23) gradually increases the importance of the multi-scale spatio-temporal enhancement subnet, and the detection error of the SETD-Net network shows a downward trend. When λ = 0.7, the SETD-Net obtains the optimal weight and achieves the best detection performance on the AVEC2013 and AVEC2014 datasets.
[0188] See Figure 9 , Figure 9 (a) Schematic diagram of the influence of the number of 3D-IR stacking layers on the network performance on the AVEC2013 dataset. Figure 9 (b) Schematic diagram of the influence of the number of 3D-IR stacking layers on the network performance on the AVEC2014 dataset. When N = 2, the feature extraction ability of the shallow stacking module is limited, resulting in the network being difficult to capture the deep semantics related to depression. As the number of 3D-IR stacking layers increases, its feature extraction ability gradually enhances, and the network detection error shows a downward trend. When N = 10, the SETD-Net achieves the best detection effect on the AVEC2013 and AVEC2014 datasets. As N continues to increase, deep stacking will cause network over-parameterization, resulting in an upward trend in the detection error.
[0189] 5.3.7 Visualization
[0190] To visually display the detection effect of the SETD-Net network, the present invention uses the Gradient-weighted Class Activation Mapping (Grad-CAM) method to visualize the facial regions concerned by the baseline network, the multi-scale spatio-temporal enhancement subnet, and the multi-resolution time difference subnet, such as Figure 10 . In Figure 10Among them, Subnet 1 and Subnet 2 represent the multi-scale spatio-temporal enhancement sub-network and the multi-resolution temporal difference sub-network respectively. Compared with the baseline network, the two sub-networks pay more attention to the facial regions closely related to depressive symptoms. Specifically, the baseline network fails to fully detect the information at key positions (such as eyes and mouth). In contrast, Subnet 1 and Subnet 2 can capture subtle facial features from a global perspective and focus on the movement changes in areas such as the corners of the mouth and eyes. This also shows that the proposed SETD-Net network adopts a multi-scale enhancement structure to capture hidden facial behaviors and uses a multi-resolution difference mechanism to extract scattered motion information, which can effectively improve the performance of automatic depression detection.
[0191] 5.3.8 Computational Performance Analysis of SETD-Net Network
[0192] To provide a comprehensive analysis, the present invention separately counts the parameters, FLOPs, and computational overheads of the baseline network, spatio-temporal enhancement sub-network, temporal difference sub-network, adaptive fusion sub-network, and SETD-Net network on two datasets, as shown in Table 6 (“-” indicates that the module is not trained or tested separately). Compared with the baseline network, the SETD-Net network embedded with the 3D-IR module significantly reduces the number of parameters and FLOPs. In terms of test performance, the inference duration per epoch (processing more than 4000 video segments) is less than 4 minutes, and the processing duration of a single video frame is less than 3.8 milliseconds, which indicates that the SETD-Net network can meet the real-time performance requirements in general scenarios.
[0193] Table 6. Computational Performance of SETD-Net Network
[0194]
[0195] In summary, to overcome the problem of difficult semantic extraction caused by the concealment of facial behavior features and the dispersion of key motion time series, the present invention proposes a depression detection network based on multi-scale spatio-temporal enhancement and multi-resolution time difference. The proposed network mainly includes a multi-scale spatio-temporal enhancement subnet, a multi-resolution time difference subnet, and an adaptive fusion subnet. In the multi-scale spatio-temporal enhancement subnet, a multi-scale enhancement structure is constructed to capture fine-grained global spatio-temporal semantics, thereby mining concealed facial behaviors; in the multi-resolution time difference subnet, a multi-resolution difference mechanism is proposed to capture multi-resolution dynamic difference semantics, and then attention is paid to the dispersed motion information; in the adaptive fusion subnet, a two-stage attention strategy is designed to model the correlation between global spatio-temporal semantics and dynamic difference semantics, thereby realizing the adaptive fusion of the two types of semantics. In addition, to reduce network parameters to adapt to the small sample characteristics of the depression dataset, a plug-and-play 3D inverted residual module is constructed to replace 3D-CNN and embedded in the network. The depression detection effect of the proposed network is analyzed on the datasets AVEC2013 and AVEC2014. The effectiveness of the MES structure, MDM mechanism, and DAS strategy is illustrated through ablation experiments and visualization, and the influence of parameters such as the balance factor and the number of stacked layers of 3D-IR in the network on the network detection performance is discussed. The experimental results show that the MAE and RMSE of the SETD-Net network on the AVEC2013 dataset are 6.06 and 7.60 respectively, which are 0.74 and 1.30 lower than those of the baseline network; on the AVEC2014 dataset, the MAE and RMSE are 6.03 and 7.59 respectively, which are 0.69 and 1.34 lower than those of the baseline network.
[0196] Embodiment 2
[0197] Based on Embodiment 1, Embodiment 2 of the present invention further provides an automatic depression detection system, including:
[0198] A first feature extraction module, configured to construct a multi-scale spatio-temporal enhancement subnet for global spatio-temporal semantic feature extraction;
[0199] A second feature extraction module, configured to construct a multi-resolution time difference subnet for dynamic difference semantic feature extraction;
[0200] A result output module, configured to fuse the global spatio-temporal semantic features and the dynamic difference semantic features to obtain adaptive fusion features, and input them into a fully connected layer to implement depression detection.
[0201] Specifically, before the first feature extraction module, a 3D inverted residual module is further included, and the construction process of the 3D inverted residual module is as follows:
[0202] Perform a pointwise convolution operation on the input feature map and output a channel-expanded feature map through a ReLU activation function;
[0203] Perform a depth convolution operation on the channel-expanded feature map and pass it through another ReLU activation function to obtain a depth convolution feature map;
[0204] Perform a pointwise convolution operation on the depth convolution feature map and pass it through a linear activation function to obtain a channel-compressed feature map;
[0205] Add the channel-compressed feature map to the input feature map to obtain an output feature map.
[0206] More specifically, the construction process of the multi-scale spatio-temporal enhancement subnet is as follows:
[0207] S11. The input feature tensor X is equally divided into l feature sub-blocks along the time dimension
[0208] S12. Input the feature sub-blocks into the 3D inverted residual module to perform layer-by-layer convolution operations along the forward and reverse directions respectively to obtain spatio-temporal multi-scale features and i = 1, 2,..., l, 3D-IR is the abbreviation of the 3D inverted residual module;
[0209] S13. Concatenate the spatio-temporal multi-scale features and along the time dimension, and perform a residual connection with the input feature tensor X to obtain multi-scale enhanced semantic features;
[0210] S14. Input the multi-scale enhanced semantic features into N layers of 3D-IR, and then perform global average pooling to obtain global spatio-temporal semantic features.
[0211] More specifically, S12 includes:
[0212]
[0213] Among them, 3D-IR(·) represents a single-layer 3D inverted residual module.
[0214] More specifically, the process of constructing the multi-resolution time difference subnet is as follows:
[0215] S21. The input feature tensor X is divided into T frame feature maps and perform a first-order difference operation with a step size of m to obtain a first-order difference map with a step size of m, m ∈ [1, M], M represents the resolution, that is, the total number of different step sizes;
[0216] S22. Perform a second-order difference operation on each frame of the first-order difference map with a step size of m to obtain a second-order difference map with a step size of m;
[0217] S23. Concatenate the first-order difference map and the second-order difference map along the channel dimension, and use single-layer pointwise convolution to fuse the channel features, thereby obtaining the difference map E with a step size of m. m ;
[0218] S24. Input the difference map E m into the N-layer 3D-IR, and then perform global average pooling to obtain the dynamic difference semantic features with a step size of m.
[0219] More specifically, S21 includes:
[0220] Split the input feature tensor X into T-frame feature maps and perform a first-order difference operation on :
[0221]
[0222] where m ∈ [1, M] is the difference step size; represents the result of the first-order difference operation for the t-th frame with a step size of m; U m represents the first-order difference map with a step size of m.
[0223] More specifically, S22 includes:
[0224] Perform a second-order difference operation on the basis of the first-order difference map U m :
[0225]
[0226] where represents the output of the second-order difference operation for the t-th frame with a step size of m; G m represents the second-order difference map with a step size of m.
[0227] More specifically, the result output module is further used for:
[0228] S31. Concatenate the dynamic difference semantic features of each step size along the token dimension to obtain the mapping matrix F, and multiply the mapping matrix F with three different first weight matrices respectively to obtain the query matrix, key matrix, and value matrix of each attention head of the multi-head self-attention mechanism;
[0229] S32. Perform scaled dot-product attention calculation on the query matrix, key matrix, and value matrix of the attention heads of the multi-head self-attention mechanism to obtain the calculation result of each attention head, and perform a multi-head concatenation operation on the calculation results of each attention head to obtain the self-attention difference features;
[0230] S33. Preset three second weight matrices. Multiply the global spatio-temporal semantic features by one second weight matrix to construct the query matrix of each attention head of the multi-head cross-attention mechanism. Multiply the self-attention differential features by the other two second weight matrices respectively to obtain the key matrix and value matrix of each attention head of the multi-head cross-attention mechanism;
[0231] S34. Perform scaled dot-product attention calculation on the query matrix, key matrix, and value matrix of the attention head of the multi-head cross-attention mechanism to obtain the calculation result of this attention head. Perform a multi-head concatenation operation on the calculation results of each attention head to obtain the cross-attention differential features;
[0232] S35. Concatenate the cross-attention differential features and the global spatio-temporal semantic features along the channel dimension, and input them into the fully connected layer to predict the depression level.
[0233] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An automatic depression detection method, characterized in that, Including: S1. Construct a multi-scale spatio-temporal enhancement subnet to extract global spatio-temporal semantic features; S2. Construct a multi-resolution temporal difference subnet to extract dynamic difference semantic features; S3. Fuse the global spatio-temporal semantic features and the dynamic difference semantic features to obtain an adaptive fusion feature, and input it into a fully connected layer to achieve depression detection.
2. The automatic depression detection method according to claim 1, characterized in that Before S1, it also includes constructing a 3D inverted residual module, and the construction process of the 3D inverted residual module is as follows: Perform a pointwise convolution operation on the input feature map and output a channel-expanded feature map through a ReLU activation function; Perform a depthwise convolution operation on the channel-expanded feature map and obtain a depthwise convolution feature map through another ReLU activation function; Perform a pointwise convolution operation on the depthwise convolution feature map and obtain a channel-compressed feature map through a linear activation function; Add the channel-compressed feature map to the input feature map to obtain an output feature map.
3. An automatic depression detection method according to claim 2, characterized in that, The construction process of the multi-scale spatio-temporal enhancement subnet is as follows: S11. The input feature tensor X is equally divided into l feature sub - blocks along the time dimension S12. Input the feature sub-block into the 3D inverted residual module to perform layer-by-layer convolutional operations along the forward and reverse directions respectively to obtain spatio-temporal multi-scale features and 3D-IR is the abbreviation of the 3D inverted residual module; S13. Concatenate the spatio-temporal multi-scale features and along the time dimension, and perform a residual connection with the input feature tensor X to obtain multi-scale enhanced semantic features; S14. Input the multi-scale enhanced semantic features into N layers of 3D-IR, and then perform global average pooling to obtain global spatio-temporal semantic features.
4. An automatic depression detection method according to claim 3, characterized in that, S12 includes: Among them, 3D-IR(·) represents a single-layer 3D inverted residual module.
5. The automatic depression detection method according to claim 3, characterized in that, The process of the multi-resolution temporal difference subnet is as follows: S21. The input feature tensor X is divided into T frame feature maps And for Perform a first-order difference operation with a step size of m to obtain a first-order difference map with a step size of m, where m ∈ [1, M], and M represents the resolution, that is, the total number of different step sizes; S22. Perform a second-order difference operation on each frame of the first-order difference map with a step size of m to obtain a second-order difference map with a step size of m; S23. Concatenate the first-order difference map and the second-order difference map along the channel dimension, and use single-layer pointwise convolution to fuse the channel features, thereby obtaining the difference map E with a stride of m m ; S24. Input the difference graph E m into the N-layer 3D-IR, and then perform global average pooling to obtain the dynamic difference semantic feature with a stride of m.
6. The automatic depression detection method according to claim 5, wherein S21 includes: Divide the input feature tensor X into T-frame feature maps And perform First-order difference operation on: Among them, m ∈ [1, M] is the difference step size; represents the first-order difference operation result of the t-th frame with step size m; U m represents the first-order difference graph with step size m.
7. An automatic depression detection method according to claim 6, characterized in that S22 includes: On the basis of the first-order difference diagram U m perform a second-order difference operation: Among them, represents the output of the t-th frame of the second-order difference operation with a step size of m; G m represents the second-order difference graph with a step size of m.
8. An automatic depression detection method according to claim 5, characterized in that S3 Including: S31. Concatenate the dynamic difference semantic features of each step along the token dimension to obtain a mapping matrix F, and multiply the mapping matrix F by three different first weight matrices respectively to obtain the query matrix, key matrix, and value matrix of each attention head of the multi-head self-attention mechanism; S32. Perform scaled dot-product attention calculation on the query matrix, key matrix, and value matrix of the attention head of the multi-head self-attention mechanism to obtain the calculation result of this attention head, and perform a multi-head concatenation operation on the calculation results of each attention head to obtain a self-attention difference feature; S33. Preset three second weight matrices, multiply the global spatio-temporal semantic features by a second weight matrix to construct the query matrix of each attention head of the multi-head cross-attention mechanism, and multiply the self-attention difference feature by the other two second weight matrices respectively to obtain the key matrix and value matrix of each attention head of the multi-head cross-attention mechanism; S34. Perform scaled dot-product attention calculation on the query matrix, key matrix, and value matrix of the attention head of the multi-head cross-attention mechanism to obtain the calculation result of this attention head, and perform a multi-head concatenation operation on the calculation results of each attention head to obtain a cross-attention difference feature; S35. Concatenate the cross-attention difference feature and the global spatio-temporal semantic feature along the channel dimension, and input it into a fully connected layer to predict the depression level.
9. An automatic depression detection system, characterized in that, Including: The first feature extraction module is used to construct a multi-scale spatio-temporal enhancement subnet to extract global spatio-temporal semantic features; The second feature extraction module is used to construct a multi-resolution temporal difference subnet to extract dynamic difference semantic features; The result output module is used to fuse the global spatio-temporal semantic features and the dynamic difference semantic features to obtain an adaptive fusion feature, and input it into a fully connected layer to achieve depression detection.
10. An automatic depression detection system according to claim 9, characterized in that, Before the first feature extraction module, there is also a 3D inverted residual module, and the construction process of the 3D inverted residual module is as follows: Perform a pointwise convolution operation on the input feature map and output a channel-expanded feature map through a ReLU activation function; Perform a depthwise convolution operation on the channel-expanded feature map and obtain a depthwise convolution feature map through another ReLU activation function; Perform a pointwise convolution operation on the depthwise convolution feature map and obtain a channel-compressed feature map through a linear activation function; Add the channel-compressed feature map to the input feature map to obtain the output feature map.
Citation Information
Patent Citations
Depression assessment method and device based on audio-visual multi-modal data fusion
CN118173267A