Depression Intensity Recognition Method Based on Spatiotemporal Feature Integration and Global-Local Feature Fusion

Through the integration of space-time features and global-local feature fusion methods, the problem of difficulty in extracting global and local information related to depression intensity in the prior art is solved, the accuracy of depression intensity recognition is improved, and the semantic consistency description between global and local features and the effective extraction of dynamic information on different time scales is achieved.

CN118506421BActive Publication Date: 2025-06-24XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410609548.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-06-24
Estimated Expiration
2044-05-16

AI Technical Summary

Technical Problem

In the recognition of depression intensity, it is difficult to effectively use the depth model to learn semantic consistency between global context information and local information from facial expressions, and time domain information modeling is difficult to extract dynamic information from different time scales, affecting the accuracy of recognition of depression intensity.

Method used

Using the methods of spatiotemporal feature integration and global-local feature fusion, the global and local features of face images are extracted and fused through the LSTIA-PLEGDF global feature learning branch model, local feature learning branch model and global-local semantic correlation feature fusion module, and combined with short- and long-term information aggregation modules, the model's spatiotemporal perception ability is enhanced.

Benefits of technology

It improves the accuracy of recognition of depression intensity in facial images, effectively utilizes semantic consistency information of global and local features, and improves the model's ability to extract dynamic information on different time scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118506421B_ABST
    Figure CN118506421B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features, and the steps are as follows: 1. Construct a global feature learning branch model, and input the global face image sequence of patients with depression into the model to obtain global face features; 2. Construct a local feature learning branch model, and input the image sequence of the eye region of patients with depression into the model to obtain local features; 3. Construct a semantic correlation feature fusion module for global face features and local features, calculate the correlation weight of global face features and local features, and fuse the semantic consistency information of global features and local features to obtain fused features; 4. Input the fused features into a fully connected layer, and then connect them to a fully connected layer with only one neuron to output the depression intensity score, and identify the depression intensity through the depression intensity score. The method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features in the present invention is reasonable and effective, and can effectively improve the accuracy of identifying depression intensity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of depression intensity recognition, and specifically relates to a depression intensity recognition method integrating spatio-temporal features and fusing global-local features. Background Art

[0002] Depressed patients will show melancholy and sad expressions on their faces, and behaviors such as reduced eye contact and less facial movement changes will occur. These manifestations contain rich depression information. Therefore, diagnosing depression through the facial expressions of patients has become one of the current research hotspots.

[0003] In the field of depression intensity recognition, due to the characteristics of facial videos such as easy acquisition and non-contact, and the fact that facial expressions are closely related to depression, visual facial features can be used as an important basis for judging depression intensity. Although extracting visual facial features is very valuable for depression recognition, how to effectively encode the spatio-temporal information of such visual facial features is still extremely challenging. In the existing spatial domain information modeling methods, the fact that the depression state is usually manifested in multiple local regions is ignored during the extraction of overall facial depression information, and the model cannot well model the relationship between the depression information in different local regions. Secondly, in the methods of combining global and local information, separate network branches are usually set to separately learn the features of local regions rich in depression information such as eyes and mouths and the global face, and then directly splice the semantic information of the global and local regions, and then perform depression intensity recognition, ignoring the relevance between the global and local semantic information. In the time domain information modeling methods, the dynamic feature encoding method can only learn the dynamic information of a fixed time length, reducing the ability of the model to extract dynamic information of different time scales and ignoring the information relevance between frames within a fixed time range. Using recurrent neural networks and their variants to integrate dynamic features of different time lengths ignores that the importance of time features is different in different periods. As can be seen from the above, in this field, there are mainly two challenges for spatial domain information modeling: (1) how to effectively learn the global context information related to depression intensity from facial expressions using a deep model; (2) how to make full use of the semantic consistency information between the global and local information related to depression intensity to improve the performance of the depression intensity recognition model. There are mainly two challenges for time domain information modeling: (1) how to effectively model the spatio-temporal relationship between frames of short-term depression features; (2) how to effectively model the time context relationship of long-term depression features. The above challenges directly affect the performance of the current depression intensity recognition model, making it difficult for the model to have good recognition accuracy in the depression intensity recognition task based on video face images.

[0004] To this end, the present application proposes a method for identifying depression intensity, namely Long-Short-Term Information Aggregation, Enhancement of Local Perception of Global Depression Features, and Fusion of Global-Local Semantic Relevance Features (LSTIA-PLEGDF-FGLSCF), which is used to improve the accuracy of identifying the depression intensity of face images. Summary of the Invention

[0005] The object of the present invention is to provide a method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features, which can effectively improve the accuracy of identifying the depression intensity of face images.

[0006] The technical solution adopted by the present invention is a method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features, and the specific steps are as follows:

[0007] Step 1: Construct an LSTIA-PLEGDF global feature learning branch model, and input the global image sequence of the face of a depression patient into this model to obtain the global face features;

[0008] Step 2: Use a three-dimensional residual convolutional network to construct a local feature learning branch model, and input the image sequence of the eye region of a depression patient into this model to obtain local features;

[0009] Step 3: Construct a semantic relevance feature fusion module for the global and local face features, calculate the correlation weight of the global and local face features, fuse the semantic consistency information of the global and local features, and obtain a fused feature with consistent global and local semantics;

[0010] Step 4: Input the fused feature into a fully connected layer, and then connect it to a fully connected layer containing only one neuron to output the depression intensity score, and identify the depression intensity through the depression intensity score.

[0011] The characteristics of the present invention also lie in:

[0012] The LSTIA-PLEGDF global feature learning branch model in Step 1 is composed of a short-term information aggregation module, a local perception enhancement module for global depression features, and a long-term information aggregation module. In the LSTIA-PLEGDF global feature learning branch model, the short-term information aggregation module, the local perception enhancement module for global depression features, and the long-term information aggregation module are stacked, and the spatio-temporal information of the global features is effectively extracted in a hierarchical stacking manner.

[0013] The specific construction process of the LSTIA-PLEGDF global feature learning branch model is as follows:

[0014] Step 1.1: Construct a short-term information aggregation module to obtain the initial global features of the global face image sequence;

[0015] Step 1.2: Use the initial global feature as the input feature, construct a global depression feature local perception enhancement module, and obtain the high-level face global feature;

[0016] Step 1.3: Use the high-level face global feature as the input, construct a long-term information aggregation module, and obtain the global feature of the face global image sequence.

[0017] The construction process of the short-term information aggregation module in Step 1.1 is as follows:

[0018] Step 1.1.1: Obtain the face global image sequence X input , and use a sliding window of a fixed size in the time direction to obtain consecutive short-term features {X1, X2,..., X i} of the face global image sequence. Along the channel dimension, each short-term feature X i is divided into four segments respectively to obtain four channel sub-segments {X i1 , X i2 , X i3 , X i4};

[0019] Among them, X input is the face global image sequence, C represents the number of image channels, T1 represents the length of the image sequence, H represents the image height, and W represents the image width; {X1, X2,..., X i} represents consecutive short-term features; {X i1 , X i2 , X i3 , X i4} represents the channel sub-segments after the channel dimension division of X i ;

[0020] Step 1.1.2: For the first two segments {X i , X i1} of the four channel sub-segments of each short-term feature X i2 , use a temporal convolutional layer for feature learning respectively. For the last two segments {X i , X i3} of the four channel sub-segments of each short-term feature X i4 , use a spatial convolutional layer for feature learning respectively. After the feature learning of the last three segments {X i2 , X i3 , X i4}, add them to the output of the previous segment in turn, and then perform feature learning through a spatio-temporal convolutional layer to obtain the output feature The specific calculation process is as follows:

[0021]

[0022]

[0023]

[0024]

[0025] Among them, Conv 3×1×1 (·), Conv 3×3×3 (·), Conv 1×3×3 (·) respectively represent the 3D convolution operation for feature learning, + represents the superposition of feature elements, and the output feature represents the time information and spatial information;

[0026] Step 1.1.3: Aggregate the output feature to obtain the aggregated feature Then calculate and corresponding time attention weight w T and spatial attention weight w S , and the calculation process is as follows:

[0027]

[0028]

[0029]

[0030]

[0031] Among them, Concate C (·) represents concatenation in the channel dimension, SP(·) and TP(·) respectively represent max pooling in the spatial dimension and max pooling in the time dimension, Sigmoid(·) is the activation function, FC * (·) represents the fully connected layer, represents the output feature after feature learning of the channel sub - segment {X i1 , X i2 , X i3 , X i4}, represents the feature after channel concatenation, represents the feature after channel concatenation, w T represents the time attention weight of S represents the spatial attention weight of;

[0032] Step 1.1.4: Through wT and w S enhance respectively the expressiveness of time information and the expressiveness of spatial information, and aggregate spatio-temporal features. The calculation process is as follows:

[0033]

[0034]

[0035]

[0036] wherein, × represents element-wise multiplication, - represents element-wise subtraction, + represents element-wise addition, and Concate C (·) represents concatenation in the channel dimension, represents the feature after channel concatenation, represents the feature after channel concatenation, w T represents the temporal attention weight of S represents the spatial attention weight of represents the feature after weighting the features in the temporal dimension and the spatial dimension for ; represents the feature after weighting the features in the temporal dimension and the spatial dimension for ; represents and the short-term feature after channel concatenation;

[0037] Step 1.1.5, repeat Step 1.1.2 to Step 1.1.4 to obtain all short-term features X i the short-term feature after channel concatenation to obtain the complete initial global feature X G1 :

[0038]

[0039] wherein, Concate T (·) represents concatenation in the temporal dimension, represents the short-term feature after channel concatenation of each temporal sub-segment X i ; X G1 represents the initial global feature finally output by the short-term information aggregation module.

[0040] The construction process of the global depression feature local perception enhancement module in Step 1.2 is as follows:

[0041] Step 1.2.1, take the initial global feature XG1 The space dimension is evenly divided into 4 non - overlapping regions X L ={X local1 ,...,X localj}, j = {1, 2, 3, 4}, where X local1 and X local2 contain information about the eyes, and X local3 and X local4 contain information about the mouth; keeping the time and channel dimensions unchanged, the local feature X L of each local region X localj is transformed into a local feature vector in the space dimension The calculation process is as follows:

[0042]

[0043] where ReLU(·) represents the activation function, Reshape(·) represents compressing the feature from two - dimensional to one - dimensional in the space dimension, Conv 1×1×1 (·) is a 3 - D convolution operation, T2 represents the sequence length of the feature, C1 represents the number of channels of the feature, H1 represents the height of the feature, W1 represents the width of the feature, N represents the product of the height and width of the feature, and X localj represents any local feature in region X L , represents the local feature vector after spatial compression of local feature X localj ;

[0044] Step 1.2.2, calculate the correlation coefficient matrix r locajl between the vectors of each local feature X n,m , so as to obtain the correlation coefficient R localj between the vectors of each local feature X k , and obtain the set R of local feature correlation coefficient vectors. The specific calculation process is as follows:

[0045]

[0046]

[0047] R={R1, R2,..., R k} (18)

[0048] Wherein, T represents transpose, × represents matrix multiplication, + represents element-wise addition, T2 represents the sequence length of the features, N represents the product of the height and width of the features, Sigmoid(·) represents the activation function, x = {1, 2, 3, 4}, and x ≠ j, n = {1, 2,..., N}, m = {1, 2,..., N}, r n,m represents the local feature vector between the correlation coefficient matrix and represent different local feature vectors, r n,* reflects the local region for the local region the importance of, r *,m contains the local region information of and the local region the degree of correlation of, R k represents the local feature vector the k-th element in and the local feature vector the correlation coefficient of, R = {R1, R2,..., R k}, k = {1, 2,..., N} represents the set of local feature correlation coefficient vectors, and N represents the product of the height and width of the features;

[0049] Step 1.2.3, by arranging the local feature correlation coefficient vector R according to the position information in the spatial dimension, the region X L the correlation coefficient matrix between each local region in and the other 3 local regions The calculation process is as follows:

[0050]

[0051] Wherein, represents arranging the local feature correlation coefficient vector R according to the position information in the spatial dimension, T2 represents the sequence length of the features, C1 represents the number of channels of the features, H1 represents the height of the features, W1 represents the width of the features, R is the set of local feature correlation coefficient vectors, R = {R1, R2,..., R k}, represents the correlation coefficient matrix of the local features transformed from the local feature correlation coefficient vector R;

[0052] Step 1.2.4, through the correlation coefficient matrix calculate the locally relevant perception of the local feature X outj concatenate to obtain the final output feature X G2 , and the calculation process is as follows:

[0053]

[0054]

[0055] Among them, × represents matrix multiplication, + represents element superposition, T2 represents the sequence length of features, C1 represents the number of channels of features, H1 represents the height of features, W1 represents the width of features, C represents the number of image channels, T1 represents the image sequence length, H represents the image height, W represents the image width, Concat(·) represents splicing features along the spatial dimension, and X localj represents the region X L ={X local1 ,..., X localj} is any local feature among them, The correlation coefficient matrix of each local region in the region X L with the other three local regions, and X outj represents the local feature with local correlation perception, and X G2 represents the output feature of the global depression feature local perception enhancement module;

[0056] Step 1.2.5, input the output feature X G2 of the global depression feature local perception enhancement module into the short-term information aggregation module with the number of channels doubled, the global depression feature local perception enhancement module with the number of channels doubled, and a three-dimensional residual convolution module in sequence for global feature extraction to obtain the high-level face global feature X G3 .

[0057] The construction process of the long-term information aggregation module in Step 1.3 is as follows:

[0058] Step 1.3.1, perform interval sampling on the high-level face global feature X G3 ={X1', X2', X3',..., X ti '} in the time dimension to obtain the sampled time sub-features X t1 ' and X t2 ', specifically as follows:

[0059] X t1 '={X1', X3', X5',..., X ti-1 '} (22)

[0060] X t2 '={X2', X4', X6',..., X ti '} (23)

[0061] Among them, X t1 ' and X t2 ' represent the time sub-features obtained by sampling the global feature X G3 in the time dimension;

[0062] Step 1.3.2: Perform the first time-domain information interaction on the time sub-features X t1 ' and X t2 ', and calculate the time attention weight W1 after rough fusion. The specific calculation process is as follows:

[0063] X ft1 = Conv1×1×1(X t1 '+ X t2 ) (24)

[0064] W f1 = Sigmoid(Concat T (GAP(X ft1 ), GMP(X ft1 ))) (25)

[0065] W1 = Conv3×1×1(W f1 ) (26)

[0066] Among them, Conv1×1×1(·) and Conv3×1×1(·) are 3D convolution operations, + is element-wise addition, GAP(·) is global average pooling in the spatial dimension, which is used to capture the gently changing facial information, GMP(·) is global max pooling, which is used to capture the violently changing facial information, Sigmoid(·) is an activation function, Concat T (·) is feature concatenation in the time dimension, which expands the ability of the weight to perceive different facial behaviors, X t1 ' and X t2 ' represent the time sub-features obtained by downsampling the global feature X G3 in the time dimension, X ft1 represents the coarsely fused feature of X t1 ' and X t2 ', W1 represents the time attention weight after rough fusion of X t1 ' and X t2 ', W f1 represents the fused weight of X t1 ' and X t2 ' before dimensionality reduction in the time dimension of W1;

[0067] Step 1.3.3: Update the time information of the time sub-features X t1 ' and X t2 ', and generate a time relationship perception vector M during this process. The calculation process is as follows:

[0068] X t1 ” = W1 × X t1 ' (27)

[0069] X t2” = (1 - W1) × X t2 ' (28)

[0070] X′ ft1 = GMP(X t1 ” + X t2 ) (29)

[0071] M = Conv3×1×1(W f1 × X G3 ) (30)

[0072] where × is element-wise multiplication, - is element-wise subtraction, + is element-wise addition, GMP(·) is global max pooling, Conv3×1×1(·) is a 3D convolution operation for dimensionality reduction in the time dimension, X t1 ' and X t2 ' represent the temporal sub-features obtained by downsampling the global feature X G3 in the time dimension, X t1 ” represents the temporal sub-feature of X t1 ' weighted by the temporal attention weight W1, X t2 ” represents the temporal sub-feature of X t2 ' weighted by the temporal attention weight (1 - W1), X′ ft1 represents the coarse-grained fusion feature of X t1 ” and X t2 '; M represents the temporal relationship perception vector;

[0073] Step 1.3.4: Further enhance the module's ability to perceive long-term information through the temporal relationship perception vector M, the temporal sub-features X t1 ” and X t2 ”. The calculation process is as follows:

[0074] W2 = Sigmoid(GMP(X t1 ” + X t2 ”)+M) (31)

[0075] X t1 ”' = W2 × X t1 ” (32)

[0076] X t2 ”' = (1 - W2) × X t2 ” (33)

[0077] X ft2 = Conv1×1×1(Concat T (X t1 ”', X t2 ”')) (34)

[0078] Among them, × represents element-wise multiplication, - represents element-wise subtraction, + represents element-wise addition, GMP(·) is global max pooling, Conv1×1×1(·) is a 3D convolution operation, Sigmoid(·) is an activation function, and Concat T (·) is the temporal dimension feature concatenation, X t1 ” represents the temporal sub-feature of X t1 ' after being weighted by the temporal attention weight W1, X t2 ” represents the temporal sub-feature of X t2 ' after being weighted by the temporal attention weight (1 - W1), M represents the temporal relationship perception vector, and W2 represents the X t1 ” and X t2 ” after rough fusion of the temporal attention weight, X t1 ”' represents the temporal sub-feature of X t1 ” after being weighted by the temporal attention weight W2, X t2 ”' represents the temporal sub-feature of X t2 ” after being weighted by the temporal attention weight (1 - W2), X ft2 represents the X t1 ”' and X t2 ”' coarse-grained fusion feature;

[0079] Step 1.3.5: Utilize the coarse-grained fusion feature of the temporal sub-features X t1 ”' and X t2 ”' to enhance the temporal perception of the input feature and obtain the output global feature X G4 of the long-term information aggregation module. The specific calculation process is as follows:

[0080] X G4 = X ft2 × X G3 + X G3 (35)

[0081] Among them, × is element-wise multiplication, + is element-wise addition, and X ft2 represents the X t1 ”' and X t2 ”' coarse-grained fusion feature, X G3 represents the global feature of the input long-term information aggregation module, and represents the output global feature of the long-term information aggregation module;

[0082] Step 1.3.6: Use the output global feature X G4 of the long-term information aggregation module as the input, and then successively pass through a global feature local perception enhancement module with the number of channels doubled again, a 3D convolutional residual block with the number of channels doubled, and a long-term feature information aggregation module with the number of channels doubled for learning to obtain the global feature X global of the face global image sequence.

[0083] The specific process of step 2 is as follows: Input the eye region image sequence of the depression patients into a 50-layer 3D residual convolutional network for feature learning to obtain the local region feature X local '.

[0084] The construction process of the global and local semantic correlation feature fusion module in step 3 is as follows:

[0085] Step 3.1: Align the global feature X global of the face global image sequence and the local region feature X local ' obtained in step 2 using a linear layer to make the sizes of the global feature X global and the local region feature X local ' the same. The calculation process is:

[0086] F g = linear1(X global ) (35)

[0087] F l = linear2(X local ') (36)

[0088] where linear1(·) and linear2(·) represent linear layers, X global is the global feature of the face global image sequence, X local ' represents the local region feature learned by the 50-layer 3D residual convolutional network in step 2, and F g and F l represent the global feature and the local region feature after linear mapping;

[0089] Step 3.2: Calculate the correlation coefficient w g between the globally linearly mapped feature F l and the locally linearly mapped feature F gl . The calculation process is:

[0090] w gl = Sigmoid(F g × F l ) (37)

[0091] where × is element-wise multiplication, and the magnitude of w gl represents the degree of correlation between the local region feature F l and the global feature F g ;

[0092] Step 3.3: Calculate the global and local region correlation feature F gl through the correlation coefficient w g and the globally linearly mapped feature F gl, the calculation process is as follows:

[0093] F gl = w gl × F g (38)

[0094] Among them, × is element-wise multiplication, and w gl represents the global and local region feature correlation coefficient, and F g represents the global feature after linear mapping, and F gl represents the global and local region correlation feature;

[0095] Step 3.4: Fuse the global feature and local region feature after linear mapping. The specific calculation process is as follows:

[0096] w′ gl = Sigmoid((F g + F gl ) × (F l + F gl )) (39)

[0097] F′ gl = w′ gl × (F g + F gl ) + w′ gl × (F l + F gl ) (40)

[0098] Among them, × is element-wise multiplication, + is element-wise addition, and w′ gl represents the global and local correlation coefficient after fusing the global and local correlation features, and F g and F l represent the global feature and local region feature after linear mapping, and F gl represents the global and local region correlation feature, and F′ gl represents the fused feature with consistent global and local region semantics.

[0099] The specific process of Step 4 is as follows:

[0100] Step 4: Input the fused feature F′ with consistent global and local region semantics obtained in Step 3 gl into the fully connected layer, and then connect it to another fully connected layer with only one neuron to output the depression intensity score, so as to perform depression intensity recognition. The calculation process is expressed as:

[0101] F socre = F′ gl × w + b (41)

[0102] where, × represents element-wise multiplication, + represents element-wise addition, w represents weights, b represents bias, and F′ gl represents the fused feature with consistent semantics in the global and local regions, and F socre represents the recognized depression intensity score.

[0103] The global face image sequence of the depression patients in Step 1 and the eye region image sequence of the depression patients in Step 2 are obtained by cropping and face alignment for each image sample in the video data obtained from the depression video datasets AVEC2013 and AVEC2014.

[0104] The beneficial effects of the present invention are as follows:

[0105] (1) In the present invention, a short-term information aggregation module is designed to enhance the model's ability to extract facial behavior information with different durations within a fixed time range, and learn richer short-term dynamic depression features in consecutive frame images. This module uses a sliding window method to make the module only focus on short-term information, and then enhances the expressiveness of the temporal information and spatial information of some features respectively. The finally integrated channel information contains short-term features with different temporal and spatial information richness.

[0106] (2) In the present invention, a global depression feature local perception enhancement module is designed to extract the semantic correlation information between facial local regions, promote the interaction between information related to depression in different local regions, and thus enhance the expressiveness of the global depression feature driven by local depression features. Through this module, the feature correlation between different local regions is enhanced, and the correlation information between local regions is effectively used to enhance the local perception of global features.

[0107] (3) In the present invention, a long-term information aggregation module is designed to use an improved attention mechanism to perceive the temporal context relationship of input features, and better screen important temporal information in long-term features through perceptual weights. In this process, an improved temporal attention mechanism is used to calculate the temporal context relationship weights in the fused temporal sub-features, and at the same time generate a temporal relationship perception vector to ensure the integrity of the temporal information perceived by attention. Finally, the temporal information of the temporal sub-features and the input features is integrated to enhance the temporal perception ability of the input features;

[0108] (4) In the present invention, a global and local semantic correlation feature fusion module is designed to capture the relevance of global and local semantic information, and realize the semantic consistency description between global and local depression features. Through this module, the correlation coefficient between the global feature and the local feature is calculated, which can reflect the importance of the features in the eye region to the global feature. By weighting the global and local feature weights with this coefficient, the contribution degree of the local feature in the recognition stage can be adaptively adjusted;

[0109] (5) Experiments and analyses are carried out on the AVEC2013 and AVEC2014 datasets. The method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features of the present invention is reasonable and effective, and can effectively improve the accuracy of identifying depression intensity. BRIEF DESCRIPTION OF THE DRAWINGS

[0110] Figure 1 It is a schematic diagram of the LSTIA-PLEGDF-FGLSCF model architecture in the method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features of the present invention;

[0111] Figure 2 It is a schematic diagram of the structure of the short-term information aggregation module in the LSTIA-PLEGDF-FGLSCF model in the method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features of the present invention;

[0112] Figure 3 It is a schematic diagram of the structure of the global depression feature local perception enhancement module in the LSTIA-PLEGDF-FGLSCF model in the method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features of the present invention;

[0113] Figure 4 It is a schematic diagram of the structure of the long-term information aggregation module in the LSTIA-PLEGDF-FGLSCF model in the method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features of the present invention;

[0114] Figure 5 It is a schematic diagram of the structure of the global and local semantic correlation feature fusion module in the LSTIA-PLEGDF-FGLSCF model in the method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0115] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0116] Example 1:

[0117] In the method for identifying depression intensity by integrating spatio-temporal features and fusing global-local features of the present invention, the global face image sequence of depression patients and the eye region image sequence of depression patients are obtained from the depression datasets AVEC2013 and AVEC2014. First, each video in the depression datasets AVEC2013 and AVEC2014 is preprocessed. For the AVEC2013 dataset, one image is taken every 60 frames starting from the first frame. For the videos in the AVEC2014 dataset, the video sampling interval is 6 frames. The multi-task cascaded convolutional network is used to obtain the coordinates of face key points and the coordinates of the bounding box, and the face position is adjusted according to the face key points through affine transformation. Finally, the face images are adjusted to a unified size, and the images of the eye regions are cropped. To ensure the correctness of face cropping, each image is checked one by one to see if the cropping is correct, and the images with incorrect cropping are deleted. For the case where the number of correctly cropped images in some videos does not meet the requirements of the experiment, the sampling interval of the corresponding videos is reduced until the number of correctly cropped images meets the requirements of all experiments. The global face image sequence input to the network can be expressed as where C represents the number of image channels, T1 represents the sequence length, H represents the image height, and W represents the image width.

[0118] The specific implementation steps are as follows:

[0119] Step 1. Construct the LSTIA-PLEGDF global feature learning branch model, and input the global face image sequence of depression patients into this model to obtain the global face features;

[0120] The architecture of the LSTIA-PLEGDF global feature learning branch model is as Figure 1 shown:

[0121] The LSTIA-PLEGDF global feature learning branch model consists of a short-term information aggregation module, a global depression feature local perception enhancement module, and a long-term information aggregation module. In the LSTIA-PLEGDF global feature learning branch model, the short-term information aggregation module, the global depression feature local perception enhancement module, and the long-term information aggregation module are stacked, and the spatio-temporal information of the global features is effectively extracted in a hierarchical stacking manner.

[0122] The structure of the short-term information aggregation module is as Figure 2 shown, and its construction process is as follows:

[0123] Step 1.1.1. Obtain the global face image sequence X input , and use a sliding window with a fixed size in the time direction to obtain continuous short-term features {X1, X2,..., X i} of the global face image sequence. Along the channel dimension, each short-term feature X iDivided into four segments respectively to obtain four channel sub-segments {X i1 , X i2 , X i3 , X i4};

[0124] Among them, X input is the global face image sequence, C represents the number of image channels, T1 represents the sequence length, H represents the image height, and W represents the image width; {X1, X2,..., X i} represents continuous short-term features; {X i1 , X i2 , X i3 , X i4} represents the channel sub-segments after the channel dimension of X i is segmented;

[0125] Step 1.1.2: For the first two segments of the four channel sub-segments of each short-term feature X i {X i1 , X i2}, use the temporal convolutional layer for feature learning respectively. For the last two segments of the four channel sub-segments of each short-term feature X i {X i3 , X i4}, use the spatial convolutional layer for feature learning respectively. After the last three segments {X i2 , X i3 , X i4} are used for feature learning, add them to the output of the previous segment in sequence, and then perform feature learning through the spatio-temporal convolutional layer to obtain the output feature The specific calculation process is as follows:

[0126]

[0127]

[0128]

[0129]

[0130] Among them, Conv 3×1×1 (·), Conv 3×3×3 (·), Conv 1×3×3 (·) respectively represent 3D convolutional operations for feature learning, + represents the superposition of feature elements, and the output feature represents the temporal information and spatial information;

[0131] Step 1.1.3: Aggregate the output feature to obtain the aggregated feature Next, calculate and the corresponding temporal attention weight w T and spatial attention weight w S , and the calculation process is as follows:

[0132]

[0133]

[0134]

[0135]

[0136] Among them, Concate C (·) represents concatenation in the channel dimension, SP(·) and TP(·) respectively represent max pooling in the spatial dimension and max pooling in the temporal dimension, Sigmoid(·) is the activation function, FC * (·) represents the fully connected layer, represents the output feature after feature learning of the channel sub - segments {X i1 , X i2 , X i3 , X i4}, represents the feature after channel concatenation, represents the feature after channel concatenation, w T represents 's temporal attention weight, w S represents 's spatial attention weight;

[0137] Step 1.1.4, enhance the temporal information expressiveness of T and the spatial information expressiveness of S respectively through w and aggregate spatio - temporal features, and the calculation process is as follows: Among them, × represents element - wise multiplication, - represents element - wise subtraction, + represents element - wise addition, Concate

[0138]

[0139]

[0140]

[0141] C (·) represents concatenation in the channel dimension, represents the feature after channel concatenation, denote the feature after channel concatenation, w T denote the temporal attention weight of, w S denote the spatial attention weight of denote the feature after weighting the features in the temporal dimension and the spatial dimension, denote the feature after weighting the features in the temporal dimension and the spatial dimension, denote and the short-term features after channel concatenation;

[0142] Step 1.1.5, repeat Step 1.1.2 to Step 1.1.4 to obtain all short-term features X i the short-term features after channel concatenation obtain the complete initial global feature X G1 :

[0143]

[0144] where, Concate T (·) represents concatenation in the temporal dimension, denote each temporal sub-segment X i the short-term features after channel concatenation, X G1 denote the initial global feature finally output by the short-term information aggregation module.

[0145] The structure of the global depression feature local perception enhancement module is as Figure 3 shown, and its construction process is as follows:

[0146] Step 1.2.1, divide the initial global feature X G1 into 4 non-overlapping regions X L ={X local1 ,...,X localj} on average in the spatial dimension, j = {1, 2, 3, 4}, where, X local1 and X local2 contain information about the eyes, X local3 and X local4 contain information about the mouth; keeping the temporal and channel dimensions unchanged, convert the local feature X L of each local region X localj into a local feature vector in the spatial dimension This calculation process is:

[0147]

[0148] Among them, ReLU(·) represents the activation function, Reshape(·) represents compressing the features from two-dimensional to one-dimensional in the spatial dimension, Conv 1×1×1 (·) is a three-dimensional convolution operation, T2 represents the sequence length of the features, C1 represents the number of channels of the features, H1 represents the height of the features, W1 represents the width of the features, N represents the product of the height and width of the features, X localj represents any local feature in region X L ; represents the local feature X localj local feature vector after spatial compression;

[0149] Step 1.2.2, calculate the correlation coefficient matrix r locajl between the vectors of each local feature X n,m , so as to obtain the correlation coefficient R localj between the vectors of each local feature X k , and obtain the local feature correlation coefficient vector set R. The specific calculation process is as follows:

[0150]

[0151]

[0152] R = {R1, R2,..., R k} (18)

[0153] where T represents transpose, × represents matrix multiplication, + represents element superposition, T2 represents the sequence length of the features, N represents the product of the height and width of the features, Sigmoid(·) represents the activation function, x = {1, 2, 3, 4}, and x ≠ j, n = {1, 2,..., N}, m = {1, 2,..., N}, r n,m represents the correlation coefficient matrix between the local feature vectors , and represent different local feature vectors, r n,* reflects the importance degree of the local region to the local region , r *,m contains the correlation degree of the information of the local region with the local region , R k represents the correlation coefficient of the kth element in the local feature vector with the local feature vector , R = {R1, R2,..., R k}, where \(k = \{1, 2, \ldots, N\}\) represents the set of local feature correlation coefficient vectors;

[0154] Step 1.2.3: Arrange the local feature correlation coefficient vector \(R\) in the spatial dimension according to the position information to obtain region \(X\) L The correlation coefficient matrix between each local region in The calculation process is as follows:

[0155]

[0156] Among them, represents arranging the local feature correlation coefficient vector \(R\) in the spatial dimension according to the position information, \(T2\) represents the sequence length of the feature, \(C\) represents the number of channels of the feature, \(H1\) represents the height of the feature, \(W1\) represents the width of the feature, \(N\) represents the product of the height and width of the feature, \(R\) is the set of local feature correlation coefficient vectors, \(R=\{R1, R2, \ldots, R\) k} represents the correlation coefficient matrix of the local features transformed from the local feature correlation coefficient vector \(R\);

[0157] Step 1.2.4: Calculate the locally correlated perception local feature \(X\) through the correlation coefficient matrix Perform splicing to obtain the final output feature \(X\) outj , and the calculation process is as follows: G2

[0158]

[0159]

[0160] Among them, \(\times\) represents matrix multiplication, \(+\) represents element superposition, \(T2\) represents the sequence length of the feature, \(C1\) represents the number of channels of the feature, \(H1\) represents the height of the feature, \(W1\) represents the width of the feature, \(N\) represents the product of the height and width of the feature, \(Concat(\cdot)\) represents splicing the features in the spatial dimension, \(X\) localj represents region \(X\) L =\{X local1 , \ldots, X localj \} represents any local feature in it, Region \(X\) L The correlation coefficient matrix between each local region and the other 3 local regions in outj represents the local feature with local correlation perception, \(X\) G2 represents the output feature of the global depression feature local perception enhancement module;

[0161] Step 1.2.5: The output feature \(X\) of the global depression feature local perception enhancement module G2 ​Input into a short-term information aggregation module with doubled number of channels, a global depression feature local perception enhancement module with doubled number of channels, and a three-dimensional residual convolution module in sequence for global feature extraction to obtain the high-level face global feature X G3 。

[0162] The structure of the long-term information aggregation module is as Figure 4 shown, and its construction process is as follows:

[0163] Step 1.3.1: Perform interval sampling on the high-level face global feature X G3 ={X1', X2', X3',..., X ti '} in the time dimension to obtain the sampled time sub-features X t1 ' and X t2 '. As a special form of the sequence, modeling time series data with time intervals can obtain more intuitive global information, specifically as follows:

[0164] X t1 '={X1', X3', X5',..., X ti-1 '} (22)

[0165] X t2 '={X2', X4', X6',..., X ti '} (23)

[0166] Among them, X t1 ' and X t2 ' represent the time sub-features obtained by sampling the global feature X G3 in the time dimension;

[0167] Step 1.3.2: Perform the first time-domain information interaction on the time sub-features X t1 ' and X t2 ' to calculate the time attention weight W1 after rough fusion. The specific calculation process is as follows:

[0168] X ft1 =Conv1×1×1(X t1 '+X t2 ') (24)

[0169] W f1 =Sigmoid(Concat T (GAP(X ft1 ), GMP(X ft1 ))) (25)

[0170] W1 = Conv3×1×1(W f1 ) (26)

[0171] Among them, Conv1×1×1(·) and Conv3×1×1(·) are 3D convolution operations, + is element-wise addition, GAP(·) is global average pooling in the spatial dimension, which is used to capture the gently changing facial information, GMP(·) is global max pooling, which is used to capture the violently changing facial information, Sigmoid(·) is an activation function, and Concat T (·) is feature concatenation in the time dimension, expanding the ability of the weights to perceive different facial behaviors, X t1 ' and X t2 ' represent the global feature X G3 for the time sub-features obtained by downsampling in the time dimension, and X ft1 represents X t1 ' and X t2 ' are the coarsely fused features, and W1 represents the time attention weights after the coarse fusion of X t1 ' and X t2 '; W f1 represents the fused weights of X t1 ' and X t2 ' before the dimensionality reduction in the time dimension;

[0172] Step 1.3.3: Update the time information of the time sub-features X t1 ' and X t2 ' using the time attention weight W1. In this process, a time relationship perception vector M is generated, and the calculation process is as follows:

[0173] X t1 ” = W1 × X t1 ' (27)

[0174] X t2 ” = (1 - W1) × X t2 ' (28)

[0175] X′ ft1 = GMP(X t1 ” + X t2 ”) (29)

[0176] M = Conv3×1×1(W f1 × X G3 ) (30)

[0177] Among them, × is element-wise multiplication, - is element-wise subtraction, + is element-wise addition, GMP(·) is global max pooling, and Conv3×1×1(·) is a 3D convolution operation, which is used to reduce the dimensionality in the time dimension. X t1 ' and X t2 ' represent the global feature X G3The time sub-features for time dimension downsampling. It should be noted that in the weight assignment stage, a gating mechanism is used to reassign weights, enhancing the selectivity of the attention coefficient and facilitating better optimization of the model's parameters. X t1 ” represents X t1 'The time sub-features weighted by the time attention weight W1, X t2 ” represents X t2 'The time sub-features weighted by the time attention weight (1 - W1), X′ ft1 represents X t1 ” and X t2 'The coarse-grained fusion feature, M represents the time relationship perception vector;

[0178] Step 1.3.4, through the time relationship perception vector M, the time sub-features X t1 ” and X t2 ” further enhance the module's ability to perceive long-term information, and the calculation process is as follows:

[0179] W2 = Sigmoid(GMP(X t1 ” + X t2 ”)+M) (31)

[0180] X t1 ”' = W2 × X t1 ” (32)

[0181] X t2 ”' = (1 - W2) × X t2 ” (33)

[0182] X ft2 = Conv1×1×1(Concat T (X t1 ”', X t2 ”')) (34)

[0183] Among them, × is element-wise multiplication, - is element-wise subtraction, + is element-wise addition, GMP(·) is global max pooling, Conv1×1×1(·) is a 3D convolution operation, Sigmoid(·) is an activation function, Concat T (·) is time dimension feature concatenation, X t1 ” represents X t1 'The time sub-features weighted by the time attention weight W1, X t2 ” represents X t2 'The time sub-features weighted by the time attention weight (1 - W1), M represents the time relationship perception vector, W2 represents the time attention weight after coarse fusion of X t1 ” and X t2 ”', X t1”' represents X t1 ” The time sub - feature after being weighted by the time attention weight W2, X t2 ”' represents X t2 ” The time sub - feature after being weighted by the time attention weight (1 - W2), X ft2 represents X t1 ”' and X t2 ”' Coarse - grained fusion feature;

[0184] Step 1.3.5: Utilize the coarse - grained fusion feature of the time sub - features X t1 ”' and X t2 ”' to enhance the time perception of the input feature and obtain the output global feature X of the long - term information aggregation module G4 , and the specific calculation process is as follows:

[0185] X G4 = X ft2 ×X G3 +X G3 (35)

[0186] Among them, × is element - wise multiplication, + is element - wise addition, X ft2 represents X t1 ”' and X t2 ”' Coarse - grained fusion feature, X G3 represents the global feature of the input long - term information aggregation module, and represents the output global feature of the long - term information aggregation module;

[0187] Step 1.3.6: Take the output global feature X of the long - term information aggregation module G4 as the input, and then successively pass through a global feature local perception enhancement module with the number of channels doubled again, a 3 - D convolutional residual block with the number of channels doubled, and a long - term feature information aggregation module with the number of channels doubled for learning, to obtain the global feature X of the face global image sequence global .

[0188] The network parameter structure of the LSTIA - PLEGDF global feature learning branch model constructed in this embodiment is shown in Table 1.

[0189] Table 1 Network structure parameters of the LSTIA - PLEGDF model

[0190]

[0191]

[0192] Step 2: Use a three - dimensional residual convolutional network to construct a local feature learning branch model, and input the image sequence of the eye region of the depression patient into this model to obtain local features;

[0193] The specific process is as follows: Input the image sequence of the eye region of the depression patients into a 50-layer three-dimensional residual convolutional network for feature learning to obtain the local region feature X local '.

[0194] Step 3: Construct a semantic correlation feature fusion module for the global and local features of the face and train the model, calculate the correlation weight of the global and local features of the face, fuse the semantic consistency information of the global and local features, and obtain the fused features with consistent global and local semantics;

[0195] The structure of the global and local semantic correlation feature fusion module is as Figure 5 shown, and its construction process is as follows:

[0196] Step 3.1: Align the global feature X global of the global face image sequence, and the local region feature X local ' obtained in step 2, and use a linear layer to align the dimensions of the global feature X global and the local region feature X local '. The calculation process is as follows:

[0197] F g = linear1(X global ) (35)

[0198] F l = linear2(X local ') (36)

[0199] where linear1(·) and linear2(·) represent linear layers, X global is the global feature learned in step 1, X local ' represents the local region feature learned through a 50-layer 3D residual convolutional network in step 2, and F g and F l represent the global feature and local region feature after linear mapping;

[0200] Step 3.2: Calculate the correlation coefficient w g between the globally linearly mapped feature F l and the locally linearly mapped feature F gl . The calculation process is as follows:

[0201] w gl = Sigmoid(F g × F l ) (37)

[0202] where × is element-wise multiplication, and the magnitude of w gl represents the correlation degree between the local region feature F l and the global feature F g ;

[0203] Step 3.3: Calculate the global and local region correlation feature F through the correlation coefficient w gl and the globally mapped global feature F g , the calculation process is as follows: gl The calculation process is:

[0204] F gl = w gl × F g (38)

[0205] where × is element-wise multiplication, w gl represents the global and local region feature correlation coefficient, F g represents the globally mapped global feature, and F gl represents the global and local region correlation feature;

[0206] Step 3.4: Fuse the globally mapped global feature and the local region feature. The specific calculation process is as follows:

[0207] w' gl = Sigmoid((F g + F gl ) × (F l + F gl )) (39)

[0208] F' gl = w' gl × (F g + F gl ) + w' gl × (F l + F gl ) (40)

[0209] where × is element-wise multiplication, + is element-wise addition, w' gl represents the global and local correlation coefficient after fusing the global and local correlation features, F g and F l represent the globally mapped global feature and the local region feature, F gl represents the global and local region correlation feature, and F' gl represents the fused feature with consistent global and local region semantics.

[0210] Step 4: Input the fused feature into the fully connected layer, and then connect it to the fully connected layer with only one neuron to output the depression intensity score, and identify the depression intensity through the depression intensity score.

[0211] The specific steps are as follows:

[0212] Input the fused feature F' with consistent global and local region semantics obtained in Step 3gl Input into the fully connected layer, and then connected to another fully connected layer with only one neuron to output the depression intensity score, so as to identify the depression intensity. The calculation process is expressed as:

[0213] F socre = F' gl × w + b (41)

[0214] Among them, × is element-wise multiplication, + is element-wise addition, w represents the weight, b represents the bias, and F' gl represents the fused feature with consistent semantics in the global and local regions, and F socre represents the identified depression intensity score.

[0215] A dropout layer is embedded before the last fully connected layer, and some neurons are randomly discarded during the training process to train the model and improve the generalization of the model.

[0216] Example 2:

[0217] Compared with Example 1, in this example, an LSTIA-PLEGDF-FGLSCF model with a network depth of 35 is built for testing. The specific steps refer to Example 1, and the following tests are carried out based on the saved 35-layer LSTIA-PLEGDF-FGLSCF model:

[0218] Step A: Use the face sequence images and the corresponding eye region images cropped by the multi-task cascaded convolutional network as the input of the LSTIA-PLEGDF-FGLSCF model, and its label is the Beck depression score;

[0219] Step B: Build the LSTIA-PLEGDF global feature learning branch to learn the global spatio-temporal depression feature X global .

[0220] Step C: Use the entire three-dimensional residual network as the local feature learning branch to learn the local feature X local ' included in the eye region.

[0221] Step D: Extract the semantic consistency information of the global feature X global and the local feature X local ' through the global and local semantic correlation feature fusion module. First, align the dimensions of the global feature X global and the local feature X local ' through the linear layer; secondly, calculate the correlation coefficient w gl of the global feature and the local region feature; then, calculate the global and local region correlation feature F gl through w gl , and finally, take F glFuse the global and local features with the alignment dimensions to obtain the fused feature F′ with global and local region semantic consistency gl 。

[0222] Step E: Input the fused feature F′ gl into the fully connected layer, and then connect it to another fully connected layer with only one neuron to output the depression intensity score. If the output depression intensity score is consistent with the video label, it indicates that the model has successfully completed the depression intensity recognition task.

[0223] Example 3:

[0224] A large number of experiments were conducted and analyzed on the AVEC2013 and AVEC2014 depression datasets of the present invention to evaluate the performance of the present invention in various aspect indicators.

[0225] The evaluation indicators and results used in the experiments are compared as follows:

[0226] For the AVEC2013 and AVEC2014 datasets, the present invention treats them as a regression task. The label value range of each video is the Beck depression score from 0 to 63 points, and the label values are continuous in the real number domain. In addition, the baseline models given in the AVEC2013 and AVEC2014 datasets use the root mean square error (RMSE) and mean absolute error (MAE) as the evaluation indicators of the model. Other works evaluate the performance of the model through these two indicators. Therefore, the present invention uses these two indicators to evaluate the performance of the model based on the global and local channel attention fusion network.

[0227] The performance comparison results of different network models on the AVEC2013 dataset are shown in Table 2, and the performance comparison results of different network models on the AVEC2014 dataset are shown in Table 3.

[0228] Table 2 Comparison of schemes for identifying depression intensity on the AVEC2013 dataset

[0229]

[0230]

[0231] Table 3 Comparison of schemes for identifying depression intensity on the AVEC2014 dataset

[0232] Whether the model is pre-trained Method RMSE MAE Not pre-trained Valstar et al. 10.86 8.86 Pre-trained Zhu et al. 9.55 7.47 Pre-trained Jazaery et al. 9.20 7.22 Pre-trained Melo et al. <![CDATA 7.61 > 5.82 Pre-trained Melo et al. 7.65 6.06 Pre-trained Pan et al. 7.75 6.00 Pre-trained Zhou et al. 8.39 6.21 Pre-trained Melo et al. 8.23 6.15 Pre-trained Sun Haohao et al. 8.56 6.65 Not pre-trained Uddin et al. 8.78 6.86 Not pre-trained He et al. 8.30 6.51 Not pre-trained Shang et al. 7.84 6.08 Not pre-trained LSTIA-PLEGDF-FGLSCF 7.47 <![CDATA 5.90 >

[0233] As can be seen from Table 2 and Table 3, for the method of identifying depression intensity by integrating spatio-temporal features and fusing global-local features in the present invention, the values of the root mean square error (RMSE) and mean absolute error (MAE) metrics are 7.71 / 5.88 and 7.47 / 5.90 respectively, and the depression intensity of the participants can be accurately identified. Through the short-term information aggregation module, short-term features with different temporal and spatial information richness are effectively integrated; the global depression feature local perception enhancement module effectively enhances the feature correlation between different local regions, and effectively utilizes the correlation information between local regions to enhance the local perception of global features; the long-term information aggregation module uses an improved attention mechanism to perceive the temporal context relationship of input features, and enhances the temporal perception of long-term features through temporal relationship perception weights, which is semantically complementary to the short-term features learned by the short-term information aggregation module, ensuring that the extracted dynamic depression information is more complete and sufficient; finally, the global and local semantic correlation feature fusion module captures the relevance of global and local semantic information, realizing the semantic consistency description between global and local depression features.

[0234] In summary, the method of identifying depression intensity by integrating spatio-temporal features and fusing global-local features in the present invention is overall superior to most existing methods, verifying the effectiveness of this method in the task of identifying depression intensity, being able to effectively utilize the correlation information between local regions to enhance the local perception of facial global features, and fully extracting the dynamic depression information of facial global short-term and long-term. Finally, the correlation between the global feature containing the entire facial information and the local feature containing only the eye region information is learned, and the semantic consistency information between global and local depression features is obtained.

Claims

1. A depression intensity recognition method based on spatiotemporal feature integration and global-local feature fusion, characterized in that: The specific steps are as follows: Step 1: construct the LSTIA-PLEGDF global feature learning branch model, and input the global face image sequence of patients with depression into the model to obtain the global face features; The LSTIA-PLEGDF global feature learning branch model consists of a short-term information aggregation module, a global depression feature local perception enhancement module and a long-term information aggregation module. In the LSTIA-PLEGDF global feature learning branch model, the short-term information aggregation module, the global depression feature local perception enhancement module and the long-term information aggregation module are stacked to effectively extract the spatiotemporal information of the global features in a layered stacking manner; The specific construction process of the LSTIA-PLEGDF global feature learning branch model is as follows: Step 1.1, construct a short-term information aggregation module to obtain the initial global features of the global face image sequence; The construction process of the short-term information aggregation module is as follows: Step 1.1.1, obtain the global face image sequence X input , a fixed-size sliding window is used in the time direction to obtain continuous short-term features of the global face image sequence {X1,X2,...,X i }, along the channel dimension, each short-term feature X i Divide into four segments, and obtain four channel sub-segments {X i1 ,X i2 ,X i3 ,X i4 }; Among them, X input is a global face image sequence, C represents the number of image channels, T1 represents the length of the image sequence, H represents the image height, and W represents the image width; {X1,X2,...,X i } represents continuous short-term features; {X i1 ,X i2 ,X i3 ,X i4 } represents X i Channel sub-segments after segmentation along the channel dimension; Step 1.1.2: For each short-term feature X i The first two fragments of the four channel sub-fragments {X i1 ,X i2 }, respectively use the time convolution layer for feature learning, for each short-term feature X i The last two fragments of the four channel sub-fragments {X i3 ,X i4 }, respectively, using spatial convolutional layers for feature learning. The last three fragments {X i2 ,X i3 ,X i4 After feature learning, the output of the previous segment is added in turn, and then feature learning is performed through the spatiotemporal convolution layer to obtain the output feature The specific calculation process is as follows: Among them, Conv 3×1×1 (·), Conv 3×3×3 (·), Conv 1×3×3 (·) indicates 3D convolution operation for feature learning, + indicates feature element superposition, and output feature The temporal and spatial information represented; Step 1.1.3: Output features Perform aggregation to obtain the post-aggregation features Then calculate and The corresponding time attention weight w T and the spatial attention weight w S , the calculation process is as follows: Among them, Concate C (·) represents channel dimension concatenation, SP(·) and TP(·) represent maximum pooling in spatial dimension and maximum pooling in temporal dimension respectively, Sigmoid(·) is the activation function, FC * (·) represents a fully connected layer, represents a channel sub-fragment {X i1 ,X i2 ,X i3 ,X i4 } Output features after feature learning, express The characteristics after channel splicing, express The feature after channel splicing, w T express The temporal attention weight, w S express The spatial attention weights of Step 1.1.4, through w T and w S Enhance The time information expression and The spatial information expression of , and aggregate the spatiotemporal features, the calculation process is as follows: Among them, × represents element multiplication, - represents element subtraction, + represents element addition, Concate C (·) represents channel dimension concatenation, express The characteristics after channel splicing, express The feature after channel splicing, w T express The temporal attention weight, w S express The spatial attention weights, Express The features after weighting the time dimension and space dimension features, Express The features after weighting the time dimension and space dimension features, express and Short-term characteristics after channel splicing; Step 1.1.5: Repeat steps 1.1.2 to 1.1.4 to obtain all short-term features X i Short-term characteristics after channel splicing Get the complete initial global feature X G1 : Among them, Concate T (·) indicates time dimension splicing, Represents each time sub-segment X i Short-term characteristics after channel splicing, X G1 Represents the initial global features finally output by the short-term information aggregation module; Step 1.2, using the initial global features as input features, constructing a global depression feature local perception enhancement module to obtain high-level global face features; Step 1.3, taking the advanced global face features as input, constructing a long-term information aggregation module, and obtaining the global features of the global face image sequence; The construction process of the long-term information aggregation module is as follows: Step 1.3.1: Advanced face global feature X G3 ={X1',X2',X3',...,X ti '} Perform interval sampling in the time dimension to obtain the sampling time sub-feature X t1 ' and X t2 ', as follows: X t1 '={X1',X3',X5',...,X ti-1 '} (22) X t2 '={X2',X4',X6',...,X ti '} (23) Among them, X t1 ' and X t2 ' represents the global feature X G3 Temporal sub-features for downsampling the temporal dimension; Step 1.3.2: Time sub-feature X t1 ' and X t2 'Perform the first time domain information interaction and calculate the temporal attention weight W1 after rough fusion. The specific calculation process is as follows: X ft1 =Conv1×1×1(X t1 '+X t2 ') (24) W f1 =Sigmoid(Concat T (GAP(X ft1 ),GMP(X ft1 ))) (25) W1=Conv3×1×1(W f1 ) (26) Among them, Conv1×1×1(·) and Conv3×1×1(·) are 3D convolution operations, + is element-wise addition, GAP(·) is global average pooling in the spatial dimension, which is used to capture facial information with gentle changes, GMP(·) is global maximum pooling, which is used to capture facial information with drastic changes, Sigmoid(·) is the activation function, and Concat T (·) is the splicing of time dimension features, which expands the ability of weights to perceive different facial behaviors. t1 ' and X t2 ' represents the global feature X G3 The time sub-features for downsampling in the time dimension, X ft1 Represents X t1 ' and X t2 'Coarse-grained fusion features, W1 represents X t1 ' and X t2 'The temporal attention weight after coarse fusion, W f1 Indicates that W1 performs time dimension reduction before X t1 ' and X t2 'The fusion weight; Step 1.3.3: Update the temporal sub-feature X using the temporal attention weight W1 t1 ' and X t2 ' ... X t1 ”=W1×X t1 ' (27) X t2 ”=(1-W1)×X t2 ' (28) X′ ft1 =GMP(X t1 ”+X t2 ”) (29) M=Conv3×1×1(W f1 ×X G3 ) (30) Among them, × is element-wise multiplication, - is element-wise subtraction, + is element-wise addition, GMP(·) is the global maximum pooling, Conv3×1×1(·) is a 3D convolution operation, which is used to reduce the time dimension, X t1 ' and X t2 ' represents the global feature X G3 The time sub-features for downsampling in the time dimension, X t1 " indicates X t1 'The time sub-feature after weighting by the time attention weight W1, X t2 " indicates X t2 'The time sub-feature after weighting by the time attention weight (1-W1), X f ' t1 Represents X t1 ” and X t2 'Coarse-grained fusion features, M represents the temporal relationship perception vector; Step 1.3.4: Perceive vector M and temporal sub-feature X through temporal relationship t1 ” and X t2 "To further enhance the module's ability to perceive long-term information, the calculation process is as follows: W2=Sigmoid(GMP(X t1 ”+X t2 ”)+M) (31) X t1 ”'=W2×X t1 ” (32) X t2 ”'=(1-W2)×X t2 ” (33) X ft2 =Conv1×1×1(Concat T (X t1 ”',X t2 ”')) (34) Among them, × is element-wise multiplication, - is element-wise subtraction, + is element-wise addition, GMP(·) is global maximum pooling, Conv1×1×1(·) is a 3D convolution operation, Sigmoid(·) is an activation function, Concat T (·) is the time dimension feature concatenation, X t1 " indicates X t1 'The time sub-feature after weighting by the time attention weight W1, X t2 " indicates X t2 'The time sub-feature after weighting by the time attention weight (1-W1), M represents the time relationship perception vector, W2 represents X t1 ” and X t2 "The temporal attention weight after coarse fusion, X t1 "' indicates X t1 "The temporal sub-feature after weighting by the temporal attention weight W2, X t2 "' indicates X t2 "The time sub-feature after weighting by the time attention weight (1-W2), X ft2 Represents X t1 ”' and X t2 ”'Coarse-grained fusion features; Step 1.3.5: Use the time sub-feature X t1 ”' and X t2 The coarse-grained fusion features of "' improve the temporal perception of the input features and obtain the output global features X of the long-term information aggregation module G4 , the specific calculation process is as follows: X G4 =X ft2 ×X G3 +X G3 (35) Among them, × is element-wise multiplication, + is element-wise addition, and X ft2 Represents X t1 ”' and X t2 ”'Coarse-grained fusion features, X G3 represents the global features of the input long-term information aggregation module, represents the output global features of the long-term information aggregation module; Step 1.3.6: Output global feature X of the long-term information aggregation module G4 As input, it is then learned in sequence through a global feature local perception enhancement module with doubled number of channels, a 3D convolution residual block with doubled number of channels, and a long-term feature information aggregation module with doubled number of channels to obtain the global feature X of the global face image sequence. global ; Step 2: Use a three-dimensional residual convolutional network to build a local feature learning branch model, and input the eye area image sequence of patients with depression into the model to obtain local features; Step 3: construct a feature fusion module for the semantic correlation between global and local features of the face, calculate the correlation weights between global and local features of the face, fuse the semantic consistency information of global and local features, and obtain fused features with global and local semantic consistency; Step 4: Input the fused features into the fully connected layer, and then connect it to the fully connected layer containing only one neuron to output the depression intensity score, and identify the depression intensity through the depression intensity score.

2. The depression intensity identification method based on spatiotemporal feature integration and global-local feature fusion according to claim 1 is characterized in that: The construction process of the global depression feature local perception enhancement module in step 1.2 is as follows: Step 1.2.1: Initial global feature X G1 The spatial dimension is evenly divided into 4 non-overlapping regions X L ={X local1 ,...,X localj }, j = {1, 2, 3, 4}, where X local1 and X local2 Contains information about the eyes, X local3 and X local4 Contains the information of the mouth; keeps the time and channel dimensions unchanged, and converts each local area X L The local feature X localj Transformed into local feature vectors in spatial dimensions The calculation process is: Among them, ReLU(·) represents the activation function, Reshape(·) represents compressing the feature from two dimensions to one dimension in the spatial dimension, and Conv 1×1×1 (·) is a 3D convolution operation, T2 represents the sequence length of the feature, C1 represents the number of channels of the feature, H1 represents the height of the feature, W1 represents the width of the feature, N represents the product of the height and width of the feature, and X localj Indicates area X L Any local feature in Represents the local feature X localj Local eigenvector after spatial compression; Step 1.2.2: Calculate each local feature X localj Vector The correlation coefficient matrix r n,m , so as to obtain each local feature X localj Vector The correlation coefficient R k , and obtain the local feature correlation coefficient vector set R. The specific calculation process is as follows: R={R1,R2,...,R k } (18) Where T represents transpose, × represents matrix multiplication, + represents element superposition, T2 represents the sequence length of the feature, N represents the product of the height and width of the feature, Sigmoid(·) represents the activation function, x={1,2,3,4}, and x≠j, n={1,2,...,N}, m={1,2,...,N}, r n,m Represents the local feature vector The correlation coefficient matrix between and Represents different local eigenvectors, r n,* Reflects the local area For local areas The importance of r *,m Includes local area Information and local areas The correlation, R k Represents the local feature vector The kth element in and the local eigenvector The correlation coefficient, R = {R1, R2, ..., R k }, k = {1, 2, ..., N} represents the set of local feature correlation coefficient vectors, and N represents the product of the height and width of the feature; Step 1.2.3: Arrange the local feature correlation coefficient vector R according to the position information in the spatial dimension to obtain the region X L The correlation coefficient matrix between each local area and the other three local areas The calculation process is: in, It means that the local feature correlation coefficient vector R is arranged according to the position information in the spatial dimension, T2 represents the sequence length of the feature, C1 represents the number of channels of the feature, H1 represents the height of the feature, W1 represents the width of the feature, and R is the set of local feature correlation coefficient vectors, R = {R1, R2, ..., R k }, The correlation coefficient matrix of the local features represented by the local feature correlation coefficient vector R; Step 1.2.4: Through the correlation coefficient matrix Compute local correlation-aware local features X outj Concatenate to get the final output feature X G2 , the calculation process is as follows: Among them, × represents matrix multiplication, + represents element superposition, T2 represents the sequence length of the feature, C1 represents the number of channels of the feature, H1 represents the height of the feature, W1 represents the width of the feature, C represents the number of image channels, T1 represents the image sequence length, H represents the image height, W represents the image width, Concat(·) represents concatenating the features according to the spatial dimension, X represents localj Indicates area X L ={X local1 ,...,X localj Any local feature in}, Region X L The correlation coefficient matrix between each local area and the other three local areas, X outj represents the local feature with local correlation awareness, X G2 Represents the output features of the local perception enhancement module of the global depression feature; Step 1.2.5: Transform the output feature X of the local perception enhancement module of the global depression feature G2 The high-level face global feature X is obtained by sequentially inputting the short-term information aggregation module with doubled channels, the local perception enhancement module with doubled channels and the global depression feature. G3 .

3. The depression intensity identification method based on spatiotemporal feature integration and global-local feature fusion according to claim 1 is characterized in that: The specific process of step 2 is as follows: input the eye region image sequence of the depression patient into a 50-layer 3D residual convolutional network for feature learning, and obtain the local region feature X local '.

4. The depression intensity identification method based on spatiotemporal feature integration and global-local feature fusion according to claim 1 is characterized in that: The construction process of the global and local semantic correlation feature fusion module in step 3 is as follows: Step 3.1: The global feature X of the global face image sequence global , and the local area feature X obtained in step 2 local ', use the linear layer to align the global feature X global and local area features X local ', the calculation process is: F g =linear1(X global ) (35) F l =linear2(X local ') (36) Among them, linear1(·) and linear2(·) represent linear layers, X global is the global feature of the global face image sequence, X local ' represents the local region features learned by the 50-layer 3D residual convolutional network in step 2, F g and F l Represents the global features and local area features after linear mapping; Step 3.2: Calculate the global feature F after linear mapping g and local region features F l The correlation coefficient w gl , the calculation process is: w gl =Sigmoid(F g ×F l ) (37) Among them, × is element-wise multiplication, w gl The size indicates the local area feature F l and the global feature F g ; Step 3.3, through the correlation coefficient w gl And the global feature F after linear mapping g , calculate the global and local region correlation features F gl , the calculation process is: F gl =w gl ×F g (38) Among them, × is element-wise multiplication, w gl represents the correlation coefficient between global and local region features, F g represents the global feature after linear mapping, F gl Represents global and local region correlation features; Step 3.4: Fuse the global features and local area features after linear mapping. The specific calculation process is as follows: w g ′ l =Sigmoid((F g +F gl )×(F l +F gl )) (39) F g ′ l =w g ′ l ×(F g +F gl )+w g ′ l ×(F l +F gl ) (40) Among them, × is element-wise multiplication, + is element-wise addition, and w g ' l represents the global and local correlation coefficient after fusing the global and local correlation features, F g and F l represents the global features and local region features after linear mapping, F gl represents the global and local region correlation features, F g ' l Indicates the fusion features with consistent global and local region semantics.

5. The depression intensity identification method based on spatiotemporal feature integration and global-local feature fusion according to claim 1 is characterized in that: The specific process of step 4 is as follows: Step 4: The fusion feature F with consistent global and local region semantics obtained in step 3 g ' l The input is sent to the fully connected layer, and then connected to another fully connected layer containing only one neuron to output the depression intensity score, thereby performing depression intensity recognition. The calculation process is expressed as: F socre =F g ′ l ×w+b (41) Among them, × is element-wise multiplication, + is element-wise addition, w represents weight, b represents bias, and F g ' l Indicates the fusion features with consistent global and local semantics, F socre Represents the identified depression intensity score.

6. The depression intensity identification method based on spatiotemporal feature integration and global-local feature fusion according to claim 1 is characterized in that: The global face image sequence of the depression patient in step 1 and the eye region image sequence of the depression patient in step 2 are obtained by cropping and face aligning each image sample in the video data obtained from the depression video datasets AVEC2013 and AVEC2014.