Depression recognition method combining multi-level feature dependency and multi-scale contextual features

By combining multi-level feature dependence and multi-scale context features, multi-level emotion representation and multi-scale context features in visual emotion analysis are extracted, and depression recognition is performed through feature fusion, which solves the problem of combining scale perception and emotion representation and noise influence in the prior art, and achieves more accurate depression recognition.

CN114494770BActive Publication Date: 2025-05-06NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210035384.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2025-05-06
Estimated Expiration
2042-01-12

AI Technical Summary

Technical Problem

In the visual emotion analysis, it is difficult for the prior art to combine the perception of objects at different scales and the extraction and utilization of emotional representations at different levels, and there is an impact of image noise on classification.

Method used

Using a depression recognition method combining multi-level feature dependence and multi-scale context features, multi-level emotion representation is extracted through the ResNet101 network, combined with a multi-layer bidirectional gating network and a multi-scale adaptive context module, the dependence relationship between features and multi-scale context features are extracted, and significant emotional areas are emphasized through the convolutional attention module, and feature fusion and depression recognition are finally carried out.

Benefits of technology

Deeper and more comprehensive emotional acquisition is achieved, reducing the impact of noise and improving the accuracy of depression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494770B_ABST
    Figure CN114494770B_ABST
Patent Text Reader

Abstract

The present invention discloses a depression recognition method combining multi-level feature dependencies and multi-scale context features, including: building and improving ResNet101 to obtain four emotion representations C<supgt;2< / supgt>, C<supgt;3< / supgt>, C<supgt;4< / supgt> and C<supgt;5< / supgt> representing different levels of semantic information and detail information; inputting the multi-level emotion representations into a multi-layer bidirectional gated network to extract the dependency relationships between multi-level features; sending the multi-level emotion representations to a multi-scale adaptive context module to extract corresponding multi-scale context features, and obtaining multi-scale emotion representations O<supgt;3< / supgt>, O<supgt;4< / supgt> and O<supgt;5< / supgt> at each level through cross-level fusion; importing the C<supgt;5< / supgt> level emotion representation into a convolutional attention module to obtain an emotion representation D with significant emotion regions; obtaining a multi-cue emotion feature E through feature fusion, and performing depression recognition through a classification network. The present invention combines both intra-level and inter-level associations to obtain deeper and more comprehensive emotions, and is robust to noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sentiment analysis, and in particular to a depression recognition method combining multi-level feature dependencies with multi-scale context features. Background Art

[0002] With the development of short videos and the accelerated pace of life, sharing feelings through images and short video content has become a major way of expression on social networks. At the same time, for people with depression, as time goes by and the symptoms worsen, they prefer to use abstract and simple visual descriptions instead of intuitive and complex text descriptions. In addition, the emotions contained in image content are more diverse and in-depth. The above situation has made the task of analyzing the emotional information contained in image content more and more concerned, and it is necessary to invent a method that can automatically analyze the influence in social networks and make accurate depression judgments.

[0003] Compared with many traditional visual tasks, visual emotion analysis has two important directions in image content understanding: one is the perception of objects of different scales, and the other is the extraction and utilization of emotional representations at different levels. Current research mostly considers breakthroughs in one direction, without combining the two directions. In addition, when combining overall and local emotional features, existing methods have the influence of image noise on classification. Summary of the invention

[0004] In view of the deficiencies in the prior art, the present invention provides a depression recognition method that combines multi-level feature dependencies with multi-scale contextual features, and combines intra-level correlations with inter-level correlations to obtain deeper and more comprehensive emotions, and is robust to noise.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The embodiment of the present invention proposes a depression recognition method combining multi-level feature dependency and multi-scale context features, and the depression recognition method includes the following steps:

[0007] S1, build the ResNet101 network as the model feature extraction layer, use ELU as the activation function and add the BN layer to obtain C 2 , C 3 , C 4 and C 5 Four sentiment representations representing different levels of semantic and detail information;

[0008] S2, input the multi-level emotion representation obtained in step S1 into a multi-layer bidirectional gating network, extract the dependency between multi-level features, connect the bidirectional outputs, and obtain the association representation V of the image;

[0009] S3, the multi-level emotion representation obtained in step S1 is sent to the multi-scale adaptive context module to extract the corresponding multi-scale context features, and the multi-scale emotion representation of each level is obtained through cross-level fusion. 3 , O 4 and O 5 ;

[0010] S4, C 5 The first-level emotion representation is imported into the convolutional attention module to obtain the emotion representation D with significant emotion areas;

[0011] S5, multi-scale emotion representation O 3 , O 4 , O 5 The multi-cue emotional feature E is obtained by performing feature fusion with the emotional representation D with significant emotional regions and the associated representation V of the image;

[0012] S6, the multi-cue emotional feature E is sent to the global average pool, and two labels, whether it is depressed and the severity of depression, are obtained through two fully connected layers.

[0013] Furthermore, in step S6, the multi-cue emotion feature E is classified by two fully connected classification networks, wherein one classification network is used to generate a label indicating whether the person is depressed, and the other classification network is used to predict a depression severity score.

[0014] Furthermore, in step S1, the model feature extraction layer includes a plurality of residual blocks connected in sequence, a convolution layer of each residual block is connected to a BN layer, and an activation function used by the residual block is an ELU activation function.

[0015] Further, the residual block includes a first convolutional layer, a first BN layer, a second convolutional layer, a second BN layer, a third convolutional layer, a third BN layer and an addition layer connected in sequence;

[0016] The first convolution layer uses 64 1*1 convolution kernels to convolve the imported original feature matrix with a depth of 256 dimensions, so that the depth of the feature matrix is ​​reduced to 64 dimensions, and the output result of the first convolution layer is normalized by the first BN layer and then activated by the ELU activation function to output the first feature matrix; the second convolution layer uses 64 3*3 convolution kernels to convolve the first feature matrix with a depth of 64 dimensions output by the first convolution layer, and the output result of the second convolution layer is normalized by the second BN layer and then activated by the ELU activation function to output the second feature matrix; the third convolution layer uses 256 1*1 convolution kernels to convolve the second feature matrix with a depth of 64 dimensions output by the second convolution layer, so that the depth of the feature matrix is ​​increased to 256 dimensions, and the output result of the third convolution layer is normalized by the third BN layer and directly outputs the third feature matrix;

[0017] The addition layer is used to add the original feature matrix with a depth of 256 dimensions imported into the first convolutional layer and the third feature matrix with a dimension increased to 256 dimensions output by the third convolutional layer, and output the corresponding emotion representation result after activation with the ELU activation function.

[0018] Furthermore, in step S2, the features from low-level to high-level and from high-level to low-level are taken as two kinds of sequential data, and a multi-layer bidirectional gating network is used to fuse different dependencies.

[0019] Further, in step S3, the multi-scale adaptive context module includes a backbone layer, a sub-region branch, a regional sentiment affinity coefficient branch, a single-scale context feature calculation unit and a combination unit;

[0020] The sub-region branch is used to learn the feature map X of the backbone layer. l The local sub-region representation under different scale divisions, for each scale S k , the feature map X of the backbone layer l Divided into S k ×S k Sub-areas k = 1, 2...n, n is a positive integer, and features are extracted from each sub-region through average pooling and 1×1 convolution;

[0021] The regional sentiment affinity coefficient branch is used to learn the affinity coefficient weights between sub-regions at the same scale.

[0022]

[0023] in, is the regional sentiment affinity coefficient, S k It is the division scale. is the calculation function, M l Yes X l After resizing, l is the feature level, g(M l ) is the global information feature, j is the sub-region position; This is achieved through 1×1 convolution and sigmoid activation function. l Using global average pooling, we get the global information feature g(M l ), M l and g(M l ) and calculate the sentiment affinity coefficient of each local position j in the global context;

[0024] The single-scale context feature calculation unit is used to calculate the single-scale context feature:

[0025]

[0026] in, It is size S k The following context features, is a sub-region;

[0027] The combination unit is used to combine context features of different scales, perform batch normalization after 1×1 convolution, and obtain multi-scale context features Z l :

[0028]

[0029] Among them, Z l is a multi-scale context feature, l is the feature level, f is the channel-by-channel combination function, It is a single-scale context feature at different scales.

[0030] Furthermore, in step S3, the multi-scale emotion representation of each level is obtained by cross-level fusion according to the following formula: 3 , O 4 and O 5 :

[0031]

[0032] Among them, l is the characteristic order, O l is the l-level multi-scale sentiment representation, Z l is the multi-scale context feature, M l is the resized representation of the feature map, M l+1 It is the resized representation of the upper feature map.

[0033] Furthermore, in step S4, the emotion representation D with the significant emotion region is calculated according to the following formula:

[0034]

[0035]

[0036] Among them, D1 is the intermediate feature map, G CAM is the feature map obtained by the channel attention module, C 5 It is the fifth level of emotional representation. is the element-by-element multiplication of matrices, G CSM It is the feature map obtained by the spatial attention module.

[0037] The present invention combines the multi-level feature association extracted by the multi-layer bidirectional gating mechanism with the intra-level context features of the multi-scale division to solve the problem that only limited scales can be captured in local emotion extraction, and combines the intra-level association with the inter-level association to obtain deeper and more comprehensive emotions; in addition, the ELU activation function is used to replace the ReLU activation function in the residual block of the ResNet101 network and the BN layer is added after each convolution layer of the residual structure to accelerate the learning of the feature extraction layer and make the network robust to noise; in the final fusion, the features in the attention optimization feature map generated by the attention module are added to emphasize the areas with significant emotions in the image. Through the above measures, the emotional analysis of the image content is finally realized, and the diagnosis of depression and the determination of severity are completed.

[0038] The beneficial effects of the present invention are:

[0039] The depression recognition method proposed in the present invention combines multi-level feature dependency and multi-scale contextual features. The ELU activation function is used to replace the ReLU activation function in the residual block in the ResNet101 network, and a BN layer is added after each convolution layer of the residual structure, so that the feature extraction layer learns faster, accelerates convergence and is robust to noise, thereby improving the network accuracy.

[0040] The depression recognition method proposed in the present invention combines multi-level feature dependencies with multi-scale contextual features, adopts cross-layer and multi-layer feature fusion strategies, and adds the features in the attention optimization feature map generated by the attention module during the second fusion, emphasizing the areas with significant emotions in the image, further reducing noise and improving recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flow chart of a depression identification method combining multi-level feature dependency and multi-scale contextual features according to an embodiment of the present invention.

[0042] Figure 2 Schematic diagram of the structure of the residual block according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The present invention will now be described in further detail with reference to the accompanying drawings.

[0044] It should be noted that the terms such as "upper", "lower", "left", "right", "front", "back", etc. cited in the invention are only for the convenience of description and are not used to limit the scope of implementation of the present invention. Changes or adjustments in their relative relationships should be regarded as the scope of implementation of the present invention without substantially changing the technical content.

[0045] Figure 11 is a flow chart of a depression identification method combining multi-level feature dependency and multi-scale context features according to an embodiment of the present invention. The depression identification method comprises the following steps:

[0046] S1, build the ResNet101 network as the model feature extraction layer, use ELU as the activation function and add the BN layer to obtain C 2 , C 3 , C 4 and C 5 Four sentiment representations representing different levels of semantic and detail information;

[0047] S2, input the multi-level emotion representation obtained in step S1 into a multi-layer bidirectional gating network, extract the dependency between multi-level features, connect the bidirectional outputs, and obtain the association representation V of the image;

[0048] S3, the multi-level emotion representation obtained in step S1 is sent to the multi-scale adaptive context module to extract the corresponding multi-scale context features, and the multi-scale emotion representation of each level is obtained through cross-level fusion. 3 , O 4 and O 5 ;

[0049] S4, C 5 The first-level emotion representation is imported into the convolutional attention module to obtain the emotion representation D with significant emotion areas;

[0050] S5, multi-scale emotion representation O 3 , O 4 , O 5 The multi-cue emotional feature E is obtained by performing feature fusion with the emotional representation D with significant emotional regions and the associated representation V of the image;

[0051] S6, the multi-cue emotional feature E is sent to the global average pool, and two labels, whether it is depressed and the severity of depression, are obtained through two fully connected layers.

[0052] The depression identification method of this embodiment is described in detail below with reference to the accompanying drawings.

[0053] Step 1: Extract multi-level sentiment representation

[0054] The ResNet101 network is a commonly used and effective model in visual tasks, with enough convolutional layers to extract multi-level features. This example uses the ResNet101 network to obtain C 2 , C 3 , C 4 and C 5 Four multi-level sentiment representations representing different levels of semantic information and detail information. Considering the amount of computation, only C 3 , C 4 and C5 Preferably, the ELU activation function is used to replace the ReLU activation function in the residual block of the ResNet101 network. At the same time, in order to make the feature map during training satisfy the standard normal distribution, a BN layer is added after each convolution layer of the residual structure to accelerate the learning of the feature extraction layer and make the network robust to noise. The modified residual block is as follows Figure 2 Specifically, the depth of the feature matrix in the input residual block is 256 dimensions, and 64 1*1 convolution kernels are used to convolve it, so that the depth of the feature matrix is ​​reduced to 64 dimensions, and then convolved by 64 3*3 convolution kernels, and then convolved by 256 1*1 convolution kernels, so that the depth of the feature matrix is ​​increased to 256 dimensions, and finally the input feature matrix is ​​added to the feature matrix after three convolutions; three ELU activations are used in the entire residual block, respectively after two convolutions and after the final addition operation; three BN layers are added, respectively after two convolutions and before the final addition operation.

[0055] Step 2: Extract dependencies between features at different levels

[0056] Features at different levels are input into a multi-layer bidirectional gating network (SBi-GRU), the dependencies between features at different levels are extracted, and the outputs of the bidirectional networks are connected as the association representation V of the image.

[0057] Due to the complexity of emotional stimuli, there is a symbiotic relationship between multi-level features and emotional categories. In order to explore the dependencies between features at different levels, the features from low-level to high-level and from high-level to low-level are taken as two sequential data, and an SBi-GRU network is used to fuse different dependencies.

[0058] Step 3: Get multi-scale emotion representation at each level

[0059] The features at different levels are sent to the corresponding multi-scale adaptive context module (MACM) to obtain the corresponding multi-scale context features and perform cross-level fusion to obtain O 3 , O 4 and O 5 , as a multi-scale sentiment representation at various levels.

[0060] The process of obtaining multi-scale context features by the Multi-scale Adaptive Context Module (MACM) is as follows:

[0061] The multi-scale adaptive context module includes a backbone layer, a sub-region branch, a regional sentiment affinity coefficient branch, a single-scale context feature calculation unit and a combination unit.

[0062] The sub-region branch is used to learn the feature map X of the backbone layer. lThe local sub-region representation under different scale divisions, for each scale S k , the feature map X of the backbone layer l Divided into S k ×S k Sub-areas k=1, 2...n, n is a positive integer, and features are extracted from each sub-region through average pooling and 1×1 convolution.

[0063] The regional sentiment affinity coefficient branch is used to learn the affinity coefficient weights between sub-regions at the same scale.

[0064]

[0065] in, is the regional sentiment affinity coefficient, S k It is the division scale. is the calculation function, M l Yes X l After resizing, l is the feature level, g(M l ) is the global information feature, j is the sub-region position; This is achieved through 1×1 convolution and sigmoid activation function. l Using global average pooling, we get the global information feature g(M l ), M l and g(M l ) and calculate the sentiment affinity coefficient of each local position j in the global context;

[0066] The single-scale context feature calculation unit is used to calculate the single-scale context feature:

[0067]

[0068] in, It is size S k The following context features, is a sub-region;

[0069] The combination unit is used to combine context features of different scales, and perform batch normalization after 1×1 convolution to obtain multi-scale context features Z l :

[0070]

[0071] Among them, Z l is a multi-scale context feature, l is the feature level, f is the channel-by-channel combination function, It is a single-scale context feature at different scales.

[0072] The cross-level fusion process is as follows:

[0073]

[0074] Among them, l is the characteristic order, O l is the l-level multi-scale sentiment representation, Z l is the multi-scale context feature, M l is the resized representation of the feature map, M l+1 It is the resized representation of the upper feature map.

[0075] In order to reduce the noise from low-level features, the obtained multi-scale context feature Z l And the M obtained in the previous layer l +1 Combine them to obtain a multi-scale emotion representation at this level.

[0076] Step 4: Obtain the emotional representation D with significant emotional regions

[0077] C 5 The layer features are fed into the convolutional attention module (CBAM) to obtain the sentiment representation D with significant sentiment regions. The process of the convolutional attention module to obtain the sentiment representation with significant sentiment regions is as follows:

[0078] The CBAM module includes a channel attention module (CAM) and a spatial attention module (CSM), which perform attention allocation respectively. The calculation formula is as follows:

[0079]

[0080]

[0081] Among them, D1 is the intermediate feature map, G CAM is the feature map obtained by the channel attention module, C 5 It is the fifth level of emotional representation. is the element-by-element multiplication of matrices, G CSM It is the feature map obtained by the spatial attention module.

[0082] Step 5: Get the multi-cue emotional feature E

[0083] The multi-scale context feature O 3 , O 4 , O 5 Combined with the emotion representation D with significant emotion areas and the associated representation V of the image, the multi-cue emotion feature E is obtained.

[0084] Step 6: Identify depression

[0085] The multi-cue emotional features E are fed into a global average pool and classified by two fully connected networks, one of which generates a label identifying depression and the other predicts the depression severity score.

[0086] The present invention not only considers how to perceive objects of different scales and different emotional intensities in complex scenes, but also introduces the association between emotional representations at different levels, including color, texture, scene, object, composition, etc., to achieve breakthroughs in emotional analysis of image content from two directions. The present invention improves the noise problem and robustness by replacing the ReLU activation function in the residual block of the ResNet101 network with the ELU activation function, adding a BN layer after each convolutional layer of the residual structure, and adding the features in the attention-optimized feature map generated by the attention module during the final fusion.

[0087] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.

Claims

1. A depression recognition method combining multi-level feature dependency and multi-scale contextual features, characterized in that: The depression identification method comprises the following steps: S1, build the ResNet101 network as the model feature extraction layer, use ELU as the activation function and add the BN layer to obtain C 2 , C 3 , C 4 and C 5 Four sentiment representations representing different levels of semantic and detail information; S2, input the multi-level emotion representation obtained in step S1 into a multi-layer bidirectional gating network, extract the dependency between multi-level features, connect the bidirectional outputs, and obtain the association representation V of the image; S3, the multi-level emotion representation obtained in step S1 is sent to the multi-scale adaptive context module to extract the corresponding multi-scale context features, and the multi-scale emotion representation of each level is obtained through cross-level fusion. 3 , O 4 and O 5 ; S4, C 5 The first-level emotion representation is imported into the convolutional attention module to obtain the emotion representation D with significant emotion areas; S5, multi-scale emotion representation O 3 , O 4 , O 5 The multi-cue emotional feature E is obtained by performing feature fusion with the emotional representation D with significant emotional regions and the associated representation V of the image; S6, the multi-cue emotional feature E is sent to the global average pool, and two labels, whether it is depressed and the severity of depression, are obtained through two fully connected layers.

2. The depression identification method combining multi-level feature dependency and multi-scale contextual features according to claim 1, characterized in that: In step S6, the multi-cue emotional feature E is classified by two fully connected classification networks, one of which is used to generate a label indicating whether the person is depressed, and the other classification network is used to predict the depression severity score.

3. The depression identification method combining multi-level feature dependency and multi-scale contextual features according to claim 1, characterized in that: In step S1, the model feature extraction layer includes a plurality of residual blocks connected in sequence, a BN layer is connected after the convolution layer of each residual block, and the activation function used by the residual block is the ELU activation function.

4. The depression identification method combining multi-level feature dependency and multi-scale contextual features according to claim 3, characterized in that: The residual block includes a first convolutional layer, a first BN layer, a second convolutional layer, a second BN layer, a third convolutional layer, a third BN layer and an addition layer connected in sequence; The first convolution layer uses 64 1*1 convolution kernels to convolve the imported original feature matrix with a depth of 256 dimensions, so that the depth of the feature matrix is ​​reduced to 64 dimensions, and the output result of the first convolution layer is normalized by the first BN layer and then activated by the ELU activation function to output the first feature matrix; the second convolution layer uses 64 3*3 convolution kernels to convolve the first feature matrix with a depth of 64 dimensions output by the first convolution layer, and the output result of the second convolution layer is normalized by the second BN layer and then activated by the ELU activation function to output the second feature matrix; the third convolution layer uses 256 1*1 convolution kernels to convolve the second feature matrix with a depth of 64 dimensions output by the second convolution layer, so that the depth of the feature matrix is ​​increased to 256 dimensions, and the output result of the third convolution layer is normalized by the third BN layer and directly outputs the third feature matrix; The addition layer is used to add the original feature matrix with a depth of 256 dimensions imported into the first convolutional layer and the third feature matrix with a dimension increased to 256 dimensions output by the third convolutional layer, and output the corresponding emotion representation result after activation using the ELU activation function.

5. The depression identification method combining multi-level feature dependency and multi-scale contextual features according to claim 1, characterized in that: In step S2, the features from low-level to high-level and from high-level to low-level are taken as two kinds of sequential data, and a multi-layer bidirectional gating network is used to fuse different dependencies.

6. The depression identification method combining multi-level feature dependency and multi-scale contextual features according to claim 1, characterized in that: In step S3, the multi-scale adaptive context module includes a backbone layer, a sub-region branch, a regional sentiment affinity coefficient branch, a single-scale context feature calculation unit and a combination unit; The sub-region branch is used to learn the feature map X of the backbone layer. l The local sub-region representation under different scale divisions, for each scale S k , the feature map X of the backbone layer l Divided into S k ×S k Sub-areas n takes a positive integer, and features are extracted from each sub-region by average pooling and 1×1 convolution; The regional sentiment affinity coefficient branch is used to learn the affinity coefficient weights between sub-regions at the same scale. in, is the regional sentiment affinity coefficient, S k is the division scale, j is the sub-region position, is the calculation function, M l Yes X l After resizing, l is the feature level, g(M l ) is the global information feature; This is achieved through 1×1 convolution and sigmoid activation function. l Using global average pooling, we get the global information feature g(M l ), M l and g(M l ) and calculate the sentiment affinity coefficient of each local position j in the global context; The single-scale context feature calculation unit is used to calculate the single-scale context feature: in, It is size S k The following context features, is a sub-region; The combination unit is used to combine context features of different scales, perform batch normalization after 1×1 convolution, and obtain multi-scale context features Z l : Among them, Z l is a multi-scale context feature, l is the feature level, f is the channel-by-channel combination function, It is a single-scale context feature at different scales.

7. The depression identification method combining multi-level feature dependency and multi-scale contextual features according to claim 6, characterized in that: In step S3, multi-scale emotion representations of each level are obtained by cross-level fusion according to the following formula: 3 , O 4 and O 5 : Among them, l is the characteristic order, O l is the l-level multi-scale sentiment representation, Z l is the multi-scale context feature, M l is the resized representation of the feature map, M l+1 It is the resized representation of the upper feature map.

8. The depression identification method combining multi-level feature dependency and multi-scale contextual features according to claim 1, characterized in that: In step S4, the emotion representation D with significant emotion area is calculated according to the following formula: Among them, D1 is the intermediate feature map, G CAM is the feature map obtained by the channel attention module, C 5 It is the fifth level of emotional representation. is the element-by-element multiplication of matrices, G CSM It is the feature map obtained by the spatial attention module.

Citation Information

Patent Citations

  • Hierarchical phishing website detection method based on deep learning

    CN110602113A

  • assessment method and system for Parkinson's disease gait dyskinesia severity, and equipment

    CN111382679A