Depression image detection method fusing Gaussian pyramid and spatial domain

By integrating the detection method of Gaussian pyramid and spatial domain attention module, the shortcomings of multi-scale features and adaptive perception of key areas in facial expression analysis in the existing technology are solved, and high-precision depression image detection is achieved.

CN120656219APending Publication Date: 2025-09-16ANHUI IND TECH INNOVATION RES INST +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510660625.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies in facial expression analysis use a single-scale feature extraction strategy that ignores the diversity and dynamics of multi-scale features. Local and global features are extracted independently, and there is a lack of collaborative modeling, resulting in insufficient detection accuracy and robustness, and an inability to effectively identify adaptive perception of key facial areas.

Method used

A detection method that integrates Gaussian pyramid and spatial domain attention modules is adopted. Through primary feature extraction, fusion module, Gaussian pyramid attention module and spatial domain attention module, dynamic perception of multi-scale features and adaptive perception of key areas are achieved, thereby enhancing the integrity and discriminability of feature expression.

Benefits of technology

The accuracy and robustness of depression image detection are improved, and the model's ability to recognize complex changes in facial expressions is enhanced. It has strong applicability and low computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656219A_ABST
    Figure CN120656219A_ABST
Patent Text Reader

Abstract

The invention provides a depression image detection method fusing a Gaussian pyramid and a spatial domain. The method comprises the following steps: acquiring an original face image to be processed; the trained detection model is used for processing the original face image to obtain a detection result, and the detection model comprises a primary feature extraction module, a fusion module, a Gaussian pyramid attention module, a spatial domain attention module and a post-processing module. A Gaussian pyramid attention module is constructed and used for capturing the dynamics and diversity of multi-scale features, so that the perception ability of the model to different-scale information is enhanced; further mining significant features in the image by introducing a spatial domain attention module and based on a self-adaptive perception mechanism of a key region; and finally, through deep fusion of a Gaussian pyramid attention module and a spatial domain attention module, the integrity and discrimination of feature expression are improved, so that the detection precision of the detection model is effectively improved, and the model is low in calculation complexity and high in applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a depression image detection method integrating Gaussian pyramid and spatial domain. Background Art

[0002] With the widespread application of image data in mental health analysis, automated detection of depression has become a research focus. In recent years, facial detection models have made significant progress in the field of facial expression analysis and have demonstrated certain advantages. However, existing technologies still have obvious limitations.

[0003] Existing methods typically employ a single-scale feature extraction strategy, ignoring the diversity and dynamic nature of multi-scale features in facial expressions. Furthermore, the extraction processes for local and global features are relatively independent, lacking an effective collaborative modeling mechanism. This results in insufficient performance in capturing the complex variations in facial expressions. Furthermore, existing technologies have limited ability to focus on key facial regions and fail to achieve adaptive perception of these regions, thereby reducing the model's accuracy and robustness in complex situations. Therefore, there is an urgent need for a facial detection method that combines multi-scale features with adaptive perception of key regions to enhance the model's comprehensive representation and accurate recognition of complex facial expressions. Summary of the Invention

[0004] In view of the above defects of the prior art, the present invention provides a depression image detection method that integrates Gaussian pyramid and spatial domain to solve the technical problems of low detection accuracy and poor robustness in the prior art.

[0005] To achieve the above-mentioned purpose and other related purposes, the present invention provides a method for detecting depression images by fusing Gaussian pyramid and spatial domain, comprising: obtaining an original facial image to be processed; processing the original facial image using a trained detection model to obtain a detection result, wherein the detection model includes a primary feature extraction module, a fusion module, a Gaussian pyramid attention module, a spatial domain attention module and a post-processing module, and the processing steps of the detection model specifically include: using the primary feature extraction module to process the original facial image to obtain global facial features; using the fusion module to divide and fuse the global facial features to obtain fused features; using the Gaussian pyramid attention module to process the fused features to obtain a first attention feature; using the spatial domain attention module to process the fused features to obtain a second attention feature; using the post-processing module to process the first attention feature and the second attention feature to obtain a detection result.

[0006] In one embodiment of the present invention, the primary feature extraction module is a ResNet-18 network.

[0007] In one embodiment of the present invention, the global facial features are divided and fused to obtain fused features, including: dividing the global facial features into multiple local facial features; upsampling each local facial feature so that its size is the same as the size of the global facial feature; and weightedly fusing all the upsampled local facial features and the global facial features according to a first preset weight to obtain the fused features.

[0008] In one embodiment of the present invention, dividing the global facial feature into a plurality of local facial features includes dividing the global facial feature into four local facial features: upper left, upper right, lower left, and lower right.

[0009] In one embodiment of the present invention, the fused features are processed to obtain a first attention feature, including: continuously downsampling the fused features to obtain a plurality of first features; upsampling each first feature until its size is restored to the size of the fused feature to obtain a plurality of second features; applying a Sigmoid activation function and fusion to each of the second features to obtain a first attention weight; and multiplying the first attention weight by the fused feature element by element to obtain the first attention feature.

[0010] In one embodiment of the present invention, the fusion feature is continuously downsampled to obtain a plurality of first features, which are expressed as follows: x1 = AvgPool (x0); x i =AvgPool(x i-1 ), i∈{2,3,…,L}; where x0 is the fusion feature, AvgPool is the average pooling, and its pooling window size is 2×2 and the stride is 2.

[0011] In one embodiment of the present invention, the fused feature is processed to obtain a second attention feature, including: performing a depth-wise separable convolution operation and a standard convolution operation on the fused feature in sequence to obtain a third feature; applying a Sigmoid activation function to the third feature to obtain a second attention weight; and multiplying the second attention weight by the fused feature element-by-element to obtain the second attention feature.

[0012] In one embodiment of the present invention, the first attention feature and the second attention feature are processed to obtain a detection result, including: fusing the first attention feature and the second attention feature; passing the fused features through global average pooling and a fully connected layer in sequence to obtain a depression prediction score, which is the detection result.

[0013] To achieve the above-mentioned purpose and other related purposes, the present invention also provides an electronic device, including a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute the computer program stored in the memory to implement the method provided in any one of the above embodiments.

[0014] To achieve the above-mentioned object and other related objects, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used to enable a computer to execute the method provided in any one of the above-mentioned embodiments.

[0015] Beneficial effects of the present invention: The present invention proposes a depression image detection method that integrates Gaussian pyramid and spatial domain. The method constructs a Gaussian pyramid attention module to capture the dynamics and diversity of multi-scale features, thereby enhancing the model's perception of information at different scales; and by introducing a spatial domain attention module, based on the adaptive perception mechanism of key areas, further explores significant features in the image; finally, through the deep integration of the Gaussian pyramid attention module and the spatial domain attention module, the integrity and discrimination of feature expression are improved, thereby effectively improving the detection accuracy of the detection model, and the model has low computational complexity and strong applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 A processing flow chart of a detection model provided in one embodiment of the present invention;

[0018] Figure 2 This is an architectural diagram of a detection model provided by one embodiment of the present invention;

[0019] Figure 3 A detailed flow chart of step S200 provided in one embodiment of the present invention;

[0020] Figure 4 A detailed flow chart of step S300 provided in one embodiment of the present invention;

[0021] Figure 5 An architectural diagram of a Gaussian pyramid attention module according to an embodiment of the present invention;

[0022] Figure 6 A detailed flowchart of step S400 provided in one embodiment of the present invention;

[0023] Figure 7 This is an architectural diagram of a spatial domain attention module provided by one embodiment of the present invention;

[0024] Figure 8 A detailed flow chart of step S500 provided in one embodiment of the present invention;

[0025] Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present invention.

[0026] Description of the accompanying drawings: 101, processor; 102, memory. DETAILED DESCRIPTION

[0027] The following describes the embodiments of the present invention through specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. It should be noted that the following embodiments and the features in the embodiments can be combined with each other unless they conflict. In addition to the specific methods, equipment, and materials used in the embodiments, based on the understanding of the prior art by those skilled in the art and the description of the present invention, any methods, equipment, and materials of the prior art that are similar or equivalent to the methods, equipment, and materials in the embodiments of the present invention can also be used to implement the present invention.

[0028] It should be understood that the terms used in the examples of the present invention are for describing specific embodiments rather than for limiting the scope of protection of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those generally understood by those skilled in the art.

[0029] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of the embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0030] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations that may be implemented by the methods and computer program products of various embodiments disclosed in the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0031] See Figure 1 , Figure 1 A method for detecting depression images by integrating Gaussian pyramid and spatial domain is provided in one embodiment of the present invention, comprising the following steps: (1) obtaining an original facial image to be processed; and (2) processing the original facial image using a trained detection model to obtain a detection result. In step (1), a facial image of a person can be captured by a camera, and the image can be, for example, in RGB or other formats. It is understandable that the obtained original facial image needs to undergo certain preprocessing to adapt it to the input of the detection model. The preprocessing can include, for example, size adjustment, denoising, normalization, etc. When making a prediction, the detection model is trained, and its training process is roughly the same as that of a traditional neural network model, including steps such as constructing a data set, initializing the detection model, training the detection model using the training set, and evaluating the indicators of the trained detection model.

[0032] See Figure 2 In the present invention, the detection model includes a primary feature extraction module, a fusion module, a Gaussian pyramid attention module, a spatial domain attention module and a post-processing module (not shown in the figure). The processing steps of the detection model, namely the above-mentioned step (2), specifically include steps S100 to S500.

[0033] Step S100: Process the original facial image using the primary feature extraction module to obtain global facial features. This step mainly provides overall image information for subsequent processing.

[0034] In one embodiment of the present invention, the primary feature extraction module is a ResNet-18 network. As one of the lightest members of the ResNet family, the ResNet-18 network is specifically designed for computer vision tasks such as image classification. By introducing skip connections, it effectively solves the vanishing / exploding gradient problem in deep neural networks, making it possible to train deeper networks.

[0035] Step S200: Using the fusion module, the global facial features are divided and fused to obtain fused features. In this step, by dividing and fusion of the global facial features, the ability to express local detail features is significantly enhanced while maintaining global semantic information.

[0036] See Figure 3 In a specific embodiment of the present invention, the fusion model divides and fuses the global facial features through steps S201 to S203 to obtain fusion features.

[0037] Step S201: Divide the global facial features into a plurality of local facial features. There are various ways to divide the global facial features. In one embodiment of the present invention, the global facial features are divided into four local facial features: upper left, upper right, lower left, and lower right. Other methods of division, such as 2×3 or 3×3, may also be used.

[0038] Step S202: Upsample each local facial feature to make its size the same as the global facial feature size.

[0039] Step S203: performing weighted fusion on all local facial features after upsampling and the global facial features according to a first preset weight to obtain a fused feature.

[0040] Through the above steps S201 to S203, the global facial features can be easily divided and fused to obtain fusion features. Figure 2 It can also be seen in.

[0041] Step S300: Process the fused features using the Gaussian Pyramid Attention Module to obtain a first attention feature. The Gaussian Pyramid Attention Module combines multi-scale feature extraction with an attention mechanism and is commonly used to improve a model's ability to perceive multi-scale objects. Its core concept is to construct multi-scale features using a Gaussian pyramid and then dynamically fuse key information from different scales using an attention mechanism.

[0042] See Figure 4In a specific embodiment of the present invention, the Gaussian pyramid attention module processes the fusion features through steps S301 to S304 to obtain the first attention feature. Figure 5 As shown below, we will combine Figure 5 Come to Figure 4 The steps in are described below.

[0043] Step S301: Perform continuous downsampling operations on the fused features to obtain a number of first features. In this step, the feature maps x of the first, second, third, ... layers of the Gaussian pyramid are obtained through continuous downsampling operations. i , i∈{1,2,…,L}, L is the total number of layers.

[0044] In a specific embodiment of the present invention, average pooling AvgPool is used for downsampling, and its pooling window size is 2×2 and the stride is 2. For the feature map x1 of the first layer of the Gaussian pyramid, it can be expressed as follows: x1=AvgPool(x0), where x0 is the fusion feature; for the feature maps x1 of other layers of the Gaussian pyramid, i For example, it can be expressed as: i =AvgPool(x i-1 ), i∈{2,3,…,L}.

[0045] Step S302: Upsample each first feature until its size is restored to the size of the fused feature, to obtain a plurality of second features.

[0046] Step S303: Apply the Sigmoid activation function and fusion to each second feature to obtain the first attention weight.

[0047] Step S304: multiply the first attention weight by the fusion feature element by element to obtain the first attention feature.

[0048] The above steps S302 to S304 can be expressed together using the formula:

[0049]

[0050] In the formula, x0 is the input image or fusion feature, the dimension is B×C×H×W, Interp(x i ,H,W) represents the pyramid x i Upsample to size H×W, Conv i and BN i are the convolutional layer and batch normalization layer of each layer, σ is the Sigmoid activation function, represents element-by-element multiplication, and out1 is the output feature of the Gaussian pyramid attention module.

[0051] Step S400: Use the spatial attention module to process the fused features to obtain a second attention feature. The spatial attention module is a mechanism used to enhance neural networks' modeling of the importance of spatial locations in an image or feature map. It dynamically generates spatial weight maps to highlight key areas and suppress irrelevant background. It is widely used in tasks such as object detection, semantic segmentation, and image generation.

[0052] See Figure 6 In a specific embodiment of the present invention, the spatial domain attention module processes the fusion features through steps S401 to S403 to obtain the second attention feature. The architecture diagram of the spatial domain attention module is as follows: Figure 7 As shown below, we will combine Figure 7 Come to Figure 6 The steps in are described below.

[0053] Step S401: Perform depthwise separable convolution and standard convolution on the fused features to obtain a third feature. The convolution kernel of the depthwise separable convolution is 3×3, and the convolution kernel of the standard convolution is also 3×3.

[0054] Step S402: Apply a Sigmoid activation function to the third feature to map the attention value to between [0, 1] to obtain a second attention weight.

[0055] Step S403: Multiply the second attention weight by the fusion feature element by element to obtain the second attention feature.

[0056] The above steps S401 to S403 can be expressed as follows:

[0057]

[0058] Where DepthwiseConv is the depthwise separable convolution, Conv is the standard convolution, f is the Sigmoid activation function, represents element-by-element multiplication, and out2 is the output feature of the spatial domain attention module.

[0059] Step S500: Use the post-processing module to process the first attention feature and the second attention feature to obtain a detection result. Figure 2 It is not illustrated in the figure. It is mainly used to fuse the first attention feature and the second attention feature, and process these features to obtain the detection result.

[0060] See Figure 8In one embodiment of the present invention, the post-processing module performs processing through the following steps: S501: Fusion of the first and second attention features; S502: Processing the fused features sequentially through global average pooling and a fully connected layer to obtain a depression prediction score, which is the detection result. In this step, the fully connected layer outputs a score corresponding to the depression prediction score. The score represents the degree of depression; for example, a higher score indicates a more severe depression.

[0061] This step can be expressed as: score =F(GAP(M(out1,out2))); where y score That is, the depression prediction score output by the post-processing module, F is the fully connected layer, GAP is the global average pooling operation, and M is the Concat operation.

[0062] It should be noted that the step division of the various methods above is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they contain the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0063] See Figure 9 , Figure 9 An electronic device provided in one embodiment of the present invention includes a processor 101, a memory 102 and a communication bus; the communication bus is used to connect the processor 101 and the memory 102; the processor 101 is used to execute a computer program stored in the memory 102 to implement the above-mentioned detection method.

[0064] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is used to enable a computer to execute the above detection method.

[0065] In general, the present invention aims to provide a non-invasive intelligent detection method for depression as a preliminary screening for depression. The method first performs primary feature extraction on the original facial image to obtain global facial features, and divides them into multiple local area features. After upsampling to restore the original size, they are weighted fused with the global facial features according to preset weights; then, the multi-scale characteristics of the Gaussian pyramid and the spatial domain attention module are combined to achieve adaptive perception and multi-scale feature extraction of key facial areas, thereby capturing subtle emotional changes related to depression. By passing the fused feature map to the global average pooling layer and the fully connected layer, accurate prediction of the depression detection score is achieved. This method not only overcomes the shortcomings of traditional single-scale attention in dealing with image diversity and dynamics, but also provides new technical support and theoretical basis for early diagnosis of depression and intelligent mental health monitoring.

[0066] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. A depression image detection method integrating Gaussian pyramid and spatial domain, characterized in that: include: Obtaining the original facial image to be processed; The original facial image is processed using the trained detection model to obtain a detection result, wherein the detection model includes a primary feature extraction module, a fusion module, a Gaussian pyramid attention module, a spatial domain attention module, and a post-processing module. The processing steps of the detection model specifically include: Processing the original facial image using the primary feature extraction module to obtain global facial features; Using the fusion module, dividing and fusing the global facial features to obtain fused features; Using the Gaussian pyramid attention module, the fused feature is processed to obtain a first attention feature; Using the spatial domain attention module, the fused features are processed to obtain a second attention feature; The post-processing module is used to process the first attention feature and the second attention feature to obtain a detection result.

2. The depression image detection method integrating Gaussian pyramid and spatial domain according to claim 1, characterized in that: The primary feature extraction module is a ResNet-18 network.

3. The depression image detection method integrating Gaussian pyramid and spatial domain according to claim 1, characterized in that: The global facial features are divided and fused to obtain fused features, including: dividing the global facial features into a plurality of local facial features; Upsampling each local facial feature so that its size is the same as the size of the global facial feature; All local facial features after upsampling processing and the global facial features are weightedly fused according to a first preset weight to obtain the fused feature.

4. The depression image detection method integrating Gaussian pyramid and spatial domain according to claim 3, characterized in that: Dividing the global facial features into a plurality of local facial features, including: The global facial features are divided into four local facial features: upper left, upper right, lower left and lower right.

5. The depression image detection method integrating Gaussian pyramid and spatial domain according to claim 1, characterized in that: The fused features are processed to obtain a first attention feature, including: Performing a continuous downsampling operation on the fused features to obtain a plurality of first features; Upsampling each first feature until its size is restored to the size of the fused feature, thereby obtaining a plurality of second features; Applying a Sigmoid activation function and fusion to each of the second features to obtain a first attention weight; Multiply the first attention weight by the fusion feature element by element to obtain the first attention feature.

6. The depression image detection method integrating Gaussian pyramid and spatial domain according to claim 5, characterized in that: The fusion features are continuously downsampled to obtain several first features, which can be expressed as follows: x1 = AvgPool(x0); x i =AvgPool(x i-1 ),i∈{2,3,…,L}; Wherein, x0 is the fusion feature, AvgPool is the average pooling, and its pooling window size is 2×2 and the stride is 2.

7. The depression image detection method integrating Gaussian pyramid and spatial domain according to claim 1, characterized in that: The fused features are processed to obtain a second attention feature, including: Performing a depth-wise separable convolution operation and a standard convolution operation on the fused features in sequence to obtain a third feature; Applying a Sigmoid activation function to the third feature to obtain a second attention weight; Multiply the second attention weight by the fusion feature element by element to obtain the second attention feature.

8. The depression image detection method integrating Gaussian pyramid and spatial domain according to claim 1, characterized in that: Processing the first attention feature and the second attention feature to obtain a detection result includes: fusing the first attention feature and the second attention feature; The fused features are sequentially passed through the global average pooling and fully connected layers to obtain a depression prediction score, which is the detection result.

9. An electronic device, characterized in that: The system comprises a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program is used to enable a computer to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Dynamic space and spiral Mama fused depression image detection method

    CN121482048A