Depression degree assessment method based on space-time relevance of facial region
By introducing spatial and temporal correlation evaluation technology of facial area in depression recognition method, dynamically capturing the change patterns and context information of local facial areas is solved, and the problem of low accuracy in the prior art is achieved, and more efficient depression recognition is achieved.
Patent Information
- Application Number
- CN202510515792.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-06-20
AI Technical Summary
The existing depression recognition method based on facial images is not accurate, and ignores the changing patterns of local facial areas and the spatial and temporal context information between facial areas.
A depression degree assessment method based on the spatial and temporal correlation of facial areas is adopted. Through the feature extraction module, the fusion time and spatial attention mechanism, and the temporal feature fusion module, the change pattern of local facial areas is captured dynamically, and depression identification is combined with spatial and temporal context information.
By capturing the spatiotemporal correlation patterns of facial muscles in fine-grained size, the accuracy of depression recognition is improved, and excellent performance on the AVEC2013 and AVEC2014 datasets were achieved.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and medical diagnosis, and particularly relates to a method for evaluating the degree of depression based on the spatio-temporal correlation of facial regions. Background Art
[0002] Depression is a common mental illness, and patients usually show symptoms such as low mood, loss of interest, and sleep disorders. Traditional methods for detecting depression mainly rely on interviews with depression detection scales, which are highly subjective and rely on the cooperation and honest answers of patients, limiting their application scope and diagnostic efficiency. In recent years, researchers have begun to explore automatic depression assessment methods based on objective indicators, especially methods that combine facial expressions and deep learning techniques.
[0003] Existing methods for recognizing depression based on facial images usually process the entire facial image as a whole, ignoring the change patterns of local facial regions and the spatial and temporal context information between facial regions. This method may lead to the loss of local information, thereby affecting the accuracy of depression recognition. Therefore, how to effectively capture the change patterns of local facial regions and combine spatial and temporal context information has become a key issue in improving the accuracy of depression recognition. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for recognizing depression based on facial images to solve the problem of low accuracy of existing depression recognition methods.
[0005] To achieve the above purpose, the present invention provides a method for evaluating the degree of depression based on the spatio-temporal correlation of facial regions, including the following steps: Including the following steps: S1. Preprocess the facial image of the subject, remove the image background and extract the facial region, obtain the facial image sequence of the subject, and divide the facial image sequence into multiple facial sub-regions; S2. Establish a recognition model, which consists of a feature extraction module, a fusion of temporal and spatial attention mechanisms, and a temporal feature fusion module; The feature extraction module is used to extract feature maps from the facial image sequence; the fusion of temporal and spatial attention mechanisms is used to extract the importance of each facial sub-region; the temporal feature fusion module is used to fuse temporal and spatial features; S3. Extract feature maps from the facial image sequence through the feature extraction module, calculate the spatial importance, spatial context, temporal context, and spatial similarity of the extracted feature maps through the fusion of temporal and spatial attention mechanisms, fuse the calculated features through the temporal feature fusion module, and finally perform depression recognition to generate a depression score.
[0006] Furthermore, the fusion of temporal and spatial attention mechanisms includes the following steps: S21. Enhance the facial feature image using the channel attention mechanism. By performing average pooling and max pooling on the feature map two feature vectors are obtained Extract the importance information in the channel dimension, and are respectively input into the multi-layer perceptron MLP to generate channel attention weights where: represents the i-th frame feature map input passing through the j-th feature extraction module, AvgPool c () and MaxPool c () respectively represent global average pooling and max pooling on the feature map in the channel direction, MLP() represents the feature passing through the multi-layer perceptron, and σ represents the Sigmoid activation function; Multiply the channel attention weight by the original feature map to obtain the enhanced feature map S22. Extract the spatial importance information and calculate the spatial importance of each sub-region through convolution operations; Divide the enhanced feature map into R facial sub-regions, R = P×P, and the size of each facial sub-region is Perform average pooling on each facial sub-region to obtain the aggregated features of all facial sub-regions of the feature map
[0007]
[0008] where: represents average pooling with a kernel size of ; Concatenate the maximum and minimum values of the feature map in the channel direction into a 2-channel feature map, and then perform a convolution operation using a 7×7 convolution kernel to obtain the spatial importance weight
[0009] where: AvgPool s () and MaxPool s () respectively represent global average pooling and max pooling on the feature map in the spatial direction, Convcat[] represents the concatenation operation on vectors, and Conv 7×7 () represents a convolution operation with a convolution kernel size of 7×7, and σ represents the Sigmoid activation function; S23. Extract spatial context information and calculate the correlation between different facial sub-regions using the self-attention mechanism; Aggregate the facial sub-region features Reshape them into P×P×d, where d represents the feature dimension of each facial sub-region; Use a fully connected layer to generate the query (Q) and key (K):
[0010] where: FC q () and FC k () represent the fully connected networks for generating the query (Q) and key (K) respectively, and represent the generated query (Q) and key (K) respectively; Calculate the inter-region correlation
[0011] where: Softmax() represents the Softmax function; Use a fully connected layer to generate the spatial context importance
[0012] S24. Extract temporal context information and calculate the importance of each facial sub-region in the entire time series; Combine and reshape the queries (Q) and keys (K) of all time nodes into Q j and K j with the shape of T×P×P×d; Calculate the temporal context correlation
[0013] Use a fully connected layer to generate the temporal context importance TCI j , TCI j =FC(TC j )
[0014] S25. Extract spatial similarity information and calculate the similar facial regions to remove the fixed background and regions with less change; Calculate the similarity matrix Sim_i^j of each facial sub-region,
[0015] where: represents the aggregated features of all facial sub-regions, The transpose of the representative feature matrix , represents obtaining the norm of the feature matrix ; Use a fully connected layer to generate the importance of spatial similarity
[0016] S26. Fuse the spatial importance, spatial context, temporal context, and spatial similarity information to generate the facial attention map Φ j , Φ j = σ(Conv 7×7 (Concat[SI j , SC j , TCI j , ASI j ))
[0017] Multiply the attention map Φ j by the original feature map x j to obtain the enhanced feature map x' j , x' j = x j · Φ j
[0018] Furthermore, the working process of the temporal feature fusion module is as follows: S31. Spatial feature fusion: Fuse the features between different facial regions at the same time node as the emotional representation of the current image. This process is implemented by the self-attention mechanism: Φ sfusion = SelfAttention(x' J )
[0019] S32. Temporal feature fusion: Further fuse the feature Φ sfusion to obtain the representation Φ tfusion of the current image sequence and perform depression assessment. This process is implemented by the self-attention mechanism: Φ tfusion = SelfAttention(Φ sfusion )
[0020] S33. After completing the spatial and temporal feature fusion, fuse the two features and perform feature redirection through a feed-forward network: Add the results of the spatial feature fusion and the temporal feature fusion Φ tfusion to the enhanced feature map x' J to obtain the fused feature x fusion : x fusion = x' J + Φ tfusion
[0021] The fused features are redirected using a two - layer feed - forward network to obtain further enhanced features x'. fusion : x' fusion = x fusion + FC(ReLU(FC(x fusion )))
[0022] Where: ReLU() represents the ReLU activation function; S34. Prediction stage: A convolution with the same size as the convolutional kernel, the width, and the height of the feature map aggregates the features and obtains a unique representation of the time node. A fully - connected network is used to obtain the depression score s of each time node, and the last fully - connected network is used to obtain the prediction y of the final score. pre : y pre = FC(s).
[0023] The complexity of facial expressions stems from the characteristic of multi - region co - variation. To systematically analyze this process, researchers often use the Facial Action Coding System (FACS) to deconstruct and analyze expressions. The Facial Action Coding System (FACS) decomposes facial muscle movements into multiple independent action units (AUs), such as AU12 (raising the corners of the mouth) and AU25 (opening the lips). Each AU corresponds to a muscle activity pattern in a specific facial region (such as the eyebrow - eye region, the perioral region). Research shows that the combination of facial AUs in patients with depression presents significant specificity. For example, the correlation between the inner - eyebrow raising (AU1) and the corner - of - mouth drooping (AU15) is enhanced. Such differences can be quantified by the activation intensity and combination pattern of AUs and become important markers for emotion recognition. Existing research usually regards the data of one time node as a whole for the attention mechanism operation, that is, regarding the whole image as a whole and extracting temporal features. However, this method has two main problems: 1. The local change pattern may be masked by the overall pattern; 2. The spatial and temporal context information between facial regions is ignored. In this case, the importance of the region is completely replaced by the overall importance, which may lead to the loss of local information.
[0024] Regarding the regional characteristics of AUs, the face segmentation technology of the present invention is introduced into the depression recognition model (HARDF) to improve the research accuracy. Specifically: Independent feature extraction: By dividing the face into multiple sub - regions and combining with a convolutional neural network, the AU features of each region can be isolated and extracted, avoiding cross - region signal interference.
[0025] Dynamic association analysis: After segmentation, researchers can more accurately capture the synergistic or antagonistic relationships of AUs under different emotions (such as the co-activation of AU12 and AU6 when smiling), and identify the abnormal AU combinations unique to patients with depression (such as the abnormal association between AU4 and AU17).
[0026] Generally speaking, the face segmentation method used by the depression recognition model (HARDF) provides a high-precision and interpretable technical path for emotion research and pathological diagnosis, and is of great value for the objective assessment of mental diseases such as depression.
[0027] The beneficial effects of the present invention are as follows: The depression recognition model constructed by introducing a feature extraction module, fusing temporal and spatial attention mechanisms, and a temporal feature fusion module dynamically captures the change patterns in local facial regions, and combines spatial and temporal context information, achieving recognition performances of MAE 5.79 (RMSE 8.01) and MAE 5.76 (RMSE 7.62) on the AVEC2013 and AVEC2014 datasets respectively, which are better than the prior art. The present invention provides an efficient and objective solution for automatic depression recognition by finely capturing the spatio-temporal association patterns of facial muscles, improving the accuracy of depression recognition. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 is a schematic diagram of the architecture of the recognition model in the present invention; Figure 2 is the experimental result of different positions of the fused temporal and spatial attention mechanism on the AVEC2013 dataset in Embodiment 1 of the present invention; Figure 3 is the experimental result of different positions of the fused temporal and spatial attention mechanism on the AVEC2014 dataset in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS The present invention will be further described in detail below with reference to the drawings. The embodiments described are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0030] As Figure 1 shown in the architecture diagram of the recognition model, the recognition model is named HARDF (Hierarchical Attention-based Regional Depression Framework), and this model is composed of a feature extraction module, a fused temporal and spatial attention mechanism, which is named FSTA (Fused Spatial and Temporal Attention), and a temporal feature fusion module.
[0031] The feature extraction module is one of the components of the HARDF model. Its main task is to extract effective feature representations from the input facial images for subsequent use by the attention mechanism and the feature fusion module. In the present invention, a deep convolutional neural network (CNN) is used as the basic architecture, and specifically, the ResNet series (such as ResNet18, ResNet34, ResNet50) is used as the Backbone network.
[0032] FSTA is the core module of the HARDF model, aiming to dynamically capture the importance of each sub-region by comprehensively considering the spatial importance, spatial context, temporal context, and spatial similarity of the facial regions, thereby improving the accuracy of depression recognition.
[0033] The temporal fusion module is an important part of the HARDF model. Its main task is to effectively fuse the spatial and temporal features extracted from the FSTA module to generate the final facial expression feature representation for depression recognition. This module captures the dynamic change patterns of facial expressions in the temporal and spatial dimensions by combining the self-attention mechanism and the multi-level feature fusion strategy, thereby improving the recognition performance of the model. The core idea of the temporal fusion module is to fuse the features through the following steps.
[0034] Example 1 Data preprocessing stage The deep learning recognition model HARDF constructed in the present invention is trained and tested based on two public datasets, AVEC2013 and AVEC2014. The OpenFace 2.0 toolbox is used to process all video data. OpenFace 2.0 preprocesses all faces in the video by removing the background, centering the face, and cropping it to a resolution of 224×224. Subsequently, 15 frames are evenly sampled from every 60 frames of the AVEC2013 and AVEC2014 datasets as the input of HARDF.
[0035] Model construction In HARDF, the input X ∈ R T×C×H×W is a combination of multiple-frame facial images, where T = 15 represents the number of input image frames, C = 3 represents the number of image channels, and H = 224, W = 224 represent the height and width of the image respectively. The process of feature extraction of the image features through the BackBone Block can be expressed as: where represents the j-th Backbone Block, represents the feature x corresponding to the i-th frame of the face iThe features obtained through the j-th Backbone Block, where j ∈ {0, 1, ..., J}.
[0036] Fuse the temporal and spatial attention mechanisms The attention intensity of the facial region should be jointly determined by spatial attention, spatial context relationship, temporal context relationship, and facial region similarity. Therefore, FSTA is added in HARDF to extract enhanced facial features. The input of FSTA is the feature after being extracted by the Backbone.
[0037] Inspired by CBAM, the channel attention mechanism is first used to enhance individual features. The channel attention mechanism first performs average pooling and max pooling on the feature map, and then the feature is fed into a multi-layer perceptron for attention extraction. The multi-layer perceptron first compresses the number of channels to the original and then restores it to the original size. For the map the process of extracting its channel attention is as follows:
[0038] where AvgPool s (.) and MaxPool s (.) represent taking the average value and the maximum value of the feature channels respectively, MLP represents the multi-layer perceptron, and σ(.) represents the Sigmoid activation function. Subsequently, the feature is enhanced on the channel:
[0039] Subsequently, spatial importance, spatial context information, temporal context information, and spatial similarity features are extracted simultaneously. During the process of using the Backbone for feature extraction, its low-order feature map has a large size. If pixel-level features are used to extract the above information, it will greatly increase the model parameters and computational complexity. Therefore, inspired by VIT, the facial region is divided into R = P × P patches, where P represents the number of divisions in the feature width and height, and the aggregated feature of the patch is used as the feature of the current region. For the aggregation process is as follows:
[0040] where represents average pooling with a pooling window size of .
[0041] Spatial importance information is used to extract the more important regional features at a certain time node. In FSTA, first, the maximum and minimum values of the feature map in the channel direction are concatenated into a feature map spectrum of 2 channels, and then the spatial importance is obtained by using a convolution with a convolution kernel size of 7×7.
[0042] For The specific operations are as follows:
[0043] Among them, AvgPool s (.) and MaxPool s (.) respectively represent taking the average value and the maximum value of the feature map spectrum, and Conv 7×7 (·) represents the convolution operation with a convolution kernel of 7×7.
[0044] Spatial context information is mainly used to obtain the correlation relationships between regions. The Query (Q) and Key (K) in the self-attention mechanism are used to calculate the correlation between two regions. First, it is reset to R (p×p)×d . First is projected into a and :
[0045]
[0046] Among them, FC q (·) and FC k (·) respectively represent the fully connected neural networks for generating Query and Key. Subsequently, the spatial context correlation relationship is calculated:
[0047] Among them, the following d generations are size factors, and Softmax(·) represents the Softmax function. Finally, a fully connected layer is used to generate the regional importance SCI∈R from the spatial context p×p :
[0048] Temporal context information is used to obtain the importance of a certain region in the entire time series. For the data of the entire time series, there are
[0049] First, Q j and K j are respectively reset to a size of T×P×P×d, and subsequently, the temporal context information of all regions is calculated:
[0050] Finally, a fully connected layer is used to generate the region importance TCI ∈ R from the temporal context P×P : TCI j = FC(TC j )
[0051] Spatial similarity information is used to obtain regions with similar faces to remove fixed backgrounds and regions with less variation. For features The method for calculating the similarity of each region can be expressed as:
[0052] Subsequently, a fully connected layer is used to generate the importance map of the region from the region similarity:
[0053] Finally, all importance maps are fused to obtain the face attention map. SCI, TCI, and ASI are first reset to feature maps of P×P, and then a convolution with a convolutional kernel set to 7×7 is used to obtain the region attention map and redefine the features: Φ j = σ(conv 7×7 (concat[SI j , SC j , TCI j , ASI j ))
[0054] x’ j = x j * Φ j
[0055] Feature Fusion In this embodiment, two strategies of feature fusion and decision fusion are used for temporal feature fusion and final decision respectively. Feature fusion is divided into two parts: temporal fusion and spatial fusion. Since the width and height of the feature map x′ J ∈ R (T×C×H×W) generated by the last Backbone are small enough, fine-grained features are directly used for feature fusion. For the fusion of spatial information, x′ J is first reset to a size of T×(H×W)×C to obtain the region feature representation, and then the self-attention mechanism is used for spatial feature fusion: Φ sfusion = SelfAttention(x′ J )
[0056] For Φ sfusion The self-attention mechanism is further used for temporal feature fusion to obtain the temporal feature representation: Φ tfusion = SelfAttention(Φ sfusion )
[0057] x fusion = x' J + Φ tfusion
[0058] Finally, a two-layer feed-forward network is used for feature redirection: x' fusion = x fusion + FC(ReLU(FC(x fusion )))
[0059] In the prediction stage, a convolution with the same size as the convolutional kernel and the width and height of the feature map aggregates the features and obtains a unique representation of the time node, and then a fully connected network is used to obtain the depression score s ∈ R for each time node T Finally, a fully connected network is used to obtain the final prediction y pre : y pre = FC(s)
[0060] Loss function The SmoothL1Loss is used to optimize the model. For the true label y true and the predicted label y pre , the calculation method of SmoothL1Loss is as follows:
[0061] By default, β = 1
[0062] Experimental verification The present invention is verified on the AVEC2013 and AVEC2014 data sets, and the experimental results are as Figures 2 - 3As shown, for the AVEC2013 and AVEC2014 datasets, 15 frames are evenly sampled from every 60 frames as the input of HARDF, and ResNet18, ResNet34, and ResNet50 are used as the Backbone modules to extract features respectively. The HARDF model is constructed under the Pytorch framework. The pre-trained Backbone modules adopt the ImageNet-1K pre-trained weights provided by Pytorch. All models are trained on Tesla V100 GPUs, optimized using the Adam optimizer, and the learning rate is set to 1e-4. For the batch size, ResNet18 and ResNet34 are set to 32, ResNet50 is set to 8, and they are trained on 4 GPU cards. The HARDF using the non-pre-trained model is trained for 100 epochs, while the HARDF using the pre-trained model is trained for 20 epochs.
[0063] The experimental results are shown in Table 1: the experimental results of HARDF on the AVEC2013 and AVEC2014 datasets. For HARDF, the results of MAE and RMSE are 5.89 and 7.69 respectively on the AVEC2013 dataset, and 5.69 and 7.50 on the AVEC2014 dataset.
[0064] At the same time, the position of the FSTA module in the model of the present invention also affects the experimental results. The present invention adjusts the insertion position of FSTA in the model and conducts an analysis. The experimental results are as Figure 2 and Figure 3 shown. When combining two FSTAs, the HARDF containing FSTA2 and FSTA4 achieves the optimal results (MAE: 5.79, RMSE: 8.01) on the AVEC2013 dataset, while the HARDF containing FSTA1 and FSTA4 achieves the optimal results (MAE: 5.78, RMSE: 7.87) on the AVEC2014 dataset.
[0065] Table 1
Claims
1. A method for assessing depression based on spatiotemporal correlation of facial regions, characterized in that: The following steps are involved: S1, preprocessing the facial image of the subject, removing the image background and extracting the facial region, obtaining the facial image sequence of the subject, and dividing the facial image sequence into multiple facial sub-regions; S2. Establish a recognition model, which consists of a feature extraction module, a fusion time and space attention mechanism, and a temporal feature fusion module; The feature extraction module is used to extract feature maps from facial image sequences; the fusion of temporal and spatial attention mechanisms is used to extract the importance of each facial sub-region; The temporal feature fusion module is used to fuse temporal and spatial features; S3. Feature extraction module is used to extract feature maps from facial image sequences. The spatial importance, spatial context, temporal context and spatial similarity of the extracted feature maps are calculated by fusing the temporal and spatial attention mechanisms. The calculated features are fused by the temporal feature fusion module. Finally, depression is identified and a depression score is generated.
2. The method for evaluating depression level based on spatiotemporal correlation of facial regions according to claim 1, characterized in that: The fusion time and space attention mechanism includes the following steps: S21. Use the channel attention mechanism to enhance the facial feature image. Perform average pooling and maximum pooling to obtain two feature vectors Extract the importance information on the channel dimension and Input the multi-layer perceptron MLP respectively to generate channel attention weights in: Indicates that the input i-th frame feature map passes through the j-th feature extraction module, AvgPool c () and MaxPool c () respectively represent the global average pooling and maximum pooling of the feature map in the channel direction, MLP() represents the feature passing through the multi-layer perceptron, σ represents the Sigmoid activation function; the channel attention weight With the original feature map Multiply to get the enhanced feature map S22, extract the spatial importance information, calculate the spatial importance of each sub-region through convolution operation; Divided into R facial sub-regions, R = P × P, the size of each facial sub-region is Perform average pooling on each facial sub-region to obtain a feature map Aggregate features of all facial sub-regions in: The kernel size is Average pooling; the maximum and minimum values of the feature map in the channel direction are concatenated into a feature map of two channels, and then a 7×7 convolution kernel is used for convolution operation to obtain the spatial importance weight Where: AvgPool s () and MaxPool s () respectively represent the global average pooling and maximum pooling of the feature map in the spatial direction, Concat[] represents the concatenation operation of the vector, Conv 7×7 () represents the convolution operation with a kernel size of 7×7, and σ represents the Sigmoid activation function; S23, extract spatial context information and use the self-attention mechanism to calculate the correlation between different facial sub-regions; Aggregate features into facial sub-regions Reshape into P×P×b, where d represents the feature dimension of each facial sub-region; use a fully connected layer to generate the query (Q) and key (K): Among them: FC q () and FC k () represent the fully connected networks used to generate queries (Q) and keys (K), and Represent the generated query (Q) and key (K) respectively; calculate the correlation between regions Where: Softmax() represents the Softmax() function; a fully connected layer is used to generate spatial context importance S24, extract temporal context information, calculate the importance of each facial sub-region in the entire time series; combine and reshape the query (Q) and key (K) of all time nodes into a Q with a shape of T×P×P×d j and K j ; Calculate temporal context relevance Generate temporal context importance TCI using fully connected layers j , TCI j =FC(TC j ); S25, extracting spatial similarity information, calculating facial similarity regions to remove fixed background and regions with small changes; calculating the similarity matrix Sim_i^j of each facial sub-region, in: represents the aggregated features of all facial sub-regions, Representative feature matrix The transpose of Represents the feature matrix The modulus of ; using a fully connected layer to generate spatial similarity importance S26, integrate spatial importance, spatial context, temporal context and spatial similarity information to generate facial attention map Φ j , Φ j =σ(Conv 7×7 (Concat[SI j , S.C. j , TCI j , ASI j ])) The attention map Φ j With the original feature map x j Multiply to get the enhanced feature map x′ j , x′ j =x j ·Φ j .
3. The method for evaluating depression level based on spatiotemporal correlation of facial regions according to claim 1, characterized in that: The working process of the temporal feature fusion module is as follows: S31, spatial feature fusion: The features of different facial regions at the same time node are fused as the emotional representation of the current image. This process is implemented by the self-attention mechanism: Φ sfusion =SelfAttention(x′ J ); S32, time feature fusion: further integrate the feature Φ sfusion Fusion is performed to obtain the representation Φ of the current image sequence tfusion And depression assessment is performed, which is implemented by the self-attention mechanism: Φ tfusion =SelfAttention(Φ sfusion ); S33. After completing the fusion of spatial and temporal features, the two features are fused and feature redirection is performed through a feedforward network: The result of spatial feature fusion and temporal feature fusion Φ tfusion And the enhanced feature map x′ J Add together to get the fused feature x fusion : x fusion = x′ j +Φ tfusion Use a two-layer feedforward network to redirect the fused features to obtain further enhanced features x′ fusion :x′ fusion =x fusion +FC(ReLU(FC(x fusion )))Where: ReLU() represents the ReLU activation function; S34, prediction stage: A convolution with the same size as the convolution kernel and the width and height of the feature map aggregates the features and obtains a unique representation of the time node. A fully connected network is used to obtain the depression score s at each time node. The last fully connected network is used to obtain the prediction y of the final score. pre : y pre = FC(s).
Citation Information
Cited By
Dynamic space and spiral Mama fused depression image detection method
CN121482048A