Multi-modal data fused depression detection method

Through the DBGNet model combining scale text and heart rate variability data, the information limitations of traditional depression detection methods are solved, efficient and objective auxiliary diagnosis of depression is achieved, and the detection accuracy and efficiency are improved.

CN120236741APending Publication Date: 2025-07-01HEBEI AGRICULTURAL UNIV.
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302090.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Traditional depression detection methods rely on subjective scales or single physiological indicators, which have information limitations, resulting in low diagnostic accuracy, especially in low- and middle-income countries without diagnosis or high treatment rates.

Method used

The multimodal depression detection model was adopted, combined with the scale text data and heart rate variability data, and the features were extracted through the BMGNet and DIRNet models, and the cross-attention mechanism was used to fusion to construct a multimodal depression detection system.

Benefits of technology

It significantly improves the accuracy of depression detection, reduces the dependence on the subjective experience of professional physicians, provides technical support for patients' self-screening and clinical accurate diagnosis, and improves the objectivity and efficiency of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236741A_ABST
    Figure CN120236741A_ABST
Patent Text Reader

Abstract

The invention discloses a depression detection method fusing multi-modal data, and belongs to the technical field of medical information processing, and the depression detection method comprises the following steps: S1, collecting scale text data and heart rate variability data of a depression subject; s2, labeling the data collected in the S1 by adopting a label to obtain a label data set; s3, preprocessing the label data set; and S4, constructing a DBGNet multi-mode depression detection model to detect the preprocessed label data set, and outputting a depression detection result. The problems that a traditional scale depends on subjectivity and single-mode data information is limited are solved, the detection precision is remarkably improved, and efficient and reliable technical support is provided for clinical auxiliary diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical information processing, and particularly relates to a depression detection method that fuses multi-modal data. Background Art

[0002] The "Major Depression Report of the World Psychiatric Association" published in The Lancet pointed out that depression has become a global health crisis, especially among adolescents and middle-aged people, with a significant increase in the prevalence rate. According to the latest epidemiological survey results of mental diseases in China, it is estimated that at least 30 million people are troubled by depressive disorders every year.

[0003] In daily life, stress and mood swings are common, but they are essentially different from the persistent low mood and lack of pleasure in patients with depression. The symptoms of depression include persistent sadness, loss of interest, fatigue, sleep problems, appetite changes, and inattention, etc., and may lead to suicide in severe cases. Globally, the prevalence rate and fatality rate of depression are high. About half of the patients in high-income countries are undiagnosed or untreated, and this proportion is as high as 80% to 90% in low- and middle-income countries. In China, the number of teenagers taking leave from school or even committing suicide due to depression has increased, highlighting the severe situation of depression. However, the investment in mental health in China is less than 2% of the health budget, reflecting the insufficient attention of society to depression.

[0004] With the development of artificial intelligence technology, emerging technologies such as deep learning have gradually been applied to the medical field, including depression detection. Traditional diagnosis relies on scales or questionnaires, but patients may conceal their true state due to memory bias or fear of discrimination, affecting early diagnosis. Therefore, an automatic depression detection system has emerged, and researchers have achieved remarkable results by analyzing multi-modal data such as text, voice, facial expressions, and physiological signals.

[0005] Based on this, a depression detection method that fuses multi-modal data is proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a depression detection method that fuses multi-modal data to solve the problems in the background art.

[0007] To achieve the above purpose, the present invention provides a depression detection method that fuses multi-modal data, including the following steps:

[0008] S1. Collect the scale text data and heart rate variability data of depression subjects;

[0009] S2. According to the conversion score of the scale text data and the characteristics of the heart rate variability data, use labels to annotate the data collected in S1 to obtain a labeled data set;

[0010] The labels include two categories: depression and non-depression;

[0011] S3. Preprocess the labeled dataset;

[0012] S4. Construct the DBGNet multimodal depression detection model to detect the preprocessed labeled dataset and output the depression detection result.

[0013] Preferably, in S1, the scale text data is collected through questionnaires, social media data, or interview content; the heart rate variability data is collected in real time through wearable devices.

[0014] Preferably, the conversion score in S2 is expressed as:

[0015]

[0016] In the formula, C is the conversion score, M is the raw score, min(M) is the minimum value of the raw score, and max(M) is the maximum value of the raw score.

[0017] Preferably, the raw score is expressed as:

[0018]

[0019] In the formula, P i is the score of the i-th question, and j is the total number of questions.

[0020] Preferably, the preprocessing in S3 includes:

[0021] 1) Word segmentation, stop word removal, and text vectorization processing of the scale text data;

[0022] 2) Denoising and normalization processing of the heart rate variability data.

[0023] Preferably, in S4, the specific detection process of the DBGNet multimodal depression detection model is as follows:

[0024] S41. Respectively construct the BMGNet and DIRNet models for the scale text data and heart rate variability data in S3 to extract features, and obtain text features and heart rate variability features;

[0025] S42. Input the text features and heart rate variability features into the multimodal interaction module based on the cross-attention mechanism for feature interaction, splice the two interacted features, and then classify them through the Softmax layer to obtain the depression detection result.

[0026] Preferably, in S41, the specific process of using the BMGNet model to extract text features is as follows: For the scale text data in S3, use the Bert pre-trained model with the full-word masking technique to generate word vectors, and then use Bi-MGRU to extract the text features.

[0027] Preferably, in S41, the DIRNet model includes: a convolutional layer, a depthwise separable convolutional layer, a Block residual stacking module, a Mish activation function, and an ECA attention mechanism; the Block residual stacking module is a combination of asymmetric convolution and grouped convolution.

[0028] Preferably, the convolutional layer is three 3×3 convolutions;

[0029] The Mish activation function is expressed as:

[0030] Mish(x) = x × Tanh(Softplus(x));

[0031] In the formula, Mish(x) represents the Mish activation function, x is the input, Tanh(·) represents the hyperbolic tangent function, and Softplus(·) represents the Softplus function.

[0032] Therefore, the depression detection method of the present invention that fuses multimodal data has the following beneficial effects:

[0033] (1) By combining scale text data and heart rate variability data, a DBGNet multimodal depression detection model is constructed to achieve deep interaction of multimodal features; in the DBGNet model, the text data uses the BERT pre-trained model and Bi-MGRU to extract semantic features, and the heart rate variability data extracts physiological features through the DIRNet model. Combining the cross-attention mechanism for feature fusion to achieve complementarity, enhancing the generalization ability of the model to complex depressive features, effectively overcoming the information limitations of traditional single-modal methods (such as relying only on subjective scales or single physiological indicators), and significantly improving the accuracy of depression detection.

[0034] (2) In the DIRNet model, by using depthwise separable convolution, asymmetric convolution, and residual stacking module, by reducing redundant residual layer stacking, the number of model parameters is greatly reduced, the training speed and inference efficiency are improved, and the overfitting problem is avoided; at the same time, the DIRNet model adopts a multi-scale convolution design, combining grouped convolution and residual structure, which can effectively capture the spatial multi-scale features of heart rate variability data.

[0035] (3) By introducing the Mish activation function, the non-linear expression ability of the network is enhanced, and the utilization rate of features is improved; a lightweight ECA attention mechanism is embedded after the Block residual stacking module. Through the local cross-channel interaction strategy, without increasing the network complexity, the weight of the key feature channels is strengthened, and the model accuracy is significantly improved.

[0036] (4) Data is collected through wearable devices and scales. Combining with the DBGNet multi-modal depression detection model, an efficient and objective depression auxiliary diagnosis system is constructed, reducing the dependence on the subjective experience of professional physicians, providing technical support for patients' self-screening and clinical accurate diagnosis, and having important social application value.

[0037] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic flowchart of an embodiment of the present invention;

[0039] Figure 2 It is a structural diagram of the BMGNet model in an embodiment of the present invention;

[0040] Figure 3 It is a structural diagram of the DIRNet model in an embodiment of the present invention;

[0041] Figure 4 It is a schematic diagram of the ECA attention mechanism in an embodiment of the present invention;

[0042] Figure 5 It is an overall structural diagram of the DBGNet multi-modal depression detection model in an embodiment of the present invention;

[0043] Figure 6 It is a structural diagram of the multi-modal interaction module in an embodiment of the present invention;

[0044] Figure 7 It is a comparative diagram of the structures of three multi-modal fusion technologies in an embodiment of the present invention; where (a) is feature fusion; (b) is the cross-attention mechanism feature fusion adopted in this solution; (c) is decision fusion;

[0045] Figure 8 It is a schematic diagram of the residual stack design in an embodiment of the present invention;

[0046] Figure 9 It is a comparative diagram of different positions of ECA in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0049] Embodiment

[0050] As shown Figures 1 - 7 in the figure, the present invention provides a method for detecting depression by fusing multi-modal data. For this method, the heart rate data and text data of subjects from the psychology department of a hospital in Baoding are used to test the influence effect of multi-modal data fusion on the model performance. The steps are as follows:

[0051] S1. Collect the scale text data of depression subjects through questionnaires, social media data or interview content, and collect the heart rate variability data of depression patients in real time through wearable devices. Specifically:

[0052] Conduct the experiment in a warm and comfortable environment to facilitate the subjects to better let go of their psychological defenses. Use a heart rate monitoring bracelet to record the heart rate data of the subjects under calm breathing for 5 minutes, and then complete the depression self-rating questionnaire under the interpretation of a doctor. As shown in Table 1, the questionnaire in this embodiment mainly involves ten aspects. The subjects need to answer all the questions in each aspect, and each question is scored on a 5-point scale (1 is "completely disagree", 5 is "completely agree").

[0053] Table 1 - Depression Theme Questions

[0054]

[0055]

[0056] S2. According to the conversion score of the scale text data and the characteristics of the heart rate variability data, use labels to annotate the data collected in S1 to obtain a labeled data set; the labels include two categories: depression and non-depression. Specifically:

[0057] Calculate the conversion score from the original score of the depression self-rating questionnaire to determine the physical constitution type;

[0058] Among them, the original score M is expressed as:

[0059]

[0060] In the formula, P i is the score of the i-th question, and k is the total number of questions.

[0061] The conversion score C is expressed as:

[0062]

[0063] In the formula, min(M) is the minimum value of the original score, and max(M) is the maximum value of the original score.

[0064] The determination criterion in this embodiment is: when the conversion score C is below 40 points, it is "non-depression", and when it is above 40 points, it is "depression".

[0065] After screening, a total of 563 subjects' heart rate data and text data were collected in this embodiment; among these data, there were 403 groups of data of non-depressed individuals and 160 groups of data of depressed individuals.

[0066] S3. Preprocess the labeled dataset, and the preprocessing includes:

[0067] 1) Word segmentation, stop word removal, and text vectorization processing of the scale text data;

[0068] 2) Denoising and normalization processing of the heart rate variability data. In this embodiment, the Poincaré scatter plot is used for the visualization of the heart rate variability data, effectively realizing the early identification and evaluation of depression.

[0069] S4. Construct the DBGNet multimodal depression detection model to detect the preprocessed labeled dataset;

[0070] S41. In this model, the BMGNet and DIRNet models are respectively constructed for unimodal depression detection for text and heart rate variability data;

[0071] In this solution, the BMGNet and DIRNet models are used as feature extractors (removing the Softmax layer). The extracted text features and heart rate variability features promote feature interaction through the cross-attention mechanism to enhance the information exchange between different modalities and make up for their respective deficiencies. Then the interacted features are concatenated and classified through the Softmax layer to obtain the final result (level) of depression detection, which is specifically as follows:

[0072] 1) The BMGNet model based on text:

[0073] This model first receives the preprocessed scale text data, and then uses the Bert pre-trained model introducing the whole-word masking technology to generate word vectors, which provides the model with the ability to deeply understand the nuances of language. Then the Bi-MGRU is used to extract the text features. The BMGNet model not only optimizes the representation ability of the text data through Bert, but also significantly improves the model's understanding and analysis depth of the text information through the context interaction mechanism of Bi-MGRU.

[0074] 2) The DIRNet model based on heart rate variability:

[0075] This solution meets the requirements of heart rate variability depression detection from the following aspects:

[0076] ① The convolutional layer uses three 3×3 convolutional kernels, and depthwise separable convolution is introduced to enhance the computational efficiency of the model; the performance and effect of the neural network are improved from several aspects such as increasing nonlinearity, reducing the number of parameters, increasing the receptive field, and increasing the network depth.

[0077] Among them, the increase in non-linearity is specifically as follows: by stacking multiple 3×3 convolutional kernels, each convolutional layer introduces a non-linear transformation, thereby improving the network's representation ability.

[0078] ② The Block residual stacking module uses a combination of asymmetric convolution and grouped convolution to further enhance the network's feature extraction ability;

[0079] The Block residual stacking module consists of three parts: on the far right is a 1×3, 3×1 asymmetric convolution. Asymmetric convolutions usually have fewer parameters than symmetric convolutions, allowing the convolutional kernel to have different sizes in the horizontal and vertical directions, so it can better capture features in different directions of the input data, that is, better capture the features of heart rate variability.

[0080] ③ The Block residual stacking module also enhances the network's ability in feature extraction through the stacking of residual blocks. This multi-scale residual design not only enhances the feature abstraction ability, enabling the network to learn richer feature representations, but also improves the effectiveness of feature extraction by expanding the receptive field; it also helps to reduce the number of network parameters, thereby improving the efficiency of the model.

[0081] ④ The Mish activation function is adopted to increase the non-linear ability of the model, and its formula is:

[0082] Mish(x) = x × Tanh(Softplus(x));

[0083] In the formula, Mish(x) represents the Mish activation function, x is the input, Tanh(·) represents the hyperbolic tangent function, and Softplus(·) represents the Softplus function.

[0084] ⑤ The attention mechanism is introduced: The ECA attention mechanism, as a lightweight module, is essentially an adjusted SE attention mechanism, providing a local cross-channel interaction strategy without dimensionality reduction, as well as a method for adaptively selecting the size of a one-dimensional convolutional kernel, thereby effectively improving feature selection and enhancement means.

[0085] S42. Input the text features and heart rate variability features into the multi-modal interaction module based on the cross-attention mechanism for feature interaction, splice the two interacted features, and then perform classification through the Softmax layer to obtain the depression detection result;

[0086] The multi-modal interaction module plays a crucial role in improving the accuracy, robustness, and generalization ability of the model. It integrates data from different information sources or modalities to obtain a more comprehensive understanding. By leveraging the unique information content and advantages of each modality, through data-level, feature-level, or decision-level fusion strategies, it enhances the model's ability to recognize, interpret, and predict complex phenomena.

[0087] To further strengthen the information interaction between the two modalities of text and heart rate variability, the model introduces a multi-modal interaction module to enhance communication between different modalities. Its core technology is the cross-attention mechanism, which allows the model to capture information in different representation subspaces. Through cross-attention, the model can exchange information between text and image features.

[0088] The following is illustrated through specific experiments:

[0089] To ensure the stability and effectiveness of the experimental results, this embodiment adopts a random division method, dividing the data of 563 subjects into three independent sets according to a ratio of 6:2:2: the training set, the validation set, and the test set.

[0090] The experimental platform is built on a high-performance hardware basis. The operating system selects the widely used Windows 10 to ensure software compatibility and user familiarity. The hardware configuration includes an Intel(R) Core(TM) i9-7900X 3.30GHz processor, which provides strong support for complex data processing tasks with its excellent computing power. The effectiveness of the model is verified from the following aspects:

[0091] 1) Comparison of heart rate variability models: In this embodiment, heart rate variability data is used to train in the above model, only modifying the number of classifications in the classification layer without making other modifications to the model. The experimental results are shown in Table 2:

[0092] Table 2 - Comparison of heart rate variability models

[0093] model accuracy loss VGGNet 85.7% 0.42 ResNet34 87.9% 0.25 ResNet101 86.1% 0.13 MobileNet 85.7% 0.42 EfficientNet 86.6% 0.45 RegNet 86.6% 0.39 EfficientNetV2 81.4% 0.43 DIRNet 89.6% 0.24

[0094] It can be seen that the DIRNet of this scheme has the best effect, with an accuracy rate reaching 89.6% and a loss of 0.24. It has a residual structure and a certain network depth, and has a better effect on classifying heart rate variability data.

[0095] 2) Comparison of residual stack design:

[0096] The DIRNet module designs network models with different depths through residual stack design. To find the most suitable network depth for this scheme, models with depths of 1, 2, 3, and 4 are designed respectively. The experimental results are as Figure 8As shown in the figure, as the number of stacked residual blocks increases, the accuracy of the model gradually decreases. That is, when the stacking of residuals in the model increases, the model begins to show degradation. Excessive stacking not only fails to improve the extraction of heart rate variability features but reduces it. When the number of residual blocks is 1, 1, 1, 1, the accuracy is 91.6%, which is more suitable for the experimental data. Therefore, the stacking design of 1, 1, 1, 1 was selected for subsequent experiments.

[0097] 3) Comparison of attention modules:

[0098] To verify the influence of the ECA attention module at different positions on the model, in this study, the ECA module was inserted after the convolutional layer, after the depthwise separable convolutional layer, and after the Block residual stacking module respectively. The experimental results are as Figure 9 shown. When the ECA is only inserted after the Block residual stacking module, the model has the best effect, and the accuracy reaches 95.3%.

[0099] The ECA module adopts a local cross-channel interaction strategy, which helps the model capture the correlation between channels while keeping the number of parameters low. This local interaction strategy is more effective after the Block residual stacking module because the residual blocks pass different levels of feature information through skip connections, and the ECA module can better utilize this information to strengthen the association between channels. Finally, the position of the ECA after the Block residual stacking module was selected.

[0100] 4) Comparison of different activation functions:

[0101] An activation function is a function added to an artificial neural network to help the network learn complex problems and enhance the network's representation and learning capabilities. In this study, the activation function was replaced in the Block residual block. To verify the influence of different activation functions on the experimental results, three activation functions, ReLU, LeakyReLU, and Mish, were selected for comparative experiments. The experimental results are shown in Table 3. When the activation function is Mish, the model has the best effect.

[0102] Table 3 - Comparison of different activation functions

[0103] activation function accuracy loss ReLU 95.3% 0.16 LeakyReLU 94.2% 0.31 Mish 95.7% 0.17

[0104] 5) Ablation experiment:

[0105] The ablation experiment is a key technique for evaluating the performance of a model. It assesses the specific impact of certain parts of the model on the overall network performance by systematically removing or disabling those parts. This method can effectively reveal the contributions of different network modules to the model's accuracy, thereby verifying the effectiveness of improvement measures. In this study, the ablation experiment method was adopted to evaluate the proposed strategies. As shown in Table 4, this experiment designed combinations in three aspects: convolutional replacement, Block residual construction, and addition of attention mechanism.

[0106] Table 4 - Ablation Experiment Design

[0107] sequence number 7×7 convolution 3×3 convolution ResNet residual block Block residual block asymmetric convolution group convolution attention mechanism 1 √ √ 2 √ √ 3 √ √ 4 √ √ 5 √ √ 6 √ √ √

[0108] Table 5 - Ablation Experiment Results

[0109] sequence number accuracy precision recall F1 - score 1 89.6% 89.6% 89.5% 0.895 2 92.6% 93.2% 92.9% 0.930 3 92.4% 92.7% 92.5% 0.926 4 94.6% 94.7% 94.5% 0.946 5 95.2% 94.4% 93.9% 0.941 6 95.7% 95.6% 95.7% 0.957

[0110] Table 5 shows the results of the ablation experiment, using the model's accuracy, precision, recall, and F1-score as evaluation criteria. The following conclusions can be drawn from the experimental results:

[0111] In Experiment 1, only the 7×7 convolution and ResNet residual blocks were used without other modifications, which was the ResNet50 model; in Experiment 2, this study used three 3×3 convolutions. The comparison of experimental results shows that the three 3×3 convolutions perform better than the 7×7 convolution in all aspects. By applying the 3×3 convolution kernel multiple times, the receptive field of the network can be increased while maintaining a small convolution kernel, and at the same time, the network depth is increased, better extracting the features of heart rate variability data;

[0112] In Experiment 3, the convolution in the residual block was replaced with grouped convolution. The experimental results showed a slight decrease compared to the original model. The use of grouped convolution can reduce the number of parameters and prevent overfitting, but there is also information blockage, and the heart rate variability information in different groups cannot communicate. In Experiment 4, the convolution in the residual block was replaced with asymmetric convolutions of 1×3 and 3×1. Asymmetric convolutions can better capture features in different directions in image classification and better learn the characteristics of heart rate variability images of different populations. The Block residual stacking module is a combination of asymmetric convolution and grouped convolution. The results of Experiment 5 show that although the effect of grouped convolution is slightly lower, when combined with asymmetric convolution, it can achieve better classification results. The comparison between Experiment 5 and Experiment 2 can also prove the effectiveness of the Block residual stacking module constructed in this study.

[0113] This experiment also tried different attention mechanisms. The results showed that the ECA attention mechanism was improved based on the original model. ECA can perform local cross-channel interaction without dimensionality reduction and capture heart rate variability features more comprehensively.

[0114] 6) Model comparison experiment:

[0115] In unimodal analysis, this experiment used classification accuracy to evaluate the ability of unimodal analysis to detect depression to determine the effectiveness of heart rate variability and text analysis. In multimodal analysis, this study used four other quantitative indicators in statistical analysis to evaluate its performance, including accuracy, precision, recall, and F1 score.

[0116] In view of the situation of combining two modal information in this experiment, three fusion methods were designed: decision fusion, feature fusion, and cross-attention mechanism feature fusion proposed in this scheme, and the single-modal optimal results were compared. As shown in Table 6, the experimental results show that the cross-attention mechanism feature fusion achieved better results, and the accuracy rate reached 98.0% on the CMU-MOSI dataset used in this scheme. The cross-attention mechanism feature fusion can more comprehensively obtain depression features and improve the classification accuracy.

[0117] Table 6 - Model comparison experiment

[0118]

[0119] 7) Comparison with other methods:

[0120] Table 7 - Comparison results with other methods

[0121]

[0122]

[0123] As shown in Table 7, the model in this study performed well, with a classification accuracy of 98.0%, which is significantly better than other methods. This result not only proves the effectiveness of the depression detection method designed in this scheme, but also highlights its advancement in multimodal data processing.

[0124] Therefore, a depression detection method that integrates multi-modal data according to the present invention constructs a DBGNet multi-modal depression detection model. By combining scale text data and heart rate variability data, deep interaction of multi-modal features is achieved. In the model, BERT pre-trained model and Bi-MGRU are used to extract semantic features from text data, and the DIRNet model is used to extract physiological features from heart rate variability data. The cross-attention mechanism is used for feature fusion, enhancing the generalization ability of the model for complex depression features. The DIRNet model adopts depthwise separable convolution, asymmetric convolution and residual stacking modules to reduce the number of model parameters, improve the training speed and inference efficiency. At the same time, multi-scale convolution design is used to capture the spatial multi-scale features of heart rate variability data. In addition, the Mish activation function is introduced to enhance the non-linear expression ability of the network, and a lightweight ECA attention mechanism is embedded after the Block residual stacking module to strengthen the weights of key feature channels. Finally, by combining data collected by wearable devices and scales, a set of efficient and objective depression auxiliary diagnosis system is constructed, reducing the dependence on the subjective experience of professional physicians, providing technical support for patients' self-screening and clinical precise diagnosis, and having important social application value.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for detecting depression by fusing multimodal data, characterized in that: The following steps are involved: S1. Collect scale text data and heart rate variability data of depression subjects; S2, based on the conversion score of the scale text data and the characteristics of the heart rate variability data, use labels to annotate the data collected in S1 to obtain a labeled data set; The label includes two categories: depression and non-depression; S3, preprocessing the label data set; S4. Build the DBGNet multimodal depression detection model to detect the preprocessed labeled dataset and output the depression detection results.

2. A method for detecting depression by fusing multimodal data according to claim 1, characterized in that: In S1, the scale text data is collected through questionnaires, social media data or interview content; and the heart rate variability data is collected in real time through wearable devices.

3. The method for detecting depression by fusing multimodal data according to claim 1, characterized in that: The conversion fraction in S2 is expressed as: Where C is the conversion score, M is the original score, min(M) is the minimum value of the original score, and max(M) is the maximum value of the original score.

4. The method for detecting depression by fusing multimodal data according to claim 3, characterized in that: The raw score is expressed as: Where P i is the score of the i-th question, and j is the total number of questions.

5. The method for detecting depression by fusing multimodal data according to claim 1, characterized in that: The preprocessing in S3 includes: 1) Word segmentation, stop word removal, and text vectorization of scale text data; 2) De-noising and standardization of heart rate variability data.

6. The method for detecting depression by fusing multimodal data according to claim 1, characterized in that: In S4, the specific detection process of the DBGNet multimodal depression detection model is as follows: S41, constructing BMGNet and DIRNet models for feature extraction for the scale text data and heart rate variability data in S3, respectively, to obtain text features and heart rate variability features; S42. Input the text features and heart rate variability features into the multimodal interaction module based on the cross-attention mechanism for feature interaction, concatenate the two features after the interaction, and then classify them through the Softmax layer to obtain the depression detection results.

7. The method for detecting depression by fusing multimodal data according to claim 6, characterized in that: In S41, the specific process of extracting text features using the BMGNet model is as follows: for the scale text data in S3, the Bert pre-training model that introduces the whole word masking technology is used to generate word vectors, and then the text features are extracted using Bi-MGRU.

8. The method for detecting depression by fusing multimodal data according to claim 6, characterized in that: In S41, the DIRNet model includes: a convolution layer, a depth-separable convolution layer, a Block residual stacking module, a Mish activation function, and an ECA attention mechanism; the Block residual stacking module is a combination of asymmetric convolution and grouped convolution.