Method and system for early screening of depression based on multi-modal deep learning
The multimodal deep learning method is used to fuse multiple data types and process these data using Transformer model, which solves the problem of single data and insufficient feature extraction in the existing depression screening methods, and achieves more accurate and reliable early screening of depression.
Patent Information
- Application Number
- CN202411459250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing depression screening methods have problems such as single data, insufficient feature extraction and limited pattern recognition capabilities, making it difficult to achieve large-scale and objective early screening.
Using a multimodal deep learning method, the complex interaction relationship between multimodal data such as physiological signals, speech, facial expressions and text content is used to process the complex interaction between multimodal data, and capture subtle changes in long-term behavior patterns.
It significantly improves the accuracy and reliability of early screening for depression, can more comprehensively capture the multi-dimensional manifestations of depression, reduce misjudgment, and improve the credibility of screening results.
Smart Images

Figure CN120089395A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of early screening methods for depression, in particular to an early screening method and system for depression based on multimodal deep learning. Background Art
[0002] With the acceleration of the pace of modern social life and the increase in pressure, depression has become an increasingly serious public health problem. Early identification and intervention are crucial for improving the prognosis of patients with depression. However, traditional depression screening methods mainly rely on questionnaires and clinical interviews, which are not only time-consuming and laborious, but also easily affected by subjective factors, making it difficult to achieve large-scale and objective early screening.
[0003] In recent years, with the rapid development of artificial intelligence technology, researchers have begun to attempt to apply machine learning methods to the automated screening of depression. The closest prior art usually uses single-modal data, such as text analysis or speech feature extraction, to identify potential depressive symptoms. Although these methods have achieved certain results in some scenarios, there are still significant limitations. First, single-modal data cannot comprehensively reflect the complex manifestations of depression and is prone to missing important diagnostic clues. Second, traditional machine learning algorithms such as support vector machines or random forests, although excellent in processing structured data, are insufficient in capturing the subtle changes and long-term patterns of depressive symptoms. In addition, existing methods often ignore the complex interaction relationships between different features, which is particularly important in the multi-dimensional manifestations of depression.
[0004] More critically, the prior art performs poorly in processing time-series data and long-term behavior pattern analysis. The development of depression is often a gradual process, and single-point data is difficult to reflect the dynamic changes of symptoms. At the same time, the efficiency and scalability of existing methods in processing large-scale and heterogeneous data also face challenges, which limits their application in actual clinical environments. Summary of the Invention
[0005] In view of the above problems, the present invention proposes an early screening method and system for depression based on multimodal deep learning. The method aims to comprehensively capture various manifestations of depression by integrating multi-dimensional data such as physiological signals, speech, facial expressions, and text content. By adopting an advanced Transformer model as the core algorithm, the present invention can effectively process the complex interaction relationships between multimodal data and capture the subtle changes in long-term behavior patterns.
[0006] The present invention proposes an early screening method for depression based on multimodal deep learning, including the following steps:
[0007] Obtain the multimodal data of the subject, where the multimodal data includes physiological signals, speech, facial expressions, and text content;
[0008] For each modality data in the multimodal data, use a multi-branch deep neural network to extract features respectively, and fuse the extracted features to construct the original feature vector of the subject;
[0009] Use a Transformer with a feed-forward fully connected layer as the risk assessment model, and train the Transformer using the original feature vector to obtain the risk assessment model corresponding to the subject; and
[0010] Input the data of the subject to be predicted into the risk assessment model to obtain the risk assessment result of the subject having depression.
[0011] Preferably, the step of extracting features from each modality data specifically includes:
[0012] Extract the K-dimensional time-domain feature vector corresponding to each physiological signal p 1 ,..., p i ,..., p M} in the physiological signal dataset P of the subject, and combine the K-dimensional time-domain feature vectors into a feature matrix X∈R i , where N is the number of small segments into which the physiological signal is divided, and N≤M; N×K For the speech dataset V = {v
[0013] ,..., v 1 ,..., v m ,..., v N}, use the pre-trained vector representation model VGGish to extract each speech feature vector and perform normalization processing, and then combine the normalized speech feature vectors into a speech feature matrix where K 1 is the feature dimension of the speech;
[0014] Obtain the D-dimensional time-domain feature vector corresponding to each facial expression frame f 1 ,..., f m ,..., f N} in the facial expression dataset F, and combine the D-dimensional time-domain feature vectors into a facial expression feature matrix Z∈R i , where L is the number of emotional features, and L<N; and L×D For the text content dataset T = {t
[0015] ,..., t 1 ,..., t k ,..., t N}, using the pre-trained word vector representation model Word2Vec, represent the text content as a text feature matrix composed of K 2 dimensional feature vectors where N is the number of small segments into which the text content is divided, and N < M.
[0016] Preferably, the step of constructing the original feature vector of the subject specifically includes:
[0017] Construct the original feature vector Set the initial value of each dimension to 0;
[0018] Represent each row of the feature matrix X, the speech feature matrix Y, the facial expression feature matrix Z, and the text feature matrix W into the original feature vector V'; and
[0019] Perform normalization processing on the original feature vector V'.
[0020] Preferably, the formula for performing normalization processing on the original feature vector V' is:
[0021]
[0022] where, V j represents the j-th element of the normalized feature vector, V j ' represents the j-th element of the original feature vector, min(V') and max(V') respectively represent the minimum value and the maximum value in the original feature vector V'.
[0023] Preferably, the step of training the Transformer using the original feature vector specifically includes:
[0024] Obtain the depression state marked by experts as the gold standard to train the risk assessment model. The specific formula for training the current risk assessment model is:
[0025]
[0026] where, I = (I 1 ,..., I j ,..., I N ) ∈ R 1×N , represents the one-hot encoding of the gold standard; represents the original feature vector corresponding to the j-th sample; norm(x j ) represents the vector after regularizing the original feature vector; θ represents all the parameters of the model; λ represents the regularization strength; f θThe Transformer model function is denoted as; the softmax activation function is denoted as softmax; the loss function is denoted as L(θ).
[0027] 6. The method for early screening of depression based on multimodal deep learning according to claim 5, wherein the step of training the Transformer further comprises:
[0028] Using the validation set to verify the form of the loss gradient and fine-tuning the parameters of the risk assessment model;
[0029] Inputting the test set into the fine-tuned risk assessment model, and calculating the AUC for evaluation according to the output value of the fine-tuned risk assessment model; and
[0030] If the AUC is greater than the threshold set by the expert, the risk assessment model corresponding to the subject is obtained; otherwise, continue training until the AUC ≥ the threshold set by the expert is satisfied.
[0031] Preferably, the training process of the risk assessment model is optimized using the Adam optimizer, where the learning rate is 1e -5 .
[0032] The early screening system for depression based on multimodal deep learning that executes the method according to the claim, comprising:
[0033] A data acquisition module, configured to acquire multimodal data corresponding to a subject, where the multimodal data includes physiological signals, voice, facial expressions, and text content;
[0034] A feature extraction module, configured to respectively use a multi-branch deep neural network to extract features from each modal data, and fuse the extracted features to construct an original feature vector of the subject;
[0035] A risk assessment model training module, configured to use a Transformer with a feed-forward fully connected layer as the risk assessment model, and use the original feature vector to train the Transformer to obtain a risk assessment model corresponding to the subject; and
[0036] A risk assessment module, configured to input the data of the subject to be predicted into the risk assessment model to obtain a risk assessment result of the subject having depression.
[0037] Preferably, the feature extraction module specifically includes:
[0038] A dataset collection unit, configured to construct a physiological signal dataset P = {p 1 ,..., p i ,..., p M}, a voice dataset V = {v 1,..., v m ,..., v N}, facial expression dataset F = {f 1 ,..., f m ,..., f N}, and text content dataset T = {t 1 ,..., t k ,..., t N}, and obtain the expert-annotated depression status of the corresponding subjects;
[0039] Feature matrix construction unit, for combining the K-dimensional time-domain feature vectors corresponding to each of the physiological signals p i into a feature matrix X ∈ R N×K ;
[0040] Speech feature matrix construction unit, for using the pre-trained vector representation model VGGish to extract each speech feature vector, perform normalization processing, and then combine them into a speech feature matrix
[0041] Facial expression feature matrix construction unit, for combining the D-dimensional time-domain feature vectors corresponding to each of the facial expression frames f i into a facial expression feature matrix Z ∈ R L×D ;
[0042] Text feature matrix construction unit, for using the pre-trained word vector representation model Word2Vec to represent the text content as a text feature matrix composed of K 2 -dimensional feature vectors and
[0043] Original feature vector construction unit, for constructing an original feature vector The initial value of each dimension is set to 0, and each row of the feature matrix X, the speech feature matrix Y, the facial expression feature matrix Z, and the text feature matrix W is represented into the original feature vector V', and normalization processing is performed on the original feature vector V'.
[0044] Preferably, the risk assessment model training module specifically includes:
[0045] Model training unit, for obtaining the expert-annotated depression status as the gold standard to train the risk assessment model. The specific formula for training the current risk assessment model is:
[0046]
[0047] Where, I = (I 1 ,..., I j ,..., IN ) ∈ R 1×N , which represents the one - hot encoding of the gold standard; represents the original feature vector corresponding to the j - th sample; norm(x j ) represents the vector after regularizing the original feature vector; θ represents all the parameters of the model; λ represents the regularization strength; f θ represents the Transformer model function; softmax represents the softmax activation function; L(θ) represents the loss function;
[0048] A verification unit, which is used to use the validation set as data to verify the form of the loss gradient and fine - tune the parameters of the risk assessment model. Input the test set into the fine - tuned risk assessment model, calculate the AUC for evaluation according to the output value of the fine - tuned risk assessment model. If the AUC is greater than the threshold set by the expert, construct a risk assessment model training module to continue training until AUC ≥ the threshold set by the expert; and
[0049] A risk assessment model construction unit, which is used to construct a Transformer with a feed - forward fully - connected layer as the risk assessment model, and the input of the Transformer is the original feature vector of the subject.
[0050] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0051] The innovation of the present invention lies not only in the fusion of multi - modal data, but more in its unique algorithm design. The self - attention mechanism of the Transformer model provides an ideal solution to solve the correlation problem between different modal features. This mechanism allows the model to dynamically focus on important features at different time points and different modalities, thereby capturing the subtle changes and long - term development trends of depressive symptoms while maintaining efficient computation.
[0052] In addition, through a clever feature extraction and fusion strategy, the present invention successfully unifies the two seemingly contradictory goals of high - dimensional feature representation and model complexity control. By performing targeted pre - processing and feature extraction on each modal data, we not only retain the rich information of the original data, but also greatly reduce the computational complexity of subsequent processing. This balance not only improves the training efficiency of the model, but also enhances its scalability in practical applications.
[0053] More notably, the method of the present invention achieves a significant synergistic effect between different features. For example, abnormalities in physiological signals may be corroborated in facial expressions or speech features, while text content may provide semantic explanations for these physiological and behavioral changes. This multi - dimensional cross - verification greatly improves the reliability of the screening results and reduces the misjudgment that may be caused by a single feature.
[0054] Generally speaking, through an innovative multimodal deep learning method, the present invention effectively solves the problems existing in the existing depression screening technologies, such as single data, insufficient feature extraction, and limited pattern recognition ability. It not only significantly improves the accuracy and reliability of early depression screening, but also provides a new perspective for understanding the multi-dimensional manifestations of depression. This method is expected to play an important role in clinical practice, providing strong technical support for the early identification and intervention of depression, ultimately improving the quality of life of patients and reducing the social medical burden. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is the overall system logic block diagram of the present invention.
[0056] Figure 2 It is the Transformer training logic block diagram of the present invention.
[0057] Figure 3 It is the logic block diagram of the data acquisition module of the present invention.
[0058] Figure 4 It is the logic block diagram of the feature extraction module of the present invention.
[0059] Figure 5 It is the logic block diagram of the risk assessment model training module of the present invention.
[0060] Figure 6 It is the logic block diagram of the risk assessment module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0061] In order to further elaborate on the technical means and their effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the accompanying drawings and their preferred embodiments, and details their specific implementation manners, structures, features, and their effects as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0063] Embodiment 1
[0064] Referring to Figure 1-6 , the present invention relates to its method and system for early depression screening based on multimodal deep learning. First, the present invention proposes a method for early depression screening based on multimodal deep learning.
[0065] The method includes the following steps: Obtain the multi-modal data of the subject, including physiological signals, speech, facial expressions, and text content. These multi-modal data can comprehensively reflect the physical and mental state of the subject and provide a rich information source for subsequent analysis. Then, use a multi-branch deep neural network to extract features from each modal data respectively, and fuse the extracted features to construct the original feature vector of the subject. This step can fully explore the features of each modal data and make full use of the complementary information between different modalities through fusion. Secondly, use a Transformer with a feed-forward fully connected layer as the risk assessment model, and train the Transformer using the original feature vector to obtain the risk assessment model corresponding to the subject. The Transformer model has strong feature representation and sequence modeling capabilities and is suitable for processing multi-modal time series data. Finally, input the data of the subject to be predicted into the risk assessment model to obtain the risk assessment result of the subject's depression.
[0066] In one embodiment of the present invention, the step of extracting features from each modal data specifically includes:
[0067] Extract the K-dimensional time-domain feature vector corresponding to each physiological signal p 1 ,..., p i ,..., p M} in the physiological signal dataset P = {p i of the subject, and combine the K-dimensional time-domain feature vectors into a feature matrix X ∈ R N×K , where N is the number of small segments into which the physiological signal is divided, and N ≤ M;
[0068] For the speech dataset V = {v 1 ,..., v m ,..., v N}, use the pre-trained vector representation model VGGish to extract each speech feature vector and perform normalization processing, and then combine the normalized speech feature vectors into a speech feature matrix where K 1 is the feature dimension of the speech;
[0069] Obtain the D-dimensional time-domain feature vector corresponding to each facial expression frame f 1 ,..., f m ,..., f N} in the facial expression dataset F = {f i , and combine the D-dimensional time-domain feature vectors into a facial expression feature matrix Z ∈ R L×D , where L is the number of emotional features, and L < N; and
[0070] For the text content dataset T = {t 1 ,..., tk , ..., t N}, using the pre-trained word vector representation model Word2Vec, represent the text content as a text feature matrix composed of K 2 -dimensional feature vectors where N is the number of small segments into which the text content is divided, and N < M.
[0071] First, for the physiological signal dataset P, extract the K-dimensional time-domain feature vector corresponding to each physiological signal p i and combine these vectors into a feature matrix R N×K . Here, N represents the number of small segments into which the physiological signal is divided, and N ≤ M. For example, an electrocardiogram signal can be divided into 10-second segments, and features such as the mean, standard deviation, and peak value of each segment can be extracted. Then, for the speech dataset V, use the pre-trained VGGish model to extract each speech feature vector and perform normalization processing to form a speech feature matrix The VGGish model is pre-trained on a large-scale audio dataset and can effectively capture features such as the timbre and emotion of speech. Second, for the facial expression dataset F, extract the D-dimensional time-domain feature vector of each expression frame f i to form a facial expression feature matrix Z ∈ R L×D . Here, L represents the number of emotion features, and L < N. For example, facial action units (AUs) can be used as features to capture micro-expression changes. Finally, for the text content dataset T, use the pre-trained Word2Vec model to represent the text as K2-dimensional feature vectors to form a text feature matrix Word2Vec can map words to a low-dimensional space and capture semantic information.
[0072] Preferably, the steps of constructing the original feature vector of the subject specifically include: First, construct the original feature vector with initial values all being 0. Then, represent each row of the feature matrix X, the speech feature matrix Y, the facial expression feature matrix Z, and the text feature matrix W into the original feature vector V'; and perform normalization processing on the original feature vector V'. This mapping method preserves the temporal relationship of each modality data. Finally, perform normalization processing on V' to make the feature scales of different modalities consistent.
[0073] In an embodiment of the present invention, the formula for performing normalization processing on the original feature vector V' is:
[0074]
[0075] where V j represents the j-th element of the normalized feature vector, V j′ represents the j-th element of the original feature vector, and min(V′) and max(V′) represent the minimum and maximum values in the original feature vector V′ respectively. This normalization method maps the eigenvalues to the interval [0, 1], which helps to improve the stability and convergence speed of model training.
[0076] In one embodiment of the present invention, the steps of training the Transformer using the original feature vector specifically include:
[0077] Obtain the depression state annotated by experts as the gold standard to train the risk assessment model. The specific formula for training the current risk assessment model is:
[0078]
[0079] where I = (I 1 , …, I j , …, I N ) ∈ R 1×N , representing the one-hot encoding of the gold standard; represents the original feature vector corresponding to the j-th sample; norm(x j ) represents the vector after regularizing the original feature vector; θ represents all the parameters of the model; λ represents the regularization strength; f θ represents the Transformer model function; softmax represents the softmax activation function; L(θ) represents the loss function. This loss function combines cross-entropy loss and L2 regularization, which can effectively prevent overfitting.
[0080] In one embodiment of the present invention, the steps of training the Transformer further include:
[0081] Use the validation set to verify the shape of the loss gradient and fine-tune the parameters of the risk assessment model;
[0082] Input the test set into the fine-tuned risk assessment model, and calculate the AUC for evaluation according to the output value of the fine-tuned risk assessment model; and
[0083] If the AUC is greater than the threshold set by the expert, the risk assessment model corresponding to the subject is obtained; otherwise, continue training until AUC ≥ the threshold set by the expert (e.g., 0.85). This iterative optimization strategy can ensure the generalization ability of the model.
[0084] In one embodiment of the present invention, the training process of the risk assessment model is optimized using the Adam optimizer, where the learning rate is 1e -5 . The Adam optimizer combines the advantages of the momentum method and the adaptive learning rate method, and a smaller learning rate helps the model to converge more stably in the later stage of training.
[0085] In one embodiment of the present invention, the system architecture of the present invention includes a data acquisition module 1, a feature extraction module 2, a risk assessment model training module 3, and a risk assessment module 4. This modular design makes the system structure clear and facilitates implementation and maintenance.
[0086] The data acquisition module 1 is used to acquire multimodal data corresponding to the subject, and the multimodal data includes physiological signals, voice, facial expressions, and text content;
[0087] The feature extraction module 2 is used to respectively use a multi-branch deep neural network to extract features from each modal data, and fuse the extracted features to construct an original feature vector of the subject;
[0088] The risk assessment model training module 3 is used to use a Transformer with a feed-forward fully connected layer as a risk assessment model, and use the original feature vector to train the Transformer to obtain a risk assessment model corresponding to the subject; and
[0089] The risk assessment module 4 is used to input the data of the subject to be predicted into the risk assessment model to obtain a risk assessment result of the subject having depression.
[0090] Further, the feature extraction module 2 includes a data set collection unit 21, a feature matrix construction unit 22, a voice feature matrix construction unit 23, a facial expression feature matrix construction unit 24, a text feature matrix construction unit 25, and an original feature vector construction unit 26. Each unit is responsible for processing data of a specific modality and finally constructs a fused original feature vector.
[0091] The data set collection unit 21 is used to construct a physiological signal data set P = {p 1 ,..., p i ,..., p M}, a voice data set V = {v 1 ,..., v m ,..., v N}, a facial expression data set F = {f 1 ,..., f m ,..., f N}, and a text content data set T = {t 1 ,..., t k ,..., t N}, and obtain the expert-annotated depression status of the corresponding subject;
[0092] The feature matrix construction unit 22 is used to combine the K-dimensional time-domain feature vectors corresponding to each physiological signal p i into a feature matrix X ∈ RN×K ;
[0093] The speech feature matrix construction unit 23 is used to extract each speech feature vector using the pre-trained vector representation model VGGish, perform normalization processing, and then combine them into a speech feature matrix
[0094] The facial expression feature matrix construction unit 24 is used to combine the D-dimensional time-domain feature vectors corresponding to each of the facial expression frames f i into a facial expression feature matrix Z ∈ R L×D ;
[0095] The text feature matrix construction unit 25 is used to use the pre-trained word vector representation model Word2Vec to represent the text content as a text feature matrix composed of K 2 dimensional feature vectors and
[0096] The original feature vector construction unit 26 is used to construct the original feature vector The initial value of each dimension is set to 0. Each row of the feature matrix X, the speech feature matrix Y, the facial expression feature matrix Z, and the text feature matrix W is represented into the original feature vector V', and the original feature vector V' is normalized.
[0097] Finally, the risk assessment model training module 3 includes a model training unit 31, a verification unit 32, and a risk assessment model construction unit 33. This design realizes the separation of the model training, verification, and construction processes, which is beneficial to the optimization and performance improvement of the model.
[0098] The model training unit 31 is used to obtain the depression state annotated by experts as the gold standard to train the risk assessment model. The specific formula for training the current risk assessment model is:
[0099]
[0100] where I = (I 1 , …, I j , …, I N ) ∈ R 1×N , representing the one-hot encoding of the gold standard; represents the original feature vector corresponding to the jth sample; norm(x j ) represents the vector after regularizing the original feature vector; θ represents all the parameters of the model; λ represents the regularization strength; f θ represents the Transformer model function; softmax represents the softmax activation function; L(θ) represents the loss function;
[0101] A verification unit 32, configured to use a verification set as the form of the data verification loss gradient and fine-tune the parameters of the risk assessment model, input a test set into the fine-tuned risk assessment model, calculate the AUC for evaluation according to the output value of the fine-tuned risk assessment model, and if the AUC is greater than the threshold set by an expert, construct a risk assessment model training module to continue training until the AUC≥the threshold set by the expert; and
[0102] A risk assessment model construction unit 33, configured to construct a Transformer with a feed-forward fully-connected layer as the risk assessment model, and the input of the Transformer is the original feature vector of the subject
[0103] By integrating multi-modal data, the present invention makes full use of the complementarity of different types of data and can capture various manifestations of depression more comprehensively. For example, physiological signals may reflect potential physiological abnormalities, speech features may reveal emotional states, facial expressions may show subtle emotional changes, and text content may contain semantic-level depressive tendencies. By comprehensively analyzing these features, the method of the present invention can more accurately identify early signs of depression.
[0104] In addition, the present invention uses a Transformer model as the risk assessment model, making full use of its powerful feature representation ability and long-range dependence modeling ability. The self-attention mechanism of the Transformer enables it to capture the complex relationships between different modal features, which is particularly important for understanding the multi-dimensional manifestations of depression.
[0105] The iterative optimization strategy and multi-stage training process of the present invention ensure that the model performance reaches the expected level. By setting an AUC threshold (such as 0.85), the false positive rate and false negative rate can be balanced while ensuring the screening accuracy. This is particularly important for the early screening of depression, because too high a false positive rate may lead to unnecessary psychological burdens, while too high a false negative rate may miss the opportunity for early intervention.
[0106] In summary, the early depression screening method and system based on multi-modal deep learning provided by the present invention achieve accurate early screening of depression by effectively integrating multi-source heterogeneous data. This not only provides strong support for clinical diagnosis but also opens up new ways for personalized mental health management. In practical applications, this method can be integrated into smartphones or wearable devices to achieve daily monitoring of depression risks and provide important basis for timely intervention and treatment.
[0107] To verify the superiority of the present invention, we designed a set of examples and comparative examples and conducted a detailed comparative analysis. The following are the specific contents and test results of the examples and comparative examples.
[0108] Example 1: Early Screening Method for Depression Based on Multimodal Deep Learning
[0109] In this example, we used a multimodal dataset from a large medical institution, which contains data of 1000 subjects. Among these subjects, 200 were clinically diagnosed as depression patients, and 800 were in the healthy control group. We collected the physiological signals (including electrocardiogram and skin conductance), voice samples, facial expression videos, and text content (including social media posts and diaries) of each subject.
[0110] According to the method of the present invention, we first preprocessed and extracted features from each modal data. For physiological signals, we extracted 20-dimensional features in the time domain and frequency domain; for voice samples, we used a pre-trained VGGish model to extract 128-dimensional features; for facial expressions, we used Facial Action Units (AUs) to extract 17-dimensional features; for text content, we used a pre-trained Word2Vec model to extract 300-dimensional features. Then, we fused these features into a 765-dimensional original feature vector.
[0111] Next, we used a Transformer model with 8 attention heads and 6 layers of encoders as the risk assessment model. We adopted 5-fold cross-validation, used the Adam optimizer for training, set the learning rate to 1e-5, the batch size to 32, and the number of training epochs to 100. On the validation set, we used AUC as the evaluation metric and set the threshold to 0.85.
[0112] Comparative Example 1: Traditional Machine Learning Method Based on a Single Modality (Text)
[0113] In the comparative example, we only used text data and adopted traditional machine learning methods for depression screening. We used the TF-IDF method to extract text features and then used a Support Vector Machine (SVM) as the classifier. We also adopted 5-fold cross-validation and used the grid search method to optimize the hyperparameters of the SVM.
[0114] We used accuracy, precision, recall, F1-score, and AUC as the evaluation metrics. The following table shows the detailed test results of Example 1 and Comparative Example 1:
[0115]
[0116] As can be seen from the test results, the method of the present invention (Example 1) is significantly superior to the traditional single-modal method (Comparative Example 1) in all metrics. In particular, in terms of the AUC metric, the method of the present invention reached 0.95, far exceeding 0.83 of the comparative method, which indicates that our method has a stronger ability to distinguish between patients with depression and healthy controls.
[0117] Such test results fully illustrate the superiority of the multi-modal deep learning method of the present invention in the early screening of depression. First of all, the fusion of multi-modal data enables us to capture the manifestations of depression from multiple dimensions, rather than being limited to single text analysis. For example, physiological signals may reflect abnormal autonomic nervous system function in patients with depression, voice features may reveal intonation changes of low mood, facial expressions may capture subtle emotional inhibition, and text content may show pessimistic and negative thinking tendencies. This multi-dimensional information integration greatly improves the comprehensiveness and accuracy of screening.
[0118] Secondly, the adoption of the Transformer model as the risk assessment model is also the key to the superiority of the present invention. The self-attention mechanism of Transformer can effectively capture the complex relationships between different modal features, which is crucial for understanding the multi-dimensional manifestations of depression. For example, it may discover the subtle connection between the slight changes in facial expressions and voice emotions, or the correlation between certain keywords in the text content and abnormal physiological signals. This in-depth feature correlation analysis is difficult to achieve by traditional methods.
[0119] In addition, the high recall rate (0.91) of the method of the present invention is particularly worthy of attention. In the early screening of depression, a high recall rate means that we can better identify potential patients with depression, thus creating conditions for early intervention. At the same time, the relatively high precision rate (0.89) also ensures the reliability of the screening results and avoids unnecessary psychological burdens caused by excessive false positives.
[0120] Generally speaking, this group of comparative experiments strongly proves the superiority of the present invention in the early screening of depression. It is not only significantly superior to traditional methods in statistical metrics, but more importantly, it provides a more comprehensive and in-depth understanding of the multi-dimensional manifestations of depression. This method is expected to play an important role in clinical practice and provide strong support for the early identification and intervention of depression.
[0121] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. An early screening method for depression based on multimodal deep learning, characterized in that: The following steps are involved: Acquiring multimodal data of the subject, wherein the multimodal data includes physiological signals, speech, facial expressions, and text content; For each modal data in the multimodal data, a multi-branch deep neural network is used to extract features, and the extracted features are fused to construct an original feature vector of the subject; Using a Transformer with a feed-forward fully connected layer as a risk assessment model, and training the Transformer using the original feature vector to obtain a risk assessment model corresponding to the subject; as well as The subject data to be predicted is input into the risk assessment model to obtain a risk assessment result of the subject developing depression.
2. The method for early screening of depression based on multimodal deep learning as claimed in claim 1, characterized in that: The step of extracting features from each modal data specifically includes: Extract the physiological signal data set P corresponding to the subject i ,…,p M Each physiological signal p in i The corresponding K-dimensional time domain feature vectors are combined into a feature matrix X∈R N×K , where N is the number of small segments into which the physiological signal is divided, and N≤M; For the speech dataset V = {v1,…,v m ,…,v N }, use the pre-trained vector representation model VGGish to extract each speech feature vector and normalize it, and then combine the normalized speech feature vectors into a speech feature matrix Where K1 is the characteristic dimension of speech; Obtain each facial expression frame \(f\) in the facial expression dataset \(F=\{f_1,\ldots,f\) m ,\ldots,f N \}, and combine the \(D -\)dimensional time - domain feature vectors corresponding to the facial expression frames into a facial expression feature matrix \(Z\in\mathbb{R}\) i , where \(L\) is the number of emotional features and \(L < N\); and L×D For the text content dataset T = {t1,…,t k ,…,t N }, using the pre-trained word vector representation model Word2Vec, the text content is represented as a text feature matrix consisting of K2-dimensional feature vectors Where N is the number of small segments into which the text content is divided, and N <M。 3. The method for early screening of depression based on multimodal deep learning as claimed in claim 2, characterized in that: The step of constructing the original feature vector of the subject specifically includes: Constructing the original feature vector The initial value of each dimension is set to 0; Representing each row of the feature matrix X, the voice feature matrix Y, the facial expression feature matrix Z, and the text feature matrix W into the original feature vector V'; and The original feature vector V' is normalized.
4. The method for early screening of depression based on multimodal deep learning as claimed in claim 3, characterized in that: The formula for normalizing the original feature vector V' is: Among them, V j represents the jth element of the normalized eigenvector, V′ j represents the jth element of the original feature vector, min(V′) and max(V′) represent the minimum and maximum values in the original feature vector V′, respectively.
5. The method for early screening of depression based on multimodal deep learning according to claim 1, characterized in that: The step of using the original feature vector to train the Transformer specifically includes: The depression status annotated by experts is obtained as the gold standard to train the risk assessment model. The specific formula for training the current risk assessment model is: Among them, I=(I1,…,I j ,…,I N )∈R 1×N , represents the one-hot encoding of the gold standard; represents the original feature vector corresponding to the jth sample; norm(x j ) represents the vector after regularization of the original feature vector; θ represents all parameters of the model; λ represents the regularization strength; f θ represents the Transformer model function; softmax represents the softmax activation function; L(θ) represents the loss function.
6. The method for early screening of depression based on multimodal deep learning as claimed in claim 5, characterized in that: The step of training the Transformer also includes: Use the validation set to verify the shape of the loss gradient and fine-tune the parameters of the risk assessment model; Input the test set into the fine-tuned risk assessment model, and calculate the AUC based on the output value of the fine-tuned risk assessment model for evaluation; and If the AUC is greater than the threshold set by the expert, the risk assessment model corresponding to the subject is obtained, otherwise the training is continued until AUC ≥ the threshold set by the expert is satisfied.
7. The method for early screening of depression based on multimodal deep learning according to claim 6, characterized in that: The training process of the risk assessment model is optimized using the Adam optimizer, with a learning rate of 1e -5 .
8. An early screening system for depression based on multimodal deep learning that implements the method according to any one of claims 1 to 7, characterized in that: include: A data acquisition module, used to acquire multimodal data corresponding to the subject, wherein the multimodal data includes physiological signals, voice, facial expressions and text content; The feature extraction module is used to extract features from each modality data using a multi-branch deep neural network, and fuse the extracted features to construct the original feature vector of the subject; A risk assessment model training module, used to use a Transformer with a feed-forward fully connected layer as a risk assessment model, train the Transformer using the original feature vector, and obtain a risk assessment model corresponding to the subject; as well as The risk assessment module is used to input the subject data to be predicted into the risk assessment model to obtain the risk assessment result of the subject developing depression.
9. The depression early screening system based on multimodal deep learning as claimed in claim 8, characterized in that: The feature extraction module specifically includes: Data set collection unit, used to construct physiological signal data set P = {p1, ..., p i ,…,p M }、Speech dataset V={v1,…,v m ,…,v N }、Facial expression dataset F={f1,…,f m ,…,f N } and text content dataset T = {t1,…,t k ,…,t N }, and obtain the expert-annotated depression status of the corresponding subject; A feature matrix construction unit is used to transform each of the physiological signals p i The corresponding K-dimensional time domain feature vectors are combined into a feature matrix X∈R N×K ; The speech feature matrix construction unit is used to extract each speech feature vector using the pre-trained vector representation model VGGish, normalize it, and then combine it into a speech feature matrix A facial expression feature matrix construction unit is used to transform each of the facial expression frames f i The corresponding D-dimensional time-domain feature vectors are combined into the facial expression feature matrix Z∈R L×D ; The text feature matrix construction unit is used to use the pre-trained word vector representation model Word2Vec to represent the text content as a text feature matrix composed of K2-dimensional feature vectors as well as Original feature vector construction unit, used to construct the original feature vector The initial value of each dimension is set to 0, each row of the feature matrix X, the speech feature matrix Y, the facial expression feature matrix Z and the text feature matrix W is represented in the original feature vector V′, and the original feature vector V′ is normalized.
10. The depression early screening system based on multimodal deep learning according to claim 8, characterized in that: The risk assessment model training module specifically includes: The model training unit is used to obtain the depression status annotated by experts as the gold standard to train the risk assessment model. The specific formula for training the current risk assessment model is: Among them, I=(I1,…,I j ,…,I N )∈R 1×N , represents the one-hot encoding of the gold standard; represents the original feature vector corresponding to the jth sample; norm(x j ) represents the vector after regularization of the original feature vector; θ represents all parameters of the model; λ represents the regularization strength; f θ represents the Transformer model function; softmax represents the softmax activation function; L(θ) represents the loss function; A verification unit, used to use the verification set as data to verify the form of the loss gradient and fine-tune the parameters of the risk assessment model, input the test set into the fine-tuned risk assessment model, calculate the AUC according to the output value of the fine-tuned risk assessment model for evaluation, and if the AUC is greater than the threshold set by the expert, construct a risk assessment model training module to continue training until AUC ≥ the threshold set by the expert is satisfied; and The risk assessment model construction unit is used to construct a Transformer with a feed-forward fully connected layer as a risk assessment model, wherein the input of the Transformer is the original feature vector of the subject.
Citation Information
Cited By
Depression tendency detection method and device based on dynamic three-dimensional face
CN120477780A
Mild depression dynamic prediction system based on multi-modal time series data deep learning
CN121506501A