Depression state detection system based on multiple modes
By designing a multimodal depression state detection system, using the CNN-BiLSTM model to analyze the user's depression self-evaluation scale and voice input, the problem of time-consuming and subjective judgment dependence of existing evaluation methods is solved, and a fast and accurate assessment of depression state is achieved.
Patent Information
- Application Number
- CN202510040240.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
AI Technical Summary
The existing methods for assessing depression status are time-consuming and labor-intensive, and the accuracy of the results depends on the subjective judgment of clinicians and lacks effective multimodal characteristics measurements.
A multimodal depression state detection system is designed, and the user's depression self-evaluation scale and speech input are obtained through the information collection module. The preprocessing module preprocesses the speech input. The feature extraction module extracts the speech characteristics and text characteristics. The depression state detection module uses the CNN-BiLSTM model to analyze multimodal features and outputs the depression state evaluation score.
A rapid and accurate assessment of depression status is achieved, reducing the dependence on clinician subjective judgments, providing diagnostic support for depression based on multimodal characteristics, and improving the objectivity and accuracy of evaluation results.
Smart Images

Figure CN119993479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of psychological state assessment, and in particular to a multimodal-based depression state detection system. Background Art
[0002] Due to the lack of effective features (physiological or psychological aspects) to measure the disease, the current assessment of depression is mainly performed by clinicians through scoring. With the development of technology, ADE (Automatic Depression Diagnosis System) has been introduced to assist diagnosis. The multimodal depression state detection system is a method that uses deep learning technology to extract features from patients' questionnaire scores, language emotions, and semantics to identify and diagnose depression. The core of this method is to pre-train with a large amount of unlabeled data, and then fine-tune on specific downstream tasks to achieve efficient and accurate depression detection.
[0003] Real-time assessment of the severity of depressive symptoms is of great significance for the diagnosis and treatment of patients with depression. In clinical practice, the evaluation methods are mainly based on psychological scales and doctor-patient interviews, which are time-consuming and labor-intensive. At the same time, the accuracy of the results mainly depends on the subjective judgment of clinicians. With the development of artificial intelligence technology, more and more machine learning methods are used to diagnose depression through feature recognition. In order to solve the problems of traditional depression detection methods, this paper starts with the multimodal information of users and selects the path with the best performance by comparing the auxiliary diagnosis effects of multiple deep learning models in depression. Summary of the invention
[0004] To solve the above problems, the present invention proposes a multimodal depression state detection system, which can effectively solve the problems mentioned in the background technology.
[0005] To achieve the above object, the technical solution of the present invention is:
[0006] A multimodal depression state detection system includes an information collection module; used to collect depression self-rating scales submitted by users and voice input when answering depression state-inducing questions;
[0007] Preprocessing module: Use pre-trained models to preprocess speech input;
[0008] Feature extraction module: extract features from speech information, convert the processed speech into text information, and extract the emotional features of keywords in the text information through the model;
[0009] Depression state detection module: retrieve the user's multimodal features, load them into the constructed CNN-BiLSTM neural network model for analysis, and obtain the depression state assessment score;
[0010] Depression status assessment module: weighted addition of depression scale score and depression status detection module assessment score, then converted into a percentage index, and output depression status assessment result.
[0011] The information collection module includes:
[0012] The self-rating depression scale completed by the user contains 20 questions, and the assessment is scored on a scale of 1-4. The 20 items are added together to get the raw score, which is then divided by 80 and rounded to the nearest integer to get the standard score.
[0013] Voice input will record the user's answers to 20 induced questions asked by the AI, with the recording frequency being 16kHz.
[0014] The preprocessing module comprises:
[0015] The pre-trained model is used to pre-process the voice clips input by the user. The voice signal is downsampled and wavelet transformed to remove noise, the voice is normalized to the minimum and maximum value to make it conform to the maximum voice length that the depression state detection model can handle, and the voice is operated using a resampling algorithm to make it conform to the sampling rate that the depression state detection model can handle.
[0016] The feature extraction module comprises:
[0017] The vq-wav2ve model fine-tuned for the speech recognition depression state detection task is used as a feature extractor. The extracted features are average pooled and dimensionally reduced in the time dimension, and the mean and standard deviation of the features are calculated, and the emotional label is added to the extracted speech depression state features. The voice input of the consultant is converted into text data, and the converted text data is cleaned, segmented, and stop words are removed. The bag-of-words model is used to convert the text into a vector expression, and the features that have a greater impact on sentiment analysis are selected through feature selection methods.
[0018] The depression state detection module comprises:
[0019] The extracted multimodal features are subjected to deep learning, where the main algorithms to be applied are convolutional neural networks and long short-term memory networks. The LSTM module reduces the problems of gradient vanishing and gradient exploding through its gating mechanism, while the CNN-BiLSTM further improves the propagation of gradients through information flow in both the forward and reverse directions. This enables the model to utilize both past and future information at the current time step, thereby better capturing the long-range dependencies in the sequence. When faced with noisy or incomplete data, the proposed CNN-BiLSTM model's long-range dependencies and context-awareness can help the model resist interference and make more accurate predictions by utilizing information from the entire sequence.
[0020] Each LSTM unit contains four gates (forget gate ft , input gate i t , output gate o t and candidate state gate ):
[0021] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0022] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0023] o t =σ(W o ·[h t-1 ,x t ]+b o )
[0024]
[0025] The LSTM unit is based on f t and i t The output of the build status is:
[0026]
[0027] Hide output in either direction:
[0028] h t =o t ⊙tanh(c t )
[0029] Among them, σ represents the activation function, tanh represents the hyperbolic tangent activation function, ⊙ represents the element product, W f , W i , W o , W c is the corresponding weight matrix, b f , b i , b o , b c is the bias term, [h t-1 ,x t ] is the concatenation of the hidden state of the previous time step and the current input.
[0030] The CNN-BiLSTM model consists of two LSTM layers, one that propagates forward to obtain h t , a backward propagation yields The information from both can eventually be combined and used by the subsequent layers or output layers of the model. Finally, the forward and reverse hidden states can be concatenated to enhance pattern recognition capabilities and capture richer contextual information.
[0031] Get the output of BiLSTM at time t:
[0032]
[0033] The output of the depression state detection module is a score system of 0-1, and the higher the score, the more obvious the depression tendency.
[0034] The depression status assessment module includes:
[0035] The input end of the depression state assessment module is connected to the standard score of the depression self-rating scale and the output end of the depression state detection module, wherein the calculation process of the depression state assessment module is:
[0036]
[0037] Among them, S represents the depression state assessment score, S1 represents the standard score of the self-rating depression scale, S2 represents the output score of the depression state detection module, and λ is the corresponding weight score.
[0038] The depression severity index ranges from 0.25 to 1.0, with the higher the index, the more severe the depression. A depression severity index below 0.5 indicates no depression, 0.50 to 0.59 indicates mild to mild depression, 0.60 to 0.69 indicates moderate to severe depression, and 0.70 or above indicates severe depression. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the composition architecture of the multimodal depression state detection system of the present invention. DETAILED DESCRIPTION
[0040] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.
[0041] like Figure 1 As shown, the system consists of:
[0042] Information collection module
[0043] The information collection module consists of two parts: the first part provides a self-rating scale for depression, and users submit the options that best suit their current mental state based on the questions. The second part will ask 20 provocative questions by artificial intelligence, and record the voice information of users when they answer the questions.
[0044] Preprocessing module
[0045] The preprocessing module first downsamples and wavelet transforms the voice signal input by the user to remove noise, performs minimum-maximum normalization on the voice to make it conform to the maximum voice length that the depression state detection model can handle, and uses a resampling algorithm to operate the voice to make it conform to the sampling rate that the depression state detection model can handle.
[0046] Feature extraction module
[0047] The vq-wav2ve model fine-tuned for the speech recognition depression state detection task is used as a feature extractor. The extracted features are average pooled and dimensionally reduced in the time dimension, and the mean and standard deviation of the features are calculated, and the emotional label is added to the extracted speech depression state features. The voice input of the consultant is converted into text data, and the converted text data is cleaned, segmented, and stop words are removed. The bag-of-words model is used to convert the text into a vector expression, and the features that have a greater impact on sentiment analysis are selected through feature selection methods.
[0048] Depression state detection module
[0049] The depression state detection module performs deep learning on the extracted multimodal features, and the main algorithms to be applied are convolutional neural networks and long short-term memory networks. The LSTM module reduces the problems of gradient vanishing and gradient exploding through its gating mechanism, while the CNN-BiLSTM further improves the propagation of gradients through information flow in both forward and reverse directions. This enables the model to use both past and future information at the current time step, thereby better capturing long-range dependencies in the sequence. When faced with noisy or incomplete data, the proposed CNN-BiLSTM model's long-range dependencies and context-awareness can help the model resist interference and make more accurate predictions by utilizing information from the entire sequence.
[0050] The depression state detection module will output the depression state score of the two inputs of speech and semantics. The output is a score on a 0-1 scale. The higher the score, the more obvious the depression tendency.
[0051] Depression status assessment module
[0052] The depression status assessment module contains three assessment data, the self-rating depression scale score, the voice depression status score and the semantic depression status score, and the depression status assessment score is calculated by weighting. The depression status assessment score ranges from 0.25 to 1.0. The higher the index, the more severe the depression. A depression severity index below 0.5 indicates no depression, 0.50-0.59 indicates mild to mild depression, 0.60-0.69 indicates moderate to severe depression, and 0.70 or above indicates severe depression.
[0053] The depression status assessment results will be returned to the user interface to provide the function of early self-examination of depression status and assist physician diagnosis, and at the same time serve as an important basis for physicians to grasp the changes in the patient's depression status in real time.
Claims
1. A multimodal depression detection system, characterized by: The system includes: Information collection module: collects depression self-rating scale scores submitted by users and voice input in answering depression-inducing questions; Preprocessing module: Use pre-trained models to preprocess speech input; Feature extraction module: extract features from speech information, convert the processed speech into text information, and extract the emotional features of keywords in the text information through the model; Depression state detection module: retrieve the user's multimodal features, load them into the constructed CNN-BiLSTM neural network model for analysis, and obtain the depression state assessment score; Depression status assessment module: weighted addition of depression scale score and depression status detection module assessment score, then converted into a percentage index, and output depression status assessment result.
2. The multimodal depression state detection system according to claim 1, characterized in that: The information collection module includes: The self-rating depression scale completed by the user contains 20 questions, and the assessment is scored on a scale of 1-4. The 20 items are added together to get the raw score, which is then divided by 80 and rounded to get the standard score. Voice input will record the user's answers to the 20 provoking questions asked by the artificial intelligence, with a recording frequency of 16kHz.
3. The multimodal depression state detection system according to claim 1, characterized in that: The preprocessing module comprises: The pre-trained model is used to pre-process the voice clips input by the user. The voice signal is downsampled and wavelet transformed to remove noise, the voice is normalized to the minimum and maximum value to make it conform to the maximum voice length that the depression state detection model can handle, and the voice is operated using a resampling algorithm to make it conform to the sampling rate that the depression state detection model can handle.
4. The multimodal depression state detection system according to claim 1, characterized in that: The feature extraction module comprises: The vq-wav2ve model fine-tuned for the speech recognition depression state detection task is used as a feature extractor. The extracted features are average pooled and dimensionally reduced in the time dimension, and the mean and standard deviation of the features are calculated, and the emotional label is added to the extracted speech depression state features. The voice input of the consultant is converted into text data, and the converted text data is cleaned, segmented, and stop words are removed. The bag-of-words model is used to convert the text into a vector expression, and the features that have a greater impact on sentiment analysis are selected through feature selection methods.
5. The multimodal depression state detection system according to claim 1, characterized in that: The depression state detection module comprises: The extracted multimodal features are subjected to deep learning, where the main algorithms to be applied are convolutional neural networks and long short-term memory networks. The LSTM module reduces the problems of gradient vanishing and gradient exploding through its gating mechanism, while the CNN-BiLSTM further improves the propagation of gradients through information flow in both the forward and reverse directions. This enables the model to utilize both past and future information at the current time step, thereby better capturing the long-range dependencies in the sequence. When faced with noisy or incomplete data, the proposed CNN-BiLSTM model's long-range dependencies and context-awareness can help the model resist interference and make more accurate predictions by utilizing information from the entire sequence. Each LSTM unit contains four gates (forget gate f t , input gate i t , output gate o t and candidate state gate ): f t =σ(W f ·[h t-1 ,x t ]+b f ) i t =σ(W i ·[h t-1 ,x t ]+b i ) the t =σ(W o ·[h t-1 ,x t ]+b o ) The LSTM unit is based on f t and i t The output of the build status is: Hide output in either direction: h t =o t ⊙tanh(c t ) Among them, σ represents the activation function, tanh represents the hyperbolic tangent activation function, ⊙ represents the element product, W f , W i , W o , W c is the corresponding weight matrix, b f 、b i 、b o 、b c is the bias term, [h t-1 , x t ] is the concatenation of the hidden state of the previous time step and the current input. The CNN-BiLSTM model consists of two LSTM layers, one that propagates forward to obtain h t , a backward propagation yields The information from both can eventually be combined and used by the subsequent layers or output layers of the model. Finally, the forward and reverse hidden states can be concatenated to enhance pattern recognition capabilities and capture richer contextual information. Get the output of BiLSTM at time t: The output of the depression state detection module is a score system of 0-1, and the higher the score, the more obvious the depression tendency.
6. The multimodal depression state detection system according to claim 1, characterized in that: The depression status assessment module includes: The input end of the depression state assessment module is connected to the standard score of the depression self-rating scale and the output end of the depression state detection module, wherein the calculation process of the depression state assessment module is: Among them, S represents the depression state assessment score, S1 represents the standard score of the self-rating depression scale, S2 represents the output score of the depression state detection module, and λ is the corresponding weight score. The depression severity index ranges from 0.25 to 1.0, with the higher the index, the more severe the depression. A depression severity index below 0.5 indicates no depression, 0.50 to 0.59 indicates mild to mild depression, 0.60 to 0.69 indicates moderate to severe depression, and 0.70 or above indicates severe depression.
Citation Information
Cited By
Medical data processing-based depression assessment system
CN120236744A
Depression assessment system based on medical data processing
CN120236744B
Early warning assessment and intervention system for senile depression dynamic psychological health risk
CN120581209A