Depression risk assessment system based on prompt learning in missing mode

Through the depression risk assessment system based on cue learning under the missing mode, using facial images and speech information to generate missing mode features, the problems of low accuracy and high computing resource consumption of depression risk assessment system in the prior art are solved, and efficient depression risk assessment is achieved.

CN120388722APending Publication Date: 2025-07-29WANJIANG EMERGING IND TECH DEV CENT +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510079397.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing depression risk assessment system is not very accurate in the missing mode, and fine-tuning of large models will consume a lot of computing resources.

Method used

A depression risk assessment system based on cue learning in the missing mode is adopted. By collecting facial images and speech information, the cross-modal Transformer encoder and sub-attention fusion module are used to generate the features of the missing mode, and multimodal fusion and evaluation are performed based on existing modal information.

Benefits of technology

It improves the accuracy of depression risk assessment under the missing mode, reduces computational consumption, enhances the robustness of the model, and provides a reference for early detection of depression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005248042510000021
    Figure FDA0005248042510000021
  • Figure FDA0005248042510000022
    Figure FDA0005248042510000022
  • Figure FDA0005248042510000024
    Figure FDA0005248042510000024
Patent Text Reader

Abstract

The invention discloses a depression risk assessment system based on prompt learning in a missing mode. The depression risk assessment system comprises the following steps: acquiring a facial image and voice of a user; sending the collected data to a CPU processor for data processing and storing the data in a directory of the system; reconstructing missing information by using a missing modal reconstruction module in the system; a cross-modal Transformer encoder is applied to each modal to fuse information from other modals, and the information from the other modals is fused by using the cross-modal Transformer encoder; the data of the three modes are sent to a sub attention fusion module model for multi-mode fusion; and predicting a result by using the evaluation network and printing the result in a display window. The invention relates to the field of image processing. According to the depression risk assessment system based on prompt learning in the missing mode, the limitation that an existing depression risk assessment system is not high in accuracy in the missing mode is solved, and the application range of the depression risk assessment system is widened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and specifically to a depression risk assessment system based on prompt learning in the missing modality scenario. Background Art

[0002] Depression is a common and serious disease. Patients with depression usually have low mood and are taciturn, communicating less with others and having difficulty concentrating on work, which also poses certain difficulties for doctors to diagnose depression.

[0003] In recent years, deep learning has received extensive attention. Its main advantage lies in its ability to train using a large amount of data sets, thereby learning the most obvious features presented in these data. The data forms of depression are rich and diverse, including speech, images, texts, etc., which provides rich data for the training of deep learning models. However, in the real world, due to reasons such as equipment failures, data corruption, and privacy issues, the situation of missing modalities may occur, which may lead to a decline in model performance. Currently, multi-modal models trained on complete data usually perform poorly when testing incomplete data, and fine-tuning large models for downstream tasks consumes a large amount of computing resources. Therefore, generating representations corresponding to missing modalities using existing modalities while reducing the number of parameters plays a crucial role in the assessment of depression risk. Summary of the Invention

[0004] (I) Technical Problems to be Solved

[0005] In view of the deficiencies of the prior art, the present invention provides a depression risk assessment system based on prompt learning in the missing modality scenario, which solves the limitation of existing depression risk assessment only for complete modalities and improves the accuracy of depression risk assessment in the missing modality scenario.

[0006] (II) Technical Solutions

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: A depression risk assessment system based on prompt learning in the missing modality scenario, including the following steps:

[0008] S1: Collect the facial images and voices of the user.

[0009] S2: Send the collected videos to the CPU processor for data processing.

[0010] S3: Add the optimal model obtained through training to the directory of this system.

[0011] S4: Generate corresponding missing modalities according to the missing signals, and apply cross-modal Transformer encoders to each modality to fuse information from other modalities.

[0012] S5: Feed the data of the three modalities into the sub-attention fusion module for multi-modal fusion;

[0013] S6: Use the evaluation network to predict the result and print the evaluation result in the display window;

[0014] Among them, the optimal model training described in step S3 includes the following steps:

[0015] The first step is to preprocess the collected facial images and speech: Use the OpenFace tool to extract 68 three-dimensional facial key points and 4 gaze vectors from each facial image; Calculate the Log-Mel spectrogram of the speech to obtain speech features; Convert the speech to text, and use the pre-trained Universal Sentence Encoder (USE) embedding model to convert the text into an embedding vector. According to the timestamps in the text, segment the data of the three modalities into the same segments;

[0016] The second step is optimal network training: Construct and train the network model. Use three-quarters of the total multi-modal data as training samples and input them into the network structure module to train the network structure module to obtain the trained network; Use the remaining one-quarter of the data as test samples, input the data into the trained network, and test the accuracy of the model.

[0017] As an optimization, the specific steps for training the network structure module are as follows:

[0018] (1). Construct a backbone network for extracting multi-modal features: The backbone network includes visual, speech, and text feature extraction modules, and each module consists of a convolutional layer, a pooling layer, a bidirectional LSTM layer, and a fully connected layer;

[0019] (2). Construct a missing modality generation network: This network processes the inputs of the three modalities of text T, speech A, and video V and deals with the missing situation of any modality; First, initialize a learnable three-dimensional generation prompt tensor respectively representing the generation prompts for speech, video, and text modalities, which are used to generate supplementary information in the case of missing modalities; For different missing patterns, the model adopts different strategies for data completion; For example, when the input is x=(x am ,x v ,x t ), it represents the missing speech modality, and the generated feature can be expressed as Secondly, add the generated feature and the missing signal prompt, aiming to inform the corresponding Transformer whether the specific modality information is real or generated; For each modality, there are two missing signal prompts: P MS indicating modality missing, P NMSIt indicates that there is no modal loss, and the missing signal prompt can be merged in the following way: Subsequently, a cross-modal attention mechanism is constructed to model the relationships between different modalities through a cross-modal Transformer; finally, a missing type projection matrix is introduced wherein Thus, the projected missing type projection matrix P' can be obtained MT = P MT · M P , and the missing type prompt is embedded into the Transformer for attention interaction; in addition, the model also includes a self-attention layer for capturing the feature representations within each modality; finally, the features processed by the cross-modal attention and self-attention networks are extracted as the output of the model;

[0020] (3) Construct a sub-attention fusion module: This module has three layers. In the first layer, given a feature map formed by connecting the features extracted from each modality , it is first processed by a 2D-CNN layer, which will learn to capture and detect the most critical features; then, this feature is added to the input feature map Y to obtain an intermediate feature map as the input to the second-layer attention block, where C represents the number of channels and H×W represents the size of the feature map; in the second layer, given the intermediate feature map X generated from the first layer, the output channel attention weight is calculated as follows: where is the global feature attention, L(X) = BN(PWConv2·ReLU(BN(PWConv1·X))) is the local channel attention, and after obtaining the channel attention weight ω, the complementary channel attention weight 1 - ω is also calculated. The refined feature RF and the complementary refined feature can be calculated through the following formula Finally, the output of the second layer, i.e., the transitional attention feature can be used as the sum of the two refined features; in the third layer, the previous process is repeated to further enhance and strengthen the depressive features of the multi-modal transitional attention feature X'. Therefore, the final output of the sub-attention fusion feature can be expressed as

[0021] As an optimization, in step S4, the missing modality generation module processes different missing situations according to the prompt of the missing signal. There are 6 kinds of modality missing situations and one complete modality situation, and the existing modalities are used to generate the missing modalities.

[0022] As an optimization, in step S2, the Log-Mel spectrogram of the speech information is calculated in data preprocessing as the input of the speech information; 68 three-dimensional facial key points and 4 gaze vectors are extracted from the facial image as the visual information input; the pre-trained Universal Sentence Encoder embedding model is used to convert the text into an embedding vector as the text input.

[0023] As an optimization, in step S6, the output of the prediction result is the total score of the PHQ8 scale score, the binary classification determination result of whether it is moderately depressed or above, and the prediction of the depression level.

[0024] As an optimization, the visual feature extraction module uses two-dimensional convolution, the input of which is B×C×F×T. Visual features are extracted through a two-dimensional convolutional layer, batch normalization (BN), ReLU activation function, max pooling layer, and bidirectional LSTM layer. The finally output feature dimension is B×T′×D, where T′ is the length of the time dimension after pooling, and D is the hidden layer dimension of the LSTM; the speech feature extraction module and the text feature extraction module use one-dimensional convolution, the inputs of which are B×F×T and B×F×T respectively. Time series features are extracted through one-dimensional convolution, pooling, Dropout, and bidirectional LSTM modules; the outputs of the three modules all retain the time series information, providing input support for the subsequent multi-modal fusion module.

[0025] (III) Beneficial effects

[0026] The present invention provides a depression risk assessment system based on prompt learning in the absence of modalities, having the following beneficial effects:

[0027] By collecting facial images and speech information and extracting corresponding modal features, the present invention can use the existing modalities to generate the features corresponding to the missing modalities through the trained missing modality generation module in the case of partial modality absence, improving the accuracy of depression risk assessment in the absence of modalities and enhancing the robustness of the model; using the method of prompt learning to train the network only requires training some prompt information, which can reduce the number of model parameters and computational consumption; driven by the continuous development of big data and deep learning, the present invention can improve the accuracy of depression risk assessment and provide some reference opinions for the early detection of depression. Description of the drawings

[0028] Figure 1 It is a flowchart of the depression risk assessment system of the present invention;

[0029] Figure 2 It is a framework diagram of the missing modality generation network of the present invention;

[0030] Figure 3 It is a framework diagram of the sub-attention fusion module of the present invention. Detailed implementation manners

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] Please refer to Figures 1-3 , the present invention provides a technical solution: as Figure 2 shown, the model training and depression risk assessment are completed in the following steps:

[0033] First step, preprocess the collected facial images and voices: use the OpenFace tool to extract 68 three-dimensional facial key points and 4 gaze vectors for each facial image; calculate the Log-Mel spectrum for the voice to obtain voice features; convert the voice into text, and use the pre-trained Universal Sentence Encoder (USE) embedding model to convert the text into an embedding vector, and segment the data of the three modalities into the same segments according to the timestamps in the text.

[0034] Second step, optimal network training: construct and train a network model, input three-quarters of the total multi-modal data as training samples into the network model, train the network structure module, and obtain the trained network; use the remaining one-quarter of the data as test samples, input the data into the trained network, and test the accuracy of the model. The specific steps for training the network structure module are as follows:

[0035] (1). Construct a backbone network for extracting multi-modal features: The backbone network includes visual, voice, and text feature extraction modules, and each module is composed of a convolutional layer, a pooling layer, a bidirectional LSTM layer, and a fully connected layer. Among them, the visual feature extraction module uses two-dimensional convolution, and its input is B×C×F×T. Visual features are extracted through a two-dimensional convolutional layer, batch normalization (BN), ReLU activation function, max pooling layer, and bidirectional LSTM layer. The final output feature dimension is B×T′×D, where T′ is the length of the time dimension after pooling, and D is the hidden layer dimension of the LSTM; the voice feature extraction module and the text feature extraction module use one-dimensional convolution, and their inputs are B×F×T and B×F×T respectively. Time series features are extracted through one-dimensional convolution, pooling, Dropout, and bidirectional LSTM modules. The outputs of the three modules all retain time series information, providing input support for the subsequent multi-modal fusion module.

[0036] (2), Construct the Missing Modality Generation Network: This network processes inputs in three modalities: text (T), speech (A), and video (V), and deals with the missing situation of any of these modalities. First, initialize a learnable three-dimensional generation prompt tensor representing the generation prompts for speech, video, and text modalities respectively, which are used to generate supplementary information in the case of missing modalities. For different missing patterns, the model adopts different strategies for data completion. For example, when the input is x = (x am , x v , x t ), it means the speech modality is missing, and the generated feature can be represented as Second, add the generated feature and the missing signal prompt, aiming to inform the corresponding Transformer whether the modality-specific information is real or generated. For each modality, there are two missing signal prompts: P MS indicating modality missing, and P NMS indicating modality not missing. The missing signal prompts can be combined as follows: Subsequently, construct a cross-modal attention mechanism to model the relationships between different modalities through a cross-modal Transformer; finally, introduce a missing type projection matrix where Thus, the projected missing type projection matrix P' MT = P MT ·M P , and embed the missing type prompt into the Transformer for attention interaction; in addition, the model also includes a self-attention layer to capture the feature representations within each modality. Finally, the features processed by the cross-modal attention and self-attention networks are extracted as the output of the model. These features can effectively fuse the information of the three modalities and remain robust in the case of any modality missing, thereby improving the overall performance of the model.

[0037] (3), Construct the Sub-Attention Fusion Module: This module has three layers. In the first layer, given a feature map concatenated by the features extracted from each modality , it is first processed by a 2D-CNN layer, which will learn to capture and detect the most critical features; then, this feature is added to the input feature map Y to obtain an intermediate feature map as the input to the second-layer attention block, where C represents the number of channels and H×W represents the size of the feature map; in the second layer, given the intermediate feature map X generated from the first layer, the output channel attention weight is calculated as: where is the global feature attention, L(X)=BN(PWConv2·ReLU(BN(PWConv1·X))) is the local channel attention. After obtaining the channel attention weight ω, the complementary channel attention weight 1 - ω is also calculated. The refined feature RF and the complementary refined feature can be calculated through the following formula Finally, the output of the second layer, that is, the transitional attention feature can be used as the sum of the two refined features; in the third layer, the previous process will be carried out again to further enhance and strengthen the depressive features of the multi-modal transitional attention feature X'. Therefore, the final sub-attention fusion feature output can be expressed as

[0038] The third step, the depressive risk assessment function: The feature output after being fused by the sub-attention fusion module is input into the evaluation network for result prediction. There are a total of 8 classification heads for predicting the PHQ-8 sub-scores, and the total score of PHQ-8 is statistically calculated to determine moderate to severe depression and above, and the depressive level is divided according to the total score, and it is divided into five levels according to the score range.

[0039] Based on the above steps, as Figure 1 shown, a depressive risk assessment system based on prompt learning under a missing modality, the specific operation steps are as follows

[0040] S1: Collect the facial images and voices of the users;

[0041] S2: Send the collected videos into the CPU processor for data processing;

[0042] S3: Add the optimal model obtained by training under the directory of this system;

[0043] S4: Generate the corresponding missing modality according to the missing signal, and apply the cross-modal Transformer encoder to each modality to fuse the information from other modalities;

[0044] S5: Send the data of the three modalities into the sub-attention fusion module for multi-modal fusion;

[0045] S6: Use the evaluation network to predict the result and print the evaluation result in the display window.

[0046] In the optimization process of the entire network of the present invention, the network gradually learns the parameters of the three prompts when the missing modality is generated best. This process can be regarded as the model learning the relationship between multi-modal data. The model can utilize the available modalities and generate the corresponding missing modality through the relationship between different modalities. At this time, the model has the ability to perform depressive risk assessment under the missing modality with high accuracy.

[0047] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A depression risk assessment system based on prompt learning in the missing modality, characterized in that: It includes the following steps: S1: Collect the facial images and voices of users; S2: Send the collected videos to the CPU processor for data processing; S3: Add the optimal model obtained through training to the directory of this system; S4: Generate corresponding missing modalities according to the missing signals, and apply cross-modal Transformer encoders to each modality to fuse information from other modalities; S5: Send the data of the three modalities to the sub-attention fusion module for multi-modal fusion; S6: Use the evaluation network to predict the results and print the evaluation results in the display window; Among them, the training of the optimal model described in step S3 includes the following steps: The first step is to preprocess the collected facial images and voices: Use the OpenFace tool to extract 68 three-dimensional facial key points and 4 gaze vectors for each facial image; Calculate the Log-Mel spectrogram for the voice to obtain voice features; Convert the voice to text, use the pre-trained Universal Sentence Encoder (USE) embedding model to convert the text into an embedding vector, and segment the data of the three modalities into the same segments according to the timestamps in the text; The second step is the optimal network training: Construct and train the network model, input three-quarters of the total multi-modal data as training samples into the network structure module, train the network structure module to obtain the trained network; Use the remaining one-quarter of the data as test samples, input the data into the trained network, and test the accuracy of the model.

2. The depression risk assessment system based on prompt learning in the missing modality according to claim 1, characterized in that: The specific steps for training the network structure module are as follows: (1). Construct a backbone network for extracting multi-modal features: The backbone network includes visual, voice, and text feature extraction modules, and each module is composed of a convolutional layer, a pooling layer, a bidirectional LSTM layer, and a fully connected layer; (2) Construct a missing modality generation network: This network processes inputs from three modalities, namely text T, speech A, and video V, and handles the missing situation of any of these modalities; First, initialize a learnable three-dimensional generation prompt tensor which represents the generation prompts for speech, video, and text modalities respectively, and is used to generate supplementary information in the case of missing modalities; For different missing patterns, the model adopts different strategies for data completion; Second, add the generated features and the missing signal prompts, aiming to inform the corresponding Transformer whether the modality-specific information is real or generated; For each modality, there are two missing signal cues: P MS indicating modality missing, P NMS indicating modality not missing. The missing signal cues can be combined as follows: Subsequently, a cross-modal attention mechanism is constructed to model the relationships between different modalities through a cross-modal Transformer; finally, a missing type projection matrix is introduced, where the projected missing type projection matrix P' can be obtained as MT = P MT · M P , and the missing type cues are embedded into the Transformer for attention interaction; in addition, the model also includes a self-attention layer to capture the feature representations within each modality; Finally, the features processed by the cross-modal attention and self-attention networks are extracted as the output of the model; (3) Constructor Attention Fusion Module: This module has three layers. In the first layer, given a feature map formed by concatenating the features extracted from each modality it is first processed by a 2D-CNN layer which will learn to capture and detect the most critical features. Then, this feature is added to the input feature map Y to obtain an intermediate feature map which serves as the input to the second layer attention block, where C represents the number of channels and H×W represents the size of the feature map. In the second layer, given the intermediate feature map X generated from the first layer, the output channel attention weights are calculated as follows: where is the global feature attention, L(X) = BN(PWConv2·ReLU(BN(PWConv1·X))) is the local channel attention. After obtaining the channel attention weight ω, the complementary channel attention weight 1 - ω is also calculated. The refined feature RF and the complementary refined feature can be calculated through the following formula Finally, the output of the second layer, i.e., the transitional attention feature can be the sum of the two refined features. In the third layer, the previous process is carried out again to further enhance and strengthen the depressive features of the multi-modal transitional attention feature X'. Therefore, the final output of the sub-attention fusion feature can be expressed as 3. The depression risk assessment system based on prompt learning in the missing modality according to claim 1, characterized in that: In step S4, the missing modality generation module processes different missing situations according to the prompts of the missing signals. There are a total of 6 modality missing situations and one complete modality situation, and the existing modalities are used to generate the missing modalities.

4. A depression risk assessment system based on prompt learning in the missing modality according to claim 1, characterized in that: In step S2, in the data preprocessing, calculate the Log-Mel spectrogram of the voice information as the input of the voice information; Extract 68 three-dimensional facial key points and 4 gaze vectors for the facial image as the visual information input; Use the pre-trained Universal Sentence Encoder embedding model to convert the text into an embedding vector as the text input.

5. The cross-modal footprint image retrieval system based on weakening modal differences according to claim 1, characterized in that: In step S6, the output of the prediction result is the total score of the PHQ8 scale score, the binary classification determination result of whether it is moderately depressed or above, and the prediction of the depression level.

6. The cross-modal footprint image retrieval system based on weakening modal differences according to claim 2, wherein: The visual feature extraction module uses two-dimensional convolution, and its input is B×C×F×T. Visual features are extracted through a two-dimensional convolutional layer, batch normalization (BN), ReLU activation function, max pooling layer, and bidirectional LSTM layer. The final output feature dimension is B×T′×D, where T′ is the length of the time dimension after pooling, and D is the hidden layer dimension of the LSTM; The speech feature extraction module and the text feature extraction module use one-dimensional convolution, with their inputs being B×F×T and B×F×T respectively. Through one-dimensional convolution, pooling, Dropout, and bidirectional LSTM modules, time series features are extracted; The outputs of the three modules all retain time series information, providing input support for the subsequent multi-modal fusion module.

Citation Information

Patent Citations

  • Depression state recognition method based on multi-modal fusion

    CN113674767A

  • Multi-modal depression detection method and system based on full attention mechanism

    CN114898861A

  • Method for auxiliary detection of crowd depression state based on multi-modal deep neural network

    CN116110565A

  • Multi-domain feature extraction and fusion identification method for water supply pipeline leakage

    CN116753471A

  • Multimodal dynamic attention fusion

    US20220392637A1