A Multimodal and Multifactor Depression Recognition System Incorporating Emotional Information

The system addresses the limitations of current depression detection methods by employing advanced models and fusion techniques to capture contextual information and individual differences, enhancing the accuracy and adaptability of emotion recognition.

CN119786021BActive Publication Date: 2025-07-15NORTHEASTERN UNIV CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510272051.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-15
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The existing multimodal depression detection methods have shortcomings in feature extraction, emotional feature extraction, multimodal feature fusion and individual differences considerations, resulting in insufficient emotional recognition and misjudgment. Especially in multiple rounds of dialogue, the contextual understanding cannot be fully understood and cannot effectively capture the complexity and individual differences of user emotions.

Method used

The preprocessing module, text feature extraction module, speech feature extraction module, multimodal fusion module and multi-factor fusion module are adopted to extract text and speech emotional characteristics through BERT and Wav2Vec2.0, and deeply integrate self-attention and cross-attention mechanisms, and take into account gender differences and external common sense knowledge base COMET to conduct common sense reasoning to build a classification model for men and women.

Benefits of technology

It improves the accuracy and robustness of emotion recognition, can better understand user language and emotional fluctuations, capture the dependencies between modals, adapt to emotional expressions of users of different genders, and improves the accuracy and adaptability of depression detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119786021B_ABST
    Figure CN119786021B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-modal multi-factor depression recognition system integrating emotion information, which relates to the field of multi-modal fusion technology. The present invention proposes to extract more accurate emotion features through transfer learning and multi-dataset training methods, so as to solve the problem of insufficient emotion recognition in existing methods. The present invention also extracts features in units of clauses to retain complete context information for better understanding of the user's language and emotional fluctuations. To optimize multi-modal data fusion, the present invention adopts a fusion method based on the attention mechanism, which can not only more effectively capture the dependencies between modalities, but also fully extract the key information in each modality, thereby improving the overall detection effect. At the same time, the present invention also takes into account the influence of gender differences and external knowledge, and improves the model's ability to accurately recognize and understand the emotions of users of different genders by introducing gender differences and external common sense knowledge bases, further enhancing the adaptability and accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal fusion, and more particularly to a multimodal multi-factor depression recognition system that integrates emotional information. Background Art

[0002] Depression has become the second largest health killer globally, seriously affecting the quality of life and health level of humans. Traditional depression diagnosis methods mainly rely on doctors' clinical judgments, through patients' self-reports, psychological assessment scales (such as the Beck Depression Inventory), and doctors' clinical observations. However, this method not only requires high medical costs but also has a long diagnosis cycle. At the same time, the symptoms of depression are complex and vary among individuals, which easily leads to misdiagnosis or missed diagnosis. Therefore, there is an urgent need for more efficient, accurate, and low-cost early detection methods for depression. With the development of technology, using computer-aided depression detection has become a popular research direction in academia in recent years. Among them, depression detection technology based on multimodal data (text and speech) is particularly remarkable. By analyzing patients' speech and text data, computers can identify potential depressive symptoms, thus providing important auxiliary diagnostic tools for doctors.

[0003] Currently, computer-aided depression detection methods mainly include three methods: text-based, speech-based, and multimodal fusion-based.

[0004] Text-based depression detection, especially deep learning methods, can identify depression more accurately than traditional methods by automatically extracting emotional features from text. Deep learning models such as BERT and LSTM can handle more complex language features and understand emotional changes in context. Different from traditional machine learning methods, deep learning does not require manual feature extraction but learns how to identify depression-related patterns in text through a large amount of data. BERT learns the deep semantic relationships of language through pre-training on a large-scale corpus, while LSTM can capture long-term dependencies in text, especially suitable for dealing with emotional fluctuations and complex contexts. For example, Wang et al. achieved better results than traditional machine learning methods by analyzing microblog data in combination with deep learning models.

[0005] Speech-based depression detection, especially deep learning methods, utilizes automatic feature extraction techniques to identify depression by analyzing deep emotional information in speech. Deep learning models, such as convolutional neural networks (CNNs) and long short-term memory networks (LSTMs), can extract complex emotional features from raw speech signals or spectrograms without manual feature design. CNNs are capable of processing local features of speech and are suitable for identifying short-term changes in sound, while LSTMs can capture long-term dependencies in speech and identify long-term emotional fluctuations such as rhythm and speech rate changes in the speech of depression patients. For example, He et al. combined handcrafted features and deep features and used a deep convolutional neural network for classification prediction, achieving good results.

[0006] Depression detection based on multimodal data has become a research hotspot by combining two or more data sources to identify depression. Compared with single modality, multimodal data is more robust. When a certain modality is missing, prediction can still be made through other modalities, while providing richer information to help identify depression more accurately. Currently, multimodal-based depression detection methods are mainly divided into three categories: feature-level fusion, decision-level fusion, and model fusion.

[0007] Feature-level fusion concatenates the features of multiple modalities and then conducts classification. For example, Gong et al. extracted text, speech, and video features for concatenation and used a random forest for classification prediction. Decision-level fusion combines the results after independent prediction of each modality, such as weighted summation or voting. Meng et al. weighted and fused the prediction results of speech and video, showing that multimodality performs better than unimodality. Model fusion captures the relationships between modalities by deeply fusing the features of different modalities, thereby improving the prediction accuracy. Li et al. constructed a multimodal hierarchical attention model.

[0008] Although certain research progress has been made in multimodal depression detection technology, existing methods still face many challenges in practical applications:

[0009] (1) Insufficient feature extraction: Most current multimodal depression detection methods rely on phonemes, sentences, or words as units for feature extraction. Although this method is effective in some scenarios, it cannot capture long-distance dependencies and context information in the user's language, resulting in incomplete context understanding. Especially in multi-turn conversations, it is often difficult to capture emotional fluctuations and user mood changes in the context, which may lead to deviations and misjudgments in emotion recognition.

[0010] (2) Limitations of emotion feature extraction: Most existing emotion recognition technologies rely on pre-set emotion lexicons. These emotion lexicons cover a large number of emotional words, but they usually have certain limitations. For example, these lexicons often lack the ability to capture delicate emotional changes and cannot effectively reflect the diversity and complexity of users' emotions. Especially for atypical emotion expressions (such as the emotional expressions of patients with hidden depression), existing emotion lexicons may not be able to fully recognize them. In addition, cultural differences, language differences, and individual differences in emotion lexicons may also affect the universality and accuracy of emotion feature extraction.

[0011] (3) Singularity of multimodal feature fusion: Multimodal data (text and speech) each contain unique information. Text data mainly reflects the user's language content and expression style, while speech data reveals the user's emotional fluctuations. To make full use of these two types of modal data, efficient fusion methods must be adopted. However, many existing multimodal fusion methods are still relatively simple and usually adopt rough fusion methods such as splicing or weighted averaging, resulting in the key information in different modal data not being fully extracted and integrated. For example, the emotional color in text information may not be fully fused with the emotional information in speech, thus affecting the overall emotion recognition effect.

[0012] (4) Ignoring external knowledge and individual differences: Many current depression recognition systems usually fail to consider external factors related to the user's feelings, mental state, life background, etc. For example, factors such as gender, age, and social background often play a crucial role in the manifestation of depression. Ignoring these factors will cause the model to have a deviation in understanding the user's emotions. In addition, the symptom manifestations of depression patients usually have individual differences, and it is difficult for traditional "one-size-fits-all" methods to capture these individual differences, resulting in a decline in the universality and accuracy of the model. Summary of the Invention

[0013] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to propose a multimodal and multi-factor depression recognition system that fuses emotion information, including a preprocessing module, a text feature extraction module, a speech feature extraction module, a multimodal fusion module, and a multi-factor fusion module;

[0014] The preprocessing module is used to preprocess the original speech of the user to be detected, obtain the preprocessed text as a clause, and a speech waveform diagram, where the abscissa of the speech waveform diagram is time and the ordinate is amplitude;

[0015] The text feature extraction module is used to add [CLS] at the start position of the preprocessed clause and [SEP] at the end position, segment the clause after adding the markers through a tokenizer to obtain the segmented clause; extract features from the segmented clause through BERT to obtain semantic featuresf ts ; Process the segmented clauses through a text emotion extraction model to obtain text emotion features f te , and combine the semantic features f ts and the text emotion features f te into text features f t ;

[0016] Among them, the text emotion extraction model is obtained through the following method:

[0017] Obtain the clauses in the SST-2 dataset as the first input samples, and use the emotion labels corresponding to the clauses in the SST-2 dataset as the first output samples. The first input samples and the first output samples form the first training samples. Based on multiple first training samples, train the RoBERTa model until the RoBERTa model converges to obtain the trained RoBERTa model. Retain the parameters and encoder structure of the trained RoBERTa model, and replace the output layer of the trained RoBERTa model with a fully connected layer FC to obtain the text emotion extraction model;

[0018] The speech feature extraction module is used to extract features from the speech waveform diagram to obtain deep speech features f ad , input the original speech into the speech emotion extraction model to obtain speech emotion features f ae , and splice the deep speech features f ad and the speech emotion features f ae to obtain complete speech features f a ;

[0019] Among them, the speech emotion extraction model is obtained by training the Wav2Vec2.0 model based on multiple second training samples. The second training samples include second input samples and second output samples. The second input samples are the speech samples in the IEMOCAP dataset, and the second output samples are the speech emotion features corresponding to the speech samples in the IEMOCAP dataset;

[0020] The multimodal fusion module is used to generate in-text-modal features according to the text features f t and the speech features f a , S t, Features within the speech modality S a , Features in speech associated with text S ta and features in text associated with speech S at , Based on features within the text modality S t , Features within the speech modality S a , Features in speech associated with text S ta and features in text associated with speech S at , Calculate to obtain the first fused multi-modal feature and the second fused multi-modal feature ;

[0021] The multi-factor fusion module is used to determine the gender of the user to be detected. When the user to be detected is male, process the first fused multi-modal feature through a classification model for males to obtain the depression prediction result of the user to be detected. When the user to be detected is female, process the second fused multi-modal feature through a classification model for females to obtain the depression prediction result of the user to be detected.

[0022] Optionally, the preprocessing module is specifically used to obtain the original speech of the user to be detected, determine the triple <start time, end time, clause> of the user to be detected in the original speech, use regular expressions to perform data cleaning on the clause to obtain the preprocessed clause; perform pre-emphasis processing on the original speech to obtain the pre-emphasized speech. Specifically, enhance the high-frequency part of the original speech and improve the high-frequency resolution of the original speech. Perform frame segmentation on the pre-emphasized speech, apply the Hanning window function to the segmented speech to enhance the speech signal, and obtain the enhanced speech signal. Draw a speech waveform diagram based on the enhanced speech signal.

[0023] Optionally, in the text feature extraction module, BERT is used to extract features from the segmented clauses to obtain semantic features f ts , including:

[0024] Obtain word embeddings, position embeddings, and token embeddings in the segmented clauses, add the word embeddings, position embeddings, and token embeddings to obtain the added embeddings, and input the added embeddings into the Transformer encoder for encoding to obtain semantic features f ts .

[0025] Optionally, the voice feature extraction module extracts features from the voice waveform diagram to obtain deep voice features f ad , including:

[0026] Perform grayscale processing on the voice waveform diagram to obtain the waveform diagram after grayscale processing, perform normalization and standardization on the waveform diagram after grayscale processing to obtain the processed waveform diagram, the pixel values in the processed waveform diagram are in the range of [0,1], and the pixel values have a unified mean and standard deviation, convert the processed waveform diagram into a one-dimensional matrix and input it into the TCN to obtain deep voice features f ad .

[0027] Optionally, the multimodal fusion module is specifically used to input the text features f t into the self-attention Transformer to obtain the in-text-modal features S t , which is specifically implemented through the following formula:

[0028] ;

[0029] ;

[0030] ;

[0031] ;

[0032] where W Qt , W Kt , W Vt are parameter matrices, T represents the transpose of the matrix, d k is the vector dimension, sum represents summation;

[0033] Input the voice features f a into the self-attention Transformer to obtain the in-voice-modal features S a , which is specifically implemented through the following formula:

[0034] ;

[0035] ;

[0036] ;

[0037] ;

[0038] Among them, W Qa 、 W Ka 、 W Va are parameter matrices;

[0039] Through the cross-attention mechanism, according to the features within the text modality S t and the features within the speech modality S a , calculate the features associated with the text in the speech S ta , which is represented by the following formula:

[0040] ;

[0041] Through the cross-attention mechanism, according to the features within the text modality S t and the features within the speech modality S a , calculate the features associated with the speech in the text S at , which is represented by the following formula:

[0042] ;

[0043] Perform a concatenation operation on the features within the text modality S t , the features within the speech modality S a , the features associated with the text in the speech S ta and the features associated with the speech S at to obtain the matrix U M = S t , S a , S ta , S at , and input the matrix U M = S t , S a , S ta , S at into two layers Transformer EncoderIn this process, the first fused multimodal feature is obtained , which is specifically represented by the following formula:

[0044] =Transformer ( U M );

[0045] For the features within the text modality S t , the features within the speech modality S a , the features in speech related to the text S ta and the features related to speech S at , a summation operation is performed to obtain the matrix U F = S t , S a , S ta , S at . The matrix U F = S t , S a , S ta , S at is input into two layers Transformer Encoder to obtain the second fused multimodal feature , which is specifically represented by the following formula:

[0046] =Transformer ( U F ).

[0047] Optionally, Cross-Attention ( S t , S a ) is specifically calculated by the following formula:

[0048] ;

[0049] ;

[0050] ;

[0051] ;

[0052] Similarly, Cross-Attention ( S a , S t ) is specifically calculated through the following formula:

[0053] ;

[0054] ;

[0055] ;

[0056] .

[0057] Optionally, in the multi-factor fusion module, the first fused multi-modal feature is processed through a classification model for men to obtain the depression prediction result of the user to be detected, including: Obtain the inference relationship related to emotions and feelings

[0058] , and process the preprocessed clauses and inference relationships according to the common sense knowledge base COMET to obtain common sense knowledge, which is specifically implemented through the following formula: xReact xReact

[0059] ;

[0060] s ti where s ti is the i -th word in the preprocessed clause, s ki is the common sense knowledge corresponding to the i -th word, and ⊕ represents concatenation;

[0061] Input the common sense knowledge s ki into BERT for encoding to obtain the embedded representation s ki of the common sense knowledge b i , which is specifically implemented through the following formula:

[0062] ;

[0063] The embedded representations of all common sense knowledge form a sentence vector B = { b 1, b 2,…, b m}, and input the sentence vector B into Transformer to obtain the target common sense knowledge K uis achieved through the following formula:

[0064] ;

[0065] Concatenate the target common sense knowledge K u and the first fused multi-modal feature to obtain the first feature , and input the first feature into BiGRU the classifier to obtain the feature containing forward and backward information, which is specifically represented by the following formula:

[0066] ;

[0067] Input the feature containing forward and backward information into the fully connected layer to obtain the depression prediction result y’ of the user to be detected, which is specifically represented by the following formula:

[0068] .

[0069] Optionally, in the multi-factor fusion module, the second fused multi-modal feature is processed by a female-oriented classification model to obtain the depression prediction result of the user to be detected, including:

[0070] Obtain the inference relationship related to emotions and feelings xReact , and process the preprocessed clause and inference relationship xReact according to the common sense knowledge base COMET to obtain common sense knowledge, which is specifically achieved through the following formula:

[0071] ;

[0072] Among them, s ti is the i -th word in the preprocessed clause, s ki is the common sense knowledge corresponding to the i -th word;

[0073] Input the common sense knowledge s ki into BERT for encoding to obtain the embedded representation s ki of the common sense knowledge b i , which is specifically achieved through the following formula:

[0074] ;

[0075] The embedded representation of all common sense knowledge forms the sentence vector B = { b 1, b 2, …, b m}. Input the sentence vector B into Transformer to obtain the target common sense knowledge K u , which is achieved through the following formula:

[0076] ;

[0077] Concatenate the target common sense knowledge K u and the second fused multi-modal feature to obtain the second feature . Input the second feature into the MLP classifier to obtain the depression prediction result y’ of the user to be detected, which is specifically represented by the following formula:

[0078] .

[0079] Optionally, the specific calculation is achieved through the following formula:

[0080] ;

[0081] ;

[0082] ;

[0083] where , and are parameter matrices, , and are bias parameters, is the output of the first hidden layer, is the output of the second hidden layer.

[0084] Optionally, the multi-modal multi-factor depression recognition system that fuses emotional information further includes a loss module;

[0085] The loss module uses the cross-entropy loss function as the loss function, which is specifically represented by the following formula:

[0086] ;

[0087] where is the true depression result of the j th sample, is the prediction result of depression for the j th sample;

[0088] The value of the loss function is used to update the parameters in the preprocessing module, text feature extraction module, speech feature extraction module, multimodal fusion module, and multi-factor fusion module.

[0089] The beneficial effects of adopting the above technical solutions are as follows:

[0090] The present invention proposes to extract more accurate emotion features through transfer learning and multi-dataset training methods, thereby solving the problem of insufficient emotion recognition in existing methods. In addition, the invention also extracts features in units of clauses to retain complete context information for better understanding of the user's language and emotional fluctuations. To optimize multimodal data fusion, the present invention adopts a fusion method based on the attention mechanism, which can not only more effectively capture the dependencies between modalities but also fully extract the key information in each modality, thereby improving the overall detection effect. At the same time, the present invention also takes into account the influence of gender differences and external knowledge, and improves the model's ability to accurately recognize and understand the emotions of users of different genders by introducing gender differences and external common sense knowledge bases, further enhancing the adaptability and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Figure 1 is a schematic structural diagram of a multimodal multi-factor depression recognition system integrating emotion information in an embodiment of the present invention;

[0092] Figure 2 is a schematic flow diagram of a multimodal multi-factor depression recognition system integrating emotion information in an embodiment of the present invention;

[0093] Figure 3 is a schematic flow diagram of determining a triple of a user to be detected in the original speech in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0094] The following further describes in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.

[0095] In view of the problems existing in the prior art, the present invention provides a multimodal multi-factor depression recognition system integrating emotion information, in combination with Figure 1 , including a preprocessing module, a text feature extraction module, a speech feature extraction module, a multimodal fusion module, and a multi-factor fusion module;

[0096] Among them, in the text feature extraction module, first use BERT to extract text semantic features to obtain sentence vectors containing semantic information, then use the pre-trained RoBERTa model to extract the emotional information contained in the text, and then concatenate the semantic features and emotional features to form text features; in the speech feature extraction module, first use TCN to extract deep speech features from the speech waveform diagram, then use the pre-trained Wav2Vec2.0 to extract speech emotional features from the original speech signal, and concatenate the deep speech features and speech emotional features to form speech features; in the multimodal fusion module, deeply fuse text features and speech features based on the self-attention mechanism and cross-attention mechanism to obtain deeply fused modal features; the multi-factor fusion module will build two classification models for depression recognition according to the different characteristics of men and women, introduce external knowledge, and use the common sense knowledge base COMET for common sense reasoning to obtain the user's emotional changes. Finally, the fused features are sent into different classifiers according to the differences between men and women to obtain the detection labels of the user, that is, the depression prediction result.

[0097] The purpose of the present invention is to improve the accuracy and robustness of the automatic depression detection system, especially in terms of emotional feature extraction, context understanding, multimodal data fusion, and gender differences and external knowledge integration. To improve the deficiencies of the existing technology, the present invention proposes to extract more accurate emotional features through transfer learning and multi-dataset training methods, so as to solve the problem of insufficient emotion recognition in the existing methods. In addition, the invention also extracts features in units of clauses to retain complete context information to better understand the user's language and emotional fluctuations. To optimize multimodal data fusion, the present invention adopts a fusion method based on the attention mechanism, which can not only more effectively capture the dependencies between modalities, but also fully extract the key information in each modality, thereby improving the overall detection effect. At the same time, the invention also takes into account the influence of gender differences and external knowledge, and improves the model's ability to accurately recognize and understand the emotions of users of different genders by introducing gender differences and external common sense knowledge bases, further improving the adaptability and accuracy of the model.

[0098] Combined Figure 2 , the specific content executed by each module includes:

[0099] The preprocessing module is used to preprocess the original speech of the user to be detected to obtain the preprocessed text as a clause and the speech waveform diagram, where the abscissa of the speech waveform diagram is time and the ordinate of the speech waveform diagram is amplitude;

[0100] The preprocessing module is specifically used to obtain the original speech of the user to be detected and determine the triple <start time, end time, clause> of the user to be detected in the original speech;

[0101] Specifically combinedFigure 3 For the original voice signal, first, the identity of the current speaker is discriminated, that is, it is judged whether the current speaker is the user to be detected, namely Figure 3 the participant in , when the current speaker is a robot, it is judged whether there is a next sentence; when the current speaker is the user to be detected, the triple <start time, end time, clause> is updated, and the start time and end time are adjusted according to the identity of the speaker of adjacent sentences. If the speakers of adjacent sentences are both the user to be detected, that is, the participant, only the clause is updated, and then it is judged whether there is a next sentence. If the speaker of the adjacent sentence is not the user to be detected, the start time or end time of the triple needs to be updated, and then it is judged whether there is a next sentence; if there is a next sentence, return: discriminate the identity of the current speaker, if there is no next sentence, return, and thus the triple <start time, end time, clause> of the user to be detected is obtained.

[0102] Using regular expressions, data cleaning is performed on the clause to remove irrelevant characters (such as " <an> ”" <fam> ”" <sync>", etc.) to avoid interference and obtain the preprocessed clause; perform pre-emphasis processing on the original speech to obtain the pre-emphasized speech. Specifically, enhance the high-frequency part of the original speech to remove the influence of lip radiation and improve the high-frequency resolution of the original speech. Since the speech signal is a non-stationary signal but has short-term stationarity within a short period, perform frame segmentation on the pre-emphasized speech, set the frame length to 25 ms, the frame shift to 10 ms, and use the overlapping segmentation method for transition. Apply the Hanning window function to the segmented speech to enhance the speech signal and obtain the enhanced speech signal. Draw a speech waveform diagram based on the enhanced speech signal, where the abscissa of the speech waveform diagram is time and the ordinate is amplitude.

[0103] When testing the system of the present invention, the DAIC-WOZ dataset is used. The DAIC-WOZ dataset has the following characteristics: Although there are pauses when the participants answer questions, their speech content is coherent, related to each other and complementary. Different from the previous research that only extracts features based on a single utterance, the present invention regards these continuous utterances as a complete clause to fully retain the user context.

[0104] The text feature extraction module is used to add [CLS] at the starting position of the preprocessed clause and [SEP] at the ending position, and segment the clause after adding the markers through a tokenizer to obtain the segmented clause; extract features from the segmented clause through BERT to obtain semantic features f ts ; process the segmented clause through a text emotion extraction model to obtain text emotion features f te , and combine the semantic features f ts and the text emotion features f te into text features f t ;

[0105] Among them, the text emotion extraction model is obtained through the following method:

[0106] Obtain the clauses in the SST-2 dataset as the first input samples, and use the sentiment labels corresponding to the clauses in the SST-2 dataset as the first output samples. The first input samples and the first output samples form the first training samples. Based on multiple first training samples, train the RoBERTa model until the RoBERTa model converges to obtain the trained RoBERTa model. Retain the parameters and encoder structure of the trained RoBERTa model, and replace the output layer of the trained RoBERTa model with a fully connected layer FC to obtain a text sentiment extraction model. After obtaining the text sentiment extraction model, transfer it to the segmented clauses to obtain the sentiment features of the clauses. That is, the finally obtained text sentiment extraction model can be used to extract text sentiment features from clauses.

[0107] Among them, BERT is used to extract features from the segmented clauses to obtain semantic features. f ts , including:

[0108] Obtain token embeddings, position embeddings, and segment embeddings in the segmented clauses, add the token embeddings, position embeddings, and segment embeddings to get the added embeddings, and input the added embeddings into the Transformer encoder for encoding to obtain semantic features. f ts .

[0109] The speech feature extraction module is used to extract features from the speech waveform diagram through TCN to obtain deep speech features. f ad , input the original speech into the speech sentiment extraction model to obtain speech sentiment features. f ae , concatenate the deep speech features f ad and the speech sentiment features f ae to obtain the complete speech features. f a ;

[0110] Among them, the speech sentiment extraction model is obtained by training the Wav2Vec2.0 model based on multiple second training samples. The second training samples include second input samples and second output samples. The second input samples are the speech samples in the IEMOCAP dataset, and the second output samples are the speech sentiment features corresponding to the speech samples in the IEMOCAP dataset.

[0111] The training process specifically includes:

[0112] Historical speech is obtained from the IEMOCAP dataset. A speech matrix is constructed based on the historical speech and input into the FeatureEncoder module. After convolution and activation processing, a dense representation of the speech is obtained. Then, the global information is captured through the Transformer module in the Context Network, followed by a masking operation and input into the Quantization Module to convert continuous features into discrete features. Finally, the processed features are fed into a linear layer to obtain the predicted speech emotion features. Based on the predicted speech emotion features and the speech emotion features corresponding to the historical speech in the IEMOCAP dataset, the parameters in the Wav2Vec2.0 model are updated until the model converges, resulting in a speech emotion extraction model. The optimal model parameters are saved and migrated to the original speech to obtain the emotion features of the speech. That is, the trained speech emotion extraction model can be used to obtain the emotion features of the speech from the original speech.

[0113] Since the waveform diagram contains information such as the energy, pitch, loudness, and timbre of the speech, it is more robust than the spectrogram. As the waveform diagram is time-related, traditional CNNs are not suitable for processing time-series data, while TCN can effectively process images and extract time-series features. Therefore, TCN is adopted in the present invention for feature extraction.

[0114] Thus, feature extraction is performed on the speech waveform diagram to obtain deep speech features f ad , including:

[0115] The speech waveform diagram is grayscale processed, and the calculation is simplified through weighted averaging to obtain the grayscaled waveform diagram. The grayscaled waveform diagram is normalized and standardized to obtain the processed waveform diagram. The pixel values in the processed waveform diagram are in the range of [0,1], and the pixel values have a unified mean and standard deviation. The processed waveform diagram is converted into a one-dimensional matrix and input into the TCN to obtain deep speech features f ad , specifically, to ensure that the input dimensions of each layer are consistent, a Padding operation is first performed, and then convolution calculations are carried out through a dilated causal convolution layer. After convolution, steps such as pruning, weight normalization, activation, Dropout, and residual connection are performed, and finally, deep speech features are output through a linear layer f ad .

[0116] The multi-modal fusion methods in automatic depression detection are mainly divided into three categories: feature-level fusion, decision-level fusion, and model fusion. The first two are shallow fusions. Although effective, they cannot fully capture the intra-modal and inter-modal relationships. Therefore, deep fusion methods are needed. The present invention proposes a model fusion method based on self-attention and cross-attention, which uses self-attention to capture intra-modal information, and cross-attention to strengthen the important correlation features between modalities, specifically implemented through a multi-modal fusion module.

[0117] The multi-modal fusion module is used to generate intra-text modality features f t and intra-speech modality features f a according to text features S t and speech features S a features associated with text in speech S ta and features associated with speech in text S at . According to the intra-text modality features S t , intra-speech modality features S a , features associated with text in speech S ta and features associated with speech in text S at , the first fused multi-modal feature and the second fused multi-modal feature are calculated;

[0118] Specifically, the multi-modal fusion module is used to input text features f t into a self-attention Transformer to obtain intra-text modality features S t , which can be represented by the following formula:

[0119] Transformer ( )

[0120] Specifically, the corresponding query vector (Q), key vector (K), value vector (V), and weight matrix are calculated, the weight matrix is normalized, then the weight matrix is multiplied by V to obtain the self-attention output, and then the self-attention output is summed to obtain the intra-text modality feature containing important intra-modal information S t , which is specifically implemented through the following formula:

[0121] ;

[0122] ;

[0123] ;

[0124] ;

[0125] Among them, W Qt and W Kt and W Vt are parameter matrices, T represents the transpose of the matrix, d k is the vector dimension, sum represents summation;

[0126] Input the speech feature f a into the self-attention Transformer to obtain the in-modal feature of the speech S a , which can be expressed by the following formula:

[0127] = Transformer ( );

[0128] Specifically, calculate the corresponding query vector (Q), key vector (K), value vector (V) and weight matrix, normalize the weight matrix, then multiply the weight matrix by V to obtain the self-attention output, and then sum the self-attention output to obtain the in-modal feature of the speech containing important in-modal information S a , which is specifically implemented by the following formula:

[0129] ;

[0130] ;

[0131] ;

[0132] ;

[0133] Among them, W Qa and W Ka and W Va are parameter matrices;

[0134] Secondly, considering the different impacts of speech on text and text on speech, this paper uses two cross-attention mechanisms to find the dependency relationship between text and speech. Specifically, through the cross-attention mechanism, according to the features within the text modality S t and the features within the speech modality S a , the features associated with the text in the speech are calculated S ta , which is represented by the following formula:

[0135] ;

[0136] Specifically, Cross-Attention ( S t , S a )'s calculation process includes: calculating f t according to the text features , calculating f a according to the speech features , , then normalizing the attention matrix and multiplying it with to obtain the cross-attention output, and then summing to get the features associated with the text in the speech , which is specifically calculated by the following formula:

[0137] ;

[0138] ;

[0139] ;

[0140] ;

[0141] Through the cross-attention mechanism, according to the features within the text modality S t and the features within the speech modality S a , the features associated with the speech in the text are calculated S at , which is represented by the following formula:

[0142] ;

[0143] Among them, Cross-Attention ( S a , S t )'s calculation process includes: according to the speech features f a Calculated , according to the text features f t Calculated 、 to obtain the attention matrix, which is normalized, and then multiplied by to obtain the cross-attention output, and the sum operation is performed on the cross-attention output to obtain the features related to speech in the text , specifically calculated through the following formula:

[0144] ;

[0145] ;

[0146] ;

[0147] .

[0148] Existing research shows that gender is closely associated with depression, and women have a higher risk of suffering from depression than men. At the same time, to deeply understand the feelings of users of different genders and perceive the emotional changes of users, the present invention constructs a multi-factor depression detection model that integrates gender information and common sense knowledge.

[0149] The multi-factor integrated depression detection model designs a classification model for men (M model) and a classification model for women (F model) respectively according to the different characteristics of male and female users. The main differences between the male and female models are reflected in two aspects: the first is that when obtaining the matrix U in the multi-modal fusion module, the M model performs a splicing operation on the self-attention output and the cross-attention output after summation, while the F model uses a summation operation. Therefore, after obtaining the features within the text modality S t , the features within the speech modality S a , the features related to text in speech S ta and the features related to speech S at , the M model performs a splicing operation and the F model performs a summation operation.

[0150] Specifically, perform a splicing operation on the features within the text modality S t , the features within the speech modality S a , the features related to text in speech S ta and the features related to speech S at to obtain the matrix U M = S t , S a , S ta , S at , the matrix U M = S t , S a , S ta , S at is input into two layers Transformer Encoder to obtain the first fused multi-modal feature , that is, Figure 1 the deep fusion in

[0151] =Transformer ( U M );

[0152] The features within the text modality S t , the features within the speech modality S a , the features in speech related to the text S ta and the features related to speech S at are summed to obtain the matrix U F = S t , S a , S ta , S at U F = S t , S a , S ta , S at is input into two layers Transformer Encoder to obtain the second fused multi-modal feature , that is, Figure 1 the deep fusion in

[0153] =Transformer ​( U F ).

[0154] The second main difference between the male and female models is in the classifier part. The M model uses the BiGRU classifier, while the F model uses the MLP classifier. The specific content is implemented through the multi-factor fusion module.

[0155] The multi-factor fusion module is used to determine the gender of the user to be detected. When the user to be detected is male, the first fused multimodal feature is classified by a male-oriented classification model. Processing is performed to obtain the depression prediction result of the user to be detected. When the user to be detected is a female, the second fusion multimodal feature is classified by the female-oriented classification model. Processing is performed to obtain the depression prediction result of the user to be detected.

[0156] In the multi-factor fusion module, an external common sense knowledge base COMET is introduced for common sense reasoning, giving the machine reasoning ability and understanding the causal relationship in the discourse. Among them, the common sense knowledge base COMET can be reasoned from the perspectives of user reaction (Reaction), intention (Intent), etc., and provide additional information. The present invention selects the reasoning relationship related to the user's emotions and feelings xReact , extract relevant knowledge. COMET is a Seq2Seq structure, which is divided into two parts: decoder and encoder. The encoder is based on BERT to generate corresponding feature vectors, and the decoder is based on GPT to generate discrete common sense knowledge related to clauses and reasoning relations.

[0157] Among them, the first fusion multimodal feature is Processing is performed to obtain the depression prediction results of the user to be tested, including:

[0158] Acquire inferences about emotions and feelings xReact , based on the common sense knowledge base COMET, the preprocessed clauses and reasoning relations xReact Processing is performed to obtain common sense knowledge. Specifically, in the COMET encoder, the clauses and reasoning relationships during preprocessing are first xReact After concatenation, masking, deletion, filling and other operations are performed to generate feature vectors. After encoding, the feature vectors are sent to the decoder for operations such as pooling and linear transformation to generate clauses and reasoning relationships. xReact The relevant common sense knowledge is implemented through the following formula:

[0159] ;

[0160] in, s ti The first clause in the preprocessed i word, s ki is the i common sense knowledge corresponding to the word, and ⊕ represents concatenation;

[0161] Encode the common sense knowledge s ki , input it into BERT for encoding, and obtain the embedding representation s ki of the common sense knowledge b i , which is specifically implemented through the following formula:

[0162] ;

[0163] The embedding representations of all common sense knowledge form a sentence vector B = { b 1, b 2, …, b m}, input the sentence vector B into Transformer to obtain the target common sense knowledge K u , which is implemented through the following formula:

[0164] ;

[0165] Concatenate the target common sense knowledge K u and the first fused multi-modal feature to obtain the first feature , input the first feature into BiGRU the classifier to obtain the feature containing forward and backward information, which is specifically represented by the following formula:

[0166] ;

[0167] Input the feature containing forward and backward information into the fully connected layer to obtain the depression prediction result y’ of the user to be detected, which is specifically represented by the following formula:

[0168] .

[0169] Among them, the second fused multi-modal feature is processed by a female-oriented classification model to obtain the depression prediction result of the user to be detected, including:

[0170] Obtain the inference relationship related to emotions and feelings xReact , according to the common sense knowledge base COMET, for the preprocessed clauses and inference relationships xReact Process it to obtain common sense knowledge. Specifically, in the COMET encoder, first concatenate the clauses and inference relationships during preprocessing xReact Then perform operations such as masking, deletion, and padding to generate feature vectors. After encoding, the feature vectors are fed into the decoder for operations such as pooling and linear transformation to generate common sense knowledge xReact related to the clauses and inference relationships. Specifically, it is achieved through the following formula:

[0171] ;

[0172] Among them, s ti is the i th word in the preprocessed clause, s ki is the i th word corresponding common sense knowledge;

[0173] Input the common sense knowledge s ki into BERT for encoding to obtain the embedding representation s ki of the common sense knowledge b i . Specifically, it is achieved through the following formula:

[0174] ;

[0175] The embedding representations of all common sense knowledge form the sentence vector B = { b 1, b 2,…, b m}. Input the sentence vector B into Transformer to obtain the target common sense knowledge K u . It is achieved through the following formula:

[0176] ;

[0177] Concatenate the target common sense knowledge K u and the second fused multi-modal feature to obtain the second feature . Input the second feature into the MLP classifier to obtain the depression prediction result y’ of the user to be detected. Specifically, it is represented by the following formula:

[0178] .

[0179] Among them, The specific calculation is implemented through the following formula:

[0180] ;

[0181] ;

[0182] ;

[0183] Among them, 、 and are parameter matrices, 、 and are bias parameters, is the output of the first hidden layer, is the output of the second hidden layer.

[0184] Among them, a multi-modal multi-factor depression recognition system that fuses emotional information further includes a loss module;

[0185] The loss module uses the cross-entropy loss function as the loss function, which is specifically represented by the following formula:

[0186] ;

[0187] Among them, is the true depression result of the j th sample, is the predicted depression result of the j th sample;

[0188] The value of the loss function is used to update the parameters in the preprocessing module, text feature extraction module, speech feature extraction module, multi-modal fusion module, and multi-factor fusion module.

[0189] So far, the present invention effectively improves the accuracy and adaptability of depression detection by integrating modal features and common sense knowledge features. The specific steps are as follows:

[0190] First, the algorithm takes the user's text data, speech data, and depression labels as inputs, and the goal is to output a set of depression recognition prediction results for all participants. In the data reading stage, the text data, speech data, and their corresponding labels of each user are read in sequence. Then, the number of iterations and batch size required for model training are set to ensure the parameter settings of the training process.

[0191] Subsequently, through the text feature extraction module and the speech feature extraction module, text features and speech features are respectively extracted from the user's text and speech data, and the two are fused to obtain multi-modal features.

[0192] After that, the common sense knowledge base is used to obtain the external knowledge of the text, enhance the relevant background in the multi-turn conversation, and further understand the emotions in the conversation. At the same time, the influence brought by gender factors is also considered, and the model is divided into two models for men and women.

[0193] Finally, depression prediction is carried out through a specific recognition formula to obtain the recognition result of the user.

[0194] During the training process, the model adopts the set number of iterations and batch size, processes the data in batches, calculates the model error using the loss function, and updates the model parameters through the gradient descent optimization algorithm (Adam). Finally, the algorithm outputs the set of depression recognition prediction results for all users, completing the multi-modal depression recognition task.

[0195] When detecting depression, the present invention solves the deficiencies of existing methods in feature extraction. First, the present invention constructs clauses to obtain the environment of multi-turn conversations. At the same time, BERT is used for text feature extraction. This model can well handle long-distance dependencies through the Transformer architecture and can capture context changes in multi-turn conversations. In terms of speech feature extraction, the present invention uses TCN to extract deep features from speech signals. Therefore, the present invention effectively captures long-distance dependency relationships and situational information, especially the capture of emotional fluctuations and transitions in multi-turn conversations.

[0196] In terms of emotion feature extraction, the present invention uses RoBERTa and Wav2Vec2.0 to extract emotion features in text and speech. Compared with the traditional method that relies on emotion dictionaries, it can more flexibly identify the emotional changes of users, especially for delicate emotional expressions and atypical emotions (such as the emotions of patients with hidden depression) and has better recognition ability.

[0197] In addition, in terms of multi-modal feature fusion, the present invention adopts the self-attention mechanism and the cross-attention mechanism. These methods can deeply explore the correlation between text and speech, and effectively solve the problem that information cannot be fully integrated when feature fusion is carried out by simple splicing or weighted average in existing methods. Through these mechanisms, the emotion features in text and speech can complement each other, thereby improving the accuracy and robustness of emotion recognition.

[0198] Finally, in the multi-factor fusion module, the present invention considers gender differences and introduces common sense reasoning (such as the COMET knowledge base), and further enhances the system's emotional understanding ability through reasoning about the emotional changes of users. The introduction of gender differences and external knowledge bases enables the present invention to provide personalized emotion analysis according to the different characteristics of users. Especially in depression recognition, it can better adapt to the emotional expression methods of different users, thereby improving the universality and accuracy of the model.

[0199] Based on the above scheme, the present invention conducted the following experiments:

[0200] The present invention used two types of datasets in total. The first type is the emotion dataset for pre-training the model, and the second type is the depression dataset for verifying the effectiveness of the model. The emotion dataset is divided into text and speech emotion datasets. The text emotion dataset used in the present invention is SST-2, which comes from Stanford University and contains movie reviews, conforming to the habits of daily language. SST-2 provides two types of emotion labels, positive (1) and negative (0), for each sentence. The speech emotion dataset is the IEMOCAP dataset provided by the SAIL Laboratory of the University of Southern California, which contains 151 conversations, 7,433 sentences, with a duration of 12 hours. The emotion labels are divided into six categories: neutral, happy, sad, angry, excited, and depressed. Due to data imbalance, the present invention conducted a four-classification task of angry, happy, neutral, and sad, and used 90% of the data for training and 10% for testing.

[0201] To verify the effectiveness of the automatic depression recognition model, this paper uses the publicly available dataset DAIC-WOZ. This dataset is constructed based on the conversations between a robot (Ellie) and participants, and the robot is remotely controlled. DAIC-WOZ is provided by the AVEC workshop, and the data includes video features, audio, transcribed text, and depression labels, gender, PHQ-8 scores and other information. There are 189 participants in total, divided into a training set of 107 people, a validation set of 35 people, and a test set of 47 people.

[0202] Meanwhile, to evaluate the performance of the model, the present invention uses the general evaluation indicators in the field of depression recognition - Precision, Recall, and F1 value to verify the effectiveness of the model.

[0203] First, to verify the effectiveness of the multi-modal fusion part, the present invention removed the multi-factor fusion part and only focused on the experimental results of the multi-modal depression recognition model SCAED. Next, Table 1 shows the results of the comparative experiment and ablation experiment.

[0204] Table 1 Results of the comparative experiment and ablation experiment

[0205] model Precision Recall F1 GSM 74.0% 67.6% 67.0% BERT 68.8% 78.6% 72.7% BiLSTM(t) 80.0% 80.9% 80.4% ResNet34 57.1% 66.7% 61.5% CNN AE 71.0% 72.0% 71.0% BiLSTM(a) 77.8% 78.8% 78.3% FF-DDE 82.8% 83.0% 82.4% DF-DDE 80.8% 80.9% 80.8% LSTM with Gating 80.0% 80.9% 81.0% SCAED 84.8% 85.1% 84.9%

[0206] Based on Table 1, the experimental results of the multi-modal depression recognition model SCAED can be obtained. Among them, the experimental results of the multi-modal depression recognition model SCAED are: the Precision of the multi-modal depression recognition model SCAED is 84.8%, the Recall is 80.9%, and the F1 value is 84.9%.

[0207] Afterwards, to verify the effectiveness of each modal feature and emotional feature, the present invention conducts an emotion ablation experiment and a modal ablation experiment, and the experimental results are shown in Table 2.

[0208] Among them, SCAED(t) represents that the model uses only text features and removes text emotion information for depression recognition; SCAED(t - e) represents that the model uses text features for depression recognition (including text semantic features and text emotion features); SCAED(a) represents that the model uses only voice features and removes voice emotion information for depression recognition; SCAED(a - e) represents that the model uses only voice features for depression recognition (including voice deep features and voice emotion features); SCAED( / e) represents that the model removes only emotion information for depression recognition; SCAED(full) represents that the complete model is used for depression recognition.

[0209] Table 2 Results of Emotion Ablation Experiment and Modal Ablation Experiment

[0210] model Precision Recall F1 SCAED(t) 70.0% 72.5% 71.4% SCAED(t-e) 80.0% 80.9% 80.4% SCAED(a) 73.3% 74.6% 73.9% SCAED(a-e) 77.8% 78.8% 78.3% SCAED( / e) 65.4% 70.2% 67.7% SCAED(full) 84.8% 85.1% 84.9%

[0211] The experimental results show that multi-modal data can complement information, and the effect is better than that of depression detection based on single-modal data; the addition of emotion information is helpful for both single-modal and multi-modal depression detection, and can effectively improve the model performance.

[0212] Next, to verify the effectiveness of the multi-factor fusion part, the present invention will add the multi-factor fusion module to the multi-modal depression recognition model SCAED to obtain the complete multi-modal multi-factor model GCM-SCAED. Table 3 shows the results of the comparative experiment and ablation experiment of GCM-SCAED.

[0213] Table 3 Results of Comparative Experiment and Ablation Experiment of GCM-SCAED

[0214] model Precision Recall F1 baseline model 60.0% 43.0% 50.0% FS-GSM 78.0% 63.5% 70.0% Two BiLSTM 78.0% 80.0% 80.0% LSTMwithGating 80.0% 80.9% 81.0% FF-DDE 82.8% 83.0% 82.4% DF-DDE 80.8% 80.9% 80.8% Trf+CNN 91.0% 83.0% 87.0% GRU / BiLSTM 79.0% 92.0% 85.0% SCAED 84.8% 85.1% 84.9% GCM-SCAED (M) 85.7% 93.8% 89.7% GCM-SCAED (F) 100.0% 84.7% 92.4% GCM-SCAED 92.3% 89.2% 90.3%

[0215] The results of the comparative experiment show that the GCM-SCAED model has achieved the best results in the F1 index on the DAIC-WOZ dataset. And the male and female models respectively achieve the best in precision and recall.

[0216] After that, in order to verify the effectiveness of the multi-factor part, in the ablation experiment part, the present invention will conduct experiments on male and female models respectively. The experimental results of the ablation experiment part of the male model are shown in Table 4. Among them, for the male model GCM-SCAED(M): (1) GCM-SCAED(M-t) means that the male model uses only text semantic features for depression recognition; (2) GCM-SCAED(M-t-e) means that the male model uses text features (including text semantic features and text emotion features) for depression recognition; (3) GCM-SCAED(M-a) means that the male model uses only voice deep features for depression recognition; (4) GCM-SCAED(M-a-e) means that the male model uses voice features (including voice deep features and voice emotion features) for depression recognition; (5) GCM-SCAED(M / C) means that the male model removes common sense knowledge; (6) GCM-SCAED(M) means that the model is a complete male depression recognition model.

[0217] Table 4 Experimental Results of the Ablation Experiment Part of the Male Model

[0218] model Precision Recall F1 GCM-SCAED (M-t) 33.3% 90.1% 62.0% GCM-SCAED (M-t-e) 60.0% 85.7% 70.6% GCM-SCAED (M-a) 40.0% 85.7% 54.5% GCM-SCAED (M-a-e) 45.5% 71.4% 55.6% GCM-SCAED (M / C) 82.5% 87.6% 85.1% GCM-SCAED (M) 85.7% 93.8% 89.7%

[0219] The experimental results of the ablation experiment part of the female model are shown in Table 5. Among them, for the female model GCM-SCAED(F): (1) GCM-SCAED(F-t) means that the female model uses only text semantic features for depression recognition; (2) GCM-SCAED(F-t-e) means that the female model uses text features (including text semantic features and text emotion features) for depression recognition; (3) GCM-SCAED(F-a) means that the female model uses only voice deep features for depression recognition; (4) GCM-SCAED(F-a-e) means that the female model uses voice features (including voice deep features and voice emotion features) for depression recognition; (5) GCM-SCAED(F / C) means that the female model removes common sense knowledge; (6) GCM-SCAED(F) means that the model is a complete female depression recognition model.

[0220] Table 5 Experimental Results of the Ablation Experiment Part of the Female Model

[0221] model Precision Recall F1 GCM-SCAED (F-t) 66.7% 57.1% 61.5% GCM-SCAED (F-t-e) 63.6% 78.9% 71.3% GCM-SCAED (F-a) 46.8% 68.9% 57.4% GCM-SCAED (F-a-e) 54.5% 85.7% 66.7% GCM-SCAED (F / C) 91.5% 81.4% 86.6% GCM-SCAED (F) 100.0% 84.7% 92.4%

[0222] In summary, the GCM-SCAED model achieves better results than the existing optimal models and is optimal in terms of the two metrics of F1 and Precision. At the same time, the relevant results show that men and women have different characteristics. Considering gender factors in depression detection and building different classification models for separate training helps to accurately identify depressed patients. Secondly, introducing a common sense knowledge base for reasoning in depression detection and adding knowledge related to the user's mood, feelings, etc. helps to deeply understand the user's thoughts, perceive changes in the user's mood, and thus can accurately distinguish between depressed patients and non-depressed users, improving the accuracy of the model.

[0223] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.< / sync> < / fam> < / an>

Claims

1. A multimodal multi-factor depression recognition system integrating emotional information, characterized in that, It includes a preprocessing module, a text feature extraction module, a speech feature extraction module, a multimodal fusion module, and a multi-factor fusion module; The preprocessing module is used to preprocess the original speech of the user to be detected, obtain the preprocessed text as a clause, and a speech waveform diagram, where the abscissa of the speech waveform diagram is time and the ordinate of the speech waveform diagram is amplitude; The text feature extraction module is used to add [CLS] at the starting position of the preprocessed clause and [SEP] at the ending position, segment the clause with the added tokens through a tokenizer to obtain the segmented clause; extract features from the segmented clause through BERT to obtain the semantic feature f ts ; process the segmented clause through a text emotion extraction model to obtain the text emotion feature f te , and combine the semantic feature f ts and the text emotion feature f te into the text feature f t ; Among them, the text emotion extraction model is obtained through the following method: Obtain the clauses in the SST-2 dataset as the first input samples, and use the emotion labels corresponding to the clauses in the SST-2 dataset as the first output samples. The first input samples and the first output samples form the first training samples. Based on multiple first training samples, train the RoBERTa model until the RoBERTa model converges, obtain the trained RoBERTa model, retain the parameters and encoder structure of the trained RoBERTa model, and replace the output layer of the trained RoBERTa model with a fully connected layer FC to obtain the text emotion extraction model; The speech feature extraction module is used to extract features from the speech waveform diagram to obtain deep speech features f ad , input the original speech into the speech emotion extraction model to obtain speech emotion features f ae , and splice the deep speech features f ad and the speech emotion features f ae to obtain complete speech features f a ; Among them, the speech emotion extraction model is obtained by training the Wav2Vec2.0 model based on multiple second training samples. The second training samples include second input samples and second output samples. The second input samples are the speech samples in the IEMOCAP dataset, and the second output samples are the speech emotion features corresponding to the speech samples in the IEMOCAP dataset; The multi-modal fusion module is used to generate in-text modality features S t , speech modality features S a , text-related features S in speech t and speech-related features S in text a according to the text feature f ta . According to the in-text modality features S at , speech modality features S t , text-related features S in speech a and speech-related features S in text ta , the first fused multi-modal feature at and the second fused multi-modal feature are calculated The multi-factor fusion module is used to determine the gender of the user to be detected. When the user to be detected is male, the first fused multi-modal feature is processed through a classification model for males to obtain the depression prediction result of the user to be detected. When the user to be detected is female, the second fused multi-modal feature is processed through a classification model for females to obtain the depression prediction result of the user to be detected. Among them, in the multi-factor fusion module, the first fused multi-modal feature is processed by a classification model for men to obtain the depression prediction result of the user to be detected, including: to obtain the depression prediction result of the user to be detected, including: Obtain the inference relationship xReact related to emotions and feelings, and process the preprocessed clause and the inference relationship xReact according to the common sense knowledge base COMET to obtain common sense knowledge, which is specifically implemented through the following formula: where s ti is the i-th word in the preprocessed clause, and s ki is the common sense knowledge corresponding to the i-th word, and ⊕ represents concatenation; Input the common sense knowledge s ki into BERT for encoding to obtain the embedded representation b ki of the common sense knowledge s i through the following formula specifically: b i = BERT(s ki ); The embedded representations of all common sense knowledge form the sentence vector B = {b1, b2, …, b m}, and the sentence vector B is input into the Transformer to obtain the target common sense knowledge K u , which is achieved through the following formula: K u = Transformer(B); The target common sense knowledge K u and the first fused multimodal feature are concatenated to obtain the first feature The first feature is input into the BiGRU classifier to obtain the feature containing forward and backward information Specifically, it is represented by the following formula: Features containing forward and backward information Input into the fully connected layer to obtain the depression prediction result y' of the user to be detected, which is specifically expressed by the following formula: In the multi-factor fusion module, the second fused multi-modal feature is processed through a classification model for women to obtain the depression prediction result of the user to be detected, including: to obtain the depression prediction result of the user to be detected, including: Obtain the inference relationship xReact related to emotions and feelings, and process the preprocessed clause and the inference relationship xReact according to the common sense knowledge base COMET to obtain common sense knowledge, which is specifically implemented through the following formula: where s ti is the i-th word in the preprocessed clause, and s ki is the common sense knowledge corresponding to the i-th word; Input the common sense knowledge s ki into BERT for encoding to obtain the embedded representation b ki of the common sense knowledge s i , which is specifically implemented through the following formula: b i = BERT(s ki ); The embedded representations of all common sense knowledge form a sentence vector B = {b1, b2, …, b m}, and the sentence vector B is input into the Transformer to obtain the target common sense knowledge K u , which is achieved through the following formula: K u = Transformer(B); The target common sense knowledge K u and the second fusion multimodal feature are concatenated to obtain the second feature The second feature is input into the MLP classifier to obtain the depression prediction result y’ of the user to be detected, which is specifically represented by the following formula:

2. The multimodal multi-factor depression recognition system integrating emotional information according to claim 1, wherein The preprocessing module is specifically used to obtain the original speech of the user to be detected, determine the triple <start time, end time, clause> of the user to be detected in the original speech, use a regular expression to clean the data of the clause to obtain the preprocessed clause; perform pre-emphasis processing on the original speech to obtain the pre-emphasized speech. Specifically, enhance the high-frequency part of the original speech and improve the high-frequency resolution of the original speech. Perform frame segmentation on the pre-emphasized speech, apply the Hanning window function to the segmented speech to enhance the speech signal to obtain the enhanced speech signal, and draw a speech waveform diagram according to the enhanced speech signal.

3. The multimodal multi-factor depression recognition system integrating emotional information according to claim 1, characterized in that, In the text feature extraction module, the BERT is used to extract features from the segmented clauses to obtain semantic features f ts , including: Obtain word embeddings, position embeddings, and token embeddings in the segmented clauses, add the word embeddings, position embeddings, and token embeddings to get the added embeddings, input the added embeddings into the Transformer encoder for encoding, and obtain the semantic feature f ts .

4. A multimodal and multi-factor depression recognition system integrating emotional information according to claim 1, characterized in that In the speech feature extraction module, feature extraction is performed on the speech waveform diagram to obtain the deep speech feature f ad , including: Perform grayscale processing on the speech waveform diagram to obtain the waveform diagram after grayscale processing. Then, perform normalization and standardization on the waveform diagram after grayscale processing to obtain the processed waveform diagram. The pixel values in the processed waveform diagram are within the range of [0, 1], and the pixel values have a unified mean and standard deviation. Convert the processed waveform diagram into a one-dimensional matrix and input it into the TCN to obtain the deep speech feature f ad .

5. A multimodal and multi-factor depression recognition system integrating emotional information according to claim 1, characterized in that, The multimodal fusion module is specifically used to input the text feature f t into the self-attention Transformer to obtain the in-text-modal feature S t , which is specifically implemented by the following formula: Among them, W Qt , W Kt , W Vt are parameter matrices, T represents the transpose of the matrix, d k is the vector dimension, and sum represents summation; Input the speech feature f a into the self-attention Transformer to obtain the feature S within the speech modality a , which is specifically implemented through the following formula: Among them, W Qa , W Ka , W Va are parameter matrices; Based on the feature S within the text modality through the cross-attention mechanism t and the feature S within the speech modality a , calculate the feature S associated with the text in the speech, ta which is represented by the following formula: S ta = Cross - Attention(S t , S a ); Based on the features S within the text modality through the cross-attention mechanism t and the features S within the speech modality a , calculate the features S in the text associated with the speech at , which is expressed by the following formula: S at = Cross - Attention(S a , S t ); For the feature S within the text modality t and the feature S within the speech modality a and the feature S associated with the text in the speech ta and the feature S associated with the speech at perform a concatenation operation to obtain the matrix U M = [S t , S a , S ta , S at . Input the matrix U M = [S t , S a , S ta , S at into two layers of Transformer Encoder to obtain the first fused multi-modal feature Specifically, it is represented by the following formula: For the feature S within the text modality t and the feature S within the speech modality a and the feature S associated with the text in the speech ta and the feature S associated with the speech at perform a summation operation to obtain the matrix U F = [S t , S a , S ta , S at . Input the matrix U F = [S t , S a , S ta , S at into two layers of Transformer Encoder to obtain the second fused multi-modal feature Specifically, it is represented by the following formula:

6. The multimodal multi-factor depression recognition system integrating emotional information according to claim 5, characterized in that, Cross-Attention(S t ,S a ) is specifically calculated through the following formula: Similarly, Cross-Attention(S a ,S t ) is calculated specifically by the following formula:

7. A multimodal and multi-factor depression recognition system integrating emotional information according to claim 1, characterized in that The specific calculation is achieved through the following formula: Among them, and are parameter matrices, and are bias parameters, is the output of the first hidden layer, is the output of the second hidden layer.

8. A multimodal and multi-factor depression recognition system integrating emotional information according to claim 1, characterized in that, The multimodal multi-factor depression recognition system that fuses emotion information also includes a loss module; The loss module is used to use the cross-entropy loss function as the loss function, which is specifically represented by the following formula: where y j is the true depression result of the j-th sample, and y j ′ is the predicted depression result of the j-th sample; The value of the loss function is used to update the parameters in the preprocessing module, the text feature extraction module, the speech feature extraction module, the multimodal fusion module, and the multi-factor fusion module.

Citation Information

Patent Citations

  • Multi-person dialogue emotion recognition method

    CN115329779A

  • Emotion recognition model training method, emotion recognition method and device

    CN117668224A

  • Conversation emotion recognition method based on prompt learning

    CN117972019A

  • Multi-modal emotion recognition method based on deep learning

    CN119150165A