Depression identification method based on attention of multi-instance multi-modal graph

Through the multi-instance multi-modal graph attention method, the corpus data set is constructed and semantic encoding is performed. Combined with the graph convolution network and attention network, the problem of insufficient accuracy of depression detection in the existing technology is solved, and more accurate depression recognition is achieved.

CN120354205APending Publication Date: 2025-07-22YUNXIDONG BIG DATA TECHNOLOGY (QINHUANGDAO) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510430545.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the durability and complexity of depression in depression detection, especially in contextual discourse modeling and multimodal feature fusion, resulting in insufficient identification of depression.

Method used

The multi-instance multi-modal graph attention method is used to construct the corpus dataset, use pre-trained models for semantic encoding, combine the bidirectional long and short-term memory network and graph convolution network to build a local context graph, use the relational graph convolution network and graph attention network for semantic feature encoding, and finally use a classifier for depression recognition.

Benefits of technology

Improves the accuracy of depression identification, can better capture the durability and complexity of depression, provide more accurate depression assessments, and support early detection and intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354205A_ABST
    Figure CN120354205A_ABST
Patent Text Reader

Abstract

The invention provides a depression identification method based on attention of a multi-instance multi-modal graph, and relates to the technical field of natural language processing. The method comprises the following steps: firstly, dividing a long-time sample into a plurality of instances according to utterance levels; thirdly, constructing a time feature graph with a multi-dimensional edge, and converting features in the graph by adopting a graph convolutional network and a graph attention network so as to enhance time dependence and enable the time dependence to be consistent with chronic characteristics of depression; the CCC index of the method is superior to that of an existing reference model, and a graph-based fusion method can align the multi-modal features into a unified space, so that the accuracy of depression assessment is effectively improved, powerful support is provided for early discovery and intervention of psychological problems such as depression, and the method is suitable for popularization and application. Wide application prospects are realized in the fields of mental health, medical care and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method for identifying depression based on multi-instance multi-modal graph attention. Background Art

[0002] Automatic Depression Estimation (ADE) is of great significance for improving people's quality of life by analyzing an individual's speech, voice, and behavioral characteristics to identify early depression. Multi-modal depression detection aims to assign scale scores to different observation modalities of each sentence, thereby comprehensively evaluating the depression level of the speaker. This research direction is closely related to people's daily lives. Therefore, more and more scholars have started to engage in the research of multi-modal depression detection. Context information provides crucial clues for better interpreting the current discourse. Therefore, the prediction model for depression needs to be highly sensitive to the context. Deep learning methods have been widely concerned by researchers because they can simulate the context-dependent relationship between discourses and have achieved remarkable results. The current research mainly combines RNN with the attention mechanism to further improve the accuracy of depression prediction. In terms of modal feature fusion, existing technologies enhance the multi-modal learning ability of the model by directly splicing multiple modal features or using the attention mechanism to weight and fuse the contributions of each modality. In terms of context discourse modeling, existing research mainly captures time series information through LSTM or performs bidirectional fully connected time information modeling with the help of the attention mechanism.

[0003] Although these studies have made some progress in the model's ability to capture global information or model context discourse, there is still a lack of in-depth exploration of the core characteristics of depression (i.e., its nature as a jump connection event). Although some studies have attempted to use models such as multi-scale spatio-temporal CNN, LSTM, and Transformer to capture time dependencies and simulate long-range dependencies, most existing methods still focus on identifying short-term and instantaneous depressive symptoms in videos, generally assuming that depression is a temporary manifestation rather than a persistent state. In fact, depression often manifests as long-term mood swings and a gradual deterioration of cognitive function. Therefore, existing models still have certain limitations in capturing the persistence and complexity of depression. To solve this problem, future research needs to explore more accurate model structures to comprehensively and persistently evaluate an individual's depressive symptoms and provide more effective intervention and treatment strategies. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for identifying depression based on multi-instance multi-modal graph attention in view of the deficiencies of the above-mentioned existing technologies, so as to consider the jump connection between contexts and the semantic information carried by the sentence itself.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0006] A depression recognition method based on multi-instance multi-modal graph attention, comprising the following steps:

[0007] Step 1: Construct a corpus dataset;

[0008] First, construct a corpus dataset where N s is the number of video samples in the dataset; It is modal data composed of a series of videos, including three modalities: audio A, video V, and text T, representing the audio feature, video feature, and text feature corresponding to the video respectively, and id is the index of the video; perform instance segmentation on the samples of each speaker. The specific method is to divide each video sample at the sentence level, that is & is A, V, or T, * is a, v, or t, and n is the number of utterances after segmentation; S id The label set of corresponds to the sentences in each sample; where y is the corresponding patient health questionnaire score, and n id is the number of samples;

[0009] Step 2: Perform sentence-level semantic encoding on the sentences in the corpus dataset;

[0010] Use a pre-trained model to perform semantic encoding on each utterance in the corpus and convert it into a corresponding embedding vector; for audio features Adopt the bag-of-words model eGeMAPS to process each 1-second step, 4-second length audio block and extract the audio features. The audio features include the emotional changes, speech rate, and pitch information of the speaker; for video features Input the aligned video frames into a pre-trained ResNet model to extract deep representations, which can capture the facial expressions of people and identify emotional fluctuations; for text features Use a pre-trained BERT model to convert the text into sentence embeddings. The BERT model can extract the emotional, tone, and intention information in the text through in-depth context understanding, thus providing rich semantic features for subsequent analysis;

[0011] The encoded multi-modal embedding vectors are used as the input for model training, laying a foundation for the next step of sentiment analysis and depression detection;

[0012] Step 3: Perform context-aware semantic encoding on the sentences in the dialogue;

[0013] First, a fixed-length window is used to select a continuous segment of statements from the dialogue for context modeling; the length of the window is n win , representing the range of the selected continuous dialogue;

[0014] Next, a bidirectional long short-term memory network is adopted to capture the context information of each modality within the window, perform context-aware encoding on the data of three modalities, namely audio, video, and text, so as to obtain the context-aware semantic encoding features of the continuous dialogue within the window, as shown in the following formula:

[0015]

[0016] Among them, and respectively represent the output and hidden state at the i-th time step; is the input feature at the i-th time step; is the hidden state at the previous time step;

[0017] The three modalities are concatenated, as shown in the following formula:

[0018]

[0019]

[0020] Among them, is the concatenated utterance feature node, i = 0, 1, 2,..., m represents the i-th time step; are the features of the audio, visual, and text modalities respectively;

[0021] Step 4: Construct a local context graph for the continuous dialogue within the window;

[0022] After obtaining the context-aware semantic encoding features, consider N node nodes around the sentence for which the dialogue behavior needs to be predicted currently. There are p past utterances and f future utterances within the window, and a directed graph N node = p + f; where each statement is regarded as a node of the directed graph, and the node edge represents the connection between node v j and v k , k = [j - p,..., j + f], represents the relationship type between nodes, t ∈ [0, 1] is the relationship index, represents the edge connection weight, with a value of 0 ≤ α jk ≤ 1, as shown in the following formula:

[0023]

[0024] The softmax function is used to calculate weights, where represents the transpose of the j-th utterance feature node, is the trainable weight matrix for the relationship type r, and k represents the utterance subscripts from j-p to j+f;

[0025] The connection weight α jk is obtained by using the softmax function to design a learnable adjacency matrix; first, define the connection methods between two contexts, that is, the relationship types between nodes, to initialize the adjacency matrix, namely, the forward and backward directed connections of the speaker himself; in the initialized adjacency matrix, the connection relationship between each two nodes is first filled with semantic similarity, and then the softmax function is used to process the similarity values of all nodes from the j-th node to the j-p to j+f-th nodes into connection weight values with a sum of 1;

[0026] Step 5: Perform multi-instance attention mechanism and semantic feature encoding on the local context graph within the window, and fuse the final hidden semantic features for classification output by each modality;

[0027] Adopt the Relational Graph Convolutional Network (RGCN) and Graph Attention Network (GAT) to aggregate the dependencies at the speaker level;

[0028] For aggregating the dependencies at the speaker level, first use RGCN to consider the relationship between the speaker and the local context information within the window to obtain the semantic features with speaker dependencies As shown in the following formula:

[0029]

[0030] Where, and are respectively the utterance feature nodes, and are both trainable weight matrices, and α jk and α jj are respectively the connection weights of edge r jk and edge rj j, is the neighborhood index of the j-th utterance under the relationship r, and σ is the activation function;

[0031] Based on the semantic features of speaker dependencies Then use GAT to update the weight α jk , and then use the updated weight to update the semantic features again based on the speaker dependencies, as shown in the following two formulas:

[0032]

[0033] where α′ jk is the updated weight, is the updated semantic-related feature with jump context dependence, is the learnable weight matrix, a is the parameterized weight vector, is the context-dependent semantic feature of statement k; represents the concatenation operation, and LeaktyReLU is the neural network activation function;

[0034] For the calculation process of self-attention, in the multi-head attention mechanism, each head independently calculates its own attention output, and then combines the results of all heads to obtain the following output feature representation:

[0035]

[0036] where, is the feature representation of node j in the second layer of the graph neural network, representing the updated feature vector after the attention mechanism; and are the attention weights, representing the strength of the dependence relationship between nodes; represents the self-attention weight of the j-th node, that is, the relationship between the node itself and itself; represents the attention weight of the j-th node and its neighbor node k; is the learnable weight matrix of each attention head in the second layer of the graph neural network, used to transform the input features into a new feature space; and are the feature representations of the j-th node and the k-th node in the second layer of the graph neural network, and the node features are calculated through the graph convolution or self-attention mechanism of the previous layer; represents the neighborhood index of the j-th node under a certain relationship r; the norm symbol || represents normalizing or standardizing the aggregated features; L represents the attention head in the second layer of the graph neural network;

[0037] Step 6: Based on the hidden layer semantic features in Step 5, use a classifier to classify the conversation;

[0038] By obtaining the context semantic features and speaker-related semantic features through Step 5, for each utterance i, first extract its context features and the node feature representation in the graph neural network Then, concatenate these features into a joint feature vector, and use a learnable weight matrix W c to perform a linear transformation on the concatenated features and add a bias term b c as shown in the following formula:

[0039]

[0040] Among them, represents the score prediction of utterance i; the symbol represents the feature concatenation operation, σ is the activation function, and Sigmoid or ReLU is selected;

[0041] Calculate the average value of all utterance scores; assume that there are n utterances in a conversation, then the final score

[0042]

[0043] where n is the number of utterances after segmentation;

[0044] Step 7: Iteratively execute Step 1 to Step 6, and use the loss function to calculate the loss between the predicted action label and the true action label, and update all the weights W in Step 1 to Step 6 until the set termination condition is reached;

[0045] During the training process, the concordance correlation coefficient loss and the mean squared error loss are used as the cost functions; the concordance correlation coefficient loss I CCC is shown as follows:

[0046]

[0047] where ρ is the Pearson correlation coefficient; μ f and μ y represent the means of the predicted value and the true label respectively, and σ f and σ y are the corresponding standard deviations respectively; the value range of the concordance correlation coefficient CCC is from -1 to 1, where 1 represents an ideal positive correlation and -1 represents a perfect negative correlation.

[0048] The beneficial effects of adopting the above technical solution are as follows: The depression recognition method based on multi-instance multi-modal graph attention provided by the present invention proposes a multi-instance multi-modal attention depression recognition model based on a graph network, considering the mutual connection between the speaker's utterances and analyzing the connection between various modalities of the speaker's utterance. In order to better combine the intermittency of depression to capture depressive features, it is proposed to divide the speaker's utterance into individual instances, so as to better capture depressive features. The graph neural network (GNN) is used to encode the nodes, and the context information and the graph embedding features of the nodes are combined. When predicting the PHQ-8 score, multiple levels of information are fully considered. In this way, the model can simultaneously understand the emotional information of each utterance and its context and relationship in the whole conversation, so as to generate a more accurate overall score. The method of the present invention can be applied to many fields, such as daily conversations, language learning, psychological counseling and hot topic discussions, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a flowchart of the depression recognition method based on multi-instance multi-modal graph attention provided by an embodiment of the present invention;

[0050] Figure 2 It is a comparison chart of the number of parameters and indicators with other model methods on the AVEC2019 dataset provided by an embodiment of the present invention;

[0051] Figure 3 It is an effect diagram of the hidden layer after different modality fusions provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0053] As Figure 1 shown, the method of this embodiment is described as follows.

[0054] Step 1: Construct a corpus dataset.

[0055] First, construct a corpus dataset where N s is the number of video samples in the dataset; It is modal data composed of a series of videos, including three modalities: audio A, video V, and text T. respectively represent the audio feature, video feature, and text feature corresponding to the video, and id is the index of the video. For the convenience of further processing and analysis, the samples of each speaker are instantiated and segmented. The specific method is to divide each video sample at the sentence level, such as The id is the index of the video, and n is the number of segmented utterances; S id The tag set of corresponds to each sentence in the video sample, where y is the corresponding patient health questionnaire score, and n id is the number of video samples. These tag information helps with subsequent sentiment analysis and depression detection.

[0056] Step 2: Perform sentence-level semantic encoding on the statements in the corpus dataset.

[0057] Use a pre-trained model to perform semantic encoding on each utterance in the corpus, converting it into a corresponding embedding vector for subsequent model processing.

[0058] For audio features Adopt the Bag-of-audio-words model eGeMAPS to process each 1-second step, 4-second length audio chunk, and extract the audio features. These audio features capture information such as the speaker's emotional changes, speech rate, and pitch, which is particularly important for emotion recognition.

[0059] For video features Input the aligned video frames (mainly face images) into a pre-trained ResNet model to extract deep representations that can capture subtle changes such as people's facial expressions and recognize emotional fluctuations.

[0060] For text features Use a pre-trained BERT model to convert the text into sentence embeddings. The BERT model can extract information such as emotion, tone, and intention in the text through in-depth context understanding, thereby providing rich semantic features for subsequent analysis.

[0061] Through the joint encoding of audio, video, and text three-modal data, we can comprehensively understand the emotion and semantic information of each sentence, thereby providing a more accurate and comprehensive feature representation for multi-modal depression detection. The encoded multi-modal embedding vectors are used as the input for model training, laying the foundation for the next step of sentiment analysis and depression detection.

[0062] Step 3: Perform context-aware semantic encoding on the statements in the conversation.

[0063] First, use a fixed-length window to select a continuous segment of statements from the conversation for context modeling. The length of the window is n win indicating the range of the selected continuous conversation. In this way, the context information in the conversation can be captured, further enhancing the model's understanding of the dependency relationship between statements.

[0064] Next, a Bidirectional Long Short-Term Memory Network (Bi-LSTM) is used to capture the context information of each modality within the window. Bi-LSTM is a deep learning model that can process both forward and backward information in a sequence, which is particularly effective for handling context-dependent tasks such as dialogue understanding and sentiment analysis. In this step, the Bidirectional Long Short-Term Memory Network is used to perform context-aware encoding on the data of three modalities: audio, video, and text, thereby obtaining the context-aware semantic encoding features of the continuous dialogue within the window, as shown in the following formula:

[0065]

[0066] where, and represent the output and hidden state at the i-th time step respectively; is the input feature at the i-th time step; is the hidden state of the previous time step.

[0067] The three modalities are concatenated, as shown in the following formula:

[0068]

[0069] where, is the concatenated utterance feature node, and i = 0, 1, 2,..., m represents the i-th time step; are the features of the audio, visual, and text modalities respectively.

[0070] Through this method, the context information of each sentence can be combined with the context before and after it, thereby providing a more accurate semantic representation for the subsequent depression prediction task.

[0071] Step 4: Construct a local context graph for the continuous dialogue within the window.

[0072] After obtaining the context-aware semantic encoding features, consider the N node nodes around the sentence for which the dialogue behavior needs to be predicted currently. There are p past utterances and f future utterances within the window (in this embodiment, the time step is regarded as the node of the utterance, so the two have the same meaning), and a directed graph N node = p + f; where each utterance is regarded as a node of the directed graph, and the node edge represents the connection between node v j and v k , k = [j - p,..., j + f], represents the relationship type between nodes, t ∈ [0, 1] is the relationship index, represents the edge connection weight, with a value range of 0 ≤ αjk ≤ 1, as shown in the following formula:

[0073]

[0074] The softmax function is used to calculate the weights, where represents the transpose of the j-th utterance feature node, is the trainable weight matrix for the relationship type r.

[0075] The connection weight α jk is obtained by designing a learnable adjacency matrix using the softmax function. First, define the connection methods between two contexts, that is, the relationship types between nodes, to initialize the adjacency matrix, namely, the forward and backward directed connections of the speaker himself; in the initialized adjacency matrix, the connection relationships between every two nodes are filled with semantic similarities first, and then the softmax function is used to process the similarity values of all nodes from the j-th node to the j-p-th to the j+f-th nodes into connection weight values with a sum of 1.

[0076] Step 5: Perform multi-instance attention mechanism and semantic feature encoding on the local context graph within the window, and fuse the final hidden semantic features for classification output by each modality.

[0077] Adopt the relational graph convolutional network RGCN and the graph attention network GAT to aggregate the dependencies at the speaker level.

[0078] For aggregating the dependencies at the speaker level, first use RGCN to consider the relationship between the speaker and the local context information within the window to obtain the semantic features with speaker dependencies as shown in the following formula:

[0079]

[0080] where, and are the utterance feature nodes respectively, and are both trainable weight matrices, α jk and α jj are the connection weights of edge r jk and edge r jj respectively, is the neighborhood index of the j-th node under a certain relationship r, and σ is the activation function.

[0081] Based on the semantic features with speaker dependencies then use GAT to update the weight α jk , and then use the updated weight to update the semantic features again based on the speaker dependencies, as shown in the following two formulas:

[0082]

[0083] where α′ ik is the updated weight, is the updated semantic-related feature with jump context dependence, is the learnable weight matrix, and a is the parameterized weight vector, is the context-dependent semantic feature of sentence k related to sentence j; denotes the concatenation operation, and LeakyReLU is the neural network activation function.

[0084] For the calculation process of self-attention, in the multi-head attention mechanism, each head independently calculates its own attention output, and then combines the results of all heads to obtain the following output feature representation:

[0085]

[0086] where, is the feature representation of node j in the second layer of the graph neural network, representing the updated feature vector after the attention mechanism; and are the attention weights, representing the strength of the dependence relationship between nodes; represents the attention weight of the j-th node itself, that is, the relationship between the node and itself; represents the attention weight of the j-th node and its neighbor node k; is the learnable weight matrix of each attention head in the second layer of the graph neural network, used to transform the input features into a new feature space; and are the feature representations of the j-th node and the k-th node in the second layer of the graph neural network, and the node features are calculated through the graph convolution or self-attention mechanism of the previous layer; represents the neighborhood index of the j-th node under a certain relationship r; the norm symbol || represents normalizing or standardizing the aggregated features; L represents the attention head in the second layer of the graph neural network.

[0087] Step 6: Based on the hidden layer semantic features in Step 5, use a classifier to classify the conversation.

[0088] By obtaining the context semantic features and speaker-related semantic features in Step 5, for each utterance i, first extract its context features and the node feature representation in the graph neural network Then, concatenate these features into a joint feature vector, and use a learnable weight matrix W c to perform a linear transformation on the concatenated features and add a bias term bc , as shown in the following formula:

[0089]

[0090] Among them, represents the score prediction of utterance i; the symbol represents the feature splicing operation, σ is the activation function, and Sigmoid or ReLU is selected.

[0091] To obtain the overall PHQ-8 score, calculate the average value of all utterance scores. Assume that there are n utterances in a conversation, then the final score is:

[0092]

[0093] Among them, n is the number of segmented utterances.

[0094] Step 7: Iteratively execute Step 1 to Step 6, and use the loss function to calculate the loss between the predicted behavior label and the true behavior label, and update the weights W in Step 1 to Step 6 until the set termination condition is reached.

[0095] During the training process, the concordance correlation coefficient (CCC) loss and the mean squared error loss are used as the cost functions. The concordance correlation coefficient loss L CCC is as shown in the following formula:

[0096]

[0097] Among them, ρ is the Pearson correlation coefficient; μ f and μ y represent the means of the predicted value and the true label respectively, and σ f and σ y are the corresponding standard deviations respectively; the value range of the concordance correlation coefficient CCC is from -1 to 1, where 1 represents an ideal positive correlation and -1 represents a perfect negative correlation.

[0098] The depression recognition method based on multi-instance multi-modal graph attention in this embodiment conducts experiments on the AVEC2019 dataset and strictly follows its training, development, and test set divisions. Figure 2Shows the performance of the method proposed in this embodiment on the AVEC2019 dataset and compares it with various existing model methods. As can be seen from the table, this embodiment performs excellently in the two key evaluation metrics of the concordance correlation coefficient (CCC) and the root mean square error (RMSE). Among them, the CCC of this embodiment reaches 0.554, significantly higher than other methods such as TensorFormer (0.493), DepressNet (0.457), and MT (0.466). In addition, in terms of the RMSE metric, the value of this embodiment is 4.61, which performs well among all the models providing RMSE data, second only to TensorFormer (4.31), but its actual number of parameters is several times that of the method proposed in this embodiment, reflecting the advantages of this embodiment in terms of prediction accuracy and the number of parameters.

[0099] Figure 3 Shows the distribution of this embodiment in the hidden layer feature space after different modality fusions. As can be seen from the figure, this embodiment has good effects in the fusion of text (T), audio (A), and video (V) multimodal features, especially under the action of the GCN ( Figure 3 column b) and GAT ( Figure 3 column c) models, the feature distribution is more compact and reasonable, further enhancing the expression ability of multimodal data. Specifically, subgraphs (a1) to (a4) respectively show the original feature distributions of different modality combinations. It can be observed that there is still a certain modality separation phenomenon in the individual modality combinations (such as T+A, T+V, A+V), while the three-modal fusion (T+A+V, subgraph a1) can make better use of all information, making the data distribution more uniform. After applying the GCN ( Figure 3 column b) and GAT ( Figure 3 column c) for feature extraction, it can be found that for GCN ( Figure 3 column b): Compared with the original feature distribution, GCN can effectively learn local and global relationships, making similar features more aggregated and enhancing the discrimination of dissimilar features. Especially in the T+A+V modality (b1), the overall feature space is more compact, which helps the model to effectively model the data. For GAT ( Figure 3 column c): GAT further uses the attention mechanism to strengthen important information. It can be seen that the feature distribution of the T+A+V modality (c1) is more orderly, further optimizing the feature fusion effect. Compared with only fusing two modalities (c2, c3, c4), the three-modal fusion can provide richer context information, making the final representation ability stronger. The above results show that in the tasks of emotion computing and depression recognition, this embodiment can effectively improve the prediction ability of the model and provide a better solution for the research and application in related fields.

[0100] Through in-depth experiments on various aspects such as modal combination, model component contribution, window size, and the influence of the number of attention heads, as well as visualizing and analyzing scatter plots of true and predicted values, attention weights, hidden layer representations, etc., it fully demonstrates the importance of multimodal fusion, the key role of specific components such as multi-instance learning, GNN, GAT, etc. in improving model performance, and the optimization significance of appropriate window size (20 utterances) and the number of attention heads (4) for the model effect. Generally speaking, the multi-instance multimodal graph attention method for depression recognition effectively integrates multimodal information, graph structure, and time series features, significantly improves the performance of depression detection, strongly verifies the effectiveness of the method in this embodiment, provides important technical support for the accurate recognition of the depressive state, and its application has room for improvement and transformation, which has a positive significance for the development of related fields.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.

Claims

1. A depression recognition method based on multi-instance multi-modal graph attention, characterized in that: It includes the following steps: Step 1: Construct a corpus dataset; First, construct a corpus dataset where N s is the number of video samples in the dataset; It is modal data composed of a series of videos, including three modalities: audio A, video V, and text T, representing the audio feature, video feature, and text feature corresponding to the video respectively, and id is the index of the video; perform instance segmentation on the samples of each speaker. The specific method is to divide each video sample at the sentence level, that is & is A, V, or T, * is a, v, or t, and n is the number of utterances after segmentation; S id The label set of is corresponding to the sentences in each sample; where y is the corresponding patient health questionnaire score, and n id is the number of samples; Step 2: Perform sentence-level semantic encoding on the sentences in the corpus dataset; The encoded multi-modal embedding vectors serve as the input for model training, laying the foundation for the next step of sentiment analysis and depression detection; Step 3: Perform context-aware semantic encoding on the sentences in the conversation; Step 4: Construct a local context graph for consecutive conversations within the window; Step 5: Perform multi-instance attention mechanism and semantic feature encoding on the local context graph within the window, and fuse the final hidden semantic features for classification output from each modality; Adopt the Relational Graph Convolutional Network (RGCN) and Graph Attention Network (GAT) to aggregate speaker-level dependencies; Step 6: Based on the hidden semantic features in Step 5, use a classifier to classify the conversation; Step 7: Iteratively execute Steps 1 to 6, and use a loss function to calculate the loss between the predicted action label and the true action label, and update all the weights W in Steps 1 to 6 until the set termination condition is reached; During the training process, the concordance correlation coefficient loss and the mean squared error loss are used as cost functions; the concordance correlation coefficient loss L CCC is shown as follows: Among them, ρ is the Pearson correlation coefficient; μ f and μ y represent the means of the predicted value and the true label respectively, and σ f and σ y are the corresponding standard deviations respectively; the value range of the concordance correlation coefficient CCC is from -1 to 1, where 1 represents an ideal positive correlation and -1 represents a perfect negative correlation.

2. The method for identifying depression based on multi-instance multi-modal graph attention according to claim 1, wherein: The specific method of Step 2 is as follows: Semantically encode each utterance in the corpus using a pre-trained model and convert it into the corresponding embedding vector; for audio features Adopt the bag-of-words model eGeMAPS to process each audio chunk with a 1-second step length and a 4-second length, and extract the audio features, where the audio features include the emotional changes of the speaker, speech rate, and pitch information; for video features Input the aligned video frames into a pre-trained ResNet model to extract deep representations, which can capture people's facial expressions and identify emotional fluctuations; for text features Use a pre-trained BERT model to convert the text into sentence embeddings. The BERT model can extract the emotion, tone, and intention information in the text through in-depth context understanding, thus providing rich semantic features for subsequent analysis.

3. The method for identifying depression based on multi-instance multi-modal graph attention according to claim 2, wherein: The specific method of Step 3 is as follows: First, a fixed-length window is used to select a continuous segment of statements from the conversation for context modeling; the length of the window is n win , representing the range of the selected continuous conversation Next, use a bidirectional long short-term memory network to capture the context information of each modality within the window, perform context-aware encoding on the data of the three modalities of audio, video, and text, so as to obtain the context-aware semantic encoding features of consecutive conversations within the window, as shown in the following formula: Among them, and represent the output and hidden state at the i-th time step, respectively; is the input feature at the i-th time step; is the hidden state at the previous time step; Concatenate the three modalities, as shown in the following formula: Among them, is the concatenated discourse feature node, where i = 0, 1, 2, …, m represents the i-th time step; are the features of the audio, visual, and text modalities respectively.

4. A method for identifying depression based on multi-instance multi-modal graph attention according to claim 3, characterized in that: The specific method of Step 4 is as follows: After obtaining the context-aware semantic encoding features, consider N nodes around the sentence for which the dialogue act needs to be predicted currently. node There are p past utterances and f future utterances within the window, and a directed graph is constructed. N node = p + f; where each statement is regarded as a node of the directed graph, and the node edge represents the connection between node v j and v k . represents the relationship type between nodes, t ∈ [0, 1] is the relationship index, represents the connection weight of the edge , and the value range is 0 ≤ α jk ≤ 1, as shown in the following formula: The softmax function is used to calculate weights, where represents the transpose of the j-th utterance feature node, is the trainable weight matrix for the relationship type r, and k represents the utterance subscripts from the j-p to the j+f-th; The connection weight α jk is obtained by designing a learnable adjacency matrix using the softmax function; first, define the connection method between two contexts, that is, the relationship type between nodes to initialize the adjacency matrix, namely, the forward and backward directed connections of the speaker itself; in the initialized adjacency matrix, the connection relationship between each two nodes is first filled with semantic similarity, and then the softmax function is used to process the similarity values of all nodes from the j-th node to the j-p-th to j+f-th nodes into connection weight values with a sum of 1.

5. The method for identifying depression based on multi-instance multi-modal graph attention according to claim 4, wherein: The specific method of Step 5 is as follows: For aggregating speaker-level dependencies, first use RGCN to consider the relationship between the speaker and the local context information within the window, and obtain semantic features with speaker dependencies As shown in the following formula: Among them, and are discourse feature nodes respectively, and are both learnable weight matrices, α jk and α jj are the connection weights of edge r jk and edge r jj respectively, is the neighborhood index of the j-th discourse under the relationship r, and σ is the activation function; Semantic features based on speaker dependencies Then use GAT to update the weight α jk , and then use the updated weight to update the semantic features again based on speaker dependencies, as shown in the following two equations: where α′ jk is the updated weight, is the updated semantic-related feature with jump context dependence, is the learnable weight matrix, a is the parameterized weight vector, is the context-dependent semantic feature of statement k; represents the concatenation operation, and LeakyReLU is the neural network activation function; The calculation process of self-attention. In the multi-head attention mechanism, each head independently calculates its own attention output, and then combines the results of all heads to obtain the following output feature representation: Among them, is the feature representation of node j in the second layer of the graph neural network, representing the updated feature vector after passing through the attention mechanism; and are the attention weights, representing the strength of the dependence relationship between nodes; represents the attention weight of the j-th node itself, that is, the relationship between the node and itself; represents the attention weight of the j-th node and its neighbor node k; is the learnable weight matrix of each attention head in the second layer of the graph neural network, used to transform the input features into a new feature space; and are the feature representations of the j-th node and the k-th node in the second layer of the graph neural network. The node features are calculated through graph convolution or self-attention mechanism in the previous layer; represents the neighborhood index of the j-th node under a certain relationship r; the norm symbol || represents normalizing or standardizing the aggregated features; L represents the attention head in the second layer of the graph neural network.

6. The method for identifying depression based on multi-instance multi-modal graph attention according to claim 5, characterized in that: The specific method of Step 6 is as follows: Obtained through step 5, considering context semantic features and speaker-related semantic features. For each utterance i, first extract its context features and the node feature representation in the graph neural network Then, concatenate these features into a joint feature vector, and use a learnable weight matrix W c , perform a linear transformation on the concatenated features, and add a bias term b c , as shown in the following formula: Among them, represents the score prediction of utterance i; the symbol represents the feature concatenation operation, σ is the activation function, and Sigmoid or ReLU is selected; Calculate the average value of the scores of all utterances; assume there are n utterances in a conversation, then the final score Where n is the number of segmented utterances.