Depression state prediction system based on complementary graph learning
Patent Information
- Application Number
- CN202610984767.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]本发明的技术任务是提供一种基于互补图学习的抑郁状态预测系统,来解决如何克服现有多模态抑郁症检测中跨模态特征融合不充分、结构信息利用不足以及检测性能受限的缺陷,提升抑郁症状态检测的准确性的问题
(一)本发明解决了现有抑郁症检测方法中,不同模态(如视觉和声学)之间由于数据分布和特征表达不一致导致的“语义鸿沟”问题,提出了一种基于互补图学习的多模态网络;不同于现有技术中简单、静态的多模态融合策略,本发明通过动态构建一个跨模态的互补图,利用一个模态中可靠的关联信息来引导发现另一个模态中的潜在语义邻居,从而实现了深度的跨模态信息交互和对齐,显著提升了抑郁症检测的准确率和鲁棒性;
Smart Images

Figure CN122800285A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically a depression state prediction system based on complementary graph learning. Background Technology
[0002] With the accelerating pace of society and increased public awareness of mental health, depression has become a prevalent mental illness that seriously harms human physical and mental health. It not only significantly reduces patients' quality of life but also easily triggers various social and medical problems. Therefore, efficient and accurate screening and diagnostic technologies for depression have become a current research hotspot. Traditional depression diagnosis mainly relies on manual scale assessments and clinical interviews with doctors, which are highly subjective, depend heavily on professional medical personnel, have low screening efficiency, and cannot meet the application needs of large-scale population surveys.
[0003] With the iteration of artificial intelligence and multimedia data analysis technologies, automated depression detection technologies based on visual modal data such as facial expressions and behavioral actions, and acoustic modal data such as speech rate, pitch, and rhythm, are gradually emerging. These two types of data have good complementarity, providing data support for intelligent diagnosis. Currently, a large number of deep learning-based multimodal depression detection models have emerged in the industry. Mainstream methods extract single-modal features through convolutional neural networks and recurrent neural networks, and combine them with splicing, weighted fusion, or attention mechanisms to achieve modality fusion. Some solutions also use graph neural networks to mine data structure information, effectively improving detection accuracy.
[0004] However, existing technologies still have many core shortcomings: First, modality fusion strategies are simple and rigid, making it difficult to fully explore cross-modal complementary semantic information and resulting in low feature utilization. Second, graph learning methods are mostly based on building static graph structures from a single modality, ignoring the intermodal guidance relationship and potential association between samples, thus limiting the graph's expressive power. Third, deep graph convolutional networks are prone to oversmoothing of node features, weakening feature discriminative power. Fourth, datasets generally suffer from class imbalance, causing the model to favor the majority of samples, significantly reducing the accuracy and robustness of depression detection results and hindering the practical application of intelligent detection technology.
[0005] Therefore, how to overcome the shortcomings of insufficient cross-modal feature fusion, inadequate utilization of structural information, and limited detection performance in existing multimodal depression detection, and improve the accuracy of depression state detection, is an urgent technical problem to be solved. Summary of the Invention
[0006] The technical objective of this invention is to provide a depression state prediction system based on complementary graph learning to overcome the shortcomings of existing multimodal depression detection, such as insufficient cross-modal feature fusion, inadequate utilization of structural information, and limited detection performance, thereby improving the accuracy of depression state detection.
[0007] The technical objective of this invention is achieved as follows: a system for predicting depressive states based on complementary graph learning, characterized in that the system comprises: The dataset construction unit is used to extract visual modality raw features and acoustic modality raw features from a public dataset of depression, and to preprocess the visual modality raw features and acoustic modality raw features respectively to obtain visual modality features and acoustic modality features, and to construct a depression dataset using the visual modality features and acoustic modality features; The model building unit is used to construct a multimodal depression state prediction model based on complementary graph learning. Specifically, it constructs an adjacency matrix of cross-modal complementary graphs using visual and acoustic modal features. Simultaneously, it encodes, aggregates, and performs residual connection fusion processing on the visual and acoustic modal features to obtain residual-enhanced visual and acoustic modal feature representations. It then performs joint fusion learning on the residual-enhanced visual and acoustic modal feature representations to obtain a joint graph aggregated feature representation. Finally, it inputs the joint graph aggregated feature representation into a classifier to obtain the depression state prediction result. The model training unit is used to construct a composite loss function using a weighted cross-entropy loss function and a supervised contrastive loss function. It trains a multimodal depression prediction model using a depression dataset, optimizes the parameters of the multimodal depression prediction model using the composite loss function, and obtains the trained multimodal depression prediction model.
[0008] Preferably, the model building unit includes: The adjacency matrix construction module for cross-modal complementary graphs is used to construct visual modal similarity matrices and acoustic modal similarity matrices using visual modal features and acoustic modal features respectively. It then constructs a cross-modal mutual differentiation matrix using these matrices and applies a Sigmoid function to nonlinearly enhance the cross-modal mutual differentiation matrix, obtaining a nonlinear enhancement structure matrix. A weighted fusion method is then used to fuse the cross-modal mutual differentiation matrix with the nonlinear enhancement structure matrix, obtaining a fused similarity matrix. The fused similarity matrix is then normalized to obtain the final probability transition matrix, which is used as the adjacency matrix of the cross-modal complementary graph. The modality adaptive residual learning module is used to extract the initial visual modality feature representation and the initial acoustic modality feature representation from the visual modality feature and the acoustic modality feature respectively through the corresponding modality encoder. The initial visual modality feature representation and the initial acoustic modality feature representation are processed to obtain refined visual modality feature representation and refined acoustic modality feature representation. Neighborhood information propagation and aggregation are performed on the refined visual modality feature representation and refined acoustic modality feature representation through probability transition matrix to obtain visual modality graph aggregated feature representation and acoustic modality graph aggregated feature representation. The initial visual modality feature representation is then fused with the visual modality graph aggregated feature representation and the initial acoustic modality feature representation and the acoustic modality graph aggregated feature representation respectively through residual connection to obtain residual-enhanced visual modality representation and residual-enhanced acoustic modality representation. The joint modality fusion and prediction module is used to concatenate the residual-enhanced visual modality feature representation and the residual-enhanced acoustic modality feature representation to obtain a joint modality representation. The joint modality representation is then aggregated with the probability transition matrix to obtain a joint graph aggregated feature representation. This joint graph aggregated feature representation is then input into a classifier to obtain a prediction result for the depression state.
[0009] More preferably, the adjacency matrix construction module for cross-modal complementary graphs includes: The modal feature normalization submodule performs L2 norm normalization on the feature vectors of each sample in the visual modal features and acoustic modal features respectively, to obtain the normalized visual feature matrix. and the normalized acoustic feature matrix Eliminate differences in feature amplitudes across different modalities and enhance semantic consistency; The similarity matrix construction submodule is used to utilize the normalized visual feature matrix. Calculate the cosine similarity between the feature vectors of each visual sample to obtain the visual modality similarity matrix. And using the normalized acoustic feature matrix Calculate the cosine similarity between the feature vectors of each acoustic sample to obtain the acoustic modal similarity matrix. The formula is as follows: ; ; Where N represents the number of visual or acoustic sample feature vectors in the visual or acoustic modal features of the depression dataset (N represents the number of samples in the depression dataset. Since both visual and acoustic modal features come from the same batch of samples, the number of samples corresponding to both visual and acoustic modal features is N). Represents an N x N real matrix; Represents visual modality; Represents acoustic modes; The cross-modal mutual differentiation matrix construction submodule is used to perform cross-modal weighted fusion using the visual modal similarity matrix and the acoustic modal similarity matrix to obtain the cross-modal mutual differentiation matrix. The formula is as follows: ; Among them, the cross-modal mutual differentiation matrix By fusing pairwise cosine similarity between visual modal features and acoustic modal features, a cross-modal complementarity graph is dynamically constructed. This enables reliable correlations between visual or acoustic modal features to guide them in discovering potential neighborhood relationships, thereby effectively utilizing the complementary information between visual and acoustic modal features. A hyperparameter representing a balance between the contributions of each modality to the joint semantic structure; The nonlinear enhancement structure matrix construction submodule is used to employ... Function with respect to cross-modal mutual differentiation matrix Perform nonlinear enhancement to construct the nonlinearly enhanced graph structure matrix. The formula is as follows: ; in, Represents an N-order identity matrix, used to introduce node self-connections; Represents an exponential function; The zero-centered Sigmoid transformation submodule is used to adjust the output range of the standard Sigmoid function to (-1, 1), thus enhancing the graph structure matrix. Maintaining a zero-center distribution enhances connectivity between relevant samples and suppresses noisy connectivity between weakly correlated samples. The linear and nonlinear structure fusion submodule is used to fuse cross-modal mutual differentiation matrices using a weighted fusion method. With nonlinear enhancement of the graph structure matrix Obtain the fused similarity matrix The formula is as follows: ; Among them, the fused similarity matrix By fusing linear and nonlinear structural information, the graph structure expressive power is enhanced while preserving the original sample relationships. Indicates the fusion weight parameters; The probability transition matrix construction submodule is used to construct the fused similarity matrix. Perform row normalization to obtain the final probability transition matrix. The formula is as follows: ; in, Represents the similarity matrix after fusion The Middle Line 1 Column elements; Represents the probability transition matrix The Middle Line 1 Column elements, using probability transition matrix As an adjacency matrix of a cross-modal complementary graph, it provides structural constraints for subsequent graph convolutional feature propagation and cross-modal semantic learning.
[0010] More preferably, the modal adaptive residual learning module includes, The initial modal feature representation acquisition submodule is used to encode the visual modal features and acoustic modal features respectively through a visual modal encoder and an acoustic modal encoder, and extract the initial visual modal feature representation and the initial acoustic modal feature representation, as shown in the following formula: ; ; in, This represents the weight matrix of the visual modality encoder; This represents the weight matrix of the acoustic modal encoder; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; This represents the initial feature representation of the visual modality; Represents the initial feature representation of the acoustic modes; Represents a non-linear activation function; Indicates batch normalization; The initial features are represented; both the visual modality encoder and the acoustic modality encoder include a linear transform layer, a batch normalization layer, and a ReLU activation function, which are used to extract the initial feature representations of the visual modality and the acoustic modality, respectively. The refined modal feature representation construction submodule is used to perform L2 normalization and Dropout processing on the initial visual modal feature representation and the initial acoustic modal feature representation, respectively, to obtain the refined visual modal feature representation. Harmony and acoustic refinement of modal features ; Graph aggregation feature representation constructs a submodule, which is used to utilize the probability transition matrix. Visual refinement modal feature representation respectively Harmony and acoustic refinement of modal features Neighborhood information is propagated and aggregated to obtain visual modality map aggregated feature representations and acoustic modality map aggregated feature representations, as shown in the following formula: ; ; in, This represents the aggregated feature representation of the visual modality graph; This represents the aggregated feature representation of the acoustic modality graph; This represents the aggregated weight matrix of the visual modality graph; This represents the aggregated weight matrix of the acoustic modality diagram; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; Indicates the characteristics after fusion; The residual-enhanced modal feature representation construction submodule is used to fuse the initial visual modal feature representation and the aggregated visual modal graph feature representation using residual connections to obtain the residual-enhanced visual modal feature representation, and to fuse the initial acoustic modal feature representation and the aggregated acoustic modal graph feature representation using residual connections to obtain the residual-enhanced acoustic modal feature representation. The residual-enhanced visual and acoustic modal feature representations, by preserving the original modal feature information and graph structure aggregation information, alleviate the oversmoothing problem generated during graph convolution learning and enhance the expressive ability of the visual and acoustic modalities for features related to depressive states; the formulas are as follows: ; ; in, This represents the preset fixed residual fusion coefficient; Visual modal feature representation of residual enhancement; This represents the acoustic modal characteristics of residual enhancement.
[0011] More preferably, the joint modal fusion and prediction module includes: The joint modal representation construction submodule is used to concatenate the residual-enhanced visual modal feature representation with the residual-enhanced acoustic modal feature representation to form a joint modal representation, as shown in the following formula: ; in, Represents joint modal representation; Indicates feature concatenation operation; The joint graph aggregation feature representation construction submodule is used to map the joint modal representation to the joint semantic space using a fully connected layer, and then re-aggregate structural information through the probability transition matrix to obtain the aggregated features in the joint semantic space. By reintroducing complementary dynamic graph structural information, the neighborhood relationship of the fused multimodal features is optimized, reducing the information loss during the fusion process and maintaining the structural consistency in the joint semantic space; the formula is as follows: ; in, Represents aggregated features in the joint semantic space; Represents the joint semantic space projection matrix; The depression state prediction submodule is used to input aggregated features in the joint semantic space into a classification head composed of fully connected layers to achieve the mapping of high-dimensional features to the target category space and obtain the depression state prediction result; the depression state prediction result includes depressed state and non-depressed state.
[0012] Preferably, the model training unit includes: The training optimization module employs a mini-batch stochastic gradient descent approach, randomly selecting a mini-batch of samples from the training set of the depression dataset each time as the current training batch. It then calculates a composite loss function based on the current training batch and uses this composite loss function to train and optimize the multimodal depression prediction model. The composite loss function includes a weighted cross-entropy loss function and a supervised contrastive loss function. The model optimization module is used to update the parameters of the multimodal depression prediction model using the Adam optimization algorithm. It calculates the gradient of the composite loss function with respect to the parameters of the multimodal depression prediction model through the backpropagation algorithm, and iteratively optimizes the parameters of the multimodal depression prediction model based on the gradient information of the parameters until the multimodal depression prediction model converges.
[0013] More specifically, the training and optimization module works as follows: ① Calculate the weighted cross-entropy loss function, as shown in the following formula: ; in, This represents the weighted cross-entropy loss function; Indicates the total number of categories; Indicates the number of samples in the current training batch; Indicates sample Category The true label; Indicates sample Category The predicted probability; Indicates category The corresponding weighting coefficients are used to balance the impact of imbalanced sample sizes across different classes on model training. ② Minimize the supervised contrastive loss function, as shown in the following formula: ; in, This represents the supervised comparison loss function; This represents the set of indices for all samples in the current training batch; Indicates the first in the current training batch One sample; Indicates that, excluding samples, the current training batch The index set corresponding to all other samples; This indicates that, excluding samples, the current training batch contains... In addition, with the sample A set of positive sample indices with the same category label; Represents the set of positive sample indices The number of samples included; It represents a scalar temperature parameter that controls the scale of the similarity distribution; ③ Integrate the weighted cross-entropy loss function and the supervised comparison loss function to construct a composite loss function, as shown in the following formula: ; in, Represents the total loss function; and This represents the loss weighting coefficient.
[0014] The depression state prediction system based on complementary graph learning of the present invention has the following advantages: (i) This invention addresses the “semantic gap” problem caused by inconsistent data distribution and feature representation between different modalities (such as visual and acoustic) in existing depression detection methods. It proposes a multimodal network based on complementary graph learning. Unlike the simple and static multimodal fusion strategies in existing technologies, this invention dynamically constructs a cross-modal complementary graph and uses reliable association information in one modality to guide the discovery of potential semantic neighbors in another modality. This achieves deep cross-modal information interaction and alignment, significantly improving the accuracy and robustness of depression detection. (ii) This invention first performs L2 norm normalization on visual and acoustic features to eliminate dimensional differences, and then calculates the pairwise cosine similarity matrix within each modality. By utilizing the strong correlation of one modality to dynamically guide the graph structure learning of another modality, this invention can effectively construct a dynamic graph structure that reflects the intrinsic complementary relationship between modalities without relying on any external prior knowledge or predefined rules, effectively solving the problem of asymmetric expression intensity of depressive cues in different modalities. (III) This invention designs an enhanced complementary graph construction method. To avoid the oversmoothing problem in deep graph networks, this invention innovatively applies a zero-center Sigmoid nonlinear transformation to the sparse cross-differentiation matrix and dynamically fuses the original linear similarity with the nonlinear enhancement structure through a weighted parameter. At the same time, by explicitly adding self-loops, it ensures that each node retains its own feature information during message propagation, thereby effectively maintaining the uniqueness of node representation and providing high-quality structured priors for subsequent modality-specific feature learning. (iv) In this invention, each modality first passes through an independent encoder, which includes a linear layer, a batch normalization layer, and a ReLU activation function. The original features are mapped to a unified semantic space. Then, the constructed complementary graph is used to perform graph convolution operations on each modality to achieve local feature smoothing and cross-modal geometric alignment. This not only enhances the model's robustness to noise but also ensures that single-modal features can maintain consistency with the global multimodal topology. (v) This invention introduces an innovative adaptive residual fusion strategy to solve the feature homogenization problem in deep graph networks. That is, it uses a preset scalar gating coefficient to balance the contribution ratio of the "self-feature preservation" branch and the "neighbor information aggregation" branch, so that the model can achieve a balance between preserving the individual features of nodes and utilizing the neighborhood context information according to the specific situation. This adaptive mechanism provides an efficient technical solution for suppressing over-smoothing and maintaining feature discriminative power. (vi) In this invention, visual modal features and acoustic modal features are input into the corresponding modal encoder to obtain the initial modal representation. After feature refinement of the initial modal representation, neighborhood information propagation and feature aggregation are performed by combining complementary dynamic graphs. Furthermore, adaptive residual connections are used to fuse the initial modal representation and graph aggregation feature representation, which enhances the structural semantic information while retaining the original modal feature information, thereby alleviating the oversmoothing problem in the graph convolution learning process. (vii) The present invention constructs a joint modality fusion and prediction module. After obtaining the single modality features refined by adaptive residual learning, the present invention does not simply splice or average them, but splices the two and then applies the learned complementary graph to perform the fusion operation. This can effectively make up for the information loss that may occur in the early fusion, and further reduce noise through structural consistency, ensuring that the feature expression in the joint semantic space is purer and more discriminative. (viii) This invention proposes a comprehensive objective loss function consisting of two parts to optimize the entire network in an end-to-end manner: The first part is a weighted cross-entropy loss, which effectively solves the class imbalance problem that is common in depression datasets by assigning higher weights to the minority class (depressed samples), forcing the model to pay more attention to the correct classification of depressed samples; The second part is a supervised contrastive learning loss, which optimizes the manifold geometry of the feature space by explicitly narrowing the feature distance of samples of the same class and widening the feature distance of samples of different classes, thereby enhancing the generalization ability of the model. That is, the weighted cross-entropy loss function and the supervised contrastive loss function are introduced at the same time to construct a composite loss function, optimize the training of the model, and improve the accuracy and generalization ability of the model for the depression state detection task. Attached Figure Description
[0015] The invention will be further described below with reference to the accompanying drawings.
[0016] Appendix Figure 1 The diagram shows the structure of a depression state prediction system based on complementary graph learning. Appendix Figure 2 A flowchart illustrating the working process of building the model unit; Appendix Figure 3 A flowchart illustrating the working process of the adjacency matrix construction module for cross-modal complementary graphs; Appendix Figure 4 This is a flowchart illustrating the working process of the modal adaptive residual learning module. Detailed Implementation
[0017] The following detailed description of the depression state prediction system based on complementary graph learning of the present invention is provided with reference to the accompanying drawings and specific embodiments. Example
[0018] As attached Figure 1 As shown, this embodiment provides a depression state prediction system based on complementary graph learning, the system comprising: The dataset construction unit is used to extract visual modality raw features and acoustic modality raw features from a public dataset of depression, and to preprocess the visual modality raw features and acoustic modality raw features respectively to obtain visual modality features and acoustic modality features, and to construct a depression dataset using the visual modality features and acoustic modality features; The model building unit is used to construct a multimodal depression state prediction model based on complementary graph learning. Specifically, it constructs an adjacency matrix of cross-modal complementary graphs using visual and acoustic modal features. Simultaneously, it encodes, aggregates, and performs residual connection fusion processing on the visual and acoustic modal features to obtain residual-enhanced visual and acoustic modal feature representations. It then performs joint fusion learning on the residual-enhanced visual and acoustic modal feature representations to obtain a joint graph aggregated feature representation. Finally, it inputs the joint graph aggregated feature representation into a classifier to obtain the depression state prediction result. The model training unit is used to construct a composite loss function using a weighted cross-entropy loss function and a supervised contrastive loss function. It trains a multimodal depression prediction model using a depression dataset, optimizes the parameters of the multimodal depression prediction model using the composite loss function, and obtains the trained multimodal depression prediction model.
[0019] The model building unit in this embodiment includes: The adjacency matrix construction module for cross-modal complementary graphs is used to construct visual modal similarity matrices and acoustic modal similarity matrices using visual modal features and acoustic modal features respectively. It then constructs a cross-modal mutual differentiation matrix using these matrices and applies a Sigmoid function to nonlinearly enhance the cross-modal mutual differentiation matrix, obtaining a nonlinear enhancement structure matrix. A weighted fusion method is then used to fuse the cross-modal mutual differentiation matrix with the nonlinear enhancement structure matrix, obtaining a fused similarity matrix. The fused similarity matrix is then normalized to obtain the final probability transition matrix, which is used as the adjacency matrix of the cross-modal complementary graph. The modality adaptive residual learning module is used to extract the initial visual modality feature representation and the initial acoustic modality feature representation from the visual modality feature and the acoustic modality feature respectively through the corresponding modality encoder. The initial visual modality feature representation and the initial acoustic modality feature representation are processed to obtain refined visual modality feature representation and refined acoustic modality feature representation. Neighborhood information propagation and aggregation are performed on the refined visual modality feature representation and refined acoustic modality feature representation through probability transition matrix to obtain visual modality graph aggregated feature representation and acoustic modality graph aggregated feature representation. The initial visual modality feature representation is then fused with the visual modality graph aggregated feature representation and the initial acoustic modality feature representation and the acoustic modality graph aggregated feature representation respectively through residual connection to obtain residual-enhanced visual modality representation and residual-enhanced acoustic modality representation. The joint modality fusion and prediction module is used to concatenate the residual-enhanced visual modality feature representation and the residual-enhanced acoustic modality feature representation to obtain a joint modality representation. The joint modality representation is then aggregated with the probability transition matrix to obtain a joint graph aggregated feature representation. This joint graph aggregated feature representation is then input into a classifier to obtain a prediction result for the depression state.
[0020] The adjacency matrix construction module for the cross-modal complementary graph in this embodiment includes: The modal feature normalization submodule performs L2 norm normalization on the feature vectors of each sample in the visual modal features and acoustic modal features respectively, to obtain the normalized visual feature matrix. and the normalized acoustic feature matrix Eliminate differences in feature amplitudes across different modalities and enhance semantic consistency; The similarity matrix construction submodule is used to utilize the normalized visual feature matrix. Calculate the cosine similarity between the feature vectors of each visual sample to obtain the visual modality similarity matrix. And using the normalized acoustic feature matrix Calculate the cosine similarity between the feature vectors of each acoustic sample to obtain the acoustic modal similarity matrix. The formula is as follows: ; ; Where N represents the number of visual or acoustic sample feature vectors in the visual or acoustic modal features of the depression dataset (N represents the number of samples in the depression dataset. Since both visual and acoustic modal features come from the same batch of samples, the number of samples corresponding to both visual and acoustic modal features is N). Represents an N x N real matrix; Represents visual modality; Represents acoustic modes; The cross-modal mutual differentiation matrix construction submodule is used to perform cross-modal weighted fusion using the visual modal similarity matrix and the acoustic modal similarity matrix to obtain the cross-modal mutual differentiation matrix. The formula is as follows: ; Among them, the cross-modal mutual differentiation matrix By fusing pairwise cosine similarity between visual modal features and acoustic modal features, a cross-modal complementarity graph is dynamically constructed. This enables reliable correlations between visual or acoustic modal features to guide them in discovering potential neighborhood relationships, thereby effectively utilizing the complementary information between visual and acoustic modal features. A hyperparameter representing a balance between the contributions of each modality to the joint semantic structure; The nonlinear enhancement structure matrix construction submodule is used to employ... Function with respect to cross-modal mutual differentiation matrix Perform nonlinear enhancement to construct the nonlinearly enhanced graph structure matrix. The formula is as follows: ; in, Represents an N-order identity matrix, used to introduce node self-connections; Represents an exponential function; The zero-centered Sigmoid transformation submodule is used to adjust the output range of the standard Sigmoid function to (-1, 1), thus enhancing the graph structure matrix. Maintaining a zero-center distribution enhances connectivity between relevant samples and suppresses noisy connectivity between weakly correlated samples. The linear and nonlinear structure fusion submodule is used to fuse cross-modal mutual differentiation matrices using a weighted fusion method. With nonlinear enhancement of the graph structure matrix Obtain the fused similarity matrix The formula is as follows: ; Among them, the fused similarity matrix By fusing linear and nonlinear structural information, the graph structure expressive power is enhanced while preserving the original sample relationships. Indicates the fusion weight parameters; The probability transition matrix construction submodule is used to construct the fused similarity matrix. Perform row normalization to obtain the final probability transition matrix. The formula is as follows: ; in, Represents the similarity matrix after fusion The Middle Line 1 Column elements; Represents the probability transition matrix The Middle Line 1 Column elements, using probability transition matrix As an adjacency matrix of a cross-modal complementary graph, it provides structural constraints for subsequent graph convolutional feature propagation and cross-modal semantic learning.
[0021] The modal adaptive residual learning module in this embodiment includes, The initial modal feature representation acquisition submodule is used to encode the visual modal features and acoustic modal features respectively through a visual modal encoder and an acoustic modal encoder, and extract the initial visual modal feature representation and the initial acoustic modal feature representation, as shown in the following formula: ; ; in, This represents the weight matrix of the visual modality encoder; This represents the weight matrix of the acoustic modal encoder; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; This represents the initial feature representation of the visual modality; Represents the initial feature representation of the acoustic modes; Represents a non-linear activation function; Indicates batch normalization; The initial features are represented; both the visual modality encoder and the acoustic modality encoder include a linear transform layer, a batch normalization layer, and a ReLU activation function, which are used to extract the initial feature representations of the visual modality and the acoustic modality, respectively. The refined modal feature representation construction submodule is used to perform L2 normalization and Dropout processing on the initial visual modal feature representation and the initial acoustic modal feature representation, respectively, to obtain the refined visual modal feature representation. Harmony and acoustic refinement of modal features ; Graph aggregation feature representation constructs a submodule, which is used to utilize the probability transition matrix. Visual refinement modal feature representation respectively Harmony and acoustic refinement of modal features Neighborhood information is propagated and aggregated to obtain visual modality map aggregated feature representations and acoustic modality map aggregated feature representations, as shown in the following formula: ; ; in, This represents the aggregated feature representation of the visual modality graph; This represents the aggregated feature representation of the acoustic modality graph; This represents the aggregated weight matrix of the visual modality graph; This represents the aggregated weight matrix of the acoustic modality diagram; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; Indicates the characteristics after fusion; The residual-enhanced modal feature representation construction submodule is used to fuse the initial visual modal feature representation and the aggregated visual modal graph feature representation using residual connections to obtain the residual-enhanced visual modal feature representation, and to fuse the initial acoustic modal feature representation and the aggregated acoustic modal graph feature representation using residual connections to obtain the residual-enhanced acoustic modal feature representation. The residual-enhanced visual and acoustic modal feature representations, by preserving the original modal feature information and graph structure aggregation information, alleviate the oversmoothing problem generated during graph convolution learning and enhance the expressive ability of the visual and acoustic modalities for features related to depressive states; the formulas are as follows: ; ; in, This represents the preset fixed residual fusion coefficient; Visual modal feature representation of residual enhancement; This represents the acoustic modal characteristics of residual enhancement.
[0022] The joint modality fusion and prediction module in this embodiment includes: The joint modal representation construction submodule is used to concatenate the residual-enhanced visual modal feature representation with the residual-enhanced acoustic modal feature representation to form a joint modal representation, as shown in the following formula: ; in, Represents joint modal representation; Indicates feature concatenation operation; The joint graph aggregation feature representation construction submodule is used to map the joint modal representation to the joint semantic space using a fully connected layer, and then re-aggregate structural information through the probability transition matrix to obtain the aggregated features in the joint semantic space. By reintroducing complementary dynamic graph structural information, the neighborhood relationship of the fused multimodal features is optimized, reducing the information loss during the fusion process and maintaining the structural consistency in the joint semantic space; the formula is as follows: ; in, Represents aggregated features in the joint semantic space; Represents the joint semantic space projection matrix; The depression state prediction submodule is used to input aggregated features in the joint semantic space into a classification head composed of fully connected layers to achieve the mapping of high-dimensional features to the target category space and obtain the depression state prediction result; the depression state prediction result includes depressed state and non-depressed state.
[0023] The model training unit in this embodiment includes: The training optimization module employs a mini-batch stochastic gradient descent approach, randomly selecting a mini-batch of samples from the training set of the depression dataset each time as the current training batch. It then calculates a composite loss function based on the current training batch and uses this composite loss function to train and optimize the multimodal depression prediction model. The composite loss function includes a weighted cross-entropy loss function and a supervised contrastive loss function. The model optimization module is used to update the parameters of the multimodal depression prediction model using the Adam optimization algorithm. It calculates the gradient of the composite loss function with respect to the parameters of the multimodal depression prediction model through the backpropagation algorithm, and iteratively optimizes the parameters of the multimodal depression prediction model based on the gradient information of the parameters until the multimodal depression prediction model converges.
[0024] The specific working process of the training optimization module in this embodiment is as follows: ① Calculate the weighted cross-entropy loss function, as shown in the following formula: ; in, This represents the weighted cross-entropy loss function; Indicates the total number of categories; Indicates the number of samples in the current training batch; Indicates sample Category The true label; Indicates sample Category The predicted probability; Indicates category The corresponding weighting coefficients are used to balance the impact of imbalanced sample sizes across different classes on model training. ② Minimize the supervised contrastive loss function, as shown in the following formula: ; in, This represents the supervised comparison loss function; This represents the set of indices for all samples in the current training batch; Indicates the first in the current training batch One sample; Indicates that, excluding samples, the current training batch The index set corresponding to all other samples; This indicates that, excluding samples, the current training batch contains... In addition, with the sample A set of positive sample indices with the same category label; Represents the set of positive sample indices The number of samples included; It represents a scalar temperature parameter that controls the scale of the similarity distribution; ③ Integrate the weighted cross-entropy loss function and the supervised comparison loss function to construct a composite loss function, as shown in the following formula: ; in, Represents the total loss function; and This represents the loss weighting coefficient.
[0025] The working process of this system is as follows: S1. Construct the dataset using dataset construction units: Extract raw visual and acoustic modal features from the MUD3, LMVD, or D-vlog datasets, and perform one or more preprocessing operations, including standardization, normalization, and missing feature imputation, on the raw visual and acoustic modal features to obtain visual and acoustic modal features. Construct a depression dataset using the visual and acoustic modal features. Simultaneously, divide the dataset into training, validation, and test sets for model training, parameter optimization, and model performance evaluation. S2. Construct a multimodal depression prediction model based on complementary graph learning using model building units: Construct an adjacency matrix of cross-modal complementary graphs using visual and acoustic modal features. Simultaneously, encode, aggregate, and perform residual connection fusion processing on the visual and acoustic modal features respectively to obtain residual-enhanced visual and acoustic modal feature representations. Perform joint fusion learning on the residual-enhanced visual and acoustic modal feature representations to obtain a joint graph aggregated feature representation. Input the joint graph aggregated feature representation into a classifier to obtain the depression state prediction result. S3. A composite loss function is constructed using the weighted cross-entropy loss function and the supervised contrastive loss function through the model training unit. The multimodal depression prediction model is trained using a depression dataset. The parameters of the multimodal depression prediction model are optimized through the composite loss function to obtain the trained multimodal depression prediction model.
[0026] As attached Figure 2 As shown, the working process of the model building unit in this embodiment is as follows: S201. The adjacency matrix construction module of the cross-modal complementary graph constructs visual modal similarity matrices and acoustic modal similarity matrices using visual modal features and acoustic modal features respectively. It then constructs a cross-modal mutual differentiation matrix using the visual modal similarity matrix and acoustic modal similarity matrix, and performs nonlinear enhancement on the cross-modal mutual differentiation matrix using the Sigmoid function to obtain a nonlinear enhancement structure matrix. Finally, it uses a weighted fusion method to fuse the cross-modal mutual differentiation matrix and the nonlinear enhancement structure matrix to obtain a fused similarity matrix. The fused similarity matrix is then normalized to obtain the final probability transition matrix, which is used as the adjacency matrix of the cross-modal complementary graph. S202. The modal adaptive residual learning module extracts the initial visual modal feature representation and the initial acoustic modal feature representation from the visual modal features and the acoustic modal features respectively through the corresponding modal encoder. It processes the initial visual modal feature representation and the initial acoustic modal feature representation to obtain refined visual modal feature representation and refined acoustic modal feature representation. It performs neighborhood information propagation and aggregation on the refined visual modal feature representation and refined acoustic modal feature representation respectively through the probability transition matrix to obtain the visual modal graph aggregated feature representation and the acoustic modal graph aggregated feature representation. It performs residual connection fusion on the initial visual modal feature representation, the visual modal graph aggregated feature representation, the initial acoustic modal feature representation, and the acoustic modal graph aggregated feature representation to obtain the residual enhanced visual modal representation and the residual enhanced acoustic modal representation. S203, the joint modality fusion and prediction module concatenates the residual-enhanced visual modality feature representation and the residual-enhanced acoustic modality feature representation to obtain a joint modality representation. It then performs structured information aggregation with the probability transition matrix to obtain a joint graph aggregated feature representation. Finally, it inputs the joint graph aggregated feature representation into a classifier to obtain the prediction result of the depression state.
[0027] As attached Figure 3 As shown, the specific working process of the adjacency matrix construction module for the cross-modal complementary graph in this embodiment is as follows: S20101. Normalize the visual and acoustic modal features: Perform L2 norm normalization on the feature vectors of each sample in the visual and acoustic modal features to obtain the normalized visual feature matrix. and the normalized acoustic feature matrix To eliminate differences in feature amplitudes across different modalities and enhance semantic consistency; S20102. Constructing the visual modality similarity matrix: using the normalized visual feature matrix Calculate the cosine similarity between samples to obtain the visual modal similarity matrix, as shown in the following formula: ; in, This is the number of samples in the current training batch. Represents the visual feature matrix. Let represent an N x N real matrix. Represents visual modality; S20103. Constructing the acoustic modal similarity matrix: using the normalized acoustic feature matrix The acoustic modal similarity matrix is obtained by calculating the cosine similarity between samples, as shown in the following formula: ; in, Represents the acoustic feature matrix. Let represent an N x N real matrix. Represents acoustic modes; S20104. Constructing the cross-modal mutual differentiation matrix: A cross-modal weighted fusion of the visual modal similarity matrix and the acoustic modal similarity matrix is performed to obtain the cross-modal mutual differentiation matrix. The formula is as follows: ; in, It is a cross-differentiation matrix that can fuse pairwise cosine similarities between visual and acoustic modalities, thereby dynamically constructing a cross-modal complementarity graph. This allows reliable correlations in one modality to guide another modality in discovering potential neighborhood relationships, thus effectively utilizing the complementary information between the two. It is a hyperparameter that balances the contribution of each modality to the joint semantic structure. It is the cosine similarity matrix of visual modalities. It is the cosine similarity matrix of the acoustic modes; S20105. Constructing a nonlinear enhancement structure matrix: using... Functions with mutual derivatives matrix Nonlinear enhancement is performed to obtain the nonlinear enhancement structure matrix. The formula is as follows: ; in, Represents an N-order identity matrix, used to introduce node self-connections. This represents the graph structure matrix after nonlinear enhancement. It is a cross-differential matrix. It is an exponential function; S20106. The zero-center Sigmoid transform adjusts the output range of the standard Sigmoid function to (-1,1), so that the enhanced structure matrix maintains a zero-center distribution, thereby enhancing the connection between related samples and suppressing noisy connections between weakly related samples. S20107, Fusion of Linear and Nonlinear Structures: Fusing Cross-Modal Transduction Matrices Using a Weighted Fusion Method With nonlinear enhancement structure matrix The fused similarity matrix is obtained. The formula is as follows: ; in, Indicates the fusion weight parameters; This represents the similarity matrix after fusion. It is a cross-differential matrix. Represents the graph structure matrix after nonlinear enhancement; S20108. By integrating linear and nonlinear structural information, the graph structure expressive power is enhanced while preserving the original sample relationships. S20109. Constructing the probability transition matrix: This involves processing the fused similarity matrix... After performing row normalization, the final probability transition matrix is obtained. The formula is as follows: ; in, Represents the fusion matrix The Middle Line 1 Column elements; Indicates the number of samples included in the current frame composition; Represents the probability transition matrix The Middle Line 1 Column elements; using the probability transition matrix As an adjacency matrix of a cross-modal complementary graph, it provides structural constraints for subsequent graph convolutional feature propagation and cross-modal semantic learning.
[0028] As attached Figure 4 As shown, the modal adaptive residual learning module in this embodiment works as follows: S20201. Obtaining Initial Modal Feature Representations: The visual modal features and acoustic modal features are encoded using a visual modal encoder and an acoustic modal encoder, respectively, to extract the initial visual modal feature representations and the initial acoustic modal feature representations, as shown in the following formulas: ; ; in, This represents the weight matrix of the visual modality encoder; This represents the weight matrix of the acoustic modal encoder; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; This represents the initial feature representation of the visual modality; Represents the initial feature representation of the acoustic modes; Represents a non-linear activation function; Indicates batch normalization; Indicates initial features; S20202. Constructing refined modal feature representations: Perform L2 normalization and Dropout processing on the initial visual modal feature representations and the initial acoustic modal feature representations respectively to obtain the refined visual modal feature representations. Harmony and acoustic refinement of modal features ; S20203. Constructing graph aggregation feature representation: utilizing the probability transition matrix. Visual refinement modal feature representation respectively Harmony and acoustic refinement of modal features Neighborhood information is propagated and aggregated to obtain visual modality map aggregated feature representations and acoustic modality map aggregated feature representations, as shown in the following formula: ; ; in, This represents the aggregated feature representation of the visual modality graph; This represents the aggregated feature representation of the acoustic modality graph; This represents the aggregated weight matrix of the visual modality graph; This represents the aggregated weight matrix of the acoustic modality diagram; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; Indicates the characteristics after fusion; S20204. Constructing the residual-enhanced modal feature representation: The visual modal initial feature representation and the visual modal graph aggregated feature representation are fused using residual concatenation to obtain the residual-enhanced visual modal feature representation. Similarly, the acoustic modal initial feature representation and the acoustic modal graph aggregated feature representation are fused using residual concatenation to obtain the residual-enhanced acoustic modal feature representation. The formula is as follows: ; ; in, This represents the preset fixed residual fusion coefficient. This represents the initial representation of the visual modality. This represents the initial representation of the acoustic modes. Visual modal feature representation of residual enhancement Representation of acoustic modal features for residual enhancement. The representation of visual modal features is represented by the aggregated graph. The representation of aggregated acoustic modal features. It is a non-linear activation function. It is a batch of normalization; S20205. Residual-enhanced modal feature representations alleviate the oversmoothing problem during graph convolution learning by preserving the original modal feature information and graph structure aggregation information, and enhance the expressive ability of visual and acoustic modalities for features related to depressive states.
[0029] The joint modality fusion and prediction module in this embodiment works as follows: S20301. Constructing a joint modality representation: Enhancing the obtained visual modality feature representation. Acoustic modal enhancement feature representation Feature concatenation is performed to form a joint modal representation. The formula is as follows: ; in, This indicates a feature concatenation operation. Visual modal feature representation of residual enhancement Acoustic modal feature representation of residual enhancement; S20302. Constructing a joint graph aggregated feature representation: Using a fully connected layer, the joint modality representation is mapped to the joint semantic space, and the constructed probability transition matrix is reintroduced. Structural information is aggregated to obtain the joint graph aggregated feature representation, as shown in the following formula: ; in, It is a probability transition matrix. Representing aggregate features in the joint semantic space, Represents the joint semantic space projection matrix. It is a non-linear activation function. It is a batch of normalization, It is a joint modal representation; S20303. The re-aggregation process optimizes the neighborhood relationships of the fused multimodal features by reintroducing complementary dynamic graph structural information, thereby reducing information loss during the fusion process and maintaining structural consistency in the joint semantic space. S20304. Constructing the prediction result of depression state: Input the joint graph aggregated feature representation into the classification head composed of fully connected layers to realize the mapping of high-dimensional features to the target category space and output the predicted probability of depression state.
[0030] The specific working process of the model training unit in this embodiment is as follows: S301. Constructing the Loss Function: The goal of the multimodal depression detection model based on complementary graph learning is to construct a composite loss function by integrating the weighted cross-entropy loss function and the supervised contrastive loss function, and then use the composite loss function to train and optimize the model, as follows: S30101. Calculate the weighted cross-entropy loss function, as shown in the following formula: ; in, Indicates the total number of categories; Indicates the number of training samples; Indicates sample Category The true label; Indicates sample Category The predicted probability; Indicates category The corresponding weighting coefficients; S30102. Calculate the supervised comparison loss function, as shown in the following formula: ; in, This represents the index in a mini-batch of samples. Indicates that, except for the batch The set of all external indices, The index set representing positive samples. It is its base number. It is a scalar temperature parameter that controls the scale of the similarity distribution. It is an exponential function; S30103. Integrate the weighted cross-entropy loss function and the supervised comparison loss function to construct a composite loss function, as detailed below:
[0031] in, Represents the composite loss function. and This represents the loss weighting coefficient. This represents the weighted cross-entropy loss function. This represents the supervised comparison loss function.
[0032] S302. Model Optimization: The Adam optimization algorithm is used to update the model parameters. The gradient of the loss function with respect to the model parameters is calculated through the backpropagation algorithm. The model parameters are then iteratively optimized based on the gradient information until the model converges.
[0033] This embodiment has been extensively validated on multiple large-scale, real-world multimodal depression datasets, fully demonstrating its effectiveness and superiority in depression detection tasks. Experimental results show that the multimodal network model based on complementary graph learning proposed in this invention significantly outperforms several state-of-the-art methods on all evaluation metrics, verifying the core value of the complementary graph learning mechanism proposed in this invention in cross-modal depression detection. Meanwhile, this embodiment demonstrates strong robustness and generalization ability on complex datasets with high heterogeneity and environmental noise. In such datasets, the performance of most baseline models degrades significantly, while the model of this invention maintains relatively stable performance, mainly due to its internal dynamic complementary graph strategy, which can effectively filter out noise and align semantic representations between different modalities, thus proving the reliability of this invention in detecting depression in complex and real-world scenarios. Furthermore, this embodiment demonstrates significant advantages in early depression detection tasks: by simulating real-world continuous monitoring scenarios and using specialized early detection evaluation indicators for assessment, experimental results show that the present invention not only has a higher detection accuracy rate but also extremely low decision delay, enabling reliable judgments to be made at a very early stage of user activity, providing technical support for timely intervention and prevention of disease deterioration.
[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A system for predicting depressive states based on complementary graph learning, characterized in that, The system includes: The dataset construction unit is used to extract visual modality raw features and acoustic modality raw features from a public dataset of depression, and to preprocess the visual modality raw features and acoustic modality raw features respectively to obtain visual modality features and acoustic modality features, and to construct a depression dataset using the visual modality features and acoustic modality features; The model building unit is used to construct a multimodal depression state prediction model based on complementary graph learning. Specifically, it constructs an adjacency matrix of cross-modal complementary graphs using visual and acoustic modal features. Simultaneously, it encodes, aggregates, and performs residual connection fusion processing on the visual and acoustic modal features to obtain residual-enhanced visual and acoustic modal feature representations. It then performs joint fusion learning on the residual-enhanced visual and acoustic modal feature representations to obtain a joint graph aggregated feature representation. Finally, it inputs the joint graph aggregated feature representation into a classifier to obtain the depression state prediction result. The model training unit is used to construct a composite loss function using a weighted cross-entropy loss function and a supervised contrastive loss function. It trains a multimodal depression prediction model using a depression dataset, optimizes the parameters of the multimodal depression prediction model using the composite loss function, and obtains the trained multimodal depression prediction model.
2. The depression state prediction system based on complementary graph learning according to claim 1, characterized in that, The model building units include: The adjacency matrix construction module for cross-modal complementary graphs is used to construct visual modal similarity matrices and acoustic modal similarity matrices using visual modal features and acoustic modal features respectively. It then constructs a cross-modal mutual differentiation matrix using these matrices and applies a Sigmoid function to nonlinearly enhance the cross-modal mutual differentiation matrix, obtaining a nonlinear enhancement structure matrix. A weighted fusion method is then used to fuse the cross-modal mutual differentiation matrix with the nonlinear enhancement structure matrix, obtaining a fused similarity matrix. The fused similarity matrix is then normalized to obtain the final probability transition matrix, which is used as the adjacency matrix of the cross-modal complementary graph. The modality adaptive residual learning module is used to extract the initial visual modality feature representation and the initial acoustic modality feature representation from the visual modality feature and the acoustic modality feature respectively through the corresponding modality encoder. The initial visual modality feature representation and the initial acoustic modality feature representation are processed to obtain refined visual modality feature representation and refined acoustic modality feature representation. Neighborhood information propagation and aggregation are performed on the refined visual modality feature representation and refined acoustic modality feature representation through probability transition matrix to obtain visual modality graph aggregated feature representation and acoustic modality graph aggregated feature representation. The initial visual modality feature representation is then fused with the visual modality graph aggregated feature representation and the initial acoustic modality feature representation and the acoustic modality graph aggregated feature representation respectively through residual connection to obtain residual-enhanced visual modality representation and residual-enhanced acoustic modality representation. The joint modality fusion and prediction module is used to concatenate the residual-enhanced visual modality feature representation and the residual-enhanced acoustic modality feature representation to obtain a joint modality representation. The joint modality representation is then aggregated with the probability transition matrix to obtain a joint graph aggregated feature representation. This joint graph aggregated feature representation is then input into a classifier to obtain a prediction result for the depression state.
3. The depression state prediction system based on complementary graph learning according to claim 2, characterized in that, The adjacency matrix construction module for cross-modal complementary graphs includes: The modal feature normalization submodule performs L2 norm normalization on the feature vectors of each sample in the visual modal features and acoustic modal features respectively, to obtain the normalized visual feature matrix. and the normalized acoustic feature matrix Eliminate differences in feature amplitudes across different modalities and enhance semantic consistency; The similarity matrix construction submodule is used to utilize the normalized visual feature matrix. Calculate the cosine similarity between the feature vectors of each visual sample to obtain the visual modality similarity matrix. And using the normalized acoustic feature matrix Calculate the cosine similarity between the feature vectors of each acoustic sample to obtain the acoustic modal similarity matrix. The formula is as follows: ; ; Where N represents the number of corresponding visual or acoustic sample feature vectors in the visual or acoustic modal features of the depression dataset; Represents an N x N real matrix; Represents visual modality; Represents acoustic modes; The cross-modal mutual differentiation matrix construction submodule is used to perform cross-modal weighted fusion using the visual modal similarity matrix and the acoustic modal similarity matrix to obtain the cross-modal mutual differentiation matrix. The formula is as follows: ; Among them, the cross-modal mutual differentiation matrix By fusing pairwise cosine similarity between visual modal features and acoustic modal features, a cross-modal complementarity graph is dynamically constructed. This enables reliable correlations between visual or acoustic modal features to guide them in discovering potential neighborhood relationships, thereby effectively utilizing the complementary information between visual and acoustic modal features. A hyperparameter representing a balance between the contributions of each modality to the joint semantic structure; The nonlinear enhancement structure matrix construction submodule is used to employ... Function with respect to cross-modal mutual differentiation matrix Perform nonlinear enhancement to construct the nonlinearly enhanced graph structure matrix. The formula is as follows: ; in, Represents an N-order identity matrix, used to introduce node self-connections; Represents an exponential function; The zero-centered Sigmoid transformation submodule is used to adjust the output range of the standard Sigmoid function to (-1, 1), thus enhancing the graph structure matrix. Maintaining a zero-center distribution enhances connectivity between relevant samples and suppresses noisy connectivity between weakly correlated samples. The linear and nonlinear structure fusion submodule is used to fuse cross-modal mutual differentiation matrices using a weighted fusion method. With nonlinear enhancement of the graph structure matrix Obtain the fused similarity matrix The formula is as follows: ; Among them, the fused similarity matrix By fusing linear and nonlinear structural information, the graph structure expressive power is enhanced while preserving the original sample relationships. Indicates the fusion weight parameters; The probability transition matrix construction submodule is used to construct the fused similarity matrix. Perform row normalization to obtain the final probability transition matrix. The formula is as follows: ; in, Represents the similarity matrix after fusion The Middle Line 1 Column elements; Represents the probability transition matrix The Middle Line 1 Column elements, using probability transition matrix As an adjacency matrix of a cross-modal complementary graph.
4. The depression state prediction system based on complementary graph learning according to claim 2, characterized in that, The modal adaptive residual learning module includes, The initial modal feature representation acquisition submodule is used to encode the visual modal features and acoustic modal features respectively through a visual modal encoder and an acoustic modal encoder, and extract the initial visual modal feature representation and the initial acoustic modal feature representation, as shown in the following formula: ; ; in, This represents the weight matrix of the visual modality encoder; This represents the weight matrix of the acoustic modal encoder; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; This represents the initial feature representation of the visual modality; Represents the initial feature representation of the acoustic modes; Represents a non-linear activation function; Indicates batch normalization; Indicates initial features; The refined modal feature representation construction submodule is used to perform L2 normalization and Dropout processing on the initial visual modal feature representation and the initial acoustic modal feature representation, respectively, to obtain the refined visual modal feature representation. Harmony and acoustic refinement of modal features ; Graph aggregation feature representation constructs a submodule, which is used to utilize the probability transition matrix. Visual refinement modal feature representation respectively Harmony and acoustic refinement of modal features Neighborhood information is propagated and aggregated to obtain visual modality map aggregated feature representations and acoustic modality map aggregated feature representations, as shown in the following formula: ; ; in, This represents the aggregated feature representation of the visual modality graph; This represents the aggregated feature representation of the acoustic modality graph; This represents the aggregated weight matrix of the visual modality graph; This represents the aggregated weight matrix of the acoustic modality diagram; Bias parameters representing the visual modality; Bias parameters representing acoustic modes; Indicates the characteristics after fusion; The residual-enhanced modal feature representation construction submodule is used to obtain the residual-enhanced visual modal feature representation by fusing the initial visual modal feature representation and the aggregated visual modal graph feature representation using residual connections, and to obtain the residual-enhanced acoustic modal feature representation by fusing the initial acoustic modal feature representation and the aggregated acoustic modal graph feature representation using residual connections, as shown in the following formula: ; ; in, This represents the preset fixed residual fusion coefficient; Visual modal feature representation of residual enhancement; This represents the acoustic modal characteristics of residual enhancement.
5. The depression state prediction system based on complementary graph learning according to claim 2, characterized in that, The joint modal fusion and prediction module includes: The joint modal representation construction submodule is used to concatenate the residual-enhanced visual modal feature representation with the residual-enhanced acoustic modal feature representation to form a joint modal representation, as shown in the following formula: ; in, Represents joint modal representation; Indicates feature concatenation operation; The joint graph aggregation feature representation construction submodule is used to map the joint modality representation to the joint semantic space using a fully connected layer, and then re-aggregate the structural information through the probability transition matrix to obtain the aggregated features in the joint semantic space, as shown in the following formula: ; in, Represents aggregated features in the joint semantic space; Represents the joint semantic space projection matrix; The depression state prediction submodule is used to input aggregated features in the joint semantic space into a classification head composed of fully connected layers to achieve the mapping of high-dimensional features to the target category space and obtain the depression state prediction result; the depression state prediction result includes depressed state and non-depressed state.
6. The depression state prediction system based on complementary graph learning according to claim 1, characterized in that, The model training unit includes: The training optimization module employs a mini-batch stochastic gradient descent approach, randomly selecting a mini-batch of samples from the training set of the depression dataset each time as the current training batch. It then calculates a composite loss function based on the current training batch and uses this composite loss function to train and optimize the multimodal depression prediction model. The composite loss function includes a weighted cross-entropy loss function and a supervised contrastive loss function. The model optimization module is used to update the parameters of the multimodal depression prediction model using the Adam optimization algorithm. It calculates the gradient of the composite loss function with respect to the parameters of the multimodal depression prediction model through the backpropagation algorithm, and iteratively optimizes the parameters of the multimodal depression prediction model based on the gradient information of the parameters until the multimodal depression prediction model converges.
7. The depression state prediction system based on complementary graph learning according to claim 6, characterized in that, The specific working process of the training optimization module is as follows: ① Calculate the weighted cross-entropy loss function, as shown in the following formula: ; in, This represents the weighted cross-entropy loss function; Indicates the total number of categories; Indicates the number of samples in the current training batch; Indicates sample Category The true label; Indicates sample Category The predicted probability; Indicates category The corresponding weighting coefficients are used to balance the impact of imbalanced sample sizes across different classes on model training. ② Minimize the supervised contrastive loss function, as shown in the following formula: ; in, This represents the supervised comparison loss function; This represents the set of indices for all samples in the current training batch; Indicates the first in the current training batch One sample; Indicates that, excluding samples, the current training batch The index set corresponding to all other samples; This indicates that, excluding samples, the current training batch contains... In addition, with the sample A set of positive sample indices with the same category label; Represents the set of positive sample indices The number of samples included; It represents a scalar temperature parameter that controls the scale of the similarity distribution; ③ Integrate the weighted cross-entropy loss function and the supervised comparison loss function to construct a composite loss function, as shown in the following formula: ; in, Represents the total loss function; and This represents the loss weighting coefficient.