Emotional disorder detection method and device based on multi-modal heterogeneous graph convolutional neural network
Through the multimodal heterogeneous pattern convolutional neural network combined with Transformer attention mechanism, it integrates EEG, eye movement, micro-expression and gait information, solving the limitations of the traditional single-modal detection method and achieving more accurate mood disorder detection.
Patent Information
- Application Number
- CN202411799645.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-09
AI Technical Summary
Traditional emotional disorder detection methods only use single-modal information, which has problems such as high subjectivity, limitations of physiological signaling and insufficient comprehensiveness, making it difficult to fully capture complex emotional experiences.
The detection method based on multimodal heterogeneous graph convolutional neural network is adopted. By collecting EEG, eye movement, micro-expression and gait information, heterogeneous graphs are constructed for feature extraction and fusion, and the weight of modal features is adaptively allocated using Transformer's attention mechanism.
It realizes effective fusion of multimodal data, reduces redundant information, improves the generalization ability of the model, and can more accurately identify emotional disorder types, providing a powerful tool for clinical diagnosis.
Smart Images

Figure CN119949829A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical artificial intelligence technology, and in particular relates to a method and device for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network. Background Art
[0002] Mood disorders refer to the exaggeration, confusion and reduction of normal emotional responses. They mainly include depression, anxiety, etc. According to the "Chinese Classification and Diagnostic Criteria for Mental Disorders (3rd Edition)", patients with mental disorders show symptoms such as decreased energy or fatigue, lack of concentration or distraction, and reckless behavior.
[0003] Emotional disorders can be detected from two aspects: neurophysiology and external behavior. For example, EEG, pupil changes, muscle electrical signals, facial expressions and other information can quantitatively reflect emotional states. Studies have shown that when patients with depression view negative emotional information, they will show longer gaze time and fewer glances; patients with anxiety often show tension and anxiety in their facial expressions; and patients with emotional disorders will have significant changes in their walking speed, posture and stride.
[0004] Traditional methods for detecting mood disorders usually only utilize information from a single modality and have some limitations: 1. High subjectivity: Participants may not be able to accurately describe their emotional state, or may be influenced by other factors, such as individual culture, personality, situation, etc.
[0005] 2. Limitations of physiological signals: Physiological parameters of a single modality are not sufficient to fully reflect emotional states. Physiological signals (such as skin galvanic response, heart rate, and breathing rate) only indirectly reflect emotions and cannot fully capture complex emotional experiences. (Individual differences).
[0006] 3. The comprehensiveness of a single modality is insufficient: Emotions are multidimensional: Emotions are not just physiological reactions or external expressions, but also include multiple dimensions such as cognition and subjective experience. A single modality cannot fully capture these dimensions.
[0007] For example, patent document CN118490227A discloses a method and system for extracting EEG spatiotemporal patterns for emotional disorder assessment tasks, which includes: EEG data acquisition and preprocessing, EEG time slicing and artifact removal, subsequence segmentation domain time domain dimensionality reduction, spatial relationship matrix construction, establishment of timing optimization module, calculation of state position weights, merging timing states, and evaluation model establishment and interpretable feature extraction. The present invention is aimed at emotional assessment tasks, and can obtain brain spatial patterns that can integrate complex dynamic changes in a small amount of data. Based on this pattern, more accurate and stable evaluation results can be obtained to assist doctors in diagnosing whether they are ill and distinguishing depression and bipolar disorder.
[0008] Patent document CN118070127A discloses a method for feature extraction and classification of bipolar disorder based on a high-order functional network, comprising the following steps: S1, obtaining resting-state functional magnetic resonance data of the subjects, performing preprocessing operations on the data, obtaining the BOLD time series of each subject and constructing a high-order functional network; S2, using the weights of the high-order functional network as a candidate feature set, and performing feature selection to obtain a feature set E1 with the greatest recognition ability for patients with bipolar disorder; S3, calculating the community overlap index of each subject based on the high-order functional network, and using it as the feature set E2; S4, performing feature fusion of E1 and E2 to obtain the final feature set; S5, using the final feature set E3 to train a support vector machine classification model to complete the recognition of bipolar disorder. Summary of the invention
[0009] The purpose of the present invention is to provide an emotional disorder detection method and device based on a multimodal heterogeneous graph convolutional neural network. The emotional disorder detection method realizes the identification of the type of emotional disorder by collecting multimodal data including EEG, eye movement, micro-expression and gait information.
[0010] In order to achieve the first objective of the present invention, the following technical solution is provided: a method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network, comprising the following steps: Obtain user data and corresponding emotional disorder types, including the user's EEG information, eye movement information, micro-expressions, and gait information; Perform data alignment processing on user data to obtain an initial data group with consistent data modality, and use the initial data corresponding to each type of information as nodes, and use the logic of inferring emotional disorders with medical prior knowledge as connecting lines to construct the corresponding emotional disorder inference meta-path; The initial data set, the types of mood disorders, and the corresponding mood disorder inference meta-pathways are combined into a data set; Constructing a corresponding convolutional neural network based on a multi-network framework, wherein the convolutional neural network includes a data alignment processing module, a heterogeneous feature extraction module, a feature fusion module and a prediction module; The data alignment processing module is used to perform data alignment processing on the input user data to generate an initial data group; The heterogeneous feature extraction module includes a multimodal feature extractor, which is used to extract data features of each group of data in the initial data group to output a corresponding multimodal feature group; The feature fusion module uses a low-rank multimodal fusion method to perform feature encoding for each modal feature to obtain multiple sets of feature vectors in the same dimension, and uses a pre-constructed mood disorder inference meta-path to associate each feature vector in the obtained multiple sets of feature vectors to obtain a dense representation of the multimodal features; The visualization module performs visualization based on the obtained dense representation of multimodal features to output the film and television relationship of the distribution of each data in the user data; The convolutional neural network was trained using the dataset to obtain a classification model for mood disorder detection; The user data to be analyzed is input into the classification model to output the film and television relationship of each data distribution.
[0011] The present invention performs data alignment for multimodal data to ensure that they have the same representation form; in order to reduce noise from sensors or other sources, we use low-rank matrix decomposition technology to decompose the feature representation of each modality into a low-dimensional representation. This helps to reduce redundant information and improve the generalization ability of the model, and by constructing a heterogeneous graph, in which nodes represent the features of different modalities and edges represent the associations between them, and finally the Transformer's attention mechanism adaptively assigns weights between different modal features, thereby amplifying the information of effective modalities and reducing the impact of data noise; at the same time, the attention mechanism also helps to deal with the imbalance between different modalities, ensuring that each modality receives appropriate attention. This comprehensive method has broad application prospects in the field of mood disorder detection and provides a powerful tool for clinical diagnosis and intervention.
[0012] Specifically, the data alignment process includes cross-modal alignment and group-level alignment, which facilitates the construction of a stable model and reduces the risk of overfitting.
[0013] Specifically, the cross-modal alignment refers to mapping the information of different modalities into a shared feature space and calculating the attention scores between the modalities to achieve the synchronous conversion of multimodal asynchronous time series signals, which not only improves the experimental results but also facilitates the generalization of the model to new data samples.
[0014] Specifically, the group horizontal alignment refers to calculating the unit vector of the multimodal latent space features and mapping the features on the spherical space, using the feature constraints of the Frobenius two-norm to reduce the risk of overfitting and improve the stability and generalization of the model.
[0015] Specifically, the multimodal feature extractor includes an STGCN encoder for EEG information, a ResNet encoder for expression information, an LSTM encoder for eye movement information, and a ConLSTM encoder for gait information.
[0016] Specifically, the specific process of the low-rank multimodal fusion is as follows: For a single modal feature, LSTM is used to compress the corresponding time series information, and the hidden state context vector in the compressed single modal feature is extracted, and encoding is performed based on the hidden state context vector to obtain the corresponding feature vector.
[0017] Specifically, the data of each modality is processed using an LSTM network. The role of LSTM is to capture long-term dependencies in time series and transform complex time series data into a compact representation, thereby retaining the most important dynamic changes for mood disorder detection. When processing time series, LSTM generates a series of hidden states that represent the contextual information at each time point. Finally, the information of certain key time steps (such as the last frame) can be extracted or the information of all time steps can be aggregated (such as taking the average or weighted average) to form a single context feature representation.
[0018] The context vector of each modality is further processed through a specific neural network layer (such as a fully connected layer, a convolutional layer, or a dimensionality reduction module) to map it to a unified feature space. During the encoding process, the complexity of the network can be constrained, or specific methods (such as low-rank decomposition) can be used to optimize the feature representation and reduce data redundancy. Multimodal data is optimized and fused, using methods such as tensor decomposition or attention mechanisms to ensure that each modality is fully expressed in the final feature representation while avoiding redundancy and noise.
[0019] Specifically, the mood disorder inference meta-pathway includes a depression pathway, a bipolar disorder pathway, a cognitive disorder pathway and an anxiety disorder pathway.
[0020] In order to achieve the second purpose of the present invention, the following technical method is provided: an emotional disorder detection device, which implements the above-mentioned emotional disorder detection method based on multimodal heterogeneous graph convolutional neural network.
[0021] Compared with the prior art, the present invention has the following beneficial effects: By using a multimodal heterogeneous graph neural network, information from different perceptual modalities can be effectively integrated; at the same time, the feature representation of each modality can be decomposed into a low-dimensional representation through low-rank matrix decomposition technology, which helps to reduce redundant information and improve the generalization ability of the model. By constructing a heterogeneous graph, in which nodes represent the features of different modalities and edges represent the associations between them; using the Transformer's attention mechanism, the weights between different modal features are adaptively assigned, so that the information of the effective modality can be amplified while the impact of data noise can be weakened. The attention mechanism also helps to deal with the imbalance between different modalities and ensure that each modality receives appropriate attention. This comprehensive method has broad application prospects in the field of mood disorder detection and provides a powerful tool for clinical diagnosis and intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flow chart of the emotional disorder detection method provided in this embodiment; Figure 2 A flow chart of a data alignment processing module provided in this embodiment; Figure 3 A flow chart of feature decomposition provided for this embodiment; Figure 4 A flowchart of feature encoding provided for this embodiment. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0024] like Figure 1 As shown, the emotional disorder detection method provided by this embodiment has the following steps: obtaining user data and the corresponding emotional disorder type, which includes the user's EEG information, eye movement information, micro-expression and gait information; Perform data alignment processing on user data to obtain an initial data group with consistent data modality, and use the initial data corresponding to each type of information as nodes, and use the logic of inferring emotional disorders with medical prior knowledge as connecting lines to construct the corresponding emotional disorder inference meta-path; The initial data set, the types of mood disorders, and the corresponding mood disorder inference meta-pathways are combined into a data set; Constructing a corresponding convolutional neural network based on a multi-network framework, wherein the convolutional neural network includes a data alignment processing module, a heterogeneous feature extraction module, a feature fusion module and a prediction module; The data alignment processing module is used to perform data alignment processing on the input user data to generate an initial data group; The heterogeneous feature extraction module includes a multimodal feature extractor, which is used to extract data features of each group of data in the initial data group to output a corresponding multimodal feature group; The feature fusion module uses a low-rank multimodal fusion method to perform feature encoding for each modal feature to obtain multiple sets of feature vectors in the same dimension, and uses a pre-constructed mood disorder inference meta-path to associate each feature vector in the obtained multiple sets of feature vectors to obtain a dense representation of the multimodal features; The visualization module performs visualization based on the obtained dense representation of multimodal features to output the film and television relationship of the distribution of each data in the user data; The convolutional neural network was trained using the dataset to obtain a classification model for mood disorder detection; The user data to be analyzed is input into the classification model to output the film and television relationship of each data distribution.
[0025] More specifically, if Figure 2 As shown in the figure, a multimodal feature alignment and optimization framework is provided. This framework combines Convolutional Neural Network (CNN) and Multimodal Transformer (MulT) to solve the problem of cross-modal and group-level data alignment, and is particularly suitable for efficient integration and optimization of heterogeneous modal data in time series and feature space.
[0026] In this embodiment, the framework uses four modal data as input: eye tracking, EEG, facial expression, and gait. First, the original signal of each modality is feature extracted through a specific convolutional network, and the extracted low-level modal features contain time series information and spatial distribution characteristics. Then, these features are input into the multimodal transformer (MulT), which consists of multiple directional paired cross-modal transformer modules.
[0027] In the MulT module, each cross-modal Transformer builds query, key, and value matrices and uses the attention mechanism to learn the correlation between the features of the two modalities. In particular, this process effectively captures the complex interaction between the modalities by repeatedly strengthening the low-level feature associations between the target modality and the other source modality. To this end, MulT designs multi-head self-attention and cross-attention mechanisms in each directional module to make information fusion more comprehensive and accurate.
[0028] The fused time series data is mapped into a unified high-dimensional feature space, thereby achieving alignment and optimization of multimodal features. In the high-dimensional feature space, the framework further introduces a feature alignment loss function, which not only considers the correlation between modalities, but also reduces the distribution discrepancy between modal features, thereby improving the robustness and generalization ability of the final model.
[0029] After completing the feature alignment of multimodal data, in order to further ensure the differences between different individuals and maintain the feature consistency of the same individual, the present invention introduces a contrastive learning strategy. The aligned multimodal features are further optimized through the contrastive learning mechanism, thereby improving the classification performance and the robustness of the model.
[0030] Specifically, contrastive learning ensures that the feature representation of the same group of samples is more closely related and reduces the confusion of different groups of samples by constructing positive pairs and negative pairs. In this embodiment, the temporal and spatial features of EEG signals with spatiotemporal specificity are extracted. The specific steps are as follows:
[0031] 1. Feature extraction: Spatial convolution is used to combine the amplitude characteristics of different latent spatial components to capture the differences of EEG signals in the spatial domain.
[0032] Temporal convolution is used to extract the temporal pattern of EEG signal amplitude changes, thereby obtaining a feature representation with strong temporal correlation.
[0033] 2. Comparative learning process: In contrastive learning, the model constructs a feature representation , which represents the samples from the mini-batch training data. For the positive samples in the same set of samples and , maximize their similarity; for negative pairs of samples from different groups and , minimizing their similarity.
[0034] By introducing the contrastive loss function, the model ensures that the distribution of feature representations in the latent space is consistent with the target features.
[0035] 3. Design of contrast loss function: The contrast loss aims to maximize the similarity between positive pairs while minimizing the similarity between negative pairs. The specific form is:
[0036] L =- 1 N ∑ i =1 N log exp (sim( Z i , Z i + ) / τ) ∑ j=1 2N l [j≠1] exp(sim( Z i , Z j ) / τ) in: N is the number of samples in the mini-batch; ) indicates a sample and Similarity measure (usually cosine similarity); is a temperature parameter used to adjust the sensitivity of contrast loss.
[0037] In this example, the total loss of a small batch is: In this way, the model can further optimize the feature distribution in the aligned high-dimensional feature space, ensure that the feature representation of the same individual is more concentrated, and at the same time improve the distinction between different individuals, thereby providing higher accuracy and generalization ability for multimodal data analysis.
[0038] like Figure 3 As shown, the brain areas studied in emotion disorders and the neural mechanisms of their functional tasks are the biological theoretical basis of the entire framework. Figure 3 The study mainly involves several key brain regions and their association with emotional tasks: Amygdala: As the core area of emotion processing, it is directly related to emotional tasks and participates in the generation and regulation of emotional responses. Prefrontal Cortex: It plays an important role in expression recognition, gait detection and attention regulation, and is responsible for higher-level cognitive and emotional control. Anterior Cingulate Cortex: It is closely related to emotional tasks and helps regulate the intensity and response of emotions. Hippocampus: It is directly related to emotional memory and emotional disorders such as depression. Right Temporoparietal Junction: It is related to attention tasks and is mainly responsible for the integration and allocation of multimodal information. Occipital Cortex and Frontal Eye Fields: It is mainly related to eye movement tasks and visual attention. Basal Ganglia and Cerebellum: It is involved in gait detection tasks and provides support for motor features in emotional disorders. Figure 3The functional tasks of each brain region and its corresponding disorders are also clearly pointed out in the paper: Expression recognition task: dominated by the prefrontal cortex, it is the core link of emotion regulation, and abnormalities may lead to depression or anxiety. Eye movement task: mainly controlled by the frontal eye field, abnormalities may be related to distraction or anxiety. Gait detection: through the collaboration of the motor cortex and cerebellum, it evaluates the individual's movement and behavioral characteristics, which are often abnormal in emotional disorders. Emotion task: dominated by the amygdala and anterior cingulate cortex, directly involved in emotion generation and regulation. Attention task: supported by the parieto-occipital junction and other cortical areas, abnormalities may be related to bipolar disorder or attention disorder. And the complex interactive network between brain regions, for example: the strong connection between the amygdala and the prefrontal cortex indicates the coupling of emotion processing and cognitive control. The connection between the parieto-occipital junction and the motor cortex reflects the synergy of attention and movement regulation. The interaction between the hippocampus and the amygdala further supports the importance of emotional memory in emotional disorders. Abnormal activity or disruption of connectivity in these brain regions may lead to different manifestations of mood disorders, providing a clear direction for subsequent functional network modeling and data fusion.
[0039] Figure 3 The upper right part of the figure shows the abnormal characteristics of mood disorders (such as depression, bipolar disorder, and anxiety) in brain area connection patterns through functional networks. It provides a modeling idea for understanding the neural mechanisms of different mood disorders and reveals the differences between various disorders in the form of functional abnormal networks.
[0040] Depression is often characterized by weakened or abnormally enhanced connectivity between specific brain regions. For example, the connection between the amygdala and the prefrontal cortex may be too weak, resulting in a decrease in the ability to regulate emotions. The depth of the color of the nodes in the diagram may reflect the degree of functional activity. Some dark nodes in the depression pathway represent highly active regions, which may be associated with excessive emotional burden. Compared with depression, the functional network of bipolar disorder may be more complex, characterized by excessive connectivity or unstable connection strength between different brain regions. This abnormal network feature may lead to extreme fluctuations in emotions, attention, and behavior in individuals. These features may be shown in the diagram through the density or complexity of the connection. The anxiety pathway may be characterized by excessive connectivity in regions such as the amygdala and anterior cingulate cortex, leading to an overreaction to threat signals, while abnormal connectivity in brain regions related to attention regulation may make it more difficult for individuals to shift their attention away from threatening stimuli. The network nodes and connections in these emotional disorder pathways form the basis for heterogeneous graph representation. Each disorder corresponds to a specific functional network, indicating the specificity of the dynamic interaction between brain regions in different emotional disorders.
[0041] These networks reveal the specificity of the neural mechanisms of different mood disorders and provide a theoretical basis for understanding the pathological mechanisms of mood disorders. By analyzing these pathways, different types of mood disorders can be distinguished, supporting more accurate diagnosis.
[0042] In addition, feature extraction and fusion of multimodal data, including EEG, eye movement, expression, gait and other modalities, are performed. Through message passing, multi-head mapping and heterogeneous fusion, these modal data are integrated into a unified high-dimensional feature space to provide input support for the modeling of functional networks. This multimodal fusion method enhances the relevance and expressiveness of cross-modal information. Focusing on the modeling of functional networks, it specifically demonstrates the construction of functional networks through heterogeneous graphs, and the extraction of specific brain region interaction patterns based on the generation of meta-paths. This meta-path analysis provides a data-driven tool for revealing the specific characteristics of emotional disorders, while supporting the prediction and classification of disorder types.
[0043] like Figure 4 It shows how to extract features from four modalities: EEG, facial expression, eye movement, and gait. The characteristics of each modality determine its corresponding neural network architecture and processing strategy. The following is a detailed description of each modality and the technical principle:
[0044] The Spatio-Temporal Graph Convolution Network (STGCN) is used to extract EEG features, which can efficiently capture the interaction of spatial and temporal features in EEG signals.
[0045] 1) Graph Convolution: Using EEG channels as nodes, a graph structure is established through functional connections between brain regions, and graph convolution operations are applied to learn the relationship features between nodes.
[0046] 2) Temporal modeling: Based on graph convolution, temporal convolution is added to extract dynamic change features and model the temporal dependency of signals.
[0047] The residual network (ResNet) is used to extract visual features in facial expression images, including local changes (such as frowning, smiling, etc.) and overall morphology, and is highly robust to subtle changes in expression.
[0048] 1) Convolution operation: extract low-level features (such as edges, textures) and high-level features (such as overall expression patterns) of facial expressions through multi-layer convolution operations.
[0049] 2) Residual Connection: It solves the gradient vanishing problem in deep neural network training and improves the expressiveness of the model.
[0050] 3) Pre-trained weights: It is possible to use weights pre-trained on large-scale datasets (such as ImageNet) for transfer learning to enhance the initial performance of the model.
[0051] Long Short-Term Memory (LSTM) network is used to capture the time series characteristics of eye movement trajectory data, such as changes in gaze point position and temporal correlation.
[0052] 1) Input data: Eye movement data usually includes gaze point coordinates (x, y) and time information, which are input into the LSTM model as a time series.
[0053] 2) Memory unit: Controls the update of historical states and the output of the current time step through the input gate, forget gate, and output gate to retain long-term dependency information.
[0054] 3) Output features: Generate high-dimensional feature representations containing eye movement behavior patterns for subsequent fusion.
[0055] Convolutional long short-term memory (Convolutional LSTM, ConLSTM) is used to extract the temporal dynamic features and spatial patterns of gait for modeling gait behavior. ConLSTM combines the spatial feature extraction capability of convolutional networks and the temporal modeling capability of LSTM, and is sensitive to the dynamic changes of gait data.
[0056] 1) Convolutional module: captures the local spatial features of gait videos, such as the direction and amplitude of movement.
[0057] 2) LSTM network: further perform temporal modeling on these local features to generate global dynamic features of the gait sequence.
[0058] 3) Data enhancement: Perform enhancement processing such as rotation and scaling on the input gait data to improve the generalization ability of the model.
[0059] Each modality uses the neural network architecture that best suits its characteristics to maximize feature extraction efficiency. This allows the extracted multimodal features to maintain high-dimensional information expression, laying the foundation for subsequent fusion and emotional disorder modeling. Neural networks designed with different architectures can process multimodal data in parallel and improve the overall efficiency of the system.
[0060] The multimodal low-rank matrix feature decomposition part integrates the high-dimensional data from multimodal feature extraction, using the idea of low-rank matrix decomposition to reduce computational complexity while retaining important information interactions between and within modalities. The following is a detailed description of its technical details:
[0061] In order to achieve effective fusion of multimodal data and solve the problems caused by feature differences, redundant information and high dimensionality between modalities, low-rank decomposition is used to capture the shared information between modalities while retaining the key features unique to the modality.
[0062] The high-dimensional feature representation is decomposed into a set of low-rank matrices represented as the product of modality-specific factors and shared factors.
[0063] in, is a mode-specific low-rank factor, is the modal sharing factor.
[0064] The regularization term is used to constrain the factors, eliminate redundant information, and ensure that the fusion process focuses on useful features.
[0065] First, the feature matrix extracted from each modality (such as EEG, expression, eye movement, gait) is decomposed into modality-specific and shared parts through low-rank decomposition: Modality-specific factors: capture the unique features of each modality, such as the frequency pattern of EEG and the local texture of expression. Modality shared factors: capture the common characteristics between modalities, such as temporal correlation across modalities. Modality-specific factors and shared factors are reconstructed through linear or nonlinear methods to generate a fused feature representation, and the expression power of the fused representation is further enhanced by combining activation functions (such as ReLU). Weights are assigned to each modality to ensure that the more important modality features dominate the information fusion process. Weight learning is achieved through an additional attention mechanism (Attention) to automatically adjust the importance of the modality. The problem of inconsistent representation of features of different modalities is solved by alignment operations in the time or space dimensions.
[0066] The multimodal low-rank matrix eigendecomposition has the advantages of high efficiency, that is, reducing data dimension and computational complexity through low-rank representation; robustness, that is, the decomposition of modal specific factors and shared factors enhances the anti-interference ability of redundant features; information retention, that is, taking into account the correlation between modalities and the independence within modalities to achieve effective information fusion.
[0067] Multimodal low-rank matrix feature decomposition is one of the core steps of the entire model. Through precise decomposition and reconstruction mechanisms, complex high-dimensional features from multimodal are integrated into a compact and information-rich representation.
[0068] The data of each modality is processed using an LSTM network. The role of LSTM is to capture long-term dependencies in time series and transform complex time series data into a compact representation, thereby retaining the most important dynamic changes for mood disorder detection. When processing time series, LSTM generates a series of hidden states that represent the contextual information at each time point. Finally, the information of certain key time steps (such as the last frame) can be extracted or the information of all time steps can be aggregated (such as taking the average or weighted average) to form a single context feature representation.
[0069] The context vector of each modality is further processed through a specific neural network layer (such as a fully connected layer, a convolutional layer, or a dimensionality reduction module) to map it to a unified feature space. During the encoding process, the complexity of the network can be constrained or specific methods (such as low-rank decomposition) can be used to optimize the feature representation and reduce data redundancy. Multimodal data is optimized and fused using methods such as tensor decomposition or attention mechanisms to ensure that each modality is fully expressed in the final feature representation while avoiding the introduction of redundancy and noise.
[0070] The cross-modal fusion part aims to integrate multimodal data (such as EEG, facial expressions, eye movements, gait, etc.) into a semantically consistent shared representation while retaining the diversity and relevance of modal features by introducing the Transformer architecture and attention mechanism. The fusion process first takes the features obtained from the low-rank cross-modal part as input, extracts the local dependency features within the modality through a convolution operation (Conv1D), and adds position information to the time series data through an embedding layer (Positional Embedding), thereby enhancing context perception.
[0071] The core cross-modal fusion is implemented by Transformer, the most important of which is the cross-modal attention mechanism. In this mechanism, the feature vector of each modality is regarded as a query, and the features of the other modalities are used as keys and values. The dependency between modalities is captured through attention calculation. At the same time, the self-attention mechanism further strengthens the fused feature representation, ensuring sufficient information interaction between modalities, and ultimately generating semantically consistent multi-modal features.
[0072] In order to achieve cross-modal feature alignment, this part also adopts strategies such as time synchronization, feature scale normalization and modality mapping. Dynamic Time Warping (DTW) is used to align time series modalities (such as eye movements and gait), and feature normalization adjusts the data scales of different modalities to be consistent. At the same time, the shared Transformer weight design enables different modalities to be mapped and interacted in a unified latent space.
[0073] In terms of optimization, the model combines multiple loss functions, including modality alignment loss and task-related classification or regression loss. Modality alignment loss enhances the semantic consistency between modalities by minimizing the mean and covariance differences of modality feature distributions. In addition, regularization terms and Dropout mechanisms are used to suppress overfitting and improve the generalization ability of the model.
[0074] This embodiment also provides an emotional disorder detection device, which is implemented by the emotional disorder detection method provided by the above embodiment, and its specific implementation content is as follows: 1) Multimodal input and alignment: EEG, facial expressions, eye movements, and gait data are aligned through the time axis and feature level and unified into a multimodal input format to ensure the consistency of information timing and feature alignment.
[0075] 2) Emotional disorder modeling: Use brain network to map functional areas, associate emotion-task pathways, and analyze the interaction patterns and task responses of brain regions related to emotion disorders.
[0076] 3) Low-rank matrix decomposition: Perform low-rank decomposition on the multimodal feature matrix to extract shared features and modality-specific factors, reduce data dimensions and enhance feature expression.
[0077] 4) Cross-modal fusion: Based on the interaction mechanism between Transformer and convolutional networks, modal features are integrated to enhance cross-modal information sharing and high-order semantic expression capabilities.
[0078] 5) Diagnostic prediction: By fusing feature vectors, classification tasks and performance evaluation are performed to generate emotional disorder detection results and provide model performance indicator analysis.
[0079] In addition, the terms "upper", "lower", "inner", "outer", "front", "rear" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps described in these embodiments do not limit the scope of the present invention.
[0080] Of course, the above description is only a specific embodiment of the present invention and is not intended to limit the scope of implementation of the present invention. All equivalent changes or modifications made according to the structure, characteristics and principles described in the patent application scope of the present invention should be included in the patent application scope of the present invention.
[0081] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network, characterized in that: The following steps are involved: Obtain user data and corresponding emotional disorder types, including the user's EEG information, eye movement information, micro-expression and gait information; Perform data alignment processing on user data to obtain an initial data group with consistent data modality, and use the initial data corresponding to each type of information as nodes, and use the logic of inferring emotional disorders with medical prior knowledge as connecting lines to construct the corresponding emotional disorder inference meta-path; The initial data set, the types of mood disorders, and the corresponding mood disorder inference meta-pathways are combined into a data set; Constructing a corresponding convolutional neural network based on a multi-network framework, wherein the convolutional neural network includes a data alignment processing module, a heterogeneous feature extraction module, a feature fusion module and a prediction module; The data alignment processing module is used to perform data alignment processing on the input user data to generate an initial data group; The heterogeneous feature extraction module includes a multimodal feature extractor, which is used to extract data features of each group of data in the initial data group to output a corresponding multimodal feature group; The feature fusion module uses a low-rank multimodal fusion method to perform feature encoding for each modal feature to obtain multiple sets of feature vectors in the same dimension, and uses a pre-constructed mood disorder inference meta-path to associate each feature vector in the obtained multiple sets of feature vectors to obtain a dense representation of the multimodal features; The visualization module performs visualization based on the obtained dense representation of multimodal features to output the film and television relationship of the distribution of each data in the user data; The convolutional neural network was trained using the dataset to obtain a classification model for mood disorder detection; The user data to be analyzed is input into the classification model to output the film and television relationship of each data distribution.
2. The method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network according to claim 1, characterized in that: The data alignment process includes cross-modality alignment and group level alignment.
3. The method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network according to claim 2, characterized in that: The cross-modal alignment refers to achieving synchronous conversion of multi-modal asynchronous time series signals by mapping information of different modalities into a shared feature space and calculating attention scores between modalities.
4. The method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network according to claim 2, characterized in that: The group horizontal alignment refers to calculating the unit vector of the multimodal latent space features and mapping the features on the spherical space with the help of the feature constraints of the Frobenius two-norm.
5. The method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network according to claim 1, characterized in that: The multimodal feature extractor includes an STGCN encoder for EEG information, a ResNet encoder for expression information, an LSTM encoder for eye movement information, and a ConLSTM encoder for gait information.
6. The method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network according to claim 1, characterized in that: The specific process of the low-rank multimodal fusion is as follows: For a single modal feature, LSTM is used to compress the corresponding time series information, and the hidden state context vector in the compressed single modal feature is extracted, and encoding is performed based on the hidden state context vector to obtain the corresponding feature vector.
7. The method for detecting emotional disorders based on a multimodal heterogeneous graph convolutional neural network according to claim 1, characterized in that: The mood disorder inference meta-pathway includes a depression pathway, a bipolar disorder pathway, a cognitive disorder pathway and an anxiety disorder pathway.
8. An emotional disorder detection device, characterized in that: The emotional disorder detection method based on a multimodal heterogeneous graph convolutional neural network is implemented by any one of claims 1 to 7.
Citation Information
Patent Citations
Biphasic affective disorder feature extraction and classification method based on high-order function network
CN118070127A
Electroencephalogram space-time pattern extraction method and system for emotional disorder assessment task
CN118490227A
Cognitive quantitative detection machine based on multi-modal emotion artificial intelligence
CN114504320A
Aspect-level sentiment analysis method fusing multi-modal data
CN114936623A
Multi-modal human gait emotion recognition method
CN115273236A
Cited By
Evaluation method and system based on angel syndrome phenotype detection
CN120727260A
Depression population determination system and method based on asynchronous electroencephalogram
CN120732418A
Multi-type target discrimination model training and target discrimination method for brain-computer interface
CN121388855A
Training of multi-type target discrimination model of brain-computer interface and target discrimination method
CN121388855B
Multi-modal emotion recognition method based on electroencephalogram signal and facial expression fusion
CN121861735A