Multi-modal brain network classification method based on feature decoupling and dynamic graph construction

The multimodal brain network classification method based on feature decoupling and dynamic graph construction solves the problem of information loss caused by independent processing of functional and structural data in existing technologies, achieves high-accuracy and comprehensive classification of brain network states, and enhances the adaptability and interpretability of the model.

CN120747650AActive Publication Date: 2025-10-03HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202511240103.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-03
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing multimodal brain network analysis methods usually process functional magnetic resonance imaging and structural magnetic resonance imaging data independently, resulting in the loss of key brain function and structural correlation information, and making it difficult to capture the dynamic characteristics of brain networks, limiting the model's sensitivity and accuracy to changes in brain network states.

Method used

A multimodal brain network classification method based on feature decoupling and dynamic graph construction is adopted. The unique features and common features in the functional and structural network matrices are separated by the feature decoupling module, and the dynamic graph is generated in combination with the graph attention network (GAT) to capture the dynamic change characteristics of the brain network. Information fusion and classification are performed through the cross-attention module and fusion module.

Benefits of technology

It improves the accuracy and comprehensiveness of brain network state classification, enhances the robustness and generalization ability of the model, provides the interpretability of the model decision-making process, and can better adapt to complex brain network data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747650A_ABST
    Figure CN120747650A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal brain network classification method based on feature decoupling and dynamic graph construction, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining functional magnetic resonance imaging data and structural magnetic resonance imaging data of a to-be-detected person, and obtaining multi-modal brain network data according to the functional magnetic resonance imaging data and the structural magnetic resonance imaging data, and processing the multi-modal brain network data based on a preset multi-modal brain network classification model to obtain the brain network state of the to-be-tested person. According to the method, by combining functional magnetic resonance imaging (fMRI) and structural magnetic resonance imaging (sMRI) data, the advantages of the two modes can be fully utilized, so that the accuracy and comprehensiveness of brain network state classification are improved, a dynamic graph attention module based on a graph attention network (GAT) can capture the dynamic change characteristics of a brain network, dynamic graph representation is constructed, and the classification accuracy of the brain network state is improved. And the sensitivity of the model to the dynamic connection relationship between brain regions is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal brain network classification method based on feature decoupling and dynamic graph construction. Background Art

[0002] In the field of neuroscience, multimodal brain network analysis has garnered widespread attention due to its ability to provide a comprehensive view of brain structure and function. Functional magnetic resonance imaging (fMRI) can reflect changes in blood flow during brain activity, while structural magnetic resonance imaging (sMRI) can reveal the brain's anatomical structure. Traditional brain network analysis methods typically process data from these two modalities independently, potentially missing critical information about brain function and structural connections.

[0003] Furthermore, existing multimodal fusion methods often simply linearly combine or stack data from different modalities. This approach struggles to capture the dynamic nature of brain networks, namely how inter-brain connectivity changes over time and under different conditions. This approach ignores the complexity and dynamics of brain network connectivity, limiting the model's sensitivity and accuracy to changes in brain network state. Summary of the Invention

[0004] The problem solved by the present invention is one or more of the above-mentioned related technical problems.

[0005] To solve the above problems, the present invention provides a multimodal brain network classification method, device and equipment based on feature decoupling and dynamic graph construction.

[0006] In a first aspect, the present invention provides a multimodal brain network classification method based on feature decoupling and dynamic graph construction, comprising: Acquire functional magnetic resonance imaging data and structural magnetic resonance imaging data of the person to be tested, Obtaining multimodal brain network data based on the functional magnetic resonance imaging data and the structural magnetic resonance imaging data, wherein the multimodal brain network data includes at least one functional connectivity network matrix and a structural network matrix; Based on a preset multimodal brain network classification model, the multimodal brain network data is processed to obtain the brain network state of the person to be tested; Among them, the multimodal brain network classification model includes a feature decoupling module and a dynamic graph attention module constructed based on the GAT network.

[0007] Optionally, the multimodal brain network classification model further includes a cross-attention module and a fusion module. The multimodal brain network data is processed based on the preset multimodal brain network classification model to obtain the brain network state of the person to be tested, including: Processing the functional connectivity network matrix and the structural network matrix respectively by the feature decoupling module to obtain functional feature data, shared functional feature data, structural feature data and shared structural feature data; Processing the shared functional feature data and the shared structural feature data through the cross-attention module to obtain a dynamic fusion graph adjacency matrix; Processing the dynamic fusion graph adjacency matrix, the functional feature data, and the structural feature data through the dynamic graph attention module to obtain a first dynamic graph and a second dynamic graph; The first dynamic image and the second dynamic image are fused by the fusion module to obtain fusion feature data, and the fusion features are classified to obtain the brain network state of the person to be tested.

[0008] Optionally, the feature decoupling module includes a first encoding module and a second encoding module; the feature decoupling module processes the functional connectivity network matrix and the structural network matrix respectively to obtain functional feature data, shared functional feature data, structural feature data, and shared structural feature data, including: Processing the functional connectivity network matrix through the first encoding module to obtain the functional feature data and the shared functional feature data; The structural network matrix is ​​processed by the second encoding module to obtain the structural feature data and the shared structural feature data.

[0009] Optionally, the first encoding module includes a feature extraction layer, a first linear projection space, and a second linear projection space; and processing the functional connectivity network matrix by the first encoding module to obtain the functional feature data and the shared functional feature data includes: Performing feature extraction on the functional connection network matrix through the feature extraction layer to obtain intermediate feature data; The intermediate feature data is mapped using the first linear projection space to obtain the functional feature data, and the intermediate feature data is mapped using the second linear projection space to obtain the shared functional feature data.

[0010] Optionally, the processing of the shared functional feature data and the shared structural feature data by the cross-attention module to obtain a dynamic fusion graph adjacency matrix includes: Converting the shared functional feature data and the shared structural feature data respectively to obtain a first vector matrix and a second vector matrix, wherein the first vector matrix includes a first query vector and a first key vector, and the second vector matrix includes a second query vector and a second key vector; Obtaining a second attention matrix based on the first query vector and the second key vector, and obtaining a first attention matrix based on the first key vector and the second query vector; The first attention matrix and the second attention matrix are fused to obtain the dynamic fusion graph adjacency matrix.

[0011] Optionally, the dynamic graph attention module includes a first GAT network, a second GAT network, a third GAT network, and a fourth GAT network; the dynamic graph attention module processes the dynamic fusion graph adjacency matrix, the functional features, and the structural features to obtain the first dynamic graph and the second dynamic graph, including: Processing the dynamic fusion graph adjacency matrix and the functional features through the first GAT network to obtain a first intermediate feature; Processing the dynamic fusion graph adjacency matrix through the second GAT network to obtain temporary functional features; Processing the dynamic fusion graph adjacency matrix and the structural features through the third GAT network to obtain a second intermediate feature; Processing the dynamic fusion graph adjacency matrix through the fourth GAT network to obtain temporary structural features; The first intermediate feature and the temporary functional feature are fused to obtain the first dynamic graph, and the second intermediate feature and the temporary structural feature are fused to obtain the second dynamic graph.

[0012] Optionally, the fusion module includes a feedforward neural network; and the fusing of the first dynamic image and the second dynamic image by the fusion module to obtain fused feature data includes: Splicing the first dynamic image and the second dynamic image to obtain splicing vector feature data; Processing the splicing vector feature data through the feedforward neural network to obtain weight vector data; Fused feature data is obtained according to the first dynamic graph, the second dynamic graph, and the weight vector data.

[0013] Optionally, the first vector matrix includes a first value vector, and the second vector matrix includes a second value vector; and the processing of the shared functional feature data and the shared structural feature data by the cross-attention module to obtain the dynamic fusion graph adjacency matrix further includes: Obtaining first temporary feature data according to the first attention matrix and the first value vector; Obtaining cross-modal feature data according to the first temporary feature data and the shared functional feature data; The cross-modal feature data is used to enhance the auxiliary output of the multimodal brain network classification model training.

[0014] In a second aspect, the present invention provides a multimodal brain network classification device based on feature decoupling and dynamic graph construction, comprising: An acquisition unit is used to acquire functional magnetic resonance imaging data and structural magnetic resonance imaging data of the person to be tested, a construction unit, configured to obtain multimodal brain network data based on the functional magnetic resonance imaging data and the structural magnetic resonance imaging data, wherein the multimodal brain network data includes at least one functional connectivity network matrix and one structural network matrix; A processing unit is used to process the multimodal brain network data based on a preset multimodal brain network classification model to obtain the brain network state of the person to be tested; wherein the multimodal brain network classification model includes a feature decoupling module and a dynamic graph attention module constructed based on the GAT network.

[0015] In the third aspect, the present invention provides a multimodal brain network classification device based on feature decoupling and dynamic graph construction, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the multimodal brain network classification method based on feature decoupling and dynamic graph construction as described in the first aspect.

[0016] The beneficial effects of the multimodal brain network classification method, device and equipment based on feature decoupling and dynamic graph construction of the present invention are: By combining functional magnetic resonance imaging (fMRI) and structural magnetic resonance imaging (sMRI) data, the advantages of both modalities can be fully utilized, thereby improving the accuracy and comprehensiveness of brain network state classification. The feature decoupling module can separate unique and shared features in the functional and structural network matrices, reducing information interference between modalities and thereby improving feature purity and expressiveness. The dynamic graph attention module based on the graph attention network (GAT) can capture the dynamic changes in brain networks, construct a dynamic graph representation, and enhance the model's sensitivity to dynamic connectivity relationships between brain regions. These features enhance the model's robustness and generalization ability, enabling it to better adapt to complex brain network data. Furthermore, the attention mechanism of the dynamic graph attention module provides interpretability of the model's decision-making process, helping to understand the role and contribution of different brain regions in specific brain network states. Overall, this invention demonstrates significant advantages in brain network classification tasks and can provide more reliable technical support for the diagnosis and research of brain diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of a flow chart of a multimodal brain network classification method based on feature decoupling and dynamic graph construction according to an embodiment of the present invention; Figure 2 Schematic diagram of the structure of a multimodal brain network classification model according to an embodiment of the present invention; Figure 3 Schematic diagram of the feature decoupling module processing process of an embodiment of the present invention; Figure 4 Schematic diagram of a cross-modal attention module according to an embodiment of the present invention; Figure 5 Schematic diagram of the structure of the fusion module of an embodiment of the present invention. DETAILED DESCRIPTION

[0018] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0019] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0020] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to"; the term "based on" means "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0023] In response to the problems existing in the above-mentioned related technologies, an embodiment of the present invention provides a multimodal brain network classification method based on feature decoupling and dynamic graph construction.

[0024] like Figure 1 As shown, an embodiment of the present invention provides a multimodal brain network classification method based on feature decoupling and dynamic graph construction, including: Step S100: acquiring functional magnetic resonance imaging data and structural magnetic resonance imaging data of a person to be tested.

[0025] Specifically, acquiring fMRI and sMRI data from individuals is typically accomplished using magnetic resonance imaging (MRI). fMRI primarily records brain blood oxygenation-dependent signals, reflecting the dynamics of neural activity; sMRI captures brain anatomical structure and identifies brain region boundaries. Combining these two allows multimodal brain network classification models to leverage both static anatomical and dynamic functional connectivity information, improving classification accuracy and robustness, and providing a powerful tool for brain disease research and clinical applications.

[0026] Step S200: obtaining multimodal brain network data based on the functional magnetic resonance imaging data and the structural magnetic resonance imaging data, wherein the multimodal brain network data includes at least one functional connection network matrix and one structural network matrix.

[0027] Specifically, fMRI data are first denoised to reduce interference from cardiac signals and head motion artifacts. Then, methods such as the Pearson correlation coefficient or sliding window correlation can be used to calculate the functional connectivity network matrix, which reflects the dynamic connectivity between functional brain regions. For sMRI data, diffusion tensor imaging (DTI) and fiber tract tracing algorithms can be used to capture the connectivity of white matter fiber tracts. Based on edge weight formulas from graph theory, a structural network matrix is ​​constructed to describe the physical connectivity between brain regions.

[0028] In some embodiments, such as Figure 2 As shown in Figure 2, the structural network matrix can be constructed using sMRI data or DTI data to construct the corresponding structural network matrix, and the multimodal brain network data includes a functional connection network matrix and two structural network matrices.

[0029] The functional connectivity network matrix captures the dynamic characteristics of functional activity in brain regions, helping to identify complex patterns of neural activity. The structural network matrix provides high-precision information on the physical connections between brain regions, avoiding the influence of imaging artifacts. Combining these two networks into multimodal brain network data not only enriches feature representation but also improves the performance of subsequent classification tasks, enabling the model to more comprehensively understand brain network states.

[0030] Step S300: Based on a preset multimodal brain network classification model, the multimodal brain network data is processed to obtain the brain network state of the person to be tested; wherein the multimodal brain network classification model includes a feature decoupling module and a dynamic graph attention module constructed based on the GAT network.

[0031] Specifically, the feature decoupling module first feeds multimodal brain network data (functional connectivity network matrix and structural network matrix) into separate feature extraction networks (e.g., independent encoders) to extract features from each modality. Then, through linear or nonlinear projection, the decoupled features are mapped into independent and shared feature spaces, respectively. This purifies information and separates commonalities from specific features, while also preserving potential collaborations between modalities. Based on the decoupled features, the dynamic graph attention module utilizes a dynamic graph generator to dynamically adjust the brain network graph structure based on feature similarities, generating a dynamic graph. A GAT-based attention mechanism then calculates attention weights between nodes, performing weighted aggregation of brain region features to capture dynamic connectivity relationships and key connectivity patterns between brain regions. The brain network states include both normal and abnormal states. Multi-level interpretable information (modality-level and brain connectivity-level) is also provided to reveal the underlying basis for the model's decisions.

[0032] In this process, the feature decoupling module avoids inter-modal interference and improves the accuracy of feature representation; the combination of the dynamic graph generator and the GAT network enhances the ability to capture the dynamic connection relationship between brain regions; the overall design improves the robustness and accuracy of the model in classifying brain network states, providing a more reliable solution for complex brain network analysis.

[0033] In this embodiment, by combining functional magnetic resonance imaging (fMRI) and structural magnetic resonance imaging (sMRI) data, the advantages of both modalities can be fully utilized, thereby improving the accuracy and comprehensiveness of brain network state classification. The feature decoupling module can separate the unique and shared features in the functional and structural network matrices, reducing information interference between modalities, thereby improving feature purity and expressiveness. The dynamic graph attention module based on the graph attention network (GAT) can capture the dynamic changes in brain networks, construct a dynamic graph representation, and enhance the model's sensitivity to dynamic connectivity relationships between brain regions. These features enhance the model's robustness and generalization ability, enabling it to better adapt to complex brain network data. In addition, the attention mechanism of the dynamic graph attention module provides interpretability of the model's decision-making process, helping to understand the role and contribution of different brain regions in specific brain network states. Overall, this invention demonstrates significant advantages in brain network classification tasks and can provide more reliable technical support for the diagnosis and research of brain diseases.

[0034] Optionally, the multimodal brain network classification model further includes a cross-attention module and a fusion module. The multimodal brain network data is processed based on the preset multimodal brain network classification model to obtain the brain network state of the person to be tested, including: Processing the functional connectivity network matrix and the structural network matrix respectively by the feature decoupling module to obtain functional feature data, shared functional feature data, structural feature data and shared structural feature data; Processing the shared functional feature data and the shared structural feature data through the cross-attention module to obtain a dynamic fusion graph adjacency matrix; Processing the dynamic fusion graph adjacency matrix, the functional feature data, and the structural feature data through the dynamic graph attention module to obtain a first dynamic graph and a second dynamic graph; The first dynamic image and the second dynamic image are fused by the fusion module to obtain fusion feature data, and the fusion features are classified to obtain the brain network state of the person to be tested.

[0035] Specifically, if Figure 2 As shown, the model also includes a cross-attention module and a fusion module to improve the classification accuracy and interpretability.

[0036] The feature decoupling module processes the functional connectivity network matrix (from fMRI data) and the structural network matrix (from sMRI data) separately. It extracts unique features (functional and structural) as well as shared features (shared functional and structural) from each modality. This aims to purify information and isolate the commonalities and unique characteristics of each modality, allowing for more accurate capture and utilization of key insights from multimodal data.

[0037] The shared feature data is processed across attention modules to learn the correlations between different modalities. Through this processing, a dynamic fusion graph adjacency matrix is ​​obtained, which characterizes the interactive relationships between functional and structural features.

[0038] The dynamic graph attention module uses the dynamic fusion graph adjacency matrix and proprietary feature data to generate two dynamic graphs: a first dynamic graph (based on functional features) and a second dynamic graph (based on structural features). These dynamic graphs capture the dynamic characteristics of brain networks and reflect the connectivity patterns between brain regions under different states.

[0039] The fusion module fuses the features of the first dynamic image and the second dynamic image to obtain fused feature data.

[0040] Finally, the fused features are classified to determine the brain network status of the person being tested.

[0041] Through feature decoupling and cross-attention mechanisms, the model enhances its representation of multimodal data features and improves feature differentiation. The use of dynamic graph attention and fusion modules enables the model to accurately capture brain network dynamics, enhancing classification accuracy. The cross-attention and fusion modules empower the model to account for the contributions of different modalities to classification decisions, improving interpretability. Furthermore, the model's dynamic graph construction and adaptive attention fusion mechanism enable it to adapt to different brain network states, significantly improving generalization.

[0042] Optionally, the feature decoupling module includes a first encoding module and a second encoding module; the feature decoupling module processes the functional connectivity network matrix and the structural network matrix respectively to obtain functional feature data, shared functional feature data, structural feature data, and shared structural feature data, including: Processing the functional connectivity network matrix through the first encoding module to obtain the functional feature data and the shared functional feature data; The structural network matrix is ​​processed by the second encoding module to obtain the structural feature data and the shared structural feature data.

[0043] Optionally, the first encoding module includes a feature extraction layer, a first linear projection space, and a second linear projection space; and processing the functional connectivity network matrix by the first encoding module to obtain the functional feature data and the shared functional feature data includes: Performing feature extraction on the functional connection network matrix through the feature extraction layer to obtain intermediate feature data; The intermediate feature data is mapped using the first linear projection space to obtain the functional feature data, and the intermediate feature data is mapped using the second linear projection space to obtain the shared functional feature data.

[0044] Specifically, the feature decoupling module processes the functional connection network matrix and the structural network matrix through the first encoding module and the second encoding module respectively. The first encoding module includes a feature extraction layer, a first linear projection space, and a second linear projection space. The feature extraction layer first extracts intermediate feature data from the functional connection network matrix, and then the first linear projection space maps the intermediate features into functional feature data, and the second linear projection space maps the intermediate features into shared functional feature data. Similarly, the second encoding module processes the structural network matrix to obtain structural feature data and shared structural feature data. This process is beneficial to improving the discriminability of feature representation, improving classification accuracy, and enhancing the interpretability of the model.

[0045] In some embodiments, such as Figure 3 As shown in Figure 2, a schematic diagram of the feature decoupling module processing process.

[0046] The input data includes brain network matrices of two modalities (functional connectivity network matrix) and (Structural Network Matrix), representing functional and structural modes respectively.

[0047] The input data of each modality is processed by the feature extraction module. This step may include preprocessing operations such as denoising and normalization.

[0048] Residual layer processing: The extracted feature data is further processed through the residual layer to enhance the model's ability to learn features.

[0049] Feature decoupling: functional connection network matrix through the first encoding module Processing to obtain functional feature data and shared functional feature data .

[0050] The structure network matrix is ​​encoded by the second encoding module Processing to obtain structural feature data and shared structural feature data . Among them, the second encoding module includes a corresponding feature extraction layer, a third linear projection space and a fourth linear projection space.

[0051] Feature projection: Through the first linear projection space ( Figure 3 The corresponding unique projection head) for the intermediate functional features Mapping to obtain functional feature data , and through the second linear projection space ( Figure 3 The corresponding common projection head) for the intermediate functional features Mapping to obtain shared functional feature data .

[0052] Through the third linear projection space ( Figure 3 The corresponding unique projection head) for the intermediate structural features Mapping to obtain structural feature data , through the third quadrilinear projection space ( Figure 3 The corresponding common projection head) for the intermediate structural features Mapping to obtain shared structural feature data .

[0053] During training, the maximum mean difference (MMD) loss is introduced To ensure that the shared features of different modalities have similar probability distributions, thus forcing the encoder to reach a "consensus" at the output of the common projection head. At the same time, the orthogonal loss is introduced. To ensure that shared features and unique features within each modality are as uncorrelated as possible, thereby forcing them to be geometrically orthogonal in the feature space. MMD implicitly maps features into a high-dimensional reproducing kernel Hilbert space (RKHS) using a kernel function (such as a Gaussian radial basis function kernel) and calculates the distance between two mean elements distributed in this space. By minimizing this distance, the system forces the two independent encoders to learn to reach a "consensus" at the output of their respective common projection heads (corresponding to the second and fourth linear projection spaces), that is, to output statistically indistinguishable features. Simultaneously, an orthogonality loss is used to ensure that shared features and unique features within each modality are as uncorrelated as possible. This loss is implemented by calculating and minimizing the squared cosine similarity between the shared feature vector and the unique feature vector, thereby forcing them to be geometrically orthogonal in the feature space and ensuring the "purity" of the decoupling.

[0054] Model training and optimization: The model is trained by minimizing the loss function through backpropagation and an optimization algorithm such as Adam or SGD.

[0055] This process enables the model to more fully express the characteristics of multimodal data through feature decoupling and cross-modal attention mechanism, improve the feature discrimination ability, thereby improving classification accuracy, and provide modality-level explanations to enhance the interpretability of the model.

[0056] Optionally, the processing of the shared functional feature data and the shared structural feature data by the cross-attention module to obtain a dynamic fusion graph adjacency matrix includes: Converting the shared functional feature data and the shared structural feature data respectively to obtain a first vector matrix and a second vector matrix, wherein the first vector matrix includes a first query vector and a first key vector, and the second vector matrix includes a second query vector and a second key vector; Obtaining a second attention matrix based on the first query vector and the second key vector, and obtaining a first attention matrix based on the first key vector and the second query vector; The first attention matrix and the second attention matrix are fused to obtain the dynamic fusion graph adjacency matrix.

[0057] Optionally, the first vector matrix includes a first value vector, and the second vector matrix includes a second value vector; and the processing of the shared functional feature data and the shared structural feature data by the cross-attention module to obtain the dynamic fusion graph adjacency matrix further includes: Obtaining first temporary feature data according to the first attention matrix and the first value vector; Obtaining cross-modal feature data according to the first temporary feature data and the shared functional feature data; The cross-modal feature data is used to enhance the auxiliary output of the multimodal brain network classification model training.

[0058] In some embodiments, such as Figure 4 Figure 2 shows a schematic diagram of the cross-modal attention module. This module's input is the shared feature matrix (including shared functional and structural feature data) from the decoupling module. A bidirectional cross-modal attention mechanism is employed to learn the interaction patterns between nodes. This mechanism is implemented using a standard multi-head attention layer, based on the principles of the Transformer architecture. The advantage of this multi-head mechanism is that it allows the model to learn multiple different attention patterns in parallel across different representation subspaces, thereby capturing richer inter-modal relationships.

[0059] Specifically, we first learn three independent linear projection matrices for each attention head and shared features of each modality, and transform them into queries (query vectors) respectively. , key (Key, key vector) Sum value (Value, value vector) These three different projections allow the model to independently learn “what to focus on (Query)”, “how to focus on (Key)” and “what information to provide (Value)”.

[0060] Then, a bidirectional attention computation is performed: in the first direction, a modality query (the first query vector corresponding to the shared feature data) with the key of another modality (The second key vector corresponding to the shared structural feature data) performs a scaled dot product operation (Scaled Dot-Product Attention), that is, calculates the matrix product And divided by a scaling factor (usually the square root of the key vector dimension) to prevent the gradient from being too small, and then passed through the Softmax function to obtain the attention score (the second attention matrix). This score is then used to adjust the corresponding value (second value vector) is weighted summed. This process is carried out in parallel in all attention heads, and the results are concatenated and transformed again through a linear layer to obtain a functional feature representation (first temporary feature data or second temporary feature data) containing structural modal context information. The same operation is performed in the second direction, that is, the first attention matrix is ​​obtained by sharing the first key vector corresponding to the functional feature data and the second query vector corresponding to the structural feature data. The first attention matrix is ​​then compared with the corresponding value vector. (first value vector) is weightedly summed to finally obtain the corresponding structural feature representation containing the structural modal context information.

[0061] The attention weight matrices (first attention matrix and second attention matrix) generated by the attention calculations in these two directions are then symmetrically fused (for example, taking the average of the weights in the two directions) and finally processed by row-wise Softmax normalization to generate a dynamic fusion graph adjacency matrix tailored for the current subject. .

[0062] In addition, to further enhance the learning of inter-modal representations, e.g. Figure 4 As shown, the output of the attention mechanism (the first temporary feature data or the second temporary feature data) is further combined with the original vector (the corresponding shared functional feature data or shared structural functional data) through a residual connection. After subsequent layer normalization and a feedforward network (composed of linear layers), it is used for a cross-modal feature generation task (the corresponding cross-modal feature data). This task is supervised by an auxiliary generation loss (such as the mean squared error loss), which encourages the model to learn more meaningful and generalizable cross-modal shared representations. The second temporary feature data is obtained based on the second attention matrix and the second value vector.

[0063] The cross-attention module captures dynamic interactions between modalities, generating an adjacency matrix containing task-aware information that truly reflects the dynamic characteristics of brain networks. The dynamic graph adjacency matrix adapts to dynamic connectivity changes between brain regions, avoiding the limitations of fixed modality relationships. Independent query and key / value projection mechanisms enhance the model's ability to represent features from different modalities. Generative tasks and residual connections strengthen feature learning, improving robustness and generalization, and providing reliable feature representations for subsequent classification.

[0064] Optionally, the dynamic graph attention module includes a first GAT network, a second GAT network, a third GAT network, and a fourth GAT network; the dynamic graph attention module processes the dynamic fusion graph adjacency matrix, the functional features, and the structural features to obtain the first dynamic graph and the second dynamic graph, including: Processing the dynamic fusion graph adjacency matrix and the functional features through the first GAT network to obtain a first intermediate feature; Processing the dynamic fusion graph adjacency matrix through the second GAT network to obtain temporary functional features; Processing the dynamic fusion graph adjacency matrix and the structural features through the third GAT network to obtain a second intermediate feature; Processing the dynamic fusion graph adjacency matrix through the fourth GAT network to obtain temporary structural features; The first intermediate feature and the temporary functional feature are fused to obtain the first dynamic graph, and the second intermediate feature and the temporary structural feature are fused to obtain the second dynamic graph.

[0065] Specifically, the dynamic graph attention module uses four graph attention networks (GATs) to process the dynamic fusion graph adjacency matrix as well as functional features and structural features to generate two dynamic graphs.

[0066] First GAT network: This network processes the dynamic fusion graph adjacency matrix and functional features to generate the first intermediate features. These features contain functional feature information weighted by the graph attention mechanism.

[0067] Second GAT network: This network processes the dynamic fusion graph adjacency matrix separately and generates temporary functional features. This helps to extract functional feature information based only on the graph structure.

[0068] Third GAT network: This network processes the dynamic fusion graph adjacency matrix and structural features to generate the second intermediate features. These features contain structural feature information weighted by the graph attention mechanism.

[0069] Fourth GAT network: This network processes the dynamic fusion graph adjacency matrix separately and generates temporary structural features. This helps to extract structural feature information based only on the graph structure.

[0070] Feature fusion: The first intermediate feature is fused with the temporary functional feature to generate a first dynamic graph. This dynamic graph integrates both functional features and graph structure information. Similarly, the second intermediate feature is fused with the temporary structural feature to generate a second dynamic graph, which integrates both structural features and graph structure information.

[0071] In some embodiments, such as Figure 2 As shown, the dynamic graph adjacency matrix generated in the previous step As the underlying topology of the Graph Attention Network (GAT), it defines the neighborhood relationship for information transfer. Next, the functional feature data separated from the feature decoupling module and structural feature data , respectively, as the initial node features of two independent GAT streams. In each GAT network, information aggregation is iteratively performed through a multi-layer GAT architecture containing residual connections. GAT uses its self-attention mechanism to aggregate neighbors (by When analyzing information (defined by the model), different weights are assigned to different neighbors. These weights are learned end-to-end, enabling the model to amplify important neighbor signals and suppress noise. This refines and contextualizes modality-specific features, generating a high-level, context-aware dynamic graph representation for subsequent decision-making.

[0072] By processing multiple GAT networks, the model is able to more fully extract information from the graph structure, thereby enhancing the expressive power of features. The fusion of intermediate and temporary features integrates information from different sources, generating a richer dynamic graph representation. Generating two dynamic graphs corresponding to functional and structural features, respectively, helps distinguish and understand the role of different modalities in brain networks. This more accurately captures the dynamic characteristics of brain networks and improves the performance of subsequent classification tasks. Furthermore, because the dynamic graphs represent functional and structural features separately, this increases the interpretability of the model's decision-making process.

[0073] Optionally, the fusion module includes a feedforward neural network; and the fusing of the first dynamic image and the second dynamic image by the fusion module to obtain fused feature data includes: Splicing the first dynamic image and the second dynamic image to obtain splicing vector feature data; Processing the splicing vector feature data through the feedforward neural network to obtain weight vector data; Fused feature data is obtained according to the first dynamic graph, the second dynamic graph, and the weight vector data.

[0074] Specifically, if Figure 5 The input of the fusion module is the first dynamic graph and the second dynamic graph obtained by the dynamic graph attention module. First, the graph is read out (such as global average pooling) and aggregated into a graph-level vector. and The two graph-level vectors are concatenated by the connection layer. The concatenated high-dimensional vector is fed into a small feedforward neural network (represented by a linear layer in the figure). The network acts as a weight generator and learns a two-dimensional vector (weight vector data) that represents the relative importance of the two modalities in the current sample. This vector is normalized by a Softmax function to generate two scalar values ​​between 0 and 1 whose sum is 1. , or attention weights or modality contribution scores. Ultimately, these two weights are used to perform a weighted sum of the two graph-level vectors of the original input, yielding a final, highly information-rich fused feature vector (fused feature data). This weighting mechanism is dynamic and sample-specific. This fused vector is then fed into the final classifier for diagnosis. Crucially, the attention weights output by this module are themselves a manifestation of interpretability, clearly quantifying the model's reliance on information from different modalities when diagnosing a specific sample.

[0075] By splicing feature data from the first and second dynamic graphs, the fusion module integrates information from different modalities to generate a comprehensive feature representation. This helps capture more comprehensive brain network characteristics and enhances the model's understanding of complex brain network states. Regarding weight learning, the feedforward neural network learns to generate weight vector data, automatically determining the relative importance of each modality in the current sample. This allows the model to dynamically adjust its reliance on information from different modalities, improving its adaptability and accuracy. Regarding feature enrichment, the resulting fused feature data is a highly condensed representation of information, combining key features from both dynamic graphs. This reduces data dimensionality while enhancing feature expressiveness, which is beneficial for improving classifier performance. Regarding interpretability, the output attention weights provide interpretability of the model's decision-making process. By quantifying the model's reliance on information from different modalities, users can more clearly understand the model's diagnostic basis for specific samples, enhancing the model's transparency and credibility. Regarding dynamic adaptability, the fusion process is dynamic and sample-specific, allowing the model to adjust its fusion strategy based on varying input data, making it more flexible and robust when processing diverse brain network data. In general, the fusion module significantly improves the performance and reliability of the model in brain network state classification tasks by integrating multimodal information, learning weights, concentrating features, and providing interpretability.

[0076] An embodiment of the present invention provides a multimodal brain network classification device based on feature decoupling and dynamic graph construction, comprising: An acquisition unit is used to acquire functional magnetic resonance imaging data and structural magnetic resonance imaging data of the person to be tested, a construction unit, configured to obtain multimodal brain network data based on the functional magnetic resonance imaging data and the structural magnetic resonance imaging data, wherein the multimodal brain network data includes at least one functional connectivity network matrix and one structural network matrix; A processing unit is used to process the multimodal brain network data based on a preset multimodal brain network classification model to obtain the brain network state of the person to be tested; wherein the multimodal brain network classification model includes a feature decoupling module and a dynamic graph attention module constructed based on the GAT network.

[0077] An embodiment of the present invention provides a multimodal brain network classification device based on feature decoupling and dynamic graph construction, comprising a memory and a processor; the memory is used to store a computer program; the processor is used to implement the multimodal brain network classification method based on feature decoupling and dynamic graph construction as described above when executing the computer program.

[0078] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the multimodal brain network classification method based on feature decoupling and dynamic graph construction as described above is implemented.

[0079] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the protection scope of the present invention.

Claims

1. A multimodal brain network classification method based on feature decoupling and dynamic graph construction, characterized by: include: Acquire functional magnetic resonance imaging data and structural magnetic resonance imaging data of the person to be tested, Obtaining multimodal brain network data based on the functional magnetic resonance imaging data and the structural magnetic resonance imaging data, wherein the multimodal brain network data includes at least one functional connectivity network matrix and a structural network matrix; Based on a preset multimodal brain network classification model, the multimodal brain network data is processed to obtain the brain network state of the person to be tested; Among them, the multimodal brain network classification model includes a feature decoupling module and a dynamic graph attention module constructed based on the GAT network.

2. The multimodal brain network classification method based on feature decoupling and dynamic graph construction according to claim 1 is characterized in that: The multimodal brain network classification model further includes a cross-attention module and a fusion module. The multimodal brain network classification model based on the preset multimodal brain network is used to process the multimodal brain network data to obtain the brain network state of the person to be tested, including: Processing the functional connectivity network matrix and the structural network matrix respectively by the feature decoupling module to obtain functional feature data, shared functional feature data, structural feature data and shared structural feature data; Processing the shared functional feature data and the shared structural feature data through the cross-attention module to obtain a dynamic fusion graph adjacency matrix; Processing the dynamic fusion graph adjacency matrix, the functional feature data, and the structural feature data through the dynamic graph attention module to obtain a first dynamic graph and a second dynamic graph; The first dynamic image and the second dynamic image are fused by the fusion module to obtain fusion feature data, and the fusion features are classified to obtain the brain network state of the person to be tested.

3. The multimodal brain network classification method based on feature decoupling and dynamic graph construction according to claim 2 is characterized in that: The feature decoupling module includes a first encoding module and a second encoding module; the feature decoupling module processes the functional connectivity network matrix and the structural network matrix respectively to obtain functional feature data, shared functional feature data, structural feature data, and shared structural feature data, including: Processing the functional connectivity network matrix through the first encoding module to obtain the functional feature data and the shared functional feature data; The structural network matrix is ​​processed by the second encoding module to obtain the structural feature data and the shared structural feature data.

4. The multimodal brain network classification method based on feature decoupling and dynamic graph construction according to claim 3 is characterized in that: The first encoding module includes a feature extraction layer, a first linear projection space and a second linear projection space; The processing of the functional connectivity network matrix by the first encoding module to obtain the functional feature data and the shared functional feature data includes: Performing feature extraction on the functional connection network matrix through the feature extraction layer to obtain intermediate feature data; The intermediate feature data is mapped using the first linear projection space to obtain the functional feature data, and the intermediate feature data is mapped using the second linear projection space to obtain the shared functional feature data.

5. The multimodal brain network classification method based on feature decoupling and dynamic graph construction according to claim 2 is characterized in that: The step of processing the shared functional feature data and the shared structural feature data by the cross-attention module to obtain a dynamic fusion graph adjacency matrix includes: Converting the shared functional feature data and the shared structural feature data respectively to obtain a first vector matrix and a second vector matrix, wherein the first vector matrix includes a first query vector and a first key vector, and the second vector matrix includes a second query vector and a second key vector; Obtaining a second attention matrix based on the first query vector and the second key vector, and obtaining a first attention matrix based on the first key vector and the second query vector; The first attention matrix and the second attention matrix are fused to obtain the dynamic fusion graph adjacency matrix.

6. The multimodal brain network classification method based on feature decoupling and dynamic graph construction according to claim 2 is characterized in that: The dynamic graph attention module includes a first GAT network, a second GAT network, a third GAT network and a fourth GAT network; The step of processing the dynamic fusion graph adjacency matrix, the functional features, and the structural features by the dynamic graph attention module to obtain a first dynamic graph and a second dynamic graph includes: Processing the dynamic fusion graph adjacency matrix and the functional features through the first GAT network to obtain a first intermediate feature; Processing the dynamic fusion graph adjacency matrix through the second GAT network to obtain temporary functional features; Processing the dynamic fusion graph adjacency matrix and the structural features through the third GAT network to obtain a second intermediate feature; Processing the dynamic fusion graph adjacency matrix through the fourth GAT network to obtain temporary structural features; The first intermediate feature and the temporary functional feature are fused to obtain the first dynamic graph, and the second intermediate feature and the temporary structural feature are fused to obtain the second dynamic graph.

7. The multimodal brain network classification method based on feature decoupling and dynamic graph construction according to claim 2 is characterized in that: The fusion module includes a feedforward neural network; The fusing the first dynamic image and the second dynamic image by the fusion module to obtain fusion feature data includes: Splicing the first dynamic image and the second dynamic image to obtain splicing vector feature data; Processing the splicing vector feature data through the feedforward neural network to obtain weight vector data; Fused feature data is obtained according to the first dynamic graph, the second dynamic graph, and the weight vector data.

8. The multimodal brain network classification method based on feature decoupling and dynamic graph construction according to claim 5 is characterized in that: The first vector matrix includes a first value vector, and the second vector matrix includes a second value vector; the processing of the shared functional feature data and the shared structural feature data by the cross-attention module to obtain the dynamic fusion graph adjacency matrix also includes: Obtaining first temporary feature data according to the first attention matrix and the first value vector; Obtaining cross-modal feature data according to the first temporary feature data and the shared functional feature data; The cross-modal feature data is used to enhance the auxiliary output of the multimodal brain network classification model training.

9. A multimodal brain network classification device based on feature decoupling and dynamic graph construction, characterized in that: include: An acquisition unit is used to acquire functional magnetic resonance imaging data and structural magnetic resonance imaging data of the person to be tested, a construction unit, configured to obtain multimodal brain network data based on the functional magnetic resonance imaging data and the structural magnetic resonance imaging data, wherein the multimodal brain network data includes at least one functional connectivity network matrix and one structural network matrix; A processing unit is used to process the multimodal brain network data based on a preset multimodal brain network classification model to obtain the brain network state of the person to be tested; wherein the multimodal brain network classification model includes a feature decoupling module and a dynamic graph attention module constructed based on the GAT network.

10. A multimodal brain network classification device based on feature decoupling and dynamic graph construction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the multimodal brain network classification method based on feature decoupling and dynamic graph construction according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal fusion method and system for brain magnetic resonance image data and medium

    CN118823541A

  • Parkinson subtype diagnosis model and device based on balance multi-mode brain network fusion and computer readable storage medium

    CN120411634A

Cited By

  • Work memory ability assessment method and device for dynamic brain network attention fusion

    CN120918653A