Counterfactual causal cascade graph network method and device for alzheimer's disease

CN122781584APending Publication Date: 2026-09-18SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610801612.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

这类方法主要衡量输入特征对预测结果的局部敏感度,本质上是在捕捉表面相关性,且极易受到“梯度噪声”干扰

Benefits of technology

通过分层因果图卷积网络,在拓扑结构层面,围绕基因→结构→功能的层级依赖关系,对遗传模态、结构模态与功能模态之间的有向依赖关系进行统一建模并提出了层级约束的邻接矩阵;在特征交互层面,设计分层因果级联注意力模块,通过参数化控制跨模态的信息传递过程,显式编码了自底向上的有向依赖结构,且拓扑约束与级联注意力相互配合,在结构连接与信息传播两个维度上共同刻画层级依赖关系;在解释机制方面,提出分阶段反事实积分梯度推理框架,在潜在表示空间中构建从正常认知到阿尔茨海默病的连续病理原型,并沿层级路径进行插值积分,从而重构从正常老化(NC)到阿尔茨海默病(AD)的疾病演变轨迹,从而量化不同基因及其下游结构与功能变化在疾病不同阶段(NC、EMCI、LMCI、AD)中的动态边际贡献,为AD的机制分析提供了全新的视角。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122781584A_ABST
    Figure CN122781584A_ABST
Patent Text Reader

Abstract

The application discloses an anti-fact causality cascade graph network method and device for Alzheimer's disease, relates to the technical field of neural network application, and comprises the following steps: acquiring multi-modal data of samples in an Alzheimer's disease related database, and performing preprocessing to construct a data set; creating a hierarchical causality graph convolution network, and performing gene-brain region heterogeneous graph construction, brain region position coding, hierarchical causality cascade and local information aggregation according to the constructed data set; and combining a staged anti-fact integral gradient reasoning framework to realize dynamic analysis on the evolution track of Alzheimer's disease. The application can better integrate information of different modes and provide more reasonable explainability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of neural network application technology, specifically to a counterfactual causal cascade graph network method and apparatus for Alzheimer's disease. Background Technology

[0002] Alzheimer's disease (AD) is a typical complex disease whose pathological process involves the long-term interaction of multiple scale factors, including genetic variations, brain anatomical degeneration, and brain functional network remodeling. With the rapid development of high-throughput gene sequencing and neuroimaging technologies, multimodal data such as single nucleotide polymorphisms (SNPs), structural magnetic resonance imaging (sMRI), and functional magnetic resonance imaging (rs-fMRI) can be acquired simultaneously at the individual level. This provides an important data foundation for characterizing the disease evolution process at the systemic level. Against this backdrop, how to effectively integrate information from different modalities within a unified computational framework and decode biologically consistent discriminative representations has become one of the core research problems in computational neuroscience and medical image analysis.

[0003] In recent years, a large number of studies have focused on shifting from single-modal analysis to deep multimodal collaborative modeling to improve the diagnostic performance of Alzheimer's disease (AD). Current mainstream methods can be summarized into three categories: (1) cross-modal fusion networks (such as MMANet and GM-ABC-NN), which use residual connections or hybrid models to capture statistical associations between modalities; (2) spatiotemporal graph neural networks (such as SF-GCL based on contrastive learning and BGAN with brain topology enhancement), which aim to explore local spatial dependencies and global connectivity patterns in brain regions; and (3) multimodal Transformer variants (such as ViT-TST, Aux-ViT, and EnCBAMUNet which introduces attention gating mechanisms), which use self-attention mechanisms to capture long-range dependencies. Although the above methods have made significant progress in classification accuracy, most of them follow a parallel fusion architecture at the methodological level. These methods usually treat gene, structural, and functional data as equal input streams and mix them through feature splicing or symmetrical cross-attention. Although this modeling approach can effectively capture statistical co-occurrence between modalities, it ignores the cascade effects between modalities in biology, making it difficult to provide a structured explanation of the interaction mechanisms between modalities.

[0004] Furthermore, the interpretability of existing models also faces challenges in clinical applications. Current mainstream interpretative methods mainly include gradient-based attribution methods, game-theoretic contribution methods, and posterior analysis based on discrete perturbations. Gradient-based local attribution, such as Grad-CAM and LRP, primarily measures the local sensitivity of input features to the prediction result, essentially capturing surface correlations and is highly susceptible to "gradient noise." Game-theoretic contribution measures, such as SHAP and DeepLIFT, while quantifying the contribution of a single feature, assume that features are independent or symmetrical. In the hierarchical pathology of AD, this method forcibly decomposes the transmission effects across modalities into isolated scores, losing the continuity of causal logic. Perturbation-based analysis, such as feature ablation, observes feedback by "erasing" features. This non-physical, discrete perturbation cannot simulate the continuous evolutionary trajectory of biological tissues during pathological processes, nor can it support counterfactual reasoning. In summary, these explanations fail to answer key questions with counterfactual semantics within the current cascade framework, such as assessing whether intervention in a particular gene feature alters functional networks in a specific way, thereby alleviating cognitive decline. Therefore, this kind of explanation, which only stays at the correlation level, makes it difficult to establish a conclusive link between the biomarkers identified by the model and the actual pathological mechanisms, limiting its credibility in clinical decision support. Summary of the Invention

[0005] This application provides a counterfactual causal cascade graph network method and apparatus for Alzheimer's disease, which can better integrate information from different modalities and provide more reasonable interpretability.

[0006] In a first aspect, embodiments of this application provide a counterfactual causal cascade graph network method for Alzheimer's disease, the counterfactual causal cascade graph network method for Alzheimer's disease comprising: Multimodal data of samples from Alzheimer's disease-related databases were obtained and preprocessed to construct the dataset; A hierarchical causal graph convolutional network is created, and based on the constructed dataset, gene-brain region heterogeneity map construction, brain region location encoding, hierarchical causal cascade, and local information aggregation are performed. Combined with a staged counterfactual integral gradient inference framework, dynamic analysis of the evolution trajectory of Alzheimer's disease is achieved.

[0007] In conjunction with the first aspect, in one implementation, the samples include normal controls, mild cognitive impairment, and confirmed cases, and the mild cognitive impairment includes early and late stages, and the samples are distributed according to a preset ratio.

[0008] In conjunction with the first aspect, in one implementation method, The preprocessing includes candidate gene screening, genetic data quality control and feature selection, structural image data processing, and resting-state functional magnetic resonance imaging data processing. The candidate gene screening was based on SNP sequence length and PPI network topology to screen for a target number of candidate genes, and the biological significance of the screened candidate genes was confirmed by GO function and KEGG pathway enrichment analysis performed by Metascape. The genetic data quality control and feature selection were performed using PLINK and MAGMA software, including deletion rate control, minimum allele frequency control, Hardy-Weinberg balance test, linkage disequilibrium pruning, SNP-to-gene mapping, PPI network-based gene screening, and gene encoding. The structural image data processing employs the FSL software library in conjunction with an active verification method to quantify the gray matter volume shrinkage in each brain region. Specifically, this includes non-brain tissue removal, tissue segmentation and registration, spatial registration, quality control, ROI coordinate calculation, and hybrid feature extraction. The resting-state functional magnetic resonance imaging data processing is based on the DPARSF toolbox, specifically including format conversion, initial time point deletion, time layer correction, head motion correction, spatial normalization, spatial smoothing, temporal filtering, covariate removal, and brain region time series extraction.

[0009] In conjunction with the first aspect, in one implementation method, The gene-brain region heterogeneity map is constructed based on biological priors, which includes gene, brain structure and brain function nodes. Causal connection constraints between modalities are established through a physical topological blocking strategy. For gene-brain region heterogeneity mapping, the specific components include: The rs-fMRI time series were cropped to align their length with the gene coding sequence. Constructing a gene-brain region heterogeneity map with causal constraints , Represents a block adjacency matrix. Let the node feature matrix be denoted as: ,

[0010] in, The initial feature representation of the gene node. The initial feature representation of brain structural nodes. If we represent the initial features of brain functional nodes, then the total number of nodes is the sum of the number of gene nodes, brain structure nodes, and brain functional nodes. This represents the adjacency matrix between gene nodes. This represents the adjacency matrix between gene nodes and brain structure nodes. This represents the adjacency matrix between brain structural nodes and gene nodes. This represents the adjacency matrix between nodes in the brain structure. This represents the adjacency matrix between brain structural nodes and brain functional nodes. This represents the adjacency matrix between brain functional nodes and brain structural nodes. Represents the adjacency matrix between brain functional nodes; compute nodes and The Pearson correlation coefficient between the eigenvectors is used to measure the initial weights of the edges:

[0011] in, This represents the node dimension, i.e., the number of time points. Represents a node In the Feature values ​​at each time point Represents a node In the Feature values ​​at each time point; By setting a threshold The block adjacency matrix is ​​sparsified to obtain And perform normalization to obtain the causal constraint adjacency matrix of the causal constraint graph convolution module:

[0012] in, Represents the causal constraint adjacency matrix. This represents the degree matrix used to normalize the adjacency matrix. This represents the identity matrix used to add self-loops.

[0013] In conjunction with the first aspect, in one implementation method, The brain region location coding is used to integrate spatial location information in the brain and encoded sequence location information based on a learnable location coding method; For brain region location coding, the specific components include: Constructing an absolute location encoding matrix based on brain region indexing Then for the first The brain region, in the _ ... and The encoding of the feature dimension is defined as follows:

[0014]

[0015] in, , Indicates the first The brain region in the first Encoding of feature dimensions Indicates the first The brain region in the first Encoding of feature dimensions; Introducing learnable position embedding vectors and Embedded vectors The initial value is determined by the corresponding position encoding matrix. The position-enhanced node features are obtained by element-wise addition: , , This indicates the location-enhanced features of brain structural nodes. This indicates the features of brain functional nodes after location enhancement, and... and As input to the hierarchical causal cascade attention module in the hierarchical causal cascade.

[0016] In conjunction with the first aspect, in one implementation method, The hierarchical causal cascade is used to explicitly model the hierarchical driving relationship of genes, structures, and functions by connecting attention streams, thereby achieving feature enhancement with causal constraints. For hierarchical causal cascades, the specific components include: Create a hierarchical causal cascade attention module containing a two-stage serial attention flow; The hierarchical causal cascade attention module performs gene-to-structure risk injection to mimic the unidirectional driving effect of genetic variation on brain anatomy. Mapped as query terms to find pathogenic genes associated with structural abnormalities in specific brain regions. Mapped to keys and values; Define projection transformation , , Structural representations regulated by genetic risk are computed using scaled dot product attention. :

[0017] in, This represents a query matrix derived from features of brain structural nodes. This represents the learnable projective weight matrix of the query. The key matrix represents the features derived from gene nodes. The learnable projective weight matrix represents the key matrix. The Value matrix represents the features derived from gene nodes. The learnable projective weight matrix represents the value matrix. This represents the layer normalization used for stable training and attention output. This represents the Softmax activation function. Indicates the dimension of the Key matrix; Use functional features as query terms ,Will As keys and values , Generate structurally constrained functional representations through cascaded attention. :

[0018] in, Represents the query vector. Represents the key vector. Represents a value vector; Will , , By concatenating the features, a full-modal node feature matrix is ​​obtained. , which serves as the input to the causal constraint graph convolution module.

[0019] In conjunction with the first aspect, in one implementation method, The local information aggregation is used to aggregate the local topological information of heterogeneous nodes in the causal neighborhood using the output of the hierarchical causal cascaded attention module as the initial feature, and input the learned gene and brain region representations into the fully connected network. For local information aggregation, the specific components include: The output of the hierarchical causal cascade attention module , , The nodes are concatenated to construct the initial node feature matrix of the full modality. ; Based on causal constraint adjacency matrix , define the first The graph convolution propagation rule for the layer is:

[0020] in, Represents a non-linear activation function. Indicates the first The node feature matrix of the layer, This represents the learnable weight matrix of the current layer. Indicates the first The node feature matrix of the layer; Will pass The node representation obtained after stacking layers of causal convolutions The data is then fed into a multilayer perceptron or a fully connected layer for final classification prediction, assuming... For the sample The true label, To determine the probability distribution predicted by the hierarchical causal graph convolutional network, the cross-entropy loss function is used as the optimization objective for training the network.

[0021] in, This indicates the number of layers in a graph convolutional network. This represents the total number of nodes in the graph convolutional network.

[0022] In conjunction with the first aspect, in one implementation, the phased counterfactual integral gradient inference framework is used to construct counterfactual prototypes of different pathological stages, perform integral gradient attribution along the evolution trajectory, and identify key pathogenic factors driving disease progression.

[0023] In conjunction with the first aspect, in one implementation method, the specific application process of the phased counterfactual integral gradient inference framework includes: For each stage of evolution in the evolutionary trajectory Calculate from baseline state Migrate to target state During the process, the first Features Predicting the probability of Alzheimer's disease The cumulative contribution, the phased counterfactual integral gradient inference framework accumulates the gradient along the linear interpolation path, and the calculation formula is:

[0024] in, Indicates the first Features In the stage Attribution value, Indicates the first state under the target condition Features The value, Indicates the first state under the baseline condition Features The value, This represents the interpolation parameter with a value range of [0,1]. Using Riemann approximation By multiplying the characteristic changes with the average gradient along the path, the risk of Alzheimer's disease is decomposed into driving factors in three consecutive stages, identifying specific biomarkers for the onset, progression, and deterioration stages. Based on the attribution tensors calculated at each stage, the phased counterfactual integral gradient inference framework performs independent contribution aggregation and ranking for different modalities.

[0025] Secondly, embodiments of this application provide a counterfactual causal cascade graph network device for Alzheimer's disease, the counterfactual causal cascade graph network device for Alzheimer's disease comprising: The module is used to acquire multimodal data of samples from Alzheimer's disease-related databases and preprocess them to build the dataset. The execution module is used to create a hierarchical causal graph convolutional network and, based on the constructed dataset, to construct a gene-brain region heterogeneous map, encode brain region locations, perform hierarchical causal cascades, aggregate local information, and combine a staged counterfactual integral gradient inference framework to achieve dynamic analysis of the evolution trajectory of Alzheimer's disease.

[0026] The beneficial effects of the technical solutions provided in this application include: By employing a hierarchical causal graph convolutional network, at the topological level, a unified model of directed dependencies between genetic, structural, and functional modalities is developed, based on the hierarchical dependency relationship of gene → structure → function. A hierarchical adjacency matrix with constraints is proposed. At the feature interaction level, a hierarchical causal cascaded attention module is designed. By parameterizing and controlling the cross-modal information transmission process, the bottom-up directed dependency structure is explicitly encoded. Topological constraints and cascaded attention work together to characterize hierarchical dependencies in both structural connectivity and information propagation dimensions. Regarding the explanatory mechanism, a staged counterfactual integral gradient inference framework is proposed. This framework constructs a continuous pathological prototype from normal cognition to Alzheimer's disease in the latent representation space and performs interpolation integration along the hierarchical path, thereby reconstructing the disease evolution trajectory from normal aging (NC) to Alzheimer's disease (AD). This quantifies the dynamic marginal contribution of different genes and their downstream structural and functional changes in different disease stages (NC, EMCI, LMCI, AD), providing a novel perspective for the mechanistic analysis of AD. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating the counterfactual causal cascade graph network method for Alzheimer's disease proposed in this application; Figure 2 Here are schematic diagrams of the structures of HC-GCN and SCIG; Figure 3 This is a schematic diagram of key gene nodes; Figure 4This is a schematic diagram of key brain region nodes; Figure 5 A schematic diagram of the early stage (NC→EMCI); Figure 6 This is a schematic diagram of the intermediate stage (EMCI→LMCI); Figure 7 A schematic diagram of the late stage (LMCI→AD); Figure 8 This is a schematic diagram of the functional modules of the counterfactual causal cascade network device for Alzheimer's disease proposed in this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] Firstly, this application provides a counterfactual causal cascade graph network method for Alzheimer's disease, which can better integrate information from different modalities and provide more reasonable interpretability. Specifically, this application proposes a hierarchical causal graph convolutional network (HC-GCN). The core innovation of this framework lies in: at the topological level, it uniformly models the directed dependencies between genetic modalities, structural modalities, and functional modalities around the hierarchical dependency relationship of "gene → structure → function" and proposes a hierarchical constraint adjacency matrix; at the feature interaction level, it designs a hierarchical causal cascade attention module (HCCA), which explicitly encodes bottom-up information transmission by parameterizing the cross-modal information transmission process. The directed dependency structure, with topological constraints and cascaded attention working together, jointly characterizes hierarchical dependencies in both structural connectivity and information propagation dimensions. In terms of explanatory mechanisms, this application further proposes a phased counterfactual integral gradient inference framework (SCIG), which constructs a continuous pathological prototype from normal cognition to Alzheimer's disease in the latent representation space and performs interpolation integration along the hierarchical path to reconstruct the disease evolution trajectory from normal aging (NC) to Alzheimer's disease (AD). This allows for the quantification of the dynamic marginal contribution of different genes and their downstream structural and functional changes in different stages of the disease (NC, EMCI, LMCI, AD), providing a new perspective for the mechanism analysis of AD.

[0030] First, it's important to note that the core challenge in AD diagnosis lies in effectively mapping and integrating heterogeneous modal features. Early attempts primarily relied on statistical projections, such as Multi-Kernel Learning (MKL) and Canonical Correlation Analysis (CCA). These methods sought to find the largest common correlation subspace between modalities for linear fusion, but they exhibited limitations when dealing with high-dimensional, nonlinear pathological features. With the rise of deep learning, feature concatenation strategies have gradually become mainstream. For example, MMANet uses multi-stream CNNs to extract features separately and then directly concatenates them in fully connected layers, while DPSCNN utilizes a Siamese network architecture to implicitly align modal distributions through shared weights. In recent years, attention mechanisms have been widely introduced to refine the interaction strength between features. For instance, ViT-TST and MDISPN utilize the self-attention or cross-attention mechanisms of Transformers to calculate the correlation matrix between modalities, achieving soft-weighted feature fusion. Although the aforementioned models introduce complex interaction units, their architectural design largely maps SNPs, sMRI, and fMRI to computationally equivalent latent vectors. This homogeneous vector concatenation or symmetric attention computation mathematically lacks the ability to model directed dependencies, causing the models to tend to capture high-frequency statistical co-occurrences rather than causal transmission paths. This application's HC-GCN introduces an asymmetric biological mechanism at the architectural level to break this computational equivalence.

[0031] Graph Neural Networks (GNNs) aggregate non-Euclidean data through message passing mechanisms, demonstrating excellent performance in brain network analysis. Existing graph construction strategies are mainly divided into two categories: data-driven and knowledge-driven. Data-driven methods, such as MCA-GCN and SF-GCL, typically construct adjacency matrices based on feature similarity between samples (e.g., cosine similarity or Pearson correlation coefficient). HSIA-GAN, proposed by Bi et al., goes further, utilizing hypergraphs to construct vertex-edge dual structures to capture higher-order associations. However, these statistical similarity-based graph construction methods often produce dense and noisy edges, easily leading to oversmoothing during GNN training and causing node features to become homogenized. Knowledge-driven methods rely on prior information (such as protein-protein interaction networks (PPIs) or predefined brain maps) to determine the graph's topology. While these methods have good biological interpretability, their graph structures are usually static and fixed, making it difficult to capture specific pathological changes at the individual level. More importantly, the graphs constructed by these two mainstream methods are usually undirected, meaning information can flow bidirectionally between nodes of any modality. This paper proposes a sparse directed graph construction strategy based on biological priors, which is actually a fusion approach: it utilizes prior knowledge (knowledge-driven) to physically block unreasonable cross-layer connections while retaining dynamic interactions at the feature level, thereby strictly constraining the flow of message transmission in the topological structure.

[0032] The interpretability research of medical AI is evolving from feature attribution to causal reasoning. As the mainstream paradigm, feature attribution methods widely employ gradient-based backpropagation techniques. For example, Lian et al. used Grad-CAM to visualize activation heatmaps of convolutional layers, and other research applications such as LRP or IG decompose predicted scores into the contribution values ​​of input features. The core of these methods lies in calculating the partial derivatives of input features with respect to the output, essentially still belonging to correlation analysis. To overcome the limitations of gradients, recent research has begun to shift towards counterfactual explanations. Oh et al. (TPAMI 2023) proposed the LEAR framework, which synthesizes "counterfactual images" through generative adversarial networks (GANs) to intuitively demonstrate the morphological changes from NC to AD. However, existing feature attribution methods typically assume that features are independent and identically distributed, ignoring the hierarchical influence between features; while existing counterfactual methods (such as LEAR) mainly focus on the generation of visual samples, making it difficult to analyze the evolutionary paths of multimodal numerical features (such as SNP sites). The SCIG proposed in this application combines the advantages of both methods, retaining the quantitative analysis capability of integral gradients while introducing counterfactual path interpolation, thereby solving the technical problem of continuous evolution analysis of discrete gene data.

[0033] In one embodiment, reference is made to Figure 1 , Figure 1 This is a flowchart illustrating the counterfactual causal cascade graph network method for Alzheimer's disease proposed in this application. Figure 1 As shown, the counterfactual causal cascade graph network method for Alzheimer's disease includes: S1: Obtain multimodal data of samples from Alzheimer's disease-related databases and preprocess them to construct the dataset; S2: Create a hierarchical causal graph convolutional network (HC-GCN) and, based on the constructed dataset, construct gene-brain region heterogeneity maps, encode brain region locations, perform hierarchical causal cascades, and aggregate local information. Combined with a phased counterfactual integral gradient inference framework (SCIG), it enables dynamic analysis of the evolution trajectory of Alzheimer's disease.

[0034] It should be noted that the sample includes normal controls, mild cognitive impairment, and confirmed cases, and mild cognitive impairment includes early and late stages, and each sample is distributed according to a preset ratio.

[0035] Specifically, this application uses multimodal data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database, a long-term, multicenter longitudinal study designed to systematically track the pathological evolution of Alzheimer's disease. This application includes three groups of participants: cognitively normal controls (NC), mild cognitive impairment (MCI), and patients diagnosed with Alzheimer's disease. MCI is further subdivided into early EMCI and late LMCI according to ADNI criteria.

[0036] The initial sample consisted of 869 subjects. After excluding subjects with excessive head movements (translation >2.0 mm or rotation >2.0°), 817 valid samples were retained, including 213 NC, 190 EMCI, 192 LMCI, and 222 AD. Data sources covered ADNI-GO, ADNI-2, ADNI-3, and ADNI-DOD stages. Sample information is shown in Table 1 below, including which ADNI centers and their numbers were used for each of the four sample types.

[0037] Table 1

[0038] Furthermore, in one embodiment, preprocessing includes candidate gene screening, genetic data quality control and feature selection, structural image data processing, and resting-state functional magnetic resonance imaging (fMRI) data processing. Specifically, since the HC-GCN model uses SNP, sMRI, and fMRI data as input, relevant data processing is required to ensure accuracy.

[0039] The candidate gene screening was based on SNP sequence length and PPI network topology to select candidate genes with a target number of genes. The biological significance of the selected candidate genes was confirmed by GO (gene ontology) function and KEGG pathway enrichment analysis using Metascape (a gene / protein list functional enrichment and systematic analysis tool).

[0040] Specifically, the HC-GCN model comprehensively considered SNP sequence length and PPI network topology, ultimately screening 42 candidate genes. Subsequently, GO function and KEGG pathway enrichment analysis using Metascape (Zhou et al., 2019) confirmed the biological significance of these genes. The actual results showed that these candidate genes not only dominate key physiological processes such as axonal development, neuronal recognition, and transsynaptic signal transduction, but also deeply participate in core pathological mechanisms of AD such as postsynaptic signaling pathway blockade and hippocampal signal regulation.

[0041] Among them, genetic data quality control and feature selection were carried out using PLINK and MAGMA software, including deletion rate control, minimum allele frequency control, Hardy-Weinberg balance test, linkage disequilibrium pruning, SNP-to-gene mapping, gene screening based on PPI network, and gene encoding.

[0042] Specifically, PLINK and MAGMA software were used to complete the quality control and feature selection of genetic data, mainly including quality screening, gene mapping and network filtering, specifically including: (1) Missingness filtering: In order to avoid interference from high missing samples or sites to subsequent analysis, a loose (20%) and a strict (2%) threshold were used to filter the sample-level and SNP-level missing rates to ensure data integrity; (2) Minimum allele frequency (MAF) control: Too low allele frequency may lead to unstable statistical estimation and reduce the generalization ability of the model. In this application, the MAF threshold was set to 0.05 to exclude rare variant sites and retain only common variants with sufficient statistical support in the population; (3) Hardy-Weinberg balance (HWE) test: The HWE test is an important step in the quality control of genetic data and can be used to detect potential genotyping errors or systematic biases. In this application, the HWE test threshold was set to less than (3) SNP sites that significantly deviate from equilibrium are removed; (4) Linkage disequilibrium (LD) pruning: In order to reduce the redundant information and multiple comparison problems caused by highly correlated SNPs, a sliding window strategy is adopted for LD pruning. Specifically, 200 SNPs are used as the window size, and sites with strong linkage disequilibrium with other SNPs in the window are removed based on the pairwise correlation coefficient; (5) Mapping SNPs to genes: MAGMA (refer to the Human Genome Project build-37) is used to map the physical location of SNPs to the corresponding genes, thereby realizing the aggregated representation from the site level to the gene level; (6) Gene screening based on PPI network: In order to improve the biological relevance of genetic features and network consistency, further The protein-protein interaction (PPI) network of the STRING database was used to screen genes. Only genes that had interaction edges in all four types of PPI networks (database source, experimental source, co-expression source and text mining source) were retained. Finally, 42 core genes with dense network associations were screened. (7) Gene-level encoding: The SNP genotypes in each gene were numerically encoded: homozygous common genotypes, heterozygous genotypes and homozygous rare genotypes were encoded as 0, 1 and 2 respectively (e.g. AA→0, AC→1, CC→2). Finally, each gene was represented as a SNP encoding sequence of length 90 as the input feature of gene modality in the model.

[0043] The structural image data processing employs the FSL software library in conjunction with an active verification method to quantify the gray matter volume shrinkage in each brain region. Specifically, this includes non-brain tissue removal, tissue segmentation and registration, spatial registration, quality control, ROI (region of interest) coordinate calculation, and hybrid feature extraction.

[0044] Specifically, the FSL (FMRIB Software Library) software library was used to preprocess the structural image data using an automatic preprocessing method combined with manual review, aiming to quantify the shrinkage of gray matter volume (GMV) in each brain region. The processing flow is as follows: (1) Non-brain tissue removal: The BET tool was used to automatically extract brain tissue, and robustfov was used to process excessively long necks or abnormal structures. To ensure the quality of the results, interactive parameter adjustment was designed: For samples with a Dice coefficient lower than 0.85, the BET parameters were manually corrected and the craniotomy step was run again; (2) Tissue segmentation and registration: The FAST tool was used to segment the craniotomy brain image into three probability maps: gray matter (GM), white matter (WM), and cerebrospinal fluid (CSF) using a Markov random field model; (3) Spatial registration: First, the segmented gray matter image was highly registered to the MNI152 standard space using a nonlinear registration tool (FNIRT) to eliminate the differences in anatomical morphology between individuals. At the same time, the brain mask and gray matter image were also transformed in the same way; (4) Quality Control: The registration quality is quantified by the Dice coefficient, and the images are automatically classified into high quality (Dice≥0.90), medium quality (0.85≤Dice<0.90) and unqualified (Dice<0.85). Unqualified samples are kept in the Rejected folder for further manual inspection and repair; (5) ROI coordinate calculation: The voxel coordinates of each AALROI are indexed to avoid each subject from repeatedly scanning the entire image; (6) Hybrid feature extraction: A 90-dimensional vector is extracted for each ROI, which includes an 85-dimensional gray matter density histogram to reflect the shape of gray matter distribution, focusing on the intensity range of 0.05–1.0. The other 5-dimensional statistical matrix means, standard deviation, entropy, skewness and kurtosis are used to describe the overall structural characteristics of gray matter, thereby obtaining a 116-dimensional gray matter volume feature vector for each subject.

[0045] Among them, the resting-state functional magnetic resonance imaging data processing is based on the DPARSF toolbox, which specifically includes format conversion, initial time point deletion, time layer correction, head motion correction, spatial normalization, spatial smoothing, temporal filtering, covariate removal, and brain region time series extraction.

[0046] Specifically, the preprocessing of resting-state functional magnetic resonance imaging (rs-fMRI) data was completed using the DPARSF toolbox, following standard temporal and spatial correction procedures to ensure the reliability of functional signals and cross-subject consistency. The specific steps are summarized as follows: (1) Format conversion: Convert the original DICOM format rs-fMRI data into the NIFTI format that DPARSF can recognize; (2) Initial time point deletion: To eliminate noise caused by magnetic field instability in the early stage of scanning, delete the first 10 time points of each subject to ensure that subsequent analysis is based on stable BOLD signals; (3) Slicetiming correction: Due to the differences in acquisition time of different brain slices, the time series is corrected to align all slices at the same reference time point; (4) Head motion correction: Correct the image shift caused by the subject's head motion through rigid registration to reduce the impact of motion artifacts on functional signals; (5) Spatial normalization: Normalize all rs-fMRI data to the standard EPI model. (6) Spatial smoothing: Gaussian kernels are used for spatial smoothing, with smoothing parameters set to FWHM=[6,6,6]mm to improve the signal-to-noise ratio and enhance the spatial continuity of functional modes; (7) Temporal filtering: Bandpass filtering (0.01–0.08Hz) is applied to the time series to retain low-frequency oscillations related to neural activity and remove high-frequency noise and low-frequency drift; (8) Covariate removal: Regression is used to remove potential confounding factors such as global signals, white matter and cerebrospinal fluid signals to further improve the specificity of functional connectivity analysis; (9) Brain region time series extraction: The brain is divided into 116 brain regions using the AAL-116 template, and the average BOLD signal of voxels in each brain region is calculated to obtain regional functional time series as input for subsequent functional modality modeling.

[0047] Furthermore, in one embodiment, the overall architecture of the hierarchical causal graph convolutional network (HC-GCN) and the staged counterfactual integral gradient inference framework (SCIG) of this application is as follows: Figure 2 As shown, hierarchical causal graph convolutional networks include gene-brain region heterogeneity map construction, brain region location encoding, hierarchical causal cascade, and local information aggregation.

[0048] Gene-brain region heterogeneity mapping is constructed based on biological priors, creating a gene-brain region heterogeneity map containing genes, brain structures, and brain functional nodes. A physical topological blocking strategy is used to establish causal connection constraints between modalities. Brain region location encoding integrates spatial location information in the brain and encoded sequence location information using a learnable location encoding method. Hierarchical causal cascade explicitly models the hierarchical driving relationships of genes, structures, and functions through a concatenated attention flow, achieving feature enhancement with causal constraints. Local information aggregation uses the output of the hierarchical causal cascade attention module as initial features, employs a graph convolutional network to aggregate the local topological information of heterogeneous nodes within their causal neighborhoods, and inputs the learned gene and brain region representations into a fully connected network to achieve accurate Alzheimer's disease (AD) diagnosis. A phased counterfactual integral gradient inference framework constructs counterfactual prototypes for different pathological stages, performs integral gradient attribution along the evolutionary trajectory, and identifies key pathogenic factors driving disease progression.

[0049] like Figure 2 As shown, in AD diagnosis and prediction, a gene-brain region causal network is first constructed and learnable brain region location encoding is introduced. Then, hierarchical causal cascade attention is used to achieve two-stage information transfer from genes to function. Finally, local and global information are aggregated using a genetic network (GCN) to achieve AD classification. In interpreting the continuous pathological states of Alzheimer's disease, patient data features at each stage are first input. Then, a continuous three-stage model is constructed in the NC-AD task, and staged integral gradient interpolation calculations are performed accordingly. Finally, the marginal contribution of different modalities in specific disease stages is quantified.

[0050] Furthermore, in one embodiment, the construction of a gene-brain region heterogeneity map specifically includes: S201: The rs-fMRI time series is cropped to align its length with the gene coding sequence; This application is used to achieve accurate classification of Alzheimer's disease (AD) by fusing multimodal data SNP, sMRI and rs-fMRI, and to discover key biomarkers driving disease progression. In order to accurately reflect homeostatic neural activity and reduce the influence of time inertia effect, the rs-fMRI time series is cropped so that its length is aligned with the gene coding sequence, that is, the number of time points is uniform, and the preferred value of the number of time points is 90. S202: Constructing a gene-brain region heterogeneity map with causal constraints It is used to associate different modal features within a unified topological space. Represents a block adjacency matrix. The node feature matrix is ​​composed of vertically concatenated embedding features from three modalities, where: ,

[0051] in, The initial feature representation of the gene nodes is given; the number of gene nodes is 42. The initial feature representation of brain structural nodes is given; the number of brain structural nodes is 116. The initial feature representation of brain functional nodes is given. If the number of brain functional nodes is 116, then the total number of nodes is the sum of the number of gene nodes, brain structure nodes, and brain functional nodes. This represents the adjacency matrix between gene nodes. This represents the adjacency matrix between gene nodes and brain structure nodes. This represents the adjacency matrix between brain structural nodes and gene nodes. This represents the adjacency matrix between nodes in the brain structure. This represents the adjacency matrix between brain structural nodes and brain functional nodes. This represents the adjacency matrix between brain functional nodes and brain structural nodes. Represents the adjacency matrix between brain functional nodes; Unlike fully connected multimodal graphs, HC-GCN introduces a physical topological blocking strategy that conforms to biological priors to explicitly establish the directed dependencies of "gene-structure-function". The structure is shown above. In this structure, explicit settings are used. 0 and When the value is 0, this constraint forces the information flow to be mediated by structural nodes, blocking the direct edge between genes and functional modes; S203: This measures the initial association strength between nodes by calculating the node... and The Pearson correlation coefficient between the eigenvectors is used to measure the initial weights of the edges:

[0052] in, This represents the node dimension, i.e., the number of time points. Represents a node In the Feature values ​​at each time point Represents a node In the Feature values ​​at each time point; S204: By setting a threshold The block adjacency matrix is ​​sparsified to obtain And perform normalization to obtain the causal constraint adjacency matrix of the causal constraint graph convolution module:

[0053] in, Represents the causal constraint adjacency matrix. This represents the degree matrix used to normalize the adjacency matrix. This represents the identity matrix used to add self-loops.

[0054] Furthermore, in one embodiment, the encoding of brain region locations specifically includes: S211: To provide a stable and distinguishable spatial reference in the early stages of training, an absolute position encoding matrix is ​​first constructed based on brain region indices. Then for the first The brain region, in the _ ... and The encoding of the feature dimension is defined as follows:

[0055]

[0056] in, , Indicates the first The brain region in the first Encoding of feature dimensions Indicates the first The brain region in the first Encoding of feature dimensions; The above sine and cosine encoding provides an absolute position prior that is independent of the node arrangement order, which can effectively break the permutation invariance constraint in graph neural networks; It should be noted that brain regions possess stable and distinguishable spatial identities in terms of anatomical structure and functional organization, while functional signals at different time points reflect their dynamic evolution. Relying solely on node features is insufficient to simultaneously characterize the spatial distinguishability and temporal evolution of brain regions. To address this, HC-GCN introduces a multimodal location coding mechanism in structural modality (sMRI) and functional modality (rs-fMRI) to explicitly inject spatial identity and stage-specific contextual information into brain region nodes, thereby providing a spatiotemporally consistent representational basis for subsequent cross-modal causal modeling. S212: Based on the above prior knowledge, a learnable position embedding vector is introduced. and Embedded vectors The initial value is determined by the corresponding position encoding matrix. The position-enhanced node features are obtained by element-wise addition: , , This indicates the location-enhanced features of brain structural nodes. This indicates the features of brain functional nodes after location enhancement, and... and This serves as the input to the hierarchical causal cascade attention module within the hierarchical causal cascade. This "prior initialization + joint optimization" mechanism enables the model (HC-GCN) to adaptively learn brain region spatial identity and functional stage information related to disease discrimination while preserving the original modality feature representation.

[0057] It should be noted that location encoding only applies to brain region nodes (sMRI and rs-fMRI) and is not introduced to gene nodes to avoid disrupting semantic consistency between different modalities. For the structural modality (sMRI), location encoding is used to distinguish the absolute identities of different brain regions in anatomical space, thereby maintaining the spatial distinguishability of nodes in graph modeling and attention calculation. For the functional modality (rs-fMRI), location encoding, while characterizing brain region identity, implicitly represents the relative evolutionary stage of functional signals within the pruning time window. Location-enhanced and It will be used as input to the Hierarchical Causal Cascaded Attention Module (HCCA) for subsequent cross-modal causal dependency modeling.

[0058] Furthermore, in one embodiment, hierarchical causal cascade specifically includes: S221: Create a hierarchical causal cascade attention module containing a two-stage serial attention flow; It should be noted that most existing multimodal fusion methods adopt parallel splicing or symmetrical cross attention mechanisms. This approach implicitly assumes that different modalities are at the same semantic level, ignoring the inherent directed dependency structure within biological systems. To explicitly model the biological central dogma of "genotype → brain structural phenotype → brain functional phenotype", HC-GCN proposes a hierarchical causal cascade attention module (HCCA). This module abandons the traditional flattened fusion strategy and designs a two-stage serial attention flow. By parameterizing the flow order of information, it injects genetic risks into the anatomical structure layer by layer and further constrains functional expression, thereby achieving feature enhancement with causal consistency. S222: The hierarchical causal cascade attention module performs gene-to-structure risk injection to mimic the unidirectional driving effect of genetic variations on brain anatomy (such as gray matter atrophy), thereby... Mapped as query terms to find pathogenic genes associated with structural abnormalities in specific brain regions. The mapping is done as keys and values, providing potential risk factors; S223: Define projection transformation , , Structural representations regulated by genetic risk are computed using scaled dot product attention. :

[0059] in, This represents a query matrix derived from features of brain structural nodes. This represents the learnable projective weight matrix of the query. The key matrix represents the features derived from gene nodes. The learnable projective weight matrix represents the key matrix. The Value matrix represents the features derived from gene nodes. The learnable projective weight matrix represents the value matrix. This represents the layer normalization used for stable training and attention output. This represents the Softmax activation function. This represents the dimension of the Key matrix; mathematically, this process maps the distribution characteristics of the gene space to the anatomical space, explicitly modeling... The conditional dependencies enable the generated structural representation It is no longer a simple anatomical measurement, but a structural feature that incorporates genetic susceptibility; S224: After acquiring the controlled structural representation, HCCA further performs structure-to-function modulation. Given that abnormalities in brain function are often downstream manifestations of structural damage, the second stage no longer uses the original structural features, but instead utilizes the output of the first stage. As a priori condition, for the position augmentation Modulation is performed, and functional features are used as query items. ,Will As keys and values , Generate structurally constrained functional representations through cascaded attention. :

[0060] in, Represents the query vector. Represents the key vector. Representing value vectors; the key innovation here lies in cascaded dependencies: functional features The generation of [data] depends not only on structured data, but also on [other factors]. Indirectly relying on genetic data, this transmission chain-forced model follows a causal path at the feature interaction level, avoiding spurious direct associations between genes and functions; S225: Will , , By concatenating the features, a full-modal node feature matrix is ​​obtained. , which serves as the input to the causal constraint graph convolution module.

[0061] Furthermore, in one embodiment, local information aggregation specifically includes: S231: The output of the hierarchical causal cascade attention module , , The nodes are concatenated to construct the initial node feature matrix of the full modality. ; It should be noted that although the HCCA module explicitly models the hierarchical driving relationship across modalities in the feature dimension through the cascaded attention mechanism, it mainly focuses on global semantic alignment and has not fully explored the local collaborative patterns of brain regions in the topological space. In order to map the hierarchical features output by HCCA to a graph topology structure with biological constraints, HC-GCN introduces a causal constraint graph convolution module. This module aims to use the message passing mechanism of graph convolutional networks (GCN) to strictly follow biological constraints while aggregating neighborhood information within the modality, thereby achieving deep fusion of features in both topological and semantic dimensions. S232: Adjacency Matrix Based on Causal Constraints , define the first The graph convolution propagation rule for the layer is:

[0062] in, This represents a non-linear activation function, such as ReLU. Indicates the first The node feature matrix of the layer, This represents the learnable weight matrix of the current layer. Indicates the first The node feature matrix of the layer; The key here is the matrix. The embedded physical topology blocking mechanism (i.e. =0), this hard structural prior mandates the legitimate flow of message passing: for functional nodes, their feature updates only aggregate local collaborative information from the functional neighborhood and causal driving information from the structural neighborhood, while explicitly prohibiting direct input from gene nodes. This design effectively prevents the leakage of non-causal features in deep networks and ensures that information propagation strictly follows the biological central dogma. S233: Will pass through The node representation obtained after stacking layers of causal convolutions The data is then fed into a multilayer perceptron or a fully connected layer for final classification prediction, assuming... For the sample The true label, To determine the probability distribution predicted by the hierarchical causal graph convolutional network, the cross-entropy loss function is used as the optimization objective for training the network.

[0063] in, This indicates the number of layers in a graph convolutional network. This represents the total number of nodes in the graph convolutional network. This loss function drives end-to-end joint optimization of the entire HC-GCN framework (including positional encoding, HCCA attention parameters, and GCN weights) by minimizing the difference between the predicted and true distributions, thereby achieving high-precision identification of Alzheimer's disease.

[0064] It should be further noted that while HC-GCN demonstrates excellent classification accuracy, its interpretability—that is, answering the question of "which gene mutations or brain region abnormalities drive the disease from early to late stages"—is crucial for a medical diagnostic model. Traditional attribution methods (such as Grad-CAM or primitive integral gradients) typically use an all-zero matrix as the baseline, which is biologically meaningless because the healthy control group (NC), rather than a "vacuum state," serves as the reference point for the disease. Therefore, this application proposes a novel interpretative framework—the Staged Counterfactual Integral Gradient Inference Framework (SCIG). This framework abandons the single baseline setting and quantifies the marginal contribution of different modalities in each stage of disease progression (NC→EMCI→LMCI→AD) by constructing a counterfactual trajectory of disease evolution. Specifically, firstly, prototypes of different cognitive stages are constructed based on the training set, defined as the centroids of the features of each class of samples. For a target patient diagnosed with AD, its counterfactual disease trajectory is defined as... The trajectory simulates the ideal state of the patient's physiological characteristics if the patient is in the early stages.

[0065] Specifically, the application process of the phased counterfactual integral gradient inference framework includes: S301: To accurately quantify the pathogenic drivers at each stage, SCIG employs a path integral strategy to calculate feature attribution for each evolutionary stage in the evolutionary trajectory. (For example, from NC to EMCI), calculate from the baseline state Migrate to target state During the process, the first Features Predicting the probability of Alzheimer's disease The cumulative contribution, the phased counterfactual integral gradient inference framework accumulates the gradient along the linear interpolation path, and the calculation formula is:

[0066] in, Indicates the first Features In the stage Attribution value, Indicates the first state under the target condition Features The value, Indicates the first state under the baseline condition Features The value, This represents the interpolation parameter with a value range of [0,1]. S302: In numerical implementation, the Riemann Sum approximation is used. By multiplying the characteristic changes with the average gradient along the path, the risk of Alzheimer's disease is decomposed into driving factors in three consecutive stages, and specific biomarkers for the initiation stage (NC→EMCI), the progression stage (EMCI→LMCI), and the deterioration stage (LMCI→AD) are identified. S303: Based on the attribution tensors calculated at each stage, a staged counterfactual integral gradient inference framework is used to independently aggregate and rank the contributions of different modalities. For gene modalities and brain region modalities, their importance scores are defined as the sum of absolute values ​​across feature dimensions. This scoring mechanism can automatically generate a Top-K driver report, listing the pathogenic genes (such as APOE4 variants) and damaged brain regions (such as the hippocampus and precuneus) with the highest contributions in each stage. This fine-grained staged attribution analysis not only validates that the decision-making logic of HC-GCN conforms to neuropathological priors (e.g., early stages are mainly driven by minor gene and structural changes, while later stages are dominated by functional network collapse), but also provides a dynamic list of potential biological pathogenic factors for clinical intervention.

[0067] The following example illustrates the counterfactual causal cascade graph network method for Alzheimer's disease proposed in this application.

[0068] In the experimental setup, the HC-GCN model architecture was first implemented using PyTorch version 1.12 and trained on a single NVIDIA RTX 4060 GPU (8GB). The key network parameters were set as follows: number of graph convolutional layers. Hidden layer dimensions The dropout rate is set to 0.2 to prevent overfitting. The optimizer used is Adam, with an initial learning rate of 0.001 and a weight decay rate of [value missing]. The composition threshold is determined through a grid search. Optimization was performed within the specified range, and the final selection was made as follows. The maximum number of training epochs is set to 100.

[0069] Secondly, to comprehensively evaluate the performance of the HC-GCN model, the following metrics were used to evaluate its classification performance: accuracy (ACC), sensitivity (SEN), specificity (SPE), F1 score, Matthews correlation coefficient (MCC), and area under the ROC curve (AUC). Furthermore, to assess its advantages in handling multimodal data, multiple comparative experiments were designed, covering scenarios for each method when using both dual and multimodal data. Regarding data partitioning, the dataset was randomly divided into a training set (80%) and a test set (20%). To avoid data leakage and enhance the model's generalization ability, a 5-fold cross-validation strategy was used to optimize hyperparameters on the training set, followed by final performance evaluation on a separate test set.

[0070] Finally, to quantify the specific contributions of each module in HC-GCN, ablation experiments were conducted by removing specific components, and the experimental results were analyzed in depth to verify the robustness of the model. Furthermore, HC-GCN was compared with state-of-the-art (SOTA) baseline methods.

[0071] To comprehensively evaluate the performance of HC-GCN, it was systematically compared with three representative Alzheimer's disease diagnostic methods: (1) traditional machine learning methods, including Random Forest (RF) and Support Vector Machine (SVM); (2) mainstream deep learning models, including Convolutional Neural Networks (CNN) and Transformer; and (3) advanced Graph Neural Networks and their variants, covering standard GCN, MCA-GCN, and two specific variants (EX-GCN and MA-GCN). These baseline models represent different methodological paradigms, including both fusion frameworks based on bimodal data and architectures that incorporate and do not employ attention mechanisms. Table 2 below shows the detailed comparison results and the names of each method. Experimental data show that, compared with existing state-of-the-art methods, HC-GCN achieved the highest ACC and AUC metrics in most classification tasks. The results indicate that by organically combining three modalities and graph convolution, HC-GCN can more effectively fuse heterogeneous multi-source information, thereby achieving excellent diagnostic performance.

[0072] Table 2 Performance comparison of different methods (NC vs. AD)

[0073] In the ablation experiments, to quantify the contributions of each key component in the HC-GCN framework, four variants were constructed by removing or replacing specific modules, and compared with the full model under the same dataset, training strategy, and evaluation metrics. The specific ablation settings are defined as follows: HCGCN without Hierarchical Causal Cascade Attention (w / o HCCA): The hierarchical causal cascade attention module is removed, and modality fusion is performed by feature concatenation based on parallel multi-head attention. The fused features are then fed into the GCN and predicted by the classifier. HCGCN without Causal Graph (w / o CG): Removes the physical topological blocking strategy in the adjacency matrix, allowing direct connections between gene nodes and functional nodes, thus degenerating the original hierarchical causal graph into a normal heterogeneous graph structure. HCGCN without GCN (w / o GCN): The attention fusion module is retained, but the core GCN layer is directly removed. The fused features are not aggregated with any graph topology information, but are directly flattened by MLP and fed into the classifier for prediction. HCGCN without Position Encoding (w / o PE): Removes the learnable positional encoding of brain region nodes and retains only the original features of brain region nodes for graph convolution computation.

[0074] Table 3 Ablation experiments of each component (NC and AD)

[0075] As shown in Table 3, the full version of HC-GCN outperforms the four variants in all evaluation metrics, demonstrating the complementarity and indispensability of each component. Removing different key components resulted in varying degrees of performance degradation. For example, replacing the hierarchical causal cascade attention module with ordinary feature fusion reduced ACC to 93.10% and MCC to 86.29%, indicating that this module plays a crucial role in hierarchical multimodal information transmission. Removing the causal graph structure reduced ACC from 94.25% to 92.95%, demonstrating that causal structure constraints can reduce noise interference in cross-modal connections, allowing the model to focus more on biologically significant gene-brain region networks. Removing the graph convolutional layer significantly reduced model performance, with ACC decreasing from 94.25% to 91.72% and MCC from 88.60% to 83.49%. This result indicates that graph convolutional networks can effectively model the topological relationships between brain regions and are an important component for capturing the global connectivity features of brain networks. Furthermore, when location encoding was removed, the model performance also declined, with ACC dropping to 91.26% and MCC from 88.60% to 82.59%. This indicates that spatial location information of brain regions plays an important role in identifying key lesion areas, and the lack of this information weakens the model's ability to identify changes in brain regions related to Alzheimer's disease.

[0076] Overall, the synergistic effect of the modules in HC-GCN can more comprehensively characterize the complex relationships between genes and brain regions, thereby improving the classification performance of Alzheimer's disease. Whether it's the causal graph design at the network architecture level, the cascaded attention mechanism, or the multimodal fusion at the data level, all aspects of HC-GCN have made indispensable and positive contributions to improving the final classification performance of Alzheimer's disease.

[0077] To further understand the model's decision-making mechanism on multimodal data, this application comprehensively evaluates gene features and brain region features by combining feature importance ranking and cross-stage stability analysis. Specifically, by comprehensively considering the global importance Z-score and the changing trends of feature contribution at different disease stages, core pathogenic factors that consistently play a key role in the model's decision-making process and exhibit high stability are identified.

[0078] Table 4 Global Importance Z-score

[0079] The final 10 key genes identified (WWOX, CNTN5, CDH13, RYR3, NRXN1, CACNA1C, CNTNAP2, ASIC2, RORA, and OPCML) can be broadly categorized into three biological modules that collectively reflect the progression of Alzheimer's disease from early synaptic abnormalities to late-stage neurodegenerative changes. (See...) Figure 3 As shown in the figure, genes such as NRXN1, CNTNAP2, CNTN5, CDH13, and OPCML exhibited relatively stable and consistently high importance in the model regarding synaptic structure and neural connectivity regulation. These genes are mainly related to neuronal adhesion and axonal guidance, and participate in the fine regulation of synaptic connections and the maintenance of neural circuit stability. The results indicate that they not only showed significant effects in the early stages of the disease but also suggest that synaptic connection impairment is a key starting point for disease development. Genes related to ion channels and neural excitability regulation, such as CACNA1C, ASIC2, and RYR3, mainly participate in the regulation of calcium ion channels and acid-sensitive channels. The performance of these genes showed a certain trend of change at different stages; that is, as the disease progresses, their abnormalities gradually accumulate, reflecting the increasing imbalance of neuronal excitability and calcium homeostasis, which in turn affects the function of a wider range of neural networks. Neuroprotective and degenerative stress modules, such as WWOX and RORA, significantly increased in importance in the middle and late stages. These genes are mainly involved in processes such as apoptosis regulation, oxidative stress response, and neuroprotection. In summary, these three types of gene modules exhibit a progressive relationship over time, from synaptic dysfunction to electrophysiological imbalance and then to neurodegenerative damage, reflecting the model's ability to dynamically capture the molecular mechanisms of disease.

[0080] At the brain region level, this application, based on joint analysis of structural and functional imaging, identified 10 key brain regions that significantly contributed to model decision-making and exhibited cross-modal consistency. These regions include Hippocampus_R (right hippocampus), ParaHippocampal_R (right parahippocampal gyrus), Amygdala_L (left amygdala), Insula_L / Insula_R (left / right insula), Precuneus_L (left precuneus), Fusiform_R (right fusiform gyrus), Temporal_Pole_Sup_L (left superior temporal pole), Cingulum_Post_R (right posterior cingulate cortex), Occipital_Sup_R (right supraoccipital gyrus), and Precentral_R (right precentral gyrus). See [link to relevant documentation]. Figure 4As shown, these brain regions can be summarized into the following three core networks: The limbic memory system, such as Hippocampus_R, ParaHippocampal_R, and Amygdala_L, is the most crucial driving node in model recognition. Structural atrophy and weakened functional connectivity in these regions are the most typical and earliest neuroimaging features of AD, directly affecting the model's assessment of the degree of cognitive impairment in memory encoding, emotion regulation, and contextual integration; the default mode network and higher cognitive integration systems, such as Precuneus_L, Cingulum_Post_R, Insula_L / Insula_R, and Temporal_Pole_Sup_L, reflect the effects from the medial temporal lobe... The pathological process spreads to the whole brain network. Among them, the precuneus and posterior cingulate cortex are usually closely related to the decline in self-related cognition and memory retrieval functions and are core nodes of the default mode network (DMN). The insula and temporal pole are involved in multimodal information integration and semantic processing. Damage to these areas indicates that cognitive impairment has extended to higher-order functional levels. Perception and executive-related networks such as Fusiform_R, Occipital_Sup_R, and Precentral_R are mainly involved in visual recognition, spatial perception, and motor executive functions. The importance of these areas gradually increases in the middle and late stages of the disease, suggesting that the pathological changes have further extended from the memory system to the perception and motor-related cortex, reflecting widespread degradation at the whole brain network level.

[0081] In interpreting the pathological state, to deconstruct the dynamic mechanism of AD evolution, the core driving factors of the three key transformation stages of the disease (NC→EMCI→LMCI→AD) were extracted based on the SCIG algorithm. Building upon this, this application further decomposes the feature contributions of 10 core genes and 11 key brain regions obtained through screening into pathogenic drivers (Patho+) that promote disease progression.

[0082] Early stage (NC→EMCI): such as Figure 5As shown, the driving factors identified by the model exhibit a significant "unidirectional pathogenicity" characteristic and are mainly concentrated in gene modalities. Synapse-related genes show a strong pathogenic driving effect at this stage. CNTN5 (Patho+: 7.089), CDH13 (Patho+: 7.171), and CACNA1C (Patho+: 7.177) are all ranked high in driving strength. These genes are mainly involved in processes such as neuronal adhesion, synaptic structural stability, and electrophysiological activity regulation. Their abnormal changes indicate that even relatively minor perturbations at the synaptic level can gradually amplify and further trigger systemic pathological reactions during the evolution from normal cognition (NC) to early mild cognitive impairment (EMCI). At the brain region level, the overall driving strength is weak, but some areas have shown a "pathogenic-compensatory counterbalancing" characteristic. The left amygdala (Amygdala_L) also has a high Patho(+) (3.000). This bidirectional contribution pattern indicates that this region has been pathologically affected in the early stages, while the brain attempts to maintain emotional and memory homeostasis through local functional reorganization.

[0083] Intermediate stage (EMCI→LMCI): such as Figure 6 As shown, the pathogenic driving force of synaptic-related genes, which played a dominant role in the early stages, significantly diminished in the middle stages of disease progression. The Patho(+) score of CNTN5 dropped sharply from 7.089 to 0.823, suggesting that synaptic functional abnormalities gradually solidified from being the absolute main driver of disease deterioration into a basic pathological state. In contrast, the pathogenic driving force at the brain region level was significantly enhanced. In sMRI data, the Patho(+) score of the right para-hippocampal gyrus (ParaHippocampal_R) climbed to 2.434, becoming the most critical structural driving node in this stage. The right hippocampus (Hippocampus_R) also showed a strong pathogenic association (Patho+: 2.644). From a quantitative analysis perspective, this precisely outlines the evolutionary trajectory of pathological damage gradually spreading from the synaptic level to the medial temporal lobe memory system.

[0084] Late stage (LMCI→AD): such as Figure 7As shown, the model captured more complex brain structural and functional abnormalities in the late stages of the disease. The driving mechanism at this stage is specifically manifested as synergistic lesions at the molecular and brain region levels. At the genetic level, genes involved in apoptosis and stress response once again became the dominant driving factors. The indicators of WWOX (Patho+: 6.884) and CDH13 (Patho+: 4.326) both showed a significant rebound. This characteristic secondary activation indicates that the late-stage pathological morphology has evolved from synaptic dysfunction to neuronal damage and even cell death. At the brain region level, the pathological abnormalities have now spread to the whole brain's multi-system network. The pathogenic driving force of the right posterior cingulate cortex (Cingulum_Post_R), the core hub of the default mode network, continues to rise (Patho+: 2.759), so the "functional collapse" here undoubtedly becomes a key signal for LMCI to deteriorate into AD. At the same time, significant lesions have also occurred in brain regions responsible for movement, such as the right precentral gyrus (Precentral_R) (Patho+: 2.090), indicating that the influence of the disease has penetrated into the sensorimotor system.

[0085] This application proposes HC-GCN, which combines hierarchical causal cascaded attention (HCCA) with causal constraint graph convolutional networks, for the precise diagnosis of Alzheimer's disease and the dynamic identification of its pathogenic mechanisms. This model explicitly models the hierarchical dependencies of "gene-structure-function," learning collaborative representations of multimodal features from both biological causality and topological dimensions. It also introduces learnable location encoding to capture the dynamic identity of brain regions in spatial and evolutionary processes. Furthermore, to go beyond static feature weight analysis, this application proposes a staged counterfactual integral gradient (SCIG) analysis framework. This framework can reconstruct the continuous trajectory of disease evolution and quantify the dynamic marginal contributions of different modalities at each stage of disease development.

[0086] Experimental results on the ADNI dataset demonstrate that HC-GCN outperforms existing state-of-the-art methods on multiple classification tasks, and the key features mined by SCIG are highly consistent with existing AD pathological evidence, validating the model's interpretability: the high frequency of synaptic and ion channel pathway-related genes (such as PTPRN2, ERBB4, and CACNA1C) throughout all stages reveals that synaptic dysfunction and calcium homeostasis imbalance are the core molecular drivers of AD. In particular, PTPRN2 appears first in Stage I, suggesting its potential as a very early screening target. Although hippocampal atrophy only becomes significant in the late stages, the olfactory cortex is included in the key list of structural modalities in both Stage I and Stage II. This is consistent with clinical observation that olfactory dysfunction often precedes memory impairment, making it an early and sensitive sentinel for AD. The model assigns the highest negative weights to Amygdala and Hippocampus in the final stage, further quantifying and confirming that structural collapse in the medial temporal lobe (MTL) region is the "gold standard" feature leading to the final diagnosis of AD. In summary, this model not only improves diagnostic performance, but also provides a new perspective and computational tool for studying the pathological mechanisms of Alzheimer's disease and for causal analysis of other neurodegenerative diseases.

[0087] This application proposes a Hierarchical Causal Graph Convolutional Network (HC-GCN) to model the hierarchical structure and causal dependencies in multimodal representation learning, aiming to construct a multimodal framework with causal explanatory capabilities. By embedding a Hierarchical Causal Cascade Attention (HCCA) module, it explicitly characterizes the hierarchical causal dependencies between multimodal features, following the principle of "genes determine structure, structure carries function." Furthermore, it combines graph convolution operations to model and aggregate local topological cooperative relationships within and across modal nodes, thereby achieving structured multimodal representation learning with causal constraints. At the inference level, this application further proposes a Stage-wise Counterfactual Integrated Gradients (SCIG) framework. By constructing continuous pathological state prototypes at different stages and performing counterfactual interpolation along hierarchical causal paths, it achieves fine-grained causal attribution of model decisions at different modalities and structural levels. Experimental results on the ADNI dataset show that HC-GCN outperforms several existing multimodal fusion methods in diagnostic performance. Meanwhile, SCIG can stably identify key gene-brain region combinations that drive disease progression and their cross-modal transmission pathways, providing a causal explanatory framework for the mechanism analysis and pathogenic factor identification of Alzheimer's disease.

[0088] Secondly, embodiments of this application also provide a counterfactual causal cascade graph network device for Alzheimer's disease.

[0089] In one embodiment, reference is made to Figure 8 , Figure 8 This is a schematic diagram of the functional modules of the counterfactual causal cascade network device for Alzheimer's disease proposed in this application. Figure 8 As shown, the counterfactual causal cascade graph network device for Alzheimer's disease includes: a construction module and an execution module.

[0090] The construction module is used to acquire multimodal data of samples from Alzheimer's disease-related databases and preprocess them to construct a dataset. The execution module is used to create a hierarchical causal graph convolutional network and, based on the constructed dataset, to construct gene-brain region heterogeneous maps, encode brain region locations, perform hierarchical causal cascades, aggregate local information, and combine a phased counterfactual integral gradient inference framework to achieve dynamic analysis of the evolution trajectory of Alzheimer's disease.

[0091] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An Alzheimer's disease-oriented counterfactual causal cascade graph network method, characterized in that, The counterfactual causal cascade graph network method for Alzheimer's disease includes: Multimodal data of samples from Alzheimer's disease-related databases were obtained and preprocessed to construct the dataset; A hierarchical causal graph convolutional network is created, and based on the constructed dataset, gene-brain region heterogeneity map construction, brain region location encoding, hierarchical causal cascade, and local information aggregation are performed. Combined with a staged counterfactual integral gradient inference framework, dynamic analysis of the evolution trajectory of Alzheimer's disease is achieved.

2. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 1, characterized in that: The samples include normal controls, mild cognitive impairment, and confirmed cases, with mild cognitive impairment including early and late stages, and each sample is distributed according to a preset ratio.

3. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 1, characterized in that: The preprocessing includes candidate gene screening, genetic data quality control and feature selection, structural image data processing, and resting-state functional magnetic resonance imaging data processing. The candidate gene screening was based on SNP sequence length and PPI network topology to screen for a target number of candidate genes, and the biological significance of the screened candidate genes was confirmed by GO function and KEGG pathway enrichment analysis performed by Metascape. The genetic data quality control and feature selection were performed using PLINK and MAGMA software, including deletion rate control, minimum allele frequency control, Hardy-Weinberg balance test, linkage disequilibrium pruning, SNP-to-gene mapping, PPI network-based gene screening, and gene encoding. The structural image data processing employs the FSL software library in conjunction with an active verification method to quantify the gray matter volume shrinkage in each brain region. Specifically, this includes non-brain tissue removal, tissue segmentation and registration, spatial registration, quality control, ROI coordinate calculation, and hybrid feature extraction. The resting-state functional magnetic resonance imaging data processing is based on the DPARSF toolbox, specifically including format conversion, initial time point deletion, time layer correction, head motion correction, spatial normalization, spatial smoothing, temporal filtering, covariate removal, and brain region time series extraction.

4. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 1, characterized in that: The gene-brain region heterogeneity map is constructed based on biological priors, which includes gene, brain structure and brain function nodes. Causal connection constraints between modalities are established through a physical topological blocking strategy. For gene-brain region heterogeneity mapping, the specific components include: The rs-fMRI time series were cropped to align their length with the gene coding sequence. Constructing a gene-brain region heterogeneity map with causal constraints , Represents a block adjacency matrix. Let the node feature matrix be denoted as: , in, The initial feature representation of the gene node. The initial feature representation of brain structural nodes. If we represent the initial features of brain functional nodes, then the total number of nodes is the sum of the number of gene nodes, brain structure nodes, and brain functional nodes. This represents the adjacency matrix between gene nodes. This represents the adjacency matrix between gene nodes and brain structure nodes. This represents the adjacency matrix between brain structural nodes and gene nodes. This represents the adjacency matrix between nodes in the brain structure. This represents the adjacency matrix between brain structural nodes and brain functional nodes. This represents the adjacency matrix between brain functional nodes and brain structural nodes. Represents the adjacency matrix between brain functional nodes; compute nodes and The Pearson correlation coefficient between the eigenvectors is used to measure the initial weights of the edges: in, This represents the node dimension, i.e., the number of time points. Represents a node In the Feature values ​​at each time point Represents a node In the Feature values ​​at each time point; By setting a threshold The block adjacency matrix is ​​sparsified to obtain And perform normalization to obtain the causal constraint adjacency matrix of the causal constraint graph convolution module: in, Represents the causal constraint adjacency matrix. This represents the degree matrix used to normalize the adjacency matrix. This represents the identity matrix used to add self-loops.

5. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 4, characterized in that: The brain region location coding is used to integrate spatial location information in the brain and encoded sequence location information based on a learnable location coding method; For brain region location coding, the specific components include: Constructing an absolute location encoding matrix based on brain region indexing Then for the first The brain region, in the _ ... and The encoding of the feature dimension is defined as follows: in, , Indicates the first The brain region in the first Encoding of feature dimensions Indicates the first The brain region in the first Encoding of feature dimensions; Introducing learnable position embedding vectors and Embedded vectors The initial value is determined by the corresponding position encoding matrix. The position-enhanced node features are obtained by element-wise addition: , , This indicates the location-enhanced features of brain structural nodes. This indicates the features of brain functional nodes after location enhancement, and... and As input to the hierarchical causal cascade attention module in the hierarchical causal cascade.

6. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 5, characterized in that: The hierarchical causal cascade is used to explicitly model the hierarchical driving relationship of genes, structures, and functions by connecting attention streams, thereby achieving feature enhancement with causal constraints. For hierarchical causal cascades, the specific components include: Create a hierarchical causal cascade attention module containing a two-stage serial attention flow; The hierarchical causal cascade attention module performs gene-to-structure risk injection to mimic the unidirectional driving effect of genetic variation on brain anatomy. Mapped as query terms to find pathogenic genes associated with structural abnormalities in specific brain regions. Mapped to keys and values; Define projection transformation , , Structural representations regulated by genetic risk are computed using scaled dot product attention. : in, This represents a query matrix derived from features of brain structural nodes. This represents the learnable projective weight matrix of the query. The key matrix represents the features derived from gene nodes. The learnable projective weight matrix represents the key matrix. The Value matrix represents the features derived from gene nodes. The learnable projective weight matrix represents the value matrix. This represents the layer normalization used for stable training and attention output. This represents the Softmax activation function. Indicates the dimension of the Key matrix; Use functional features as query terms ,Will As keys and values , Generate structurally constrained functional representations through cascaded attention. : in, Represents the query vector. Represents the key vector. Represents a value vector; Will , , By concatenating the features, a full-modal node feature matrix is ​​obtained. , which serves as the input to the causal constraint graph convolution module.

7. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 6, characterized in that: The local information aggregation is used to aggregate the local topological information of heterogeneous nodes in the causal neighborhood using the output of the hierarchical causal cascaded attention module as the initial feature, and input the learned gene and brain region representations into the fully connected network. For local information aggregation, the specific components include: The output of the hierarchical causal cascade attention module , , The nodes are concatenated to construct the initial node feature matrix of the full modality. ; Based on causal constraint adjacency matrix , define the first The graph convolution propagation rule for the layer is: in, Represents a non-linear activation function. Indicates the first The node feature matrix of the layer, This represents the learnable weight matrix of the current layer. Indicates the first The node feature matrix of the layer; Will pass The node representation obtained after stacking layers of causal convolutions The data is then fed into a multilayer perceptron or a fully connected layer for final classification prediction, assuming... For the sample The true label, To determine the probability distribution predicted by the hierarchical causal graph convolutional network, the cross-entropy loss function is used as the optimization objective for training the network. in, This indicates the number of layers in a graph convolutional network. This represents the total number of nodes in the graph convolutional network.

8. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 1, characterized in that: The phased counterfactual integral gradient inference framework is used to construct counterfactual prototypes for different pathological stages, perform integral gradient attribution along the evolution trajectory, and identify key pathogenic factors driving disease progression.

9. The counterfactual causal cascade graph network method for Alzheimer's disease as described in claim 8, characterized in that, The specific application process of the phased counterfactual integral gradient inference framework includes: For each stage of evolution in the evolutionary trajectory Calculate from baseline state Migrate to target state During the process, the first Features Predicting the probability of Alzheimer's disease The cumulative contribution, the phased counterfactual integral gradient inference framework accumulates the gradient along the linear interpolation path, and the calculation formula is: in, Indicates the first Features In the stage Attribution value, Indicates the first state under the target condition Features The value, Indicates the first state under the baseline condition Features The value, This represents the interpolation parameter with a value range of [0,1]. Using Riemann approximation By multiplying the characteristic changes with the average gradient along the path, the risk of Alzheimer's disease is decomposed into driving factors in three consecutive stages, identifying specific biomarkers for the onset, progression, and deterioration stages. Based on the attribution tensors calculated at each stage, the phased counterfactual integral gradient inference framework performs independent contribution aggregation and ranking for different modalities.

10. A counterfactual causal cascade graph network device for Alzheimer's disease, characterized in that, The counterfactual causal cascade graph network device for Alzheimer's disease includes: The module is used to acquire multimodal data of samples from Alzheimer's disease-related databases and preprocess them to build the dataset. The execution module is used to create a hierarchical causal graph convolutional network and, based on the constructed dataset, to construct a gene-brain region heterogeneous map, encode brain region locations, perform hierarchical causal cascades, aggregate local information, and combine a staged counterfactual integral gradient inference framework to achieve dynamic analysis of the evolution trajectory of Alzheimer's disease.