Alzheimer disease diagnosis system based on image gene data fusion

The characteristics of fMRI and DNA data are extracted through deep and shallow GCN modules, and the decoupled representation and adaptive weighted fusion method are used to solve the problem of poor multimodal data fusion effect in the prior art, achieving higher disease prediction and diagnostic accuracy.

CN120148831APending Publication Date: 2025-06-13BEIJING INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510311874.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate two types of information when fusing functional magnetic resonance imaging (fMRI) and DNA gene data, resulting in information loss and unsatisfactory prediction accuracy.

Method used

The deep GCN module and shallow GCN module are used to extract the global and local features of fMRI and DNA data respectively, and the global and local representation decoupling mechanism is combined with similarity map alignment and adaptive weighted fusion of attention mechanisms to achieve efficient fusion of multimodal data.

Benefits of technology

It significantly improves the accuracy and stability of complex disease prediction and diagnosis, solves the problems of insufficient cross-modal feature extraction and information loss, and ensures that the unique information of each data modal is retained and fully utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148831A_ABST
    Figure CN120148831A_ABST
Patent Text Reader

Abstract

The invention provides an Alzheimer's disease diagnosis system based on image gene data fusion, which innovatively introduces a decoupling representation technology, divides features of multi-modal data into global representation and local representation, retains unique information of each data modal by respectively extracting global features and local features, and improves the diagnosis accuracy of the Alzheimer's disease. Efficient fusion of shared information is realized; through the innovative multi-modal feature fusion method, the accuracy and stability of complex disease prediction and diagnosis are remarkably improved, and the problem that cross-modal features and modal internal features cannot be captured at the same time in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biomedical technologies, and particularly relates to an Alzheimer's disease diagnosis system based on the fusion of imaging and genetic data. Background Art

[0002] In order to obtain valuable knowledge from the complementary information in different perspectives of biomedical data, effective data fusion and effective pattern mining are key issues that need to be urgently solved. In this context, the fusion analysis of imaging data and genetic data has become a popular trend in studying disease mechanisms and disease diagnosis. For Alzheimer's Disease (AD), a large number of studies have shown that AD is closely related to the lesions in specific brain regions. In recent years, with the rapid development of high-throughput sequencing, researchers have found that genes related to AD can severely induce significant degradation of certain brain functions. Therefore, it is of great significance to fuse imaging and genetic data and conduct more accurate early diagnosis research.

[0003] Currently, some imaging and genetic multimodal data fusion studies on AD have been carried out at home and abroad, but effective data fusion technologies are still in the exploratory stage. Researchers first adopted classical algorithms such as random forest, principal component analysis, etc., but these methods cannot mine the deep features between data. Therefore, some improved methods have emerged. For example, an adaptive structured SCCA method can be used to detect the genetic associations of AD brain imaging phenotypes, and methods based on statistical independence and structural sparsity can obtain genes and brain regions related to brain diseases. With the development of artificial intelligence technology and the expansion of data volume, multimodal fusion methods based on deep learning have become the current research trend in recent years. Some researchers constructed a brain region-gene map and used GCN to extract multimodal complementary information. In addition, there are also studies using CNN to achieve survival prediction for brain disease patients. In summary, constructing an AD multimodal data fusion analysis method is not only a necessary way to mine the correlation mechanism of multimodal data, but also a key direction for constructing an accurate AD diagnosis model.

[0004] In the literature "Feature aggregation graph convolutional network based on imaging genetic data for diagnosis and pathogeny identification of Alzheimer's disease", the authors proposed an AD diagnosis and etiology identification method based on the Feature Aggregation Graph Convolutional Network (FAGCN). Its specific implementation includes the following key steps:

[0005] (1) Construction of the brain region-gene map

[0006] First, the fMRI data is processed into a brain region time series of n1*m (n1 is the number of brain regions, and m is the length of the time series); the gene SNP data is processed into n2 genes containing m loci. Then, a node-based brain region-gene network is constructed, where the brain regions and genes are defined as nodes in the network, and the weight of the edge is calculated using the Pearson correlation coefficient to describe the connection strength between nodes. A neighbor similarity matrix is used to determine similar nodes, and the features of brain region and gene nodes are updated through graph convolution to achieve multi-modal fusion of data.

[0007] (2) Feature Aggregation Graph Convolutional Network (FAGCN)

[0008] FAGCN realizes the information aggregation process through graph convolutional layers. The model passes through two graph convolutional layers (L1 and L2), and nodes successively aggregate the information of their first-order and second-order neighbors. Each layer includes two steps, namely edge-node (E-N) convolution and node-edge (N-E) convolution, and the node features are updated through the information transfer of neighbors and similar nodes. A weighted matrix and a similarity matrix are introduced in the graph convolutional layer to increase the interpretability of the results.

[0009] (3) Model training and feature selection

[0010] The FAGCN model is trained using the cross-entropy loss function, with the classification accuracy as the evaluation criterion, and the parameters of the model are optimized through the gradient descent algorithm. The full-gradient saliency map mechanism is applied to select features, extracting the brain region and gene features that are most discriminative for AD, and a t-test is performed on these features to ensure their statistical significance.

[0011] At present, the fusion of fMRI (functional Magnetic Resonance Imaging) and DNA gene data is widely used in medical research, especially in the prediction and diagnosis of complex diseases. These two types of data respectively provide important information about brain function and gene variation. However, due to their different characteristics, existing multimodal data fusion methods have significant deficiencies in information integration. First, fMRI data reflects the connections and functional activities between different regions of the brain, while DNA data provides detailed information about gene inheritance. Traditional multimodal data processing methods often fail to effectively integrate all the information from these two modalities when faced with such complexity. Specifically, the macroscopic connection patterns in fMRI data and the microscopic changes in gene data are typical representatives of the two types of information. How to retain these important patterns and details in the fusion process is a difficult problem in existing methods. Simple combination methods are prone to ignoring the interactions between these patterns or only focusing on the data features of a certain modality, ultimately resulting in unsatisfactory prediction accuracy and diagnostic effects. In addition, existing methods have not fully utilized the possible correlations between these two types of data. For example, the activity patterns of certain brain functions may be directly related to specific gene variations, while the functional activities of other parts may not depend on gene information. This kind of correlation and difference requires the fusion method to be able to distinguish which information is the common feature of the two modalities and which is the independent feature unique to the modality. Being able to simultaneously identify and integrate these shared information and independent features is crucial for improving the performance of the model and the accuracy of disease prediction. Summary of the Invention

[0012] To solve the problems of information loss, insufficient cross-modal feature extraction, and unsatisfactory fusion effect in the prior art, the present invention provides an Alzheimer's disease diagnosis system based on the fusion of imaging and gene data, which significantly improves the effect of multimodal data fusion, especially showing higher accuracy and stability in the prediction and diagnosis of complex diseases.

[0013] An Alzheimer's disease diagnosis system based on the fusion of imaging and gene data, comprising a deep GCN module, a first shallow GCN module, a second shallow GCN module, an RWD module, a global fusion module, a local fusion module, an aggregation module, and a multi-layer perceptron;

[0014] The deep GCN module is used to extract the fMRI global features of the graph model corresponding to the functional magnetic resonance imaging of the patient;

[0015] The first shallow GCN module is used to extract the fMRI local features of the graph model corresponding to the functional magnetic resonance imaging of the patient;

[0016] The RWD module is used to extract the DNA global features of the graph model corresponding to the gene sequence of the patient;

[0017] The second shallow GCN module is used to extract the DNA local features of the graph model corresponding to the gene sequence of the patient;

[0018] The global fusion module is used to fuse the fMRI global features and the DNA global features to obtain a global representation;

[0019] The local fusion module is used to fuse the fMRI local features and the DNA local features to obtain a local representation;

[0020] The aggregation module is used to fuse the global representation and the local representation to obtain an overall representation;

[0021] The multi-layer perceptron is used to determine the probability of the patient having Alzheimer's disease according to the overall representation.

[0022] Further, the method for obtaining the graph model corresponding to the functional magnetic resonance imaging of the patient is as follows:

[0023] The functional magnetic resonance imaging of the patient is divided into 116 brain regions through the AAL whole-brain template;

[0024] According to the blood oxygenation level-dependent signals of each brain region at 130 sampling time points, a node feature matrix X 116×130 ={x ij}, where x ij represents the blood oxygenation level-dependent signal of the i-th brain region at the j-th sampling time point, where i = 1, 2,... 116 and j = 1, 2,... 130;

[0025] According to the Pearson correlation coefficient between the blood oxygenation level-dependent signals of any two brain regions at 130 sampling time points, an adjacency matrix A 116×116 ={x ab}, x ab represents the Pearson correlation coefficient between the blood oxygenation level-dependent signals of the 130 sampling time points corresponding to the a-th brain region and the blood oxygenation level-dependent signals of the 130 sampling time points corresponding to the b-th brain region, where a = 1, 2,... 116 and b = 1, 2,... 116;

[0026] The node feature matrix X 116×130 ={x ij} and the adjacency matrix A 116×116 ={x ab} together constitute the graph model corresponding to the functional magnetic resonance imaging of the patient.

[0027] Further, the deep GCN module includes four layers of graph convolutional networks, and the convolutional operations of each layer of graph convolutional network are as follows:

[0028] H (l+1) =σ(LH (l) W (l) )

[0029] where H (l) is the node feature matrix of the l-th layer of graph convolutional network, H (l+1) is the node feature matrix of the (l + 1)-th layer of graph convolutional network, l = 0, 1, 2, 3, σ is the activation function, L is the Laplacian matrix, and L = D - A, D is the degree matrix determined according to the node feature matrix X 116×130 ={x ij}, A is the adjacency matrix A 116×116 ={x ab}, and W (l) is the node weight matrix of the l-th layer of graph convolutional network.

[0030] Further, the method for obtaining the graph model corresponding to the patient's gene sequence is as follows:

[0031] Assume that the length of the patient's gene sequence is L, and it is divided into L - K + 1 K-mer short sequences of length K. Among them, the total number of all possible types of K-mer short sequences is 4 K types;

[0032] Construct the node feature matrix corresponding to the graph model according to the frequencies of various types in the L - K + 1 K-mer short sequences where y i represents the number of times the i-th type of K-mer short sequence appears in the L - K + 1 K-mer short sequences, i = 1, 2,..., 4 K ;

[0033] Construct the adjacency matrix corresponding to the graph model according to the correlation between any two K-mer short sequences where, if y ab =n, it means that the a-th type of K-mer short sequence and the b-th type of K-mer short sequence appear in the form of consecutive K-mer short sequences n times in the L - K + 1 K-mer short sequences, and a = 1, 2,..., 4 K , b = 1, 2,..., 4 K ;

[0034] The graph model corresponding to the patient's gene sequence is jointly composed of the node feature matrix and the adjacency matrix .

[0035] Furthermore, the RWD module is used to extract the DNA global features of the graph model corresponding to the patient's gene sequence, specifically as follows:

[0036] Calculate the transition probability matrix P of random walk using the degree matrix D and adjacency matrix B of the graph model corresponding to the patient's gene sequence:

[0037] P = D -1 B

[0038] where D is the degree matrix determined according to the adjacency matrix B, and B is the adjacency matrix

[0039] Diffuse the node feature matrix The diffused node feature matrix H(t) is:

[0040] H(t) = P t Y

[0041] where P t is the transition probability matrix after t steps of random walk, and Y is the node feature matrix

[0042] Judge whether the difference between the node feature matrices obtained by diffusion in two adjacent steps is less than the set value. If so, the node feature matrix H(∞) obtained by the last step of diffusion is used as the final DNA global feature. If not, continue to diffuse.

[0043] Furthermore, the second shallow GCN module is used to extract the DNA local features of the graph model corresponding to the patient's gene sequence, specifically as follows:

[0044] Use the degree matrix D determined by the adjacency matrix B to determine the weights of each node in the adjacency matrix B:

[0045]

[0046] where B′ ab is the weight of node a corresponding to the a-th K-mer short sequence, d a is the degree of node a, d b is the degree of node b corresponding to the b-th K-mer short sequence, ∑ b is the set of adjacent nodes connected to this node a;

[0047] According to the weights B′ of each node ab Perform local weighting on the node feature matrix to obtain the final DNA local feature H gp :

[0048]

[0049] Among them, the weight matrix I is the identity matrix, σ is the activation function, and Y is the node feature matrix W is the learnable weight matrix.

[0050] Furthermore, the global fusion module is used to fuse the fMRI global feature and the DNA global feature to obtain the global representation, specifically as follows:

[0051] The fMRI global feature and the DNA global feature are respectively mapped to the same dimensional space through linear transformation to align the fMRI global feature and the DNA global feature;

[0052] The aligned fMRI global feature and DNA global feature are weighted and fused to obtain the global representation.

[0053] Furthermore, the local fusion module is used to fuse the fMRI local feature and the DNA local feature to obtain the local representation, specifically as follows:

[0054] The fMRI local feature and the DNA local feature are respectively mapped to the same dimensional space through linear transformation to align the fMRI local feature and the DNA local feature;

[0055] The aligned fMRI local feature and DNA local feature are adaptively weighted and fused to obtain the local representation; among them, the method for obtaining the adaptive weights corresponding to the fMRI local feature and the DNA local feature is:

[0056] H′ gp =W B ·H ip

[0057]

[0058] Among them, H′ gp is the adaptive weight corresponding to the DNA local feature, H ip ′ is the adaptive weight corresponding to the fMRI local feature, H ip is the aligned fMRI local feature, H gp is the aligned DNA local feature, T represents the transpose, and W B is the initial weight, and softmax(·) is the activation function.

[0059] Furthermore, the aggregation module is used to fuse the global representation and the local representation to obtain the overall representation, specifically as follows:

[0060] The fMRI global feature and the DNA global feature are respectively mapped to the same dimensional space through linear transformation to align the fMRI global feature and the DNA global feature;

[0061] Obtain the similarity between the aligned fMRI global features and DNA global features;

[0062] Fuse the similarities greater than the set threshold into the global representation, and fuse the similarities not greater than the set threshold into the local representation;

[0063] Overlay the global representation and the local representation after fusing the similarities to obtain the overall representation.

[0064] Furthermore, the loss function loss adopted during the training of the Alzheimer's disease diagnosis system is:

[0065] loss = αloss class + βlos same + γloss con

[0066] where loss class is the classification loss, α is the weight of the classification loss, loss same is the similarity loss, β is the weight of the similarity loss, loss con is the contrast loss, and γ is the weight of the contrast loss;

[0067] The calculation method of the classification loss loss class is:

[0068]

[0069] where N is the number of sample patients, y i is the true Alzheimer's disease status of the i-th sample patient, y i = 1 indicates having Alzheimer's disease, y i = 0 indicates not having Alzheimer's disease, and p i is the true Alzheimer's disease probability of the i-th sample patient output by the Alzheimer's disease diagnosis system;

[0070] The calculation method of the similarity loss loss same is:

[0071]

[0072] where μ k represents the k-th order central moment, K is the maximum order of calculation, μ k (xs) is the k-th order central moment of the local representation, μ k (xt) is the k-th order central moment of the global representation, and ||·|| 2 is the 2-norm;

[0073] The calculation method of the contrast loss loss con is:

[0074]

[0075] Among them, rep1 i is the sparse representation corresponding to the global representation of the i-th sample patient, and rep2 i is the sparse representation corresponding to the local representation of the i-th sample patient, and margin is the set spacing.

[0076] Beneficial effects:

[0077] 1. The present invention provides an Alzheimer's disease diagnosis system based on the fusion of imaging and genetic data, innovatively introducing a decoupled representation technology, which divides the features of multi-modal data into global representation and local representation. By separately extracting global features and local features, the unique information of each data modality is retained, and the efficient fusion of shared information is achieved. Through this innovative multi-modal feature fusion method, the present invention significantly improves the accuracy and stability of complex disease prediction and diagnosis, and solves the problem that the prior art cannot simultaneously capture cross-modal features and intra-modal features.

[0078] 2. The present invention provides an Alzheimer's disease diagnosis system based on the fusion of imaging and genetic data. Two independent networks are respectively constructed through an imaging map and a genetic map. The imaging map is established based on the AAL template, and the connectivity between brain regions is represented by a correlation matrix; while the genetic map uses the K-mer method to efficiently convert the long DNA sequence information into a K-mer frequency node feature matrix, and introduces the occurrence frequency between K-mers to define the edge weights. This method effectively retains the structural information in the imaging and genetic data, and provides a more refined node feature representation, laying a solid foundation for subsequent feature extraction and cross-modal fusion, and solving the bottleneck that traditional methods are difficult to retain the full information of genes.

[0079] 3. The present invention provides an Alzheimer's disease diagnosis system based on the fusion of imaging and genetic data, introducing a decoupling mechanism of global and local representations. Through deep GCN, cross-modal global shared information is extracted, and through shallow GCN, unique local features within the modality are captured. In particular, for the imaging map, four layers of GCN are designed to deeply explore the long-distance dependence between brain regions, and the community structure information of the graph is extracted by combining Laplacian spectral clustering; for the genetic map, the random walk diffusion method is used to capture long-distance global features, and local weighted sparse graph convolution is combined to extract local features. This method effectively decouples between global and local features, and ensures that each feature can play an independent role, being more meticulous in capturing the specificity of imaging and genetic data, retaining both the modality uniqueness and fully extracting the shared information between modalities, thereby improving the overall expression ability of multi-modal features.

[0080] 4. The present invention provides an Alzheimer's disease diagnosis system based on the fusion of imaging and gene data. During the multi-modal feature interaction process, first, similarity map alignment is adopted to align the common representations of genes and images in the same feature space, and the node alignment between genes and images is optimized through the similarity matrix W. At the same time, in the private representation fusion, the weights of each modality are adaptively weighted and adjusted through the attention mechanism to retain the private information of each modality, ensuring that both the shared features across modalities and the uniqueness of each modality are considered in the final feature fusion process. The alignment and weighting mechanism of the present invention significantly enhances the consistency of imaging and gene data in shared features, while effectively retaining the unique information of modalities. This method avoids feature confusion, improves the expression accuracy and diversity of fused features, and helps to improve the accuracy of disease prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 FIG. is a schematic block diagram of an Alzheimer's disease diagnosis system based on the fusion of imaging and gene data provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.

[0083] The present invention proposes a multi-modal data fusion method based on GCN (Graph Convolutional Network) and RWD (Random Walk Diffusion). This method extracts the global features of fMRI data through a deep GCN network, and at the same time extracts the global features of DNA data through RWD, effectively capturing the key information that may be shared in cross-modal data. These global features reflect the common patterns between different modal data, such as the potential association between the overall functional activities of brain regions and specific gene mutations, thus solving the problem of information loss in modal feature integration in the prior art. In addition to these shared information, each data modality also contains unique private representations, that is, the detailed information within the modality. These private representations are highly correlated with the modality characteristics and reflect the unique structure of each modality. Therefore, the present invention extracts the local features of fMRI data through a shallow GCN to capture the specific activity patterns within brain regions. At the same time, local weighted sparse graph convolution is used to extract the unique gene features in gene data. These private representations can reflect the information within the modality, enabling the fusion process to not only capture the common information across modalities but also retain the important features unique to each modality. To effectively fuse these common and private features, the present invention adopts GA (Graph Alignment) and an adaptive weighted fusion method based on AM (Attention Mechanism). GA is used to align the global features of different modalities to ensure that the shared information is fully utilized; at the same time, AM assigns appropriate weights to the private features of each modality through an adaptive weighting mechanism to ensure that both the unique characteristics of the modality and the synergistic effect of multi-modal data are fully utilized during the fusion process. It realizes the efficient fusion between fMRI and DNA data, enhances the interaction between modalities, and fully utilizes the synergistic effect of different modal data. Finally, aiming at the high-dimensional sparsity problem of DNA data, the present invention adopts a technology combining sparse graph convolution and RWD, effectively improving the data processing ability and ensuring good performance of the model on high-dimensional sparse data. Through these improvements, the present invention not only solves the key problems in the prior art but also significantly improves the effect of multi-modal data fusion, especially showing higher accuracy and stability in the prediction and diagnosis of complex diseases.

[0084] Specifically, as Figure 1 shown, an Alzheimer's disease diagnosis system based on imaging-gene data fusion includes a deep GCN module, a first shallow GCN module, a second shallow GCN module, an RWD module, a global fusion module, a local fusion module, an aggregation module, and a multi-layer perceptron;

[0085] The deep GCN module is used to extract the fMRI global features of the graph model corresponding to the functional magnetic resonance imaging of the patient;

[0086] The first shallow GCN module is used to extract the fMRI local features of the graph model corresponding to the functional magnetic resonance imaging of the patient;

[0087] The RWD module is used to extract the DNA global features of the graph model corresponding to the gene sequence of the patient;

[0088] The second shallow GCN module is used to extract the DNA local features of the graph model corresponding to the gene sequence of the patient;

[0089] The global fusion module is used to fuse the fMRI global features and the DNA global features to obtain a global representation;

[0090] The local fusion module is used to fuse the fMRI local features and the DNA local features to obtain a local representation;

[0091] The aggregation module is used to fuse the global representation and the local representation to obtain an overall representation;

[0092] The multi-layer perceptron is used to determine the probability of the patient having Alzheimer's disease according to the overall representation.

[0093] That is to say, the present invention extracts the global features in the fMRI data through the deep GCN to capture the long-range dependencies across brain regions; extracts the global features of the DNA data through the RWD to capture the global patterns of gene mutations; in addition, based on the shallow GCN and sparse graph convolution, the private features in each modality are extracted to ensure that the modality-specific detailed information is retained during the fusion process; during the fusion process, the GA is used to align the data features of different modalities, and the AM is used to achieve adaptive weighted fusion to ensure the efficient integration of shared features and private features; through this innovative multi-modal fusion method, the present invention significantly improves the accuracy and stability of complex disease prediction and diagnosis, and solves the problem that the prior art cannot capture cross-modal features and intra-modal features simultaneously.

[0094] It should be noted that since the brain is composed of multiple functional regions, its functional connections can usually be represented by a graph structure. For this reason, the present invention converts the fMRI data into a graph model with the help of the AAL whole-brain template. Specifically, the present invention preprocesses the acquired fMRI image data to obtain 130 time series of 116 brain regions divided by AAL to construct the node features (feature matrix) of the graph, and uses the correlation between brain regions to construct the edge features (adjacency matrix) of the graph. Further, the method for obtaining the graph model corresponding to the functional magnetic resonance imaging of the patient is as follows:

[0095] The functional magnetic resonance imaging of the patient is divided into 116 brain regions through the AAL whole-brain template;

[0096] According to the blood oxygenation level-dependent signals of each brain region at 130 sampling time points, a node feature matrix X of the graph model is constructed 116×130 ={x ij}, where x ij represents the blood oxygenation level-dependent signal of the i-th brain region at the j-th sampling time point, thereby representing its functional activity intensity, where i = 1, 2, …, 116 and j = 1, 2, …, 130;

[0097] According to the Pearson correlation coefficient between the blood oxygenation level-dependent signals of any two brain regions corresponding to 130 sampling time points, an adjacency matrix A of the graph model is constructed 116×116 ={x ab}, x ab represents the Pearson correlation coefficient between the blood oxygenation level-dependent signals of 130 sampling time points corresponding to the a-th brain region and the blood oxygenation level-dependent signals of 130 sampling time points corresponding to the b-th brain region, thereby representing the functional connection situation between brain regions, where a = 1, 2, …, 116 and b = 1, 2, …, 116;

[0098] The node feature matrix X 116×130 ={x ij} and the adjacency matrix A 116×116 ={x ab} jointly constitute the graph model corresponding to the patient's functional magnetic resonance imaging.

[0099] Furthermore, since DNA data consists of high-dimensional long sequences arranged by ATCG and there are significant differences in distribution and dimension from imaging data, traditional analysis methods usually have limitations. For example, most existing studies first annotate the data as genes or select fixed-length loci for analysis of disease risk genes. These methods not only rely on experience but may also lead to information loss. To solve the problem of excessive DNA data dimension, the present invention adopts the K-mer method to divide the long gene sequences into multiple short sequences of length K; specifically, the method for obtaining the graph model corresponding to the patient's gene sequence is as follows:

[0100] Assume that the length of the patient's gene sequence is L, and it is divided into L - K + 1 K-mer short sequences of length K, where the total number of all possible types of K-mer short sequences is 4 K types;

[0101] According to the frequencies of various types in the L - K + 1 K-mer short sequences, a node feature matrix corresponding to the graph model is constructed where y iDenote the number of occurrences of the \(i\)-th \(K\)-mer short sequence among \(L - K + 1\) \(K\)-mer short sequences, where \(i = 1, 2, \ldots, 4\). K ;

[0102] Construct the adjacency matrix corresponding to the graph model according to the correlation between any two \(K\)-mer short sequences. Among them, if \(y\) ab \(= n\), it means that the \(a\)-th \(K\)-mer short sequence and the \(b\)-th \(K\)-mer short sequence appear \(n\) times in the form of consecutive \(K\)-mer short sequences among the \(L - K + 1\) \(K\)-mer short sequences, and \(a = 1, 2, \ldots, 4\) K , \(b = 1, 2, \ldots, 4\) K ;

[0103] The graph model corresponding to the patient's gene sequence is jointly constituted by the node feature matrix and the adjacency matrix .

[0104] That is to say, the present invention can divide a gene sequence of length \(L\) into \(L - K + 1\) short sequences of length \(K\). Since there are four groups of bases (A, T, C, G) in genes, the total number of short sequences that can be formed is 4 k species. For the sake of convenient description, take a simple example: a gene sequence is \(\{ATCG\}\), when \(K = 2\), it can be divided into three short sequences of length 2, \(\{AT\}\), \(\{TC\}\), \(\{CG\}\), and it can be known that when \(K = 2\), there are 4 2 \(= 16\) short sequences (AA, AT, AC, AG, TA, TT, TC, TG, CA, CT, CC, CG, GA, GT, GC, GG).

[0105] By calculating the occurrence frequency of each \(K\)-mer, the original long DNA sequence can be converted into a graph structure. This method can make full use of gene data information and facilitate the subsequent fusion with image data. Specifically, 4 k species of \(K\)-mers constitute 4 k nodes of the gene graph, where the \(K\)-mer frequency value is used to constitute the node features (feature matrix) of the graph, and the frequency of two consecutive \(K\)-mers is used to constitute the edge features (adjacency matrix) of the graph. Taking the above example, this gene sequence has a total of 16 nodes, among which the frequencies of the three nodes AT, TC, and CG are 1, and the others are 0, which constitutes the 16 * 1 feature matrix corresponding to this graph; the adjacency matrix represents the relationship between nodes and is a 16 * 16 positive symmetric matrix. Since AT and TC and TC and CG are continuously connected, the connection relationship of these two groups of node pairs is 1, and the others are 0.

[0106] In summary, the specific shape of the data constituting the gene graph is: the node feature matrix Dimension is 4 K ×1, where K represents the K-mer length, and 4 K represents the number of K-mer types. Since the feature matrix only records the frequency values of K-mers, its feature dimension is 1; the adjacency matrix Dimension is 4 K ×4 K , representing the correlation between K-mers. Since many short K-mer sequences do not appear (the feature value of this node is 0) or do not appear continuously (the adjacency value of the node pair is 0), both the feature matrix and the adjacency matrix are sparse matrices. The feature matrix and the adjacency matrix of the gene graph together constitute the graph G g .

[0107] It should be noted that in the research on multi-modal fusion for disease prediction, existing methods usually only focus on the synergistic effect between different modalities, while ignoring the unique impact of each modality on the occurrence of diseases. In addition to the common features between modalities, each modality itself also contains rich information, which is crucial for disease prediction and analysis. To solve this problem, the present invention introduces a decoupled representation method, which divides the features in multi-modal data into common representation and private representation, and captures the global features shared between modalities and the unique features within each modality respectively. The common representation, also known as the global feature, is used to capture the macroscopic information of the graph structure, including the long-distance dependence relationship between nodes and the extensive graph connection pattern. It can extract the shared patterns or information across modalities, and is particularly suitable for cross-modal data fusion, because these features can usually reflect the similarity and correlation between modalities. In multi-modal fusion, the global feature is the basis for integrating different modalities. The private representation, also known as the local feature, focuses on the features of each node and its neighborhood in a single-modal graph, and contains the detailed information within the modality. The private representation can reflect the unique structure or information in a specific modality, and is usually closely related to the characteristics of this modality. These local features not only retain the uniqueness within the modality, but also effectively avoid information loss, thereby enhancing the model's ability to capture the individual impact of each modality. In the implementation process of the present invention, considering the characteristic differences between fMRI and DNA data, the present invention designs feature extraction schemes for these two types of data respectively.

[0108] (1) Feature extraction of imaging data

[0109] To extract effective features from fMRI data, the present invention adopts a GCN-based method. The present invention divides the features of the imaging data into common representations (global features) and private representations (local features), and uses deep GCN and shallow GCN for extraction respectively. The common representation is extracted by deep GCN, which can capture long-range dependencies and global information across nodes in the graph structure. Specifically, deep GCN aggregates and transmits node information layer by layer through multiple convolutional layers, enabling the model to capture global features spanning multiple brain regions. In the present invention, 4-layer GCN is adopted to accumulate more context information, especially the global dependencies existing between distant nodes. To further enhance the ability of global feature extraction, the present invention introduces SC (Spectral Clustering), and optimizes feature extraction by calculating the Laplacian matrix of the graph. The Laplacian matrix is the core topological structure of the graph, which can help identify the community structure in the graph, thereby making the global feature extraction more accurate. The deep GCN module includes four layers of graph convolutional networks, and the convolutional operations of each layer of graph convolutional network are as follows:

[0110] H (l+1) =σ(LH (l) W (l) )

[0111] where H (l) is the node feature matrix of the l-th layer of graph convolutional network, H (l+1) is the node feature matrix of the (l + 1)-th layer of graph convolutional network, l = 0, 1, 2, 3, σ is the activation function, L is the Laplacian matrix, and L = D - A, D is the degree matrix determined according to the node feature matrix X 116×130 ={x ij}, A is the adjacency matrix A 116×116 ={x ab}, W (l) is the node weight matrix of the l-th layer of graph convolutional network. In this way, deep GCN gradually accumulates global information and generates the global representation H ic .

[0112] In addition, the present invention also uses shallow GCN to extract private representations, focusing on the local features of each node and its neighborhood in the graph. Shallow GCN can focus on the feature changes of adjacent nodes, so as to better retain the details and uniqueness within the modality. In the present invention, 2-layer GCN is used for the extraction of local features. Through fewer levels of convolutional operations, it ensures the capture of the independent features of each brain region, reflects the unique structure and information within the modality, and obtains the private local feature representation H ip. By combining deep GCN and shallow GCN, the present invention can effectively extract global and local features from fMRI data, thus providing a solid foundation for subsequent multi-modal data fusion.

[0113] (2) Gene data feature extraction

[0114] To extract effective features from gene data, the present invention adopts a specific method for sparse graph structures; due to the sparsity of DNA data, the eigenvalue of many nodes in the constructed graph is 0, and there are few associations between nodes, so special processing is required.

[0115] Based on this, the specific process of the RWD module for extracting the DNA global features of the graph model corresponding to the gene sequence of the patient is as follows:

[0116] First, extract the common representation (global feature) through random walk diffusion. Random walk diffusion can simulate the diffusion process of information on the entire graph and capture the global features of long-distance dependencies. Specifically, calculate the transition probability matrix P of the random walk using the degree matrix D and the adjacency matrix B of the graph model corresponding to the gene sequence of the patient:

[0117] P = D -1 B

[0118] where D is the degree matrix of each node in the adjacency matrix determined according to the adjacency matrix B, and B is the adjacency matrix representing the connection relationship between nodes;

[0119] The transition matrix P of the random walk represents the transition probability of a node in a single random walk. During the random walk process, the feature X of the gene graph node will gradually spread to neighboring nodes over time t. Use the transition probability matrix P to diffuse the node feature matrix, and the diffused node feature matrix H(t) is:

[0120] H(t) = P t Y

[0121] where P t is the transition probability matrix after t steps of random walk, and Y is the node feature matrix

[0122] Judge whether the difference between the node feature matrices obtained by two adjacent steps of diffusion is less than the set value. If so, the node feature matrix H(∞) obtained by the last step of diffusion is used as the final DNA global feature. If not, continue the diffusion.

[0123] That is to say, this diffusion process simulates the global behavior of information diffusion in the graph, and features are gradually transmitted from one node to other nodes. When the diffusion process reaches the steady state (i.e., when t→∞), the feature matrix converges to the steady-state solution, denoted as

[0124] H gc (∞) = P ∞ X

[0125] The state vector H gc at this time contains global feature information and can capture the global dependencies in the entire sparse graph structure.

[0126] Furthermore, in order to extract the private representation (local features), the present invention uses locally weighted sparse graph convolution. Due to the sparsity of the gene graph, the sparse convolution operation only aggregates non-zero neighbor nodes, thereby effectively processing sparse connections and retaining local important features. During the convolution process, a local weighting mechanism is introduced to adjust the convolution weights based on the local importance weights of the gene graph nodes; the specific process of the second shallow GCN module for extracting the DNA local features of the graph model corresponding to the gene sequence of the patient is as follows:

[0127] Use the degree matrix D determined by the adjacency matrix B to determine the weights of each node in the adjacency matrix B:

[0128]

[0129] where B′ ab is the weight of node a corresponding to the a-th K-mer short sequence, d a is the degree of node a, d b is the degree of node b corresponding to the b-th K-mer short sequence, ∑ b is the set of adjacent nodes having a connection relationship with this node a;

[0130] According to the weights B′ of each node ab perform local weighting on the node feature matrix to obtain the final DNA local feature H gp :

[0131]

[0132] where the weight matrix I is the identity matrix, σ is the activation function, Y is the node feature matrix W is the learnable weight matrix.

[0133] It should be noted that H gpEnsure that the convolution process focuses on important neighbor nodes, preserves local information, and effectively extracts the unique features of each node. This method can capture both global shared information and retain local uniqueness, thereby enhancing the integrity and accuracy of feature extraction.

[0134] It should be noted that in addition to using deep and shallow GCNs with fixed numbers of layers for feature extraction, it is possible to consider automatically adjusting the number of GCN layers according to the data complexity. For example, when the data dimension is high or there are more complex relationships, increase the number of GCN layers; while in simple scenarios, reduce the number of layers. This can ensure that the model is more adaptable, and this alternative idea is relatively straightforward and will not significantly change the model framework.

[0135] In addition, during the feature extraction process, in addition to separately extracting global and local features, different weights can be set for the two types of features, allowing the model to flexibly adjust these weights according to the data situation. This can make feature extraction more flexible and effectively meet the requirements of different modalities.

[0136] (3) Common representation fusion

[0137] Furthermore, during the common representation fusion process, gene data and imaging data represent different types of biological information. Therefore, if they are directly fused, their global feature similarities may be ignored. To solve this problem, the present invention adopts a similarity map alignment method to align their global features by calculating the node similarities between the gene and imaging modalities. This method ensures the complementarity of cross-modal data, enabling gene data and imaging data to work together during the fusion process and making full use of their respective information. Specifically, first, feature mapping of the common representations of genes and images is required. Since the shapes of the gene graph and the image graph are different and their feature space dimensions are inconsistent, the present invention maps the common representations of genes and images into the same dimensional space through linear transformation. This can ensure that they can be compared and operated in the same feature space during the subsequent alignment and fusion processes. This is achieved by a learnable similarity matrix W in the map alignment-based method. This matrix can be automatically adjusted to maximize the alignment of the features of genes and images. Through this process, the common representation features of genes and images will be aligned.

[0138] Based on this, the operation steps for the global fusion module to fuse the fMRI global features and the DNA global features to obtain the global representation are summarized as follows:

[0139] Map the fMRI global features and the DNA global features into the same dimensional space through linear transformation respectively to achieve the alignment of the fMRI global features and the DNA global features; the aligned fMRI global features and DNA global features are respectively represented as:

[0140] H gc ' = H gc ·W c and H ic ' = H ic ·W c T

[0141] wherein, W c is a learnable alignment matrix, and H gc ' and H ic ' are respectively the common representation feature matrices of the gene and the image after feature alignment.

[0142] The aligned fMRI global features and DNA global features are weighted and fused to obtain a global representation.

[0143] Through this alignment process, the features of genes and images can be aligned within the same feature space, ensuring that the similarity of cross-modal features is maximized. Finally, after alignment, the features of the aligned genes and images need to be weighted and fused. Through weighted fusion, the present invention can combine complementary information from genes and images within a unified feature space. This not only ensures the complementarity of cross-modal data but also improves the efficiency and effect of feature fusion, making the finally generated fused features perform better in disease prediction and diagnosis tasks.

[0144] It should be noted that in addition to using a dynamic attention mechanism for feature fusion, a fixed-weight fusion method can also be directly adopted. This method only needs to assign fixed weights to data of different modalities at the initial stage and does not need to dynamically adjust the weights during the fusion process. This alternative solution is simpler and more feasible and is suitable for certain specific datasets, especially when computing resources are limited or the features are relatively clear.

[0145] In addition, a linear fusion method can also be used to fuse the global features of the two modalities; in linear fusion, a complex attention mechanism is not required, and only simple linear weighted summation of the features of different modalities is needed. Although the effect may not be as good as dynamic fusion, in many application scenarios, linear fusion can already provide sufficiently good performance.

[0146] (4) Private Representation Fusion

[0147] Furthermore, during the multi-modal fusion process, in addition to fusing the global shared features, it is also necessary to retain the private features of each modality. To achieve this goal, the present invention adopts an adaptive weighted fusion method based on the attention mechanism, which retains the uniqueness of each modality by dynamically adjusting the weights of different modality features. First, the private representations of genes and images are transformed into a feature space of the same dimension through feature mapping to ensure that these private features can be effectively compared and fused. Subsequently, the private features of genes and images are adaptively weighted and adjusted using attention weights, and this adjustment ensures that the private features do not interfere with each other during the fusion process; finally, the adaptively weighted private features of genes and images are weighted and fused to ensure that the unique information of each modality is retained and the final fused features are more accurate in disease prediction.

[0148] Based on this, the local fusion module is used to fuse the fMRI local features and the DNA local features to obtain a local representation, specifically:

[0149] The fMRI local features and the DNA local features are respectively mapped into the same dimensional space through linear transformation to achieve the alignment of the fMRI local features and the DNA local features;

[0150] The aligned fMRI local features and DNA local features are adaptively weighted and fused to obtain a local representation; among them, the method for obtaining the adaptive weights corresponding to the fMRI local features and the DNA local features is:

[0151] H′ gp =W B ·H ip

[0152]

[0153] Among them, H′ gp is the adaptive weight corresponding to the DNA local features, H ip ′ is the adaptive weight corresponding to the fMRI local features, H ip is the aligned fMRI local features, H gp is the aligned DNA local features, T represents transpose, W B is the initial weight, W B is used to represent the importance weight distribution of the private features of each modality, and softmax(·) is the activation function. Through the attention matrix, the model can allocate appropriate weights to each modality according to the contribution degree of the private features.

[0154] Furthermore, the aggregation module is used to fuse the global representation and the local representation to obtain an overall representation, specifically:

[0155] The fMRI global features and DNA global features are respectively mapped into the same dimensional space through linear transformation to achieve the alignment of the fMRI global features and DNA global features;

[0156] Obtain the similarity between the aligned fMRI global features and DNA global features;

[0157] Fuse the similarities greater than the set threshold into the global representation, and fuse the similarities not greater than the set threshold into the local representation;

[0158] Overlay the global representation and the local representation after fusing the similarities to obtain the overall representation.

[0159] That is to say, in the present invention, in order to effectively fuse image data and gene data, it is first necessary to map the original map data corresponding to each modality into the same dimension. Specifically, the global features of each modality are extracted through global average pooling, and these features are mapped to the same shape to ensure that data of different modalities can be fused and compared in the same space. Next, calculate the global graph similarity between different modalities, set a threshold according to the similarity (for example, the threshold is set to 0.5 to distinguish strong or weak similarity parts), and construct a similarity mask based on this (1 for greater than the threshold, 0 for less than). This mask is used to mark the shared or significantly different parts between modalities. Finally, these similarity information are overlaid on the previously extracted common representation and private representation. Specifically, the similarity information greater than the threshold is fused into the common representation features, and the similarity information less than the threshold is fused into the private representation features. This operation ensures that the feature interaction of multimodal data can make full use of the correlation and uniqueness between modalities, thereby improving the effect of feature fusion. In this process, both the shared information in the global features is retained, and the modality uniqueness in the private features is also ensured.

[0160] (5) Integration and downstream analysis

[0161] In the process of multimodal data fusion, the global representation and the local representation respectively capture different levels of features. The global representation reflects the shared information across modalities, while the local representation retains the unique features within the modality. In order to effectively integrate these two types of information, the present invention adopts an adaptive weighted fusion method based on the attention mechanism. This mechanism can adaptively adjust their weights in the fusion process according to the importance of the global and local features, ensuring that both the shared global information between modalities and the uniqueness of each modality are fully reflected during the fusion process, and finally generating a fused representation.

[0162] After the feature integration is completed, the globally and locally fused representations are input into an MLP (Multi-Layer Perceptron) for a binary classification task. The multi-layer perceptron model further processes these fused features to ultimately predict the probability of having Alzheimer's disease, improving the accuracy and stability of disease diagnosis.

[0163] Furthermore, during the model training process, the present invention adopts a comprehensive loss function to optimize the learning of multi-modal fusion features. The loss function loss used in the training of the Alzheimer's disease diagnosis system is:

[0164] loss = αloss class + βloss same + γloss con

[0165] where loss class is the classification loss, α is the weight of the classification loss, loss same is the similarity loss, β is the weight of the similarity loss, loss con is the contrastive loss, and γ is the weight of the contrastive loss;

[0166] The calculation method of the classification loss loss class is:

[0167]

[0168] where N is the number of sample patients, y i is the true Alzheimer's disease status of the i-th sample patient, y i = 1 indicates having Alzheimer's disease, y i = 0 indicates not having Alzheimer's disease, and p i is the true probability of having Alzheimer's disease of the i-th sample patient output by the Alzheimer's disease diagnosis system; by minimizing this loss, the model can learn an accurate classification decision boundary;

[0169] To ensure the consistency of the common representation features of multi-modal data among different modalities, we introduce a similarity loss, namely the CMD loss (Central Moment Discrepancy). The CMD loss measures the difference in high-order statistics between the source data and the target data to ensure the alignment of multi-modal data in the feature space. The calculation method of the similarity loss loss same is:

[0170]

[0171] where μ kDenote the k-th order central moment, where K is the maximum order of calculation, and μ k (xs) is the k-th order central moment of the local representation, and μ k (xt) is the k-th order central moment of the global representation, and ||·|| 2 is the L2 norm; this loss realizes the similarity alignment of cross-modal data by reducing the differences between the local representation and the global representation in different order central moments;

[0172] It should be noted that in addition to using the CMD loss and the cross-entropy loss in the loss function, a simple L2 regularization can also be added to the loss function of the model. L2 regularization is a classical method that prevents overfitting by penalizing the model parameters. This alternative is relatively straightforward and easy to implement, and can improve the generalization ability of the model.

[0173] To further enhance the separation ability of the private representation, a contrastive loss is used to ensure the difference between the common representation and the private representation of multi-modal data. The contrastive loss is divided into two parts for inter-modal and intra-modal; the inter-modal contrastive loss is used to ensure the difference between the private representations of different modalities, and the intra-modal contrastive loss is used to separate the common representation and the private representation within the modality. Based on this, the contrastive loss loss con is calculated as follows:

[0174]

[0175] where rep1 i is the sparse representation corresponding to the global representation of the i-th sample patient, and rep2 i is the sparse representation corresponding to the local representation of the i-th sample patient, and margin is the set spacing. By setting the margin (spacing), it can be ensured that the distance between the private representations of different modalities and the common representation and the private representation within the same modality in the feature space is large enough to maintain the modality uniqueness and the feature decoupling within the modality.

[0176] Through the combination of the above three loss functions, the present invention not only optimizes the classification performance during the training process, but also ensures the consistency of cross-modal features and the independence of private features of each modality.

[0177] It should be noted that in order to better adjust the weights of different loss functions, cross-validation can be introduced during the model training process. By testing the model performance on multiple data sets, the weights of the cross-entropy loss, the CMD loss, and the contrastive loss can be automatically adjusted. This method does not change the core algorithm of the model, but only increases the flexibility in the use of the loss function.

[0178] In summary, a diagnostic system for Alzheimer's disease based on the fusion of imaging and genetic data provided by the present invention has the following advantages:

[0179] (1) Optimization of Multimodal Graph Construction - Structuring of K-mer Frequencies and Image Graph Features

[0180] Although the existing FAGCN method uses both image and gene data for fusion, it constructs a single network only by directly associating the similarity or correlation between brain regions and genes, which easily ignores the complex relationships between image and gene features in high-dimensional features.

[0181] In the present invention, two independent networks are constructed respectively from an image graph and a gene graph. The image graph is established based on the AAL template, and the connectivity between brain regions is represented by a correlation matrix; while for the gene graph, the K-mer method is used to efficiently convert the long DNA sequence information into a K-mer frequency node feature matrix, and the occurrence frequencies between K-mers are introduced to define the edge weights; therefore, the present invention effectively retains the structural information in the image and gene data, and provides a more refined node feature representation, laying a solid foundation for subsequent feature extraction and cross-modal fusion, and solving the bottleneck that traditional methods are difficult to retain all gene information.

[0182] (2) Deep Decoupled Representation of Feature Extraction - Private and Shared Representations of Global and Local Features

[0183] The graph convolution operation of the existing FAGCN method cannot fully capture the local features within the modality and the global dependencies across modalities, thus possibly ignoring the unique impacts of each modality on disease prediction.

[0184] In terms of feature extraction, the present invention innovatively introduces a decoupling mechanism of shared and private representations, extracts cross-modal global shared information through deep GCN, and captures unique local features within the modality through shallow GCN. Specifically, for the image graph, four layers of GCN are designed to deeply explore the long-range dependencies between brain regions, and the community structure information of the graph is extracted by combining Laplacian spectral clustering. For the gene graph, the random walk diffusion method is used to capture long-range global features, and local weighted sparse graph convolution is combined to extract local features.

[0185] Therefore, the present invention can effectively decouple between global and local features and ensure that each feature can play an independent role; compared with FAGCN, the present invention is more meticulous in capturing the specificities of image and gene data, retains both the modality uniqueness and fully extracts the shared information between modalities, thereby improving the overall expression ability of multi-modal features.

[0186] (3) Improvement of Feature Alignment and Multi-modal Feature Fusion - Similarity Map Alignment and Adaptive Attention Weighted Fusion

[0187] Although FAGCN realizes the feature fusion of image and gene data, it does not refine and distinguish each modality feature, and lacks the optimization of the spatial alignment of modality features.

[0188] In the process of multi-modal feature interaction of the present invention, first, similarity map alignment is adopted to align the common representations of genes and images in the same feature space, and the node alignment between genes and images is optimized through the similarity matrix W; at the same time, in the private representation fusion, the weights of each modality are adaptively weighted and adjusted through the attention mechanism to retain the private information of each modality, ensuring that both cross-modal shared features and the uniqueness of each modality are considered in the final feature fusion process.

[0189] Compared with the single fusion method of FAGCN, the alignment and weighting mechanism of the present invention significantly enhances the consistency of image and gene data in shared features, and effectively retains the unique information of modalities at the same time. This method avoids feature confusion, improves the expression accuracy and diversity of fused features, and helps to improve the accuracy of disease prediction.

[0190] Of course, the present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can certainly make various corresponding changes and deformations according to the present invention. However, these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.

Claims

1. An Alzheimer's disease diagnosis system based on image gene data fusion, characterized in that: Including deep GCN module, first shallow GCN module, second shallow GCN module, RWD module, global fusion module, local fusion module, aggregation module, multi-layer perceptron; The deep GCN module is used to extract the fMRI global features of the graph model corresponding to the patient's functional magnetic resonance imaging; The first shallow GCN module is used to extract fMRI local features of the graph model corresponding to the patient's functional magnetic resonance imaging; The RWD module is used to extract the DNA global features of the graph model corresponding to the patient's gene sequence; The second shallow GCN module is used to extract DNA local features of the graph model corresponding to the patient's gene sequence; The global fusion module is used to fuse fMRI global features and DNA global features to obtain a global representation; The local fusion module is used to fuse fMRI local features and DNA local features to obtain local representation; The aggregation module is used to fuse the global representation and the local representation to obtain the overall representation; The multilayer perceptron is used to determine the probability of a patient suffering from Alzheimer's disease based on the overall representation.

2. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 1, characterized in that: The method for obtaining the graphical model corresponding to the patient's functional magnetic resonance imaging is as follows: The patients' functional magnetic resonance imaging was divided into 116 brain regions using the AAL whole-brain template; The node feature matrix X of the graph model is constructed based on the blood oxygen level dependence signal of each brain region at 130 sampling time points. 116×130 ={x ij }, where x ij represents the blood oxygen level dependence signal of the i-th brain region at the j-th sampling time point, where i = 1, 2, ... 116, j = 1, 2, ... 130; The adjacency matrix A of the graph model is constructed based on the Pearson correlation coefficient between the blood oxygen level dependent signals at 130 sampling time points corresponding to any two brain regions. 116×116 ={x ab }, x ab represents the Pearson correlation coefficient between the blood oxygen level dependence signal at 130 sampling time points corresponding to the a-th brain region and the blood oxygen level dependence signal at 130 sampling time points corresponding to the b-th brain region, where a=1,2,…116, b=1,2,…116; By node feature matrix X 116×130 ={x ij } and the adjacency matrix A 116×116 ={x ab }Together they constitute the graphical model corresponding to the patient's functional magnetic resonance imaging.

3. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 2, characterized in that: The deep GCN module consists of four layers of graph convolutional networks. The convolution operations of each layer of graph convolutional networks are as follows: H (l+1) =σ(LH (l) W (l) ) Among them, H (l) is the node feature matrix of the l-th layer graph convolutional network, H (l+1) is the node feature matrix of the l+1th layer graph convolutional network, l=0,1,2,3, σ is the activation function, L is the Laplacian matrix, and L=DA, D is the node feature matrix X 116×130 ={x ij }Determined degree matrix, A is the adjacency matrix A 116×116 ={x ab },W (l) is the node weight matrix of the l-th layer graph convolutional network.

4. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 1, characterized in that: The method for obtaining the graph model corresponding to the patient's gene sequence is: Assuming that the length of the patient's gene sequence is L, it is divided into L-K+1 K-mer short sequences of length K, where all possible types of K-mer short sequences are 4 in total. K kind; Construct the node feature matrix Y4 corresponding to the graph model according to the frequency of occurrence of various classes in L-K+1 K-mer short sequences K ×1 ={y i }, where y i Indicates the number of times the i-th K-mer short sequence appears in L-K+1 K-mer short sequences, i = 1, 2, ..., 4 K ; Construct the adjacency matrix B4 corresponding to the graph model based on the correlation between any two K-mer short sequences K ×4 K ={y ab }, where if y ab = n, indicating that the a-th K-mer short sequence and the b-th K-mer short sequence appear n times in the L-K+1 K-mer short sequences in the form of continuous K-mer short sequences, and a = 1, 2, ... 4 K , b=1,2,…4 K ; From the node feature matrix Y4 K ×1 ={y i } and adjacency matrix B4 K ×4 K ={y ab } together constitute the graph model corresponding to the patient's genetic sequence.

5. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 4, characterized in that: The RWD module is used to extract the DNA global features of the graph model corresponding to the patient's gene sequence: The transition probability matrix P of the random walk is calculated using the degree matrix D and adjacency matrix B of the graph model corresponding to the patient's gene sequence: P=D -1 B Where D is the degree matrix determined according to the adjacency matrix B, and B is the adjacency matrix B4 K ×4 K ={y ab }; Use the transition probability matrix P to transform the node feature matrix Y4 K ×1 ={y i } is diffused, and the node feature matrix H(t) after diffusion is: H(t)=P t Y Among them, P t is the transition probability matrix after t steps of random walk, Y is the node feature matrix Y4 K ×1 ={y i }; Determine whether the difference between the node feature matrices obtained in two adjacent diffusion steps is less than the set value. If yes, the node feature matrix H(∞) obtained in the last diffusion step is used as the final DNA global feature. If no, continue to diffuse.

6. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 4, characterized in that: The second shallow GCN module is used to extract the DNA local features of the graph model corresponding to the patient's gene sequence: The degree matrix D determined by the adjacency matrix B is used to determine the weight of each node in the adjacency matrix B: Among them, B′ ab is the weight of node a corresponding to the a-th K-mer short sequence, d a is the degree of node a, d b is the degree of node b corresponding to the b-th K-mer short sequence, Σ b is the set of adjacent nodes connected to the node a; According to the weight B′ of each node ab Perform local weighting on the node feature matrix to obtain the final DNA local feature H gp : Among them, the weight matrix I is the identity matrix, σ is the activation function, and Y is the node feature matrix W is a learnable weight matrix.

7. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 1, characterized in that: The global fusion module is used to fuse the fMRI global features and the DNA global features to obtain a global representation as follows: The fMRI global features and the DNA global features are mapped to the same dimensional space respectively through linear transformation to achieve the alignment of the fMRI global features and the DNA global features; The aligned fMRI global features and DNA global features are weighted fused to obtain a global representation.

8. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 1, characterized in that: The local fusion module is used to fuse fMRI local features and DNA local features to obtain a local representation specifically as follows: The fMRI local features and the DNA local features are respectively mapped to the same dimensional space through linear transformation to achieve the alignment of the fMRI local features and the DNA local features; The aligned fMRI local features and DNA local features are adaptively weighted fused to obtain local representation; wherein the method for obtaining the adaptive weights corresponding to the fMRI local features and the DNA local features is: H′ gp =W B ·H ip Among them, H′ gp is the adaptive weight corresponding to the local feature of DNA, H ip ′ is the adaptive weight corresponding to the local feature of fMRI, H ip is the aligned fMRI local feature, H gp is the local feature of the aligned DNA, T represents transposition, and W B is the initial weight and softmax(·) is the activation function.

9. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 1, characterized in that: The aggregation module is used to fuse the global representation and the local representation to obtain the overall representation as follows: The fMRI global features and the DNA global features are mapped to the same dimensional space respectively through linear transformation to achieve the alignment of the fMRI global features and the DNA global features; Obtain the similarity between the aligned fMRI global features and the DNA global features; The similarities greater than the set threshold are fused into the global representation, and the similarities less than the set threshold are fused into the local representation; The global representation and local representation after fusion of similarity are superimposed to obtain the overall representation.

10. The Alzheimer's disease diagnosis system based on image gene data fusion as claimed in claim 1, characterized in that: The loss function used in the training of the Alzheimer's disease diagnosis system is: loss=αloss class +βloss same +βloss con Among them, loss class is the classification loss, α is the weight of the classification loss, loss same is the similarity loss, β is the weight of the similarity loss, loss con is the contrast loss, γ is the weight of the contrast loss; Classification loss class The calculation method is: Where N is the number of sample patients, y i is the actual Alzheimer's disease condition of the i-th sample patient, y i =1 indicates Alzheimer's disease, y i =0 means no Alzheimer's disease, p i The true probability of Alzheimer's disease for the i-th sample patient output by the Alzheimer's disease diagnosis system; Similarity loss same The calculation method is: Among them, μ k represents the kth order central moment, K is the maximum order of calculation, μ k (xs) is the kth order central moment of the local representation, μ k (xt) is the kth order central moment of the global representation, ||·||2 is the 2-norm; Contrast loss con The calculation method is: Among them, rep1 i is the sparse representation corresponding to the global representation of the i-th sample patient, rep2 i is the sparse representation corresponding to the local representation of the i-th sample patient, and margin is the set spacing.

Citation Information

Cited By

  • Spatial transcriptome region identification method driven by multi-modal graph fusion

    CN121354671A