Alzheimer disease auxiliary classification method and system, terminal and storage medium

By combining the characteristics of PET and sMRI, using backbone networks and attention mechanisms for feature weighting and fusion, the problem of inaccurate classification of Alzheimer's disease in the prior art is solved, and a more accurate disease classification is achieved.

CN120088527APending Publication Date: 2025-06-03SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510002299.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art cannot accurately classify Alzheimer's disease in combination with PET and sMRI, mainly due to the complex complementary relationship between different imaging mechanisms.

Method used

By obtaining the target positron emission tomography and structural magnetic resonance imaging, gray matter and white matter features are extracted and these features are input into the backbone network. The local attention module and the collaborative attention module are used for feature weighting and fusion, and finally a classification prediction result is generated through a multi-layer perceptron.

Benefits of technology

Accurate classification of Alzheimer's disease is achieved, and complementary features of different modalities are mined by capturing richer image features and enhancing the interaction between different modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088527A_ABST
    Figure CN120088527A_ABST
Patent Text Reader

Abstract

The invention discloses an Alzheimer's disease auxiliary classification method and system, a terminal and a storage medium, and the method comprises the steps: obtaining a target positron emission tomography image and a structural magnetic resonance image, obtaining gray matter and white matter from the structural magnetic resonance image, inputting the gray matter and white matter into a backbone network, outputting a first positron emission tomography imaging feature, a first grey matter feature and a first white matter feature; inputting the first positron emission tomography imaging feature, the first grey matter feature and the first white matter feature into a collaborative attention module to obtain a second positron emission tomography imaging feature, a second grey matter feature and a second white matter feature; performing fusion to obtain fusion features, inputting the fusion features into an enhanced feature fusion module, and obtaining multi-modal features through normalization and a multi-layer perceptron; and generating a classification prediction result according to the multi-modal features. According to the method, an accurate classification prediction result can be obtained according to the target positron emission tomography and the structural magnetic resonance imaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural networks, and in particular, to a method, a system, a terminal and a storage medium for assisting in classifying Alzheimer's disease. Background Art

[0002] Alzheimer's disease (AD) is one of the most common neurodegenerative diseases, mainly manifested by loss of cognitive ability. Early diagnosis of AD is crucial for timely intervention in disease development and disease management. Therefore, it is crucial to obtain the current stage of Alzheimer's disease.

[0003] However, due to the complex complementary relationship between PET (Positron Emission Tomography) and sMRI (Structural Magnetic Resonance Imaging) with different imaging mechanisms, it is currently impossible to accurately classify Alzheimer's disease by combining PET and sMRI.

[0004] Therefore, the prior art still needs to be improved and developed. Summary of the Invention

[0005] The main object of the present invention is to provide a method, a system, a terminal and a computer-readable storage medium for assisting in classifying Alzheimer's disease, aiming to solve the problem in the prior art that due to the complex complementary relationship between PET (Positron Emission Tomography) and sMRI (Structural Magnetic Resonance Imaging) with different imaging mechanisms, it is currently impossible to accurately classify Alzheimer's disease by combining PET and sMRI.

[0006] To achieve the above object, the present invention provides a method for assisting in classifying Alzheimer's disease, and the method for assisting in classifying Alzheimer's disease includes the following steps:

[0007] Obtain a target positron emission tomography image and a structural magnetic resonance imaging, and obtain gray matter and white matter from the structural magnetic resonance imaging. Input the target positron emission tomography image, the gray matter and the white matter into a backbone network, and output a first positron emission tomography image imaging feature, a first gray matter feature and a first white matter feature, wherein a local attention module is set in each layer of the backbone network;

[0008] Input the first positron emission tomography (PET) image features, the first gray matter features, and the first white matter features into the co-attention module for feature weighting to obtain the second PET image features, the second gray matter features, and the second white matter features;

[0009] Fuse the second PET image features, the second gray matter features, and the second white matter features to obtain fused features, and input the fused features into the enhanced feature fusion module. After normalization and a multi-layer perceptron, obtain multi-modal features;

[0010] Generate a classification prediction result based on the multi-modal features.

[0011] Optionally, the step of inputting the target PET image, the gray matter, and the white matter into the backbone network to output the first PET image features, the first gray matter features, and the first white matter features specifically includes:

[0012] Input the target PET image, the gray matter, and the white matter into the backbone network. After preprocessing through an upsampling layer and dimensionality reduction through a dimensionality reduction layer, obtain the dimensionality-reduced PET image features, the dimensionality-reduced gray matter features, and the dimensionality-reduced white matter features respectively;

[0013] Extract features from the dimensionality-reduced PET image features, the dimensionality-reduced gray matter features, and the dimensionality-reduced white matter features through four residual layers to obtain the first PET image features, the first gray matter features, and the first white matter features;

[0014] Among them, after feature extraction in each residual layer, use a local attention module to obtain a local weight map, multiply the local weight map by the features extracted in the current residual layer, and use the multiplication result as the input for the next residual layer.

[0015] Optionally, the co-attention module includes a first co-module and a second co-module. The step of inputting the first PET image features, the first gray matter features, and the first white matter features into the co-attention module for feature weighting to obtain the second PET image features, the second gray matter features, and the second white matter features specifically includes:

[0016] Input the first positron emission tomography (PET) image features, the first gray matter features, and the first white matter features into a first collaborative module to generate first PET image collaborative features, first gray matter collaborative features, and first white matter collaborative features, and at the same time input the first PET image features, the first gray matter features, and the first white matter features into a second collaborative module to generate second PET image collaborative features, second gray matter collaborative features, and second white matter collaborative features;

[0017] Generate second PET image features based on the first PET image features, the first PET image collaborative features, and the second PET image collaborative features;

[0018] Generate second gray matter features based on the first gray matter features, the first gray matter collaborative features, and the second gray matter collaborative features;

[0019] Generate second white matter features based on the first white matter features, the first white matter collaborative features, and the second white matter collaborative features.

[0020] Optionally, the step of inputting the first PET image features, the first gray matter features, and the first white matter features into a first collaborative module to generate first PET image collaborative features, first gray matter collaborative features, and first white matter collaborative features specifically includes:

[0021] Input the first PET image features, the first gray matter features, and the first white matter features into a first collaborative module, and perform compression through global average pooling to obtain a PET image resonance compression vector, a gray matter compression vector, and a white matter compression vector;

[0022] According to the PET image resonance compression vector, the gray matter compression vector, and the white matter compression vector, obtain a fusion vector through a fully connected layer in the first collaborative module, and obtain channel weights corresponding to the PET image resonance compression vector, the gray matter compression vector, and the white matter compression vector according to the fusion vector;

[0023] Multiply the PET image resonance compression vector, the gray matter compression vector, and the white matter compression vector element-wise with the corresponding channel weights respectively to obtain the first PET image collaborative features, the first gray matter collaborative features, and the first white matter collaborative features.

[0024] Optionally, inputting the first positron emission tomography (PET) image imaging feature, the first gray matter feature, and the first white matter feature into a second collaborative module to generate a second PET image collaborative feature, a second gray matter collaborative feature, and a second white matter collaborative feature specifically includes:

[0025] Input the first PET image imaging feature, the first gray matter feature, and the first white matter feature into the first collaborative module and then into the second collaborative module, compress and connect them through convolution operations to obtain a fused volume;

[0026] According to three convolutions in the second collaborative module, process the fused volume respectively to obtain the spatial weights corresponding to the first PET image imaging feature, the first gray matter feature, and the first white matter feature;

[0027] Multiply the first PET image imaging feature, the first gray matter feature, and the first white matter feature element-wise with the corresponding spatial weights respectively to obtain the second PET image collaborative feature, the second gray matter collaborative feature, and the second white matter collaborative feature.

[0028] In addition, to achieve the above object, the present invention also provides an Alzheimer's disease auxiliary classification system, wherein the Alzheimer's disease auxiliary classification system includes:

[0029] A feature extraction module, configured to obtain a target positron emission tomography (PET) image and a structural magnetic resonance imaging (MRI), obtain gray matter and white matter from the structural MRI, input the target PET image, the gray matter, and the white matter into a backbone network, and output a first PET image imaging feature, a first gray matter feature, and a first white matter feature, wherein a local attention module is set in each layer of the backbone network;

[0030] A modality complementary module, configured to input the first PET image imaging feature, the first gray matter feature, and the first white matter feature into a collaborative attention module for feature weighting to obtain a second PET image imaging feature, a second gray matter feature, and a second white matter feature;

[0031] A fusion enhancement module, configured to fuse the second PET image imaging feature, the second gray matter feature, and the second white matter feature to obtain a fused feature, input the fused feature into an enhanced feature fusion module, and obtain a multi-modal feature through normalization and a multi-layer perceptron;

[0032] An output module, configured to generate a classification prediction result according to the multi-modal feature.

[0033] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an Alzheimer's disease auxiliary classification program stored in the memory and executable on the processor, and when the Alzheimer's disease auxiliary classification program is executed by the processor, the steps of the Alzheimer's disease auxiliary classification method as described above are implemented.

[0034] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an Alzheimer's disease auxiliary classification program, and when the Alzheimer's disease auxiliary classification program is executed by a processor, the steps of the Alzheimer's disease auxiliary classification method as described above are implemented.

[0035] In the present invention, a target positron emission tomography image and a structural magnetic resonance imaging are obtained, and gray matter and white matter are obtained from the structural magnetic resonance imaging. The target positron emission tomography image, the gray matter and the white matter are input into a backbone network, and a first positron emission tomography image imaging feature, a first gray matter feature and a first white matter feature are output, wherein a local attention module is set in each layer of the backbone network; the first positron emission tomography image imaging feature, the first gray matter feature and the first white matter feature are input into a collaborative attention module, and feature weighting is performed to obtain a second positron emission tomography image imaging feature, a second gray matter feature and a second white matter feature; the second positron emission tomography image feature, the second gray matter feature and the second white matter feature are fused to obtain a fused feature, and the fused feature is input into an enhanced feature fusion module, and after normalization and multi-layer perceptron, a multimodal feature is obtained; and a classification prediction result is generated according to the multimodal feature. In the Alzheimer's disease auxiliary classification method of the present invention, the features in PET and sMRI are combined with the advantages of convolutional networks and attention mechanisms to capture richer image features, and by introducing a local attention module, the focus on local related features is strengthened. A collaborative attention module is used to enhance the interaction between different modalities and explore the complementary features of different modalities, so that accurate classification prediction results can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flow chart of a preferred embodiment of the Alzheimer's disease auxiliary classification method of the present invention;

[0037] Figure 2 is a schematic diagram of the structure of the overall network in the auxiliary classification method for Alzheimer's disease of the present invention;

[0038] Figure 3 is a schematic diagram of the structure of the collaborative attention module in the auxiliary classification method for Alzheimer's disease of the present invention;

[0039] Figure 4 It is a structural diagram of a preferred embodiment of the Alzheimer's disease auxiliary classification system of the present invention;

[0040] Figure 5 It is a structural diagram of a preferred embodiment of the terminal of the present invention. Specific embodiments

[0041] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the present invention will be further described in detail below with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] Alzheimer's disease is one of the most common neurodegenerative diseases, mainly manifested as loss of cognitive ability. Early AD diagnosis is crucial for timely intervention in disease development and disease management. Therefore, it is crucial to obtain the current stage of Alzheimer's disease. However, due to the complex complementary relationship between PET and sMRI with different imaging mechanisms, the current fusion methods of the two images in research, such as simple addition or splicing, cannot make full use of their complementary information, resulting in the inability to accurately classify Alzheimer's disease by combining PET and sMRI at present.

[0043] To address one or more of the above problems, the present invention obtains a target positron emission tomography (PET) image and a structural magnetic resonance imaging (sMRI), obtains gray matter and white matter from the sMRI, inputs the target PET image, the gray matter and the white matter into a backbone network, and outputs a first PET image imaging feature, a first gray matter feature and a first white matter feature. Among them, a local attention module is set in each layer of the backbone network; the first PET image imaging feature, the first gray matter feature and the first white matter feature are input into a collaborative attention module for feature weighting to obtain a second PET image imaging feature, a second gray matter feature and a second white matter feature; the second PET image imaging feature, the second gray matter feature and the second white matter feature are fused to obtain a fusion feature, and the fusion feature is input into an enhanced feature fusion module, and after normalization and a multi-layer perceptron, a multi-modal feature is obtained; a classification prediction result is generated according to the multi-modal feature.

[0044] The Alzheimer's disease auxiliary classification method according to a preferred embodiment of the present invention, as Figure 1 shown, the Alzheimer's disease auxiliary classification method includes the following steps:

[0045] Step S10: Obtain a target positron emission tomography (PET) image and a structural magnetic resonance imaging (MRI), and obtain gray matter and white matter from the structural MRI. Input the target PET image, the gray matter, and the white matter into a backbone network, and output a first PET image imaging feature, a first gray matter feature, and a first white matter feature. Among them, a local attention module is set in each layer of the backbone network.

[0046] Specifically, after obtaining the corresponding structural MRI in the present invention, gray matter (GM) and white matter (WM) are obtained, and feature information is extracted through the correspondence between gray matter and white matter.

[0047] Further, the inputting the target PET image, the gray matter, and the white matter into the backbone network to output a first PET image imaging feature, a first gray matter feature, and a first white matter feature specifically includes:

[0048] Input the target PET image, the gray matter, and the white matter into the backbone network. After preprocessing through an upsampling layer, perform dimensionality reduction through a dimensionality reduction layer to respectively obtain a dimensionality-reduced PET image imaging feature, a dimensionality-reduced gray matter feature, and a dimensionality-reduced white matter feature;

[0049] Perform feature extraction on the dimensionality-reduced PET image imaging feature, the dimensionality-reduced gray matter feature, and the dimensionality-reduced white matter feature through four residual layers to obtain a first PET image imaging feature, a first gray matter feature, and a first white matter feature;

[0050] Among them, after feature extraction in each residual layer, a local attention module is used to obtain a local weight map, multiply the local weight map by the feature extracted in the current residual layer, and use the multiplied result as the input of the next residual layer.

[0051] Specifically, in the present invention, the backbone network is designed as a 3D architecture based on Resnet-18. The network architecture includes an upsampling layer, a dimensionality reduction layer, and a residual layer. Features of three modes, namely GM, WM, and PET, are extracted through the 3D ResNet-18 backbone network. And, in the present invention, to optimize feature fusion, the 3D ResNet channels are set to [16, 32, 64, 128] to adapt to the feature extraction of 3D data.

[0052] Specifically as Figure 2As shown, after the corresponding target positron emission tomography (PET) image, the gray matter, and the white matter are input into the backbone network, they are first preprocessed by an upsampling layer, and then multi-modal features are extracted through a dimensionality reduction layer and four residual layers. The dimensionality reduction layer performs corresponding dimensionality reduction, and the four residual layers are shown in the figure. As can be seen in the figure, after GM, WM, and PET are processed by the four residual layers, the first PET image imaging feature F is finally obtained. p The first gray matter feature F g and the first white matter feature F w , where where d, h, w, and c represent depth, height, width, and channels respectively, and sigmoid shown in the figure is a non-linear function, Multiplication represents multiplication, and Element-Wise Addition represents element-wise addition.

[0053] After each residual layer is processed, in order to guide the model to focus on the important information part of the input features, a local attention module (Local Attention Module, LA) is added after each residual layer to promote local feature extraction. After the residual layer is processed, for the output of the residual layer, by calculating the local weight map and applying it to the current residual layer output, important features are strengthened and unimportant features are weakened; the obtained local weight map is multiplied by the current residual layer and used as the input of the next residual layer.

[0054] Furthermore, the processing process of the local attention module can be expressed as:

[0055]

[0056] where F i represents the output of the i-th layer, i represents the i-th residual layer (i = 1, 2, 3, 4), ⊙ is element-wise multiplication, σ represents the Sigmoid function, Conv1 is a 1×1×1 convolution, ReLU represents the ReLU activation function, BN represents batch normalization, Conv3 is a 3×3×3 convolution, is the result of multiplying the corresponding local weight map by the current residual layer.

[0057] Step S20: Input the first PET image imaging feature, the first gray matter feature, and the first white matter feature into the collaborative attention module for feature weighting to obtain the second PET image imaging feature, the second gray matter feature, and the second white matter feature.

[0058] Specifically, in the present invention, the obtained first positron emission tomography (PET) image features, the first gray matter features, and the first white matter features are input into a collaborative attention module, so as to obtain corresponding second PET image features, second gray matter features, and second white matter features.

[0059] Further, the collaborative attention module includes a first collaborative module and a second collaborative module. The step of inputting the first PET image features, the first gray matter features, and the first white matter features into the collaborative attention module for feature weighting to obtain the second PET image features, second gray matter features, and second white matter features specifically includes:

[0060] Input the first PET image features, the first gray matter features, and the first white matter features into the first collaborative module to generate first PET image collaborative features, first gray matter collaborative features, and first white matter collaborative features, and at the same time input the first PET image features, the first gray matter features, and the first white matter features into the second collaborative module to generate second PET image collaborative features, second gray matter collaborative features, and second white matter collaborative features;

[0061] Generate second PET image features according to the first PET image features, the first PET image collaborative features, and the second PET image collaborative features;

[0062] Generate second gray matter features according to the first gray matter features, the first gray matter collaborative features, and the second gray matter collaborative features;

[0063] Generate second white matter features according to the first white matter features, the first white matter collaborative features, and the second white matter collaborative features.

[0064] Specifically, as Figure 2 shown, input the first PET image features, the first gray matter features, and the first white matter features into a Collaborative Attention Module (CA) for feature weighting, and the corresponding second PET image features F p_CA , second gray matter features F g_CA , and second white matter features F w_CA. In the collaborative attention module, there are a first collaborative module CFuse and a second collaborative module SFuse. The first positron emission tomography (PET) imaging feature, the first gray matter feature, and the first white matter feature are respectively input into the first collaborative module and the second collaborative module. Each module outputs the features corresponding to each feature and adds them together to obtain the second PET imaging feature, the second gray matter feature, and the second white matter feature.

[0065] Among them, the CFuse module strengthens the complementarity of different modality feature maps at the channel level through a channel-level fusion strategy, thereby highlighting the channel information that is more critical for the classification task; while the SFuse module focuses on feature fusion at the spatial level, can identify and emphasize the important features of different modalities in spatial positions, and realizes the effective integration of spatial information.

[0066] Furthermore, the input of the first PET imaging feature, the first gray matter feature, and the first white matter feature into the first collaborative module to generate the first PET collaborative feature, the first gray matter collaborative feature, and the first white matter collaborative feature specifically includes:

[0067] Input the first PET imaging feature, the first gray matter feature, and the first white matter feature into the first collaborative module, and compress them through global average pooling to obtain the PET resonance compression vector, the gray matter compression vector, and the white matter compression vector;

[0068] According to the PET resonance compression vector, the gray matter compression vector, and the white matter compression vector, through the fully connected layer in the first collaborative module, obtain the fusion vector, and obtain the channel weights corresponding to the PET resonance compression vector, the gray matter compression vector, and the white matter compression vector according to the fusion vector;

[0069] Multiply the PET resonance compression vector, the gray matter compression vector, and the white matter compression vector element-wise with the corresponding channel weights respectively to obtain the first PET collaborative feature, the first gray matter collaborative feature, and the first white matter collaborative feature.

[0070] Specifically, as Figure 3 shown, in the first collaborative module CFuse of the present invention, that is, the CFuse module as Figure 3 shown in (a) in, for the first PET imaging feature F p , the first gray matter feature F g and the first white matter feature F w, first compress through global average pooling to obtain the positron emission tomography resonance compression vector C p , the gray matter compression vector C g and the white matter compression vector C w , where After that, splice and pass through a fully connected layer to generate a fusion vector C fuse , through this fusion vector C fuse The channel weights corresponding to each channel can be obtained correspondingly, so that multiplying the corresponding channel weights and the corresponding fusion vector element by element can obtain the corresponding first positron emission tomography co-feature The first gray matter co-feature and the first white matter co-feature

[0071] Furthermore, C fuse = ω[C g , C w , C p +b;

[0072]

[0073] Among them, F m represents the corresponding first positron emission tomography imaging feature, the first gray matter feature and the first white matter feature, (where m ∈ {p, g, w}) represents the corresponding first positron emission tomography co-feature, the first gray matter co-feature and the first white matter co-feature, [·] represents the concatenation operator, ω, ω m and b, b m represent the corresponding weights and biases, and ⊙ is the element-wise multiplication.

[0074] Inputting the first positron emission tomography imaging feature, the first gray matter feature and the first white matter feature into the second co-module to generate the second positron emission tomography co-feature, the second gray matter co-feature and the second white matter co-feature specifically includes:

[0075] Input the first positron emission tomography imaging feature, the first gray matter feature and the first white matter feature into the first co-module and then into the second co-module, compress and connect through convolution operations to obtain a fusion volume;

[0076] According to the three convolutions in the second co-module, process the fusion volume respectively to obtain the spatial weights corresponding to the first positron emission tomography imaging feature, the first gray matter feature and the first white matter feature;

[0077] Multiply the first positron emission tomography (PET) image features, the first gray matter features, and the first white matter features element-wise with the corresponding spatial weights to obtain the second PET image collaborative features, the second gray matter collaborative features, and the second white matter collaborative features.

[0078] Specifically, as Figure 3 shown, in the present invention, the second collaborative module SFuse, i.e., the SFuse module as Figure 3 shown in (b) in the figure, for the first PET image features F p , the first gray matter features F g , and the first white matter features F w , they are first compressed into volumes respectively through 1×1×1 convolution operations , and then a fused volume S fuse is obtained through concatenation and further 1×1×1 convolution. The fused volume S fuse respectively passes through three 3×3×3 convolutions to generate spatial weights, and then the spatial weights are applied back to the input F m for element-wise multiplication respectively, to obtain (where m ∈ {p, g, w}) representing the corresponding second PET image collaborative features, the second gray matter collaborative features, and the second white matter collaborative features.

[0079] Furthermore, S fuse = K⊙[S g , S w , S p + B;

[0080]

[0081] where B, B m represent the corresponding weights and biases, K, K m are convolution kernels (where m ∈ {p, g, w}), and σ(·) is the sigmoid function.

[0082] Step S30: Fuse the second PET image features, the second gray matter features, and the second white matter features to obtain fused features, and input the fused features into the enhanced feature fusion module. After normalization and a multi-layer perceptron, multi-modal features are obtained.

[0083] Specifically, in the present invention, in order to effectively capture multi-modal long-distance features and generate distinctive representations, an Enhanced Feature Fusion Module (EFF) with spatial attention function is introduced. This module utilizes a lightweight Lite Transformer (LT) module, which can maintain powerful performance with fewer parameters and thus fits the classification task.

[0084] Furthermore, inputting the fusion feature into the enhanced feature fusion module, through normalization and a multi-layer perceptron, multi-modal features are obtained, specifically including:

[0085] Input the fusion feature into the enhanced feature fusion module, the normalization layer performs normalization, and two convolutional layers are used for processing to obtain the query, key, and value;

[0086] Input the query, key, and value into the multi-head attention mechanism in the enhanced feature fusion module to calculate the weighted value, and perform a residual connection with the fusion feature. After passing through normalization and a multi-layer perceptron for processing and non-linear transformation, the multi-modal features are obtained.

[0087] In the present invention, first add the second positron emission tomography imaging feature, the second gray matter feature, and the second white matter feature for fusion to obtain the fusion feature F fuse , and then input the fusion feature into the enhanced feature fusion module.

[0088] The enhanced feature fusion module uses a Layer Norm layer to process the fusion feature to achieve normalization. Then, two convolutional layers are used to generate the query (Q), key (K), and value (V), and the multi-head attention mechanism is applied to calculate the weighted value, thereby capturing the long-distance dependence relationship between features; the output obtained from the multi-head attention mechanism is combined with the fusion feature through a residual connection, and then through normalization and a multi-layer perceptron (MLP) for non-linear transformation to learn the complex relationship between features, and multi-modal features are obtained.

[0089] Furthermore, the processing process of the enhanced feature fusion module can be expressed as:

[0090]

[0091] Among them, Concat(·) represents the concatenation operation, Softmax represents the Softmax activation function, Q i , K i , V i are respectively the query, key, and value of the i-th head in the multi-head attention, d h =d / h is the dimension of the output feature of each head, h is the total number of heads, and the input is F fuseis the result of adding the CA outputs. MLP(·) represents a fully connected layer, and LN(·) represents a normalization layer. Finally, the multi-modal feature representation is obtained as F out 。

[0092] Step S40: Generate a classification prediction result based on the multi-modal feature.

[0093] The feature map output by the EFF module maps the features to the classification space through an average pooling layer (AvgPool) and a fully connected layer (FC). Finally, the network outputs the prediction probabilities for three classes, corresponding to AD, MCI (Mild Cognitive Impairment), and CN (Cognitively Normal), respectively.

[0094] Furthermore, the generating of the classification prediction result based on the multi-modal feature specifically includes:

[0095] The multi-modal feature passes through an average pooling layer and a fully connected layer, mapping the corresponding features to the classification space to obtain the prediction probabilities for three preset classes;

[0096] Output the prediction probabilities for the three classes as the classification prediction result.

[0097] Specifically, the three preset classes are AD, MCI, and CN. The present invention obtains the prediction probabilities corresponding to these three classes, and through this result, it can assist the user in judging the corresponding situation of the current target positron emission tomography image and structural magnetic resonance imaging.

[0098] In the present invention, when training the entire network model, imaging data and gene data from a database are used, such as data from the Alzheimer's Disease Neuroimaging Initiative, and the corresponding AD, MCI, and CN are used as corresponding labels to train the entire network. When the training meets the requirements, the overall network in the present invention is obtained.

[0099] In an embodiment of the present invention, data in the Alzheimer's Disease Neuroimaging Initiative database is obtained, and the corresponding sMRI and PET for each data are obtained. The dataset is preprocessed using SPM12 and CAT, and GM and WM are segmented from the sMRI. Finally, the preprocessed sMRI and PET images have the same size, i.e., 128×128×128. When training the network in the present invention using the corresponding constructed training set, it is trained using Pytorch on a single GPU with 24GB of memory, and the learning rate, batch size, and number of iterations are set to 10 -4, 16 and 40, and the Adam optimizer was used to update the model parameters; a 10-fold cross validation was used, and the mean and standard deviation of each fold were taken to quantitatively evaluate the model performance.

[0100] Among them, the evaluation indicators include accuracy, sensitivity, specificity, F1-score and area under the curve (AUC):

[0101]

[0102] Among them, TP (true positive), TN (true negative), FP (false positive) and FN (false negative) are the number of true positive, true negative, false positive and false negative samples respectively.

[0103] When the evaluation index corresponding to the result obtained after training meets the preset value, the training is terminated. Through this process, the present invention can realize the training of the network.

[0104] The present invention obtains a target positron emission tomography image and a structural magnetic resonance imaging, obtains gray matter and white matter from the structural magnetic resonance imaging, inputs the target positron emission tomography image, the gray matter and the white matter into a backbone network, and outputs a first positron emission tomography image imaging feature, a first gray matter feature and a first white matter feature, wherein a local attention module is set at each layer in the backbone network; the first positron emission tomography image imaging feature, the first gray matter feature and the first white matter feature are input into a collaborative attention module, feature weighting is performed to obtain a second positron emission tomography image imaging feature, a second gray matter feature and a second white matter feature; the second positron emission tomography image feature, the second gray matter feature and the second white matter feature are fused to obtain a fused feature, the fused feature is input into an enhanced feature fusion module, and multimodal features are obtained through normalization and a multi-layer perceptron; and a classification prediction result is generated according to the multimodal feature. In the Alzheimer's disease auxiliary classification method of the present invention, the features in PET and sMRI are combined with the advantages of convolutional networks and attention mechanisms to capture richer image features, and by introducing a local attention module, the focus on local related features is strengthened. A collaborative attention module is used to enhance the interaction between different modalities and explore the complementary features of different modalities, so that accurate classification prediction results can be obtained.

[0105] Furthermore, if Figure 4As shown above, based on the above Alzheimer's disease auxiliary classification method, the present invention also correspondingly provides an Alzheimer's disease auxiliary classification system, wherein the Alzheimer's disease auxiliary classification system includes:

[0106] A feature extraction module 41, configured to obtain a target positron emission tomography (PET) image and a structural magnetic resonance imaging (MRI), obtain gray matter and white matter from the structural MRI, input the target PET image, the gray matter, and the white matter into a backbone network, and output a first PET image imaging feature, a first gray matter feature, and a first white matter feature, wherein a local attention module is set in each layer of the backbone network;

[0107] A modality complementary module 42, configured to input the first PET image imaging feature, the first gray matter feature, and the first white matter feature into a collaborative attention module for feature weighting to obtain a second PET image imaging feature, a second gray matter feature, and a second white matter feature;

[0108] A fusion enhancement module 43, configured to fuse the second PET image imaging feature, the second gray matter feature, and the second white matter feature to obtain a fusion feature, input the fusion feature into an enhanced feature fusion module, and obtain multi-modal features through normalization and a multi-layer perceptron;

[0109] An output module 44, configured to generate a classification prediction result according to the multi-modal features.

[0110] Further, the inputting the target PET image, the gray matter, and the white matter into the backbone network and outputting a first PET image imaging feature, a first gray matter feature, and a first white matter feature specifically includes:

[0111] Inputting the target PET image, the gray matter, and the white matter into the backbone network, performing preprocessing through an upsampling layer, and performing dimensionality reduction through a dimensionality reduction layer to respectively obtain a dimensionality-reduced PET image imaging feature, a dimensionality-reduced gray matter feature, and a dimensionality-reduced white matter feature;

[0112] Performing feature extraction on the dimensionality-reduced PET image imaging feature, the dimensionality-reduced gray matter feature, and the dimensionality-reduced white matter feature through four residual layers to obtain a first PET image imaging feature, a first gray matter feature, and a first white matter feature;

[0113] Wherein, after feature extraction in each residual layer, a local attention module is used to obtain a local weight map, the local weight map is multiplied by the feature extracted by the current residual layer, and the multiplication result is used as the input of the next residual layer.

[0114] The co-attention module includes a first co-module and a second co-module. Inputting the first positron emission tomography (PET) image features, the first gray matter features, and the first white matter features into the co-attention module for feature weighting to obtain the second PET image features, the second gray matter features, and the second white matter features specifically includes:

[0115] Inputting the first PET image features, the first gray matter features, and the first white matter features into the first co-module to generate the first PET image co-features, the first gray matter co-features, and the first white matter co-features, and simultaneously inputting the first PET image features, the first gray matter features, and the first white matter features into the second co-module to generate the second PET image co-features, the second gray matter co-features, and the second white matter co-features;

[0116] Generating the second PET image features according to the first PET image features, the first PET image co-features, and the second PET image co-features;

[0117] Generating the second gray matter features according to the first gray matter features, the first gray matter co-features, and the second gray matter co-features;

[0118] Generating the second white matter features according to the first white matter features, the first white matter co-features, and the second white matter co-features.

[0119] Further, as Figure 5 shown, based on the above Alzheimer's disease assisted classification method and system, the present invention also correspondingly provides a terminal, which includes a processor 10, a memory 20, and a display 30. Figure 5 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0120] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In some other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit of the terminal and the external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code of the installed terminal, etc. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, an Alzheimer's disease assisted classification program 40 is stored on the memory 20, and the Alzheimer's disease assisted classification program 40 can be executed by the processor 10, thereby implementing the Alzheimer's disease assisted classification method in the present invention.

[0121] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing the Alzheimer's disease assisted classification method, etc.

[0122] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components 10 - 30 of the terminal communicate with each other through a system bus.

[0123] In one embodiment, when the processor 10 executes the Alzheimer's disease assisted classification program 40 in the memory 20, the steps of the above Alzheimer's disease assisted classification method are implemented.

[0124] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an Alzheimer's disease assisted classification program, and when the Alzheimer's disease assisted classification program is executed by a processor, the following steps are implemented:

[0125] Obtain a target positron emission tomography (PET) image and a structural magnetic resonance imaging (MRI), and obtain gray matter and white matter from the structural MRI. Input the target PET image, the gray matter, and the white matter into a backbone network, and output a first PET image imaging feature, a first gray matter feature, and a first white matter feature. Among them, a local attention module is set in each layer of the backbone network;

[0126] Input the first PET image imaging feature, the first gray matter feature, and the first white matter feature into a collaborative attention module for feature weighting to obtain a second PET image imaging feature, a second gray matter feature, and a second white matter feature;

[0127] Fuse the second PET image imaging feature, the second gray matter feature, and the second white matter feature to obtain a fused feature. Input the fused feature into an enhanced feature fusion module, and through normalization and a multi-layer perceptron, obtain a multi-modal feature;

[0128] Generate a classification prediction result according to the multi-modal feature.

[0129] Among them, the inputting the target PET image, the gray matter, and the white matter into the backbone network to output a first PET image imaging feature, a first gray matter feature, and a first white matter feature specifically includes:

[0130] Input the target PET image, the gray matter, and the white matter into the backbone network. After preprocessing through an upsampling layer, perform dimensionality reduction through a dimensionality reduction layer to respectively obtain a dimensionality-reduced PET image imaging feature, a dimensionality-reduced gray matter feature, and a dimensionality-reduced white matter feature;

[0131] Extract features of the dimensionality-reduced PET image imaging feature, the dimensionality-reduced gray matter feature, and the dimensionality-reduced white matter feature through four residual layers to obtain a first PET image imaging feature, a first gray matter feature, and a first white matter feature;

[0132] Among them, after extracting features in each residual layer, use a local attention module to obtain a local weight map, multiply the local weight map by the features extracted by the current residual layer, and use the multiplication result as the input of the next residual layer.

[0133] Among them, the collaborative attention module includes a first collaborative module and a second collaborative module. The inputting the first PET image imaging feature, the first gray matter feature, and the first white matter feature into the collaborative attention module for feature weighting to obtain a second PET image imaging feature, a second gray matter feature, and a second white matter feature specifically includes:

[0134] Input the first positron emission tomography (PET) image features, the first gray matter features, and the first white matter features into the first collaborative module to generate the first PET image collaborative features, the first gray matter collaborative features, and the first white matter collaborative features. At the same time, input the first PET image features, the first gray matter features, and the first white matter features into the second collaborative module to generate the second PET image collaborative features, the second gray matter collaborative features, and the second white matter collaborative features;

[0135] Generate the second PET image features based on the first PET image features, the first PET image collaborative features, and the second PET image collaborative features;

[0136] Generate the second gray matter features based on the first gray matter features, the first gray matter collaborative features, and the second gray matter collaborative features;

[0137] Generate the second white matter features based on the first white matter features, the first white matter collaborative features, and the second white matter collaborative features.

[0138] Among them, the step of inputting the first PET image features, the first gray matter features, and the first white matter features into the first collaborative module to generate the first PET image collaborative features, the first gray matter collaborative features, and the first white matter features specifically includes:

[0139] Input the first PET image features, the first gray matter features, and the first white matter features into the first collaborative module, and compress them through global average pooling to obtain the PET image resonance compression vector, the gray matter compression vector, and the white matter compression vector;

[0140] According to the PET image resonance compression vector, the gray matter compression vector, and the white matter compression vector, obtain a fusion vector through the fully connected layer in the first collaborative module, and obtain the channel weights corresponding to the PET image resonance compression vector, the gray matter compression vector, and the white matter compression vector according to the fusion vector;

[0141] Multiply the PET image resonance compression vector, the gray matter compression vector, and the white matter compression vector element-wise with the corresponding channel weights respectively to obtain the first PET image collaborative features, the first gray matter collaborative features, and the first white matter collaborative features.

[0142] Among them, inputting the first positron emission tomography (PET) image feature, the first gray matter feature, and the first white matter feature into a second collaborative module to generate a second PET image collaborative feature, a second gray matter collaborative feature, and a second white matter collaborative feature specifically includes:

[0143] Input the first PET image feature, the first gray matter feature, and the first white matter feature into the first collaborative module and then into the second collaborative module, compress and connect them through convolution operations to obtain a fused volume;

[0144] According to three convolutions in the second collaborative module, process the fused volume respectively to obtain the spatial weights corresponding to the first PET image feature, the first gray matter feature, and the first white matter feature;

[0145] Multiply the first PET image feature, the first gray matter feature, and the first white matter feature element-wise with the corresponding spatial weights respectively to obtain the second PET image collaborative feature, the second gray matter collaborative feature, and the second white matter collaborative feature.

[0146] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal including that element.

[0147] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0148] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for assisting classification of Alzheimer's disease, characterized in that: The Alzheimer's disease auxiliary classification method comprises: Obtain a target positron emission tomography image and a structural magnetic resonance imaging, and obtain gray matter and white matter from the structural magnetic resonance imaging, input the target positron emission tomography image, the gray matter and the white matter into a backbone network, and output a first positron emission tomography image feature, a first gray matter feature and a first white matter feature, wherein a local attention module is set at each layer in the backbone network; Inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a collaborative attention module, performing feature weighting, and obtaining a second positron emission tomography imaging feature, a second gray matter feature, and a second white matter feature; The second positron emission tomography imaging feature, the second gray matter feature and the second white matter feature are fused to obtain a fused feature, and the fused feature is input into an enhanced feature fusion module, and after normalization and multi-layer perceptron, a multi-modal feature is obtained; A classification prediction result is generated according to the multimodal features.

2. The Alzheimer's disease auxiliary classification method according to claim 1, characterized in that: The step of inputting the target positron emission tomography image, the gray matter, and the white matter into a backbone network and outputting a first positron emission tomography image feature, a first gray matter feature, and a first white matter feature specifically includes: The target positron emission tomography image, the gray matter and the white matter are input into the backbone network, and after being preprocessed by the upsampling layer, dimension reduction is performed by the dimension reduction layer to obtain reduced-dimensional positron emission tomography image features, reduced-dimensional gray matter features and reduced-dimensional white matter features, respectively; The reduced-dimensional positron emission tomography imaging feature, the reduced-dimensional gray matter feature, and the reduced-dimensional white matter feature are all subjected to feature extraction through four residual layers to obtain a first positron emission tomography imaging feature, a first gray matter feature, and a first white matter feature; Among them, after each residual layer extracts features, a local attention module is used to obtain a local weight map, the local weight map is multiplied by the features extracted by the current residual layer, and the multiplication result is used as the input of the next residual layer.

3. The Alzheimer's disease auxiliary classification method according to claim 1, characterized in that: The collaborative attention module includes a first collaborative module and a second collaborative module, wherein the first positron emission tomography imaging feature, the first gray matter feature and the first white matter feature are input into the collaborative attention module, and feature weighting is performed to obtain a second positron emission tomography imaging feature, a second gray matter feature and a second white matter feature, specifically including: Inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a first collaborative module to generate a first positron emission tomography collaborative feature, a first gray matter collaborative feature, and a first white matter collaborative feature, and simultaneously inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a second collaborative module to generate a second positron emission tomography collaborative feature, a second gray matter collaborative feature, and a second white matter collaborative feature; generating a second positron emission tomography imaging feature according to the first positron emission tomography imaging feature, the first positron emission tomography synergy feature and the second positron emission tomography synergy feature; generating a second gray matter feature according to the first gray matter feature, the first gray matter synergy feature, and the second gray matter synergy feature; A second white matter feature is generated according to the first white matter feature, the first white matter synergy feature, and the second white matter synergy feature.

4. The Alzheimer's disease auxiliary classification method according to claim 3, characterized in that: The step of inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a first coordination module to generate a first positron emission tomography coordination feature, a first gray matter coordination feature, and a first white matter coordination feature specifically includes: Inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a first collaborative module, and compressing them by global average pooling to obtain a positron emission tomography resonance compression vector, a gray matter compression vector, and a white matter compression vector; According to the resonance compression vector of the positron emission tomography image, the gray matter compression vector and the white matter compression vector, a fusion vector is obtained through a fully connected layer in a first collaborative module, and channel weights corresponding to the resonance compression vector of the positron emission tomography image, the gray matter compression vector and the white matter compression vector are obtained according to the fusion vector; The resonance compression vector of the positron emission tomography scan, the gray matter compression vector, and the white matter compression vector are respectively multiplied element by element with the corresponding channel weights to obtain the first positron emission tomography scan synergy feature, the first gray matter synergy feature, and the first white matter synergy feature.

5. The Alzheimer's disease auxiliary classification method according to claim 3, characterized in that: The step of inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a second coordination module to generate a second positron emission tomography coordination feature, a second gray matter coordination feature, and a second white matter coordination feature specifically includes: Inputting the first positron emission tomography imaging feature, the first gray matter feature and the first white matter feature into a first collaborative module and then into a second collaborative module, compressing and connecting them through a convolution operation to obtain a fused volume; According to the three convolutions in the second collaborative module, the fused mentions are processed respectively to obtain spatial weights corresponding to the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature; The first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature are respectively multiplied element by element with the corresponding spatial weights to obtain the second positron emission tomography synergy feature, the second gray matter synergy feature, and the second white matter synergy feature.

6. An auxiliary classification system for Alzheimer's disease, characterized in that: The Alzheimer's disease auxiliary classification system includes: a feature extraction module, used to obtain a target positron emission tomography image and a structural magnetic resonance imaging, and to obtain gray matter and white matter from the structural magnetic resonance imaging, input the target positron emission tomography image, the gray matter and the white matter into a backbone network, and output a first positron emission tomography image feature, a first gray matter feature and a first white matter feature, wherein a local attention module is set at each layer in the backbone network; a modality complementation module, configured to input the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a collaborative attention module, perform feature weighting, and obtain a second positron emission tomography imaging feature, a second gray matter feature, and a second white matter feature; A fusion enhancement module is used to fuse the second positron emission tomography imaging feature, the second gray matter feature and the second white matter feature to obtain a fusion feature, and input the fusion feature into an enhancement feature fusion module to obtain a multimodal feature through normalization and a multi-layer perceptron; An output module is used to generate a classification prediction result based on the multimodal features.

7. The Alzheimer's disease auxiliary classification system according to claim 6, characterized in that: The step of inputting the target positron emission tomography image, the gray matter, and the white matter into a backbone network and outputting a first positron emission tomography image feature, a first gray matter feature, and a first white matter feature specifically includes: The target positron emission tomography image, the gray matter and the white matter are input into the backbone network, and after being preprocessed by the upsampling layer, dimension reduction is performed by the dimension reduction layer to obtain reduced-dimensional positron emission tomography image features, reduced-dimensional gray matter features and reduced-dimensional white matter features, respectively; The reduced-dimensional positron emission tomography imaging feature, the reduced-dimensional gray matter feature, and the reduced-dimensional white matter feature are all subjected to feature extraction through four residual layers to obtain a first positron emission tomography imaging feature, a first gray matter feature, and a first white matter feature; Among them, after each residual layer extracts features, a local attention module is used to obtain a local weight map, the local weight map is multiplied by the features extracted by the current residual layer, and the multiplication result is used as the input of the next residual layer.

8. The Alzheimer's disease auxiliary classification system according to claim 6, characterized in that: The collaborative attention module includes a first collaborative module and a second collaborative module, wherein the first positron emission tomography imaging feature, the first gray matter feature and the first white matter feature are input into the collaborative attention module, and feature weighting is performed to obtain a second positron emission tomography imaging feature, a second gray matter feature and a second white matter feature, specifically including: Inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a first collaborative module to generate a first positron emission tomography collaborative feature, a first gray matter collaborative feature, and a first white matter collaborative feature, and simultaneously inputting the first positron emission tomography imaging feature, the first gray matter feature, and the first white matter feature into a second collaborative module to generate a second positron emission tomography collaborative feature, a second gray matter collaborative feature, and a second white matter collaborative feature; generating a second positron emission tomography imaging feature according to the first positron emission tomography imaging feature, the first positron emission tomography synergy feature and the second positron emission tomography synergy feature; generating a second gray matter feature according to the first gray matter feature, the first gray matter synergy feature, and the second gray matter synergy feature; A second white matter feature is generated according to the first white matter feature, the first white matter synergy feature, and the second white matter synergy feature.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and an Alzheimer's disease auxiliary classification program stored in the memory and executable on the processor. When the Alzheimer's disease auxiliary classification program is executed by the processor, the steps of the Alzheimer's disease auxiliary classification method as described in any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an Alzheimer's disease auxiliary classification program, and when the Alzheimer's disease auxiliary classification program is executed by a processor, the steps of the Alzheimer's disease auxiliary classification method according to any one of claims 1 to 5 are implemented.