Alzheimer's disease diagnosis method based on comparative learning and Mama

By combining comparative learning and Mamba, the problem of high feature extraction and computing resource consumption in Alzheimer's disease diagnosis is solved, and a high accuracy diagnosis is achieved, with good clinical application potential.

CN120339207AActive Publication Date: 2025-07-18TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510383568.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

When the existing Alzheimer's disease diagnosis method uses deep learning technology, it is difficult to effectively extract distinctive feature characteristics, and the computing resource consumption is high, which limits its deployment in clinical applications.

Method used

Using a diagnostic method based on contrast learning and Mamba, the structural magnetic resonance imaging data is characterized by fusion and slice processing, features are extracted using pre-trained convolutional neural network, and long-distance dependence and context modeling are combined with the selective state space model Mamba, and finally the diagnosis is performed using the KAN classifier.

Benefits of technology

It improves the accuracy rate of Alzheimer's disease diagnosis to 94%, reduces the dependence on computing resources, and has good clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339207A_ABST
    Figure CN120339207A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and medical science, and particularly relates to an Alzheimer's disease diagnosis method based on comparative learning and Mama, which comprises the following steps of: carrying out feature fusion on an image in three dimensions, slicing the image dimension by dimension, and carrying out feature extraction by utilizing a pre-trained convolutional neural network; the extracted features are subjected to nonlinear mapping processing, and each slice is represented as a feature vector; and then, inputting the feature vector fused with the information of multiple dimensions into a selective state space model Mamba, and carrying out modeling analysis on the complex relationship between the slices through the advantages of the feature vector in the aspects of long-distance dependency capture and context memory modeling. And finally, classifying the input data by using a classifier integrated with the KAN, and predicting whether the sample suffers from the Alzheimer's disease or not. The diagnosis precision is effectively improved, the dependence on computing resources is reduced, and the method has good clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and medical technology, and particularly relates to a method for diagnosing Alzheimer's disease based on contrastive learning and Mamba. Background Art

[0002] Alzheimer's Disease (AD) is an irreversible neurodegenerative disease and one of the most common cognitive impairment diseases in the elderly population, especially in the current aging society. It is estimated that there are more than 55 million dementia patients globally, and this number is expected to soar to 139 million by 2050. The exact cause of AD is still unclear, but once a person is affected, there is no cure. In the past, diagnosis mainly relied on doctors' rich clinical experience, which was very time-consuming and laborious. Brain scans, such as Structural Magnetic Resonance Images (sMRI), provide a non-invasive way to capture the pathological patterns of the disease. Currently, Structural Magnetic Resonance Imaging (sMRI) has become an important tool for detecting neurodegenerative diseases in clinical practice, providing valuable insights into exploring the dynamic morphological features associated with Alzheimer's Disease (AD).

[0003] The key to dementia diagnosis based on Structural Magnetic Resonance Imaging (sMRI) lies in discriminative representation learning. Since the brain atrophy caused by Alzheimer's Disease (AD) is very subtle and only occurs in a few local regions, it is challenging to extract discriminative feature representations from sMRI for accurate AD diagnosis. Inspired by the progress of deep learning, a large number of studies have focused on using powerful deep neural networks (DNNs) as sMRI feature extractors to learn discriminative disease representations, including two-dimensional convolutional neural networks (2D CNNs), three-dimensional convolutional neural networks (3D CNNs), and Transformers. 2D-CNN methods usually transfer a network pre-trained on ImageNet (such as ResNet) to classify sMRI slices, and then aggregate all slice-level predictions to generate subject-level predictions. 3D-CNN methods apply 3D convolutions to the whole-brain sMRI or some anatomically defined regions determined empirically in advance and directly make subject-level predictions. Due to the increased kernel dimension, these methods tend to customize a shallower architecture to avoid a large number of training parameters.

[0004] In recent years, with the popularity of vision transformers, some studies have also explored reconfiguring the transformer architecture to adapt to sMRI slices or whole-brain sMRI for Alzheimer's disease-related diagnostic tasks. Although these studies have reported impressive diagnostic accuracies, further improvements have been hindered by the omission of dementia-related regions. This is because the feature representations at the higher levels of deep neural networks tend to respond more to the global semantics of the entire image. On the other hand, transformers also consume a significant amount of computing resources, and given that medical institutions are unlikely to be equipped with very expensive computing devices, this also limits their deployment in clinical applications. Summary of the Invention

[0005] In view of the above technical problems existing in the traditional Alzheimer's disease diagnostic methods, the present invention provides an Alzheimer's disease diagnostic method based on contrast learning and Mamba, which focuses on enhancing the model's ability to diagnose Alzheimer's disease using sMRI while reducing computing resource consumption. The model design conforms to the diagnostic mechanism of Alzheimer's disease, and an AD and NC recognition accuracy of 94% is achieved using only sMRI data.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0007] An Alzheimer's disease diagnostic method based on contrast learning and Mamba, comprising the following steps:

[0008] S1. Selection and establishment of a dataset: Select and collect T1 structural magnetic resonance data from different public datasets to construct a dataset for diagnosing Alzheimer's disease;

[0009] S2. Data preprocessing: The data preprocessing process includes: format conversion, noise removal, bias field correction, skull stripping, image registration, resampling, intensity normalization, and secondary format conversion;

[0010] S3. Model construction: This model consists of two parts; the first part is a contrast learning framework, which aims to help the model learn the underlying features of the image; this module generates the positive sample pairs required in contrast learning by randomly shuffling the order of sMRI slices in a certain dimension; the second part is the MMK model; the MMK model consists of four parts, a feature fusion module, a multi-view multi-plane feature extraction module, a Mamba module, and a classifier module, and the model is trained using the dataset to obtain a trained model;

[0011] S4. Use the AdamW algorithm to optimize the training of the model and set the corresponding model training parameters;

[0012] S5. Training the model: Use the training set, validation set, and test set to train, validate, and test the model. Use the cross-entropy loss function. According to the evaluation metrics, save the best model during the validation process, and use the test set to conduct experimental tests on the effectiveness of the proposed model; all data is divided according to the subjects to ensure that there is no data leakage.

[0013] The method for selecting and collecting T1 structural magnetic resonance imaging (sMRI) data from different public datasets in S1 is as follows: Select and collect T1 structural magnetic resonance imaging (sMRI) data from the ADNI, AIBL, and OASIS databases, delete the unavailable data contained therein, collect T1 structural magnetic resonance imaging (sMRI) data from the ADNI database to form the ADNI dataset. This dataset is split into a training set, a validation set, and a test set. Here, the division is based on the subjects, so as to ensure that the data of the same subject will not appear in the training set and the test set or the validation set at the same time, avoiding the problem of data leakage; Use the T1 structural magnetic resonance imaging (sMRI) data collected from the AIBL dataset and the OASIS dataset as the test set to test and verify the generalization and robustness of the model.

[0014] The method for the preliminary processing of the input data in S2 is as follows: First, convert the original data from the DCM format to the NIfTI format; Then, use the N4 bias field correction algorithm to process the image to eliminate the influence of the low-frequency bias field caused by magnetic field inhomogeneity; Subsequently, use the HD-BET neural network to perform skull stripping on the sMRI data to remove the non-brain tissue part; Next, register the processed sMRI to the MNI152 standard template to ensure spatial consistency; Then, resample the image to adjust the voxel spacing; After that, use the zero-mean unit-variance normalization method to standardize the image intensity of all voxels; Finally, convert the normalized NIfTI image to the.npy format suitable for the model to read. The size of the data after preprocessing is H = W = D.

[0015] In S3, use the contrastive learning framework to pre-train the model so that the model can initially capture the underlying features of sMRI, and use the selective state space models Mamba and KAN to process the vectorized three-dimensional information.

[0016] The method for training the model is as follows:

[0017] S3.1. Use the contrastive learning framework to train on the generated positive and negative sample pairs. The positive samples are generated by randomly selecting one of the three dimensions of the structural magnetic resonance imaging sMRI. The three dimensions are the sagittal plane, the axial plane, and the coronal plane, and shuffling the order of the image slices in one of the dimensions.

[0018] S3.2. The MMK model is pre-trained using positive and negative samples to obtain the weights of a pre-trained model, and all subsequent work will be carried out on this model;

[0019] S3.3. Construct an MMK model that uses the pre-trained weights to process structural magnetic resonance imaging (sMRI) data.

[0020] The method for constructing the MMK model that uses the pre-trained weights to process sMRI data in S3.3 is as follows:

[0021] S3.3.1 Feature Encoding Module: This module is centered around a three-dimensional convolutional neural network (3D-CNN). The initial convolutional layer uses a convolutional kernel of size 5x5x5, with a stride of 1, and padding is set to 2 to ensure that the input size remains unchanged after the convolutional operation. Subsequently, batch normalization is implemented through BatchNorm3d to accelerate model training and improve the stability of the network. The activation function uses GELU to introduce a non-linear transformation to enhance the expressive power of the model. Then, the module contains a second three-dimensional convolutional layer, still using a convolutional kernel of 5x5x5, a stride of 1, and padding set to 2, while combined with BatchNorm3d for normalization processing. The entire module further enhances the ability to capture and express the features of the input image by gradually increasing the number of output channels while keeping the input size constant;

[0022] S3.3.2 Multi-View Multi-Plane Feature Embedding Module: This module includes two stages: slicing each dimension and feature extraction from the slices. In the slicing stage, first, the positions of the axial, sagittal, coronal, and channel dimensions are sequentially swapped to achieve multi-dimensional slicing of the sMRI, that is Next, the sliced data is merged in the second dimension, that is A sequence of slices containing three dimensions is formed. After entering the feature extraction stage, the pre-trained ResNet34 model is used to process the slices, and each slice is mapped into a vector representation from different dimensions through this model. Subsequently, through a non-linear mapping module: This module consists of two layers of multi-layer perceptron (MLP) and a RELU activation function, which is used to perform non-linear mapping on the feature vectors output by the previous module to adjust the shape of the feature vectors to the target dimension to meet the requirements of subsequent processing;

[0023] S3.3.3 Construction of the Mamba Module:

[0024] This module consists of a selective state space model Mamba, which is used for long-range modeling of the sequence of feature vectors output by the above modules and capturing global context information. Its calculation formula is as follows:

[0025]

[0026] Among them and The calculation formulas are as follows:

[0027]

[0028] After the data passes through this module, its shape will not change, which is convenient for subsequent processing.

[0029] S3.3.4. Classifier module: This module consists of an adaptive average pooling layer, a Flatten Layer, a Dropout layer, and a KAN layer; among them, the Adaptive Average Pooling layer is used to reduce the dimension of the input features. By adaptively adjusting the size of the pooling window, a feature map with a fixed output size is generated, thereby improving the adaptability of the model to inputs of different sizes; subsequently, the Flatten Layer is used to convert the multi-dimensional feature map into a one-dimensional vector to prepare for the subsequent processing of the fully connected layer; then, a Dropout layer is added, which is used to randomly mask some neurons during the training process, thereby effectively preventing overfitting and enhancing the generalization ability of the model; finally, the KAN layer is adopted. Through its powerful non-linear mapping ability, it models high-dimensional features and completes classification.

[0030] In S4, the AdamW algorithm is used to optimize the training of the model, and β = 0.9, β2 = 0.999 are set. It is trained for 50 epochs, the learning rate is 1e-5, and the batch size is 3.

[0031] The model training process in S5 is as follows:

[0032] After all structural magnetic resonance imaging (sMRI) is preprocessed, in the contrast learning part, the positive sample pairs required in the contrast learning are generated by randomly shuffling the order of the slices in a certain dimension of the sMRI, and the other data in a batch are used as negative samples and input into the MMK model. The class-consistent contrast loss function is used, and the parameters of each layer in the network are updated through the backpropagation of the loss and the stochastic gradient descent algorithm; this will generate a set of pre-trained weights;

[0033] In the second part of the MMK training, the data in the ADNI dataset is divided into a training set, a validation set, and a test set; the training set is input into the MMK model using the pre-trained weights generated in the contrast learning for training. The cross-entropy loss function is adopted to calculate the error between the predicted value output by the model and the label, and the parameters in the model are updated through the backpropagation and gradient descent algorithm; at the same time, the validation set is used for model selection at the end of each epoch.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] The present invention fully exploits the rich feature information contained in 3D medical images. After feature fusion in three dimensions (coronal plane, sagittal plane, and axial plane) of the images, slices are made one by one dimension and a pre-trained convolutional neural network is used for feature extraction. The extracted features are processed through non-linear mapping to represent each slice as a feature vector. Subsequently, the feature vectors integrating information from multiple dimensions are input into the selective state space model Mamba. Through its advantages in capturing long-range dependencies and modeling context memory, the complex relationships between slices are modeled and analyzed. Finally, a classifier integrating KAN (Kolmogorov–Arnold Network) is used to classify the input data to predict whether the sample has Alzheimer's disease. This method effectively improves the diagnostic accuracy, reduces the dependence on computing resources, and has good clinical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained based on the provided drawings.

[0037] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the ratio relationship, or adjustment of the size should still fall within the scope covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.

[0038] Figure 1 is a flowchart of the method of the present invention;

[0039] Figure 2 is a schematic diagram of the contrastive learning framework proposed by the present invention;

[0040] Figure 3 is a schematic diagram of the structure of the model proposed by the present invention;

[0041] Figure 4 is a curve graph of the experimental results of the present invention;

[0042] Figure 5 is a bar graph of the experimental results of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. These descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0044] The following will further describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0045] As Figures 1 to 5 shown, the present invention proposes an Alzheimer's disease diagnosis method based on contrast learning and Mamba, including the following steps:

[0046] Step 1. Selection and establishment of the dataset: Select and collect the required T1 structural magnetic resonance imaging data from the ADNI1, ADNIGO, ADNI2, ADNI3, OASIS, and AIBL on the IDA official website, and construct a dataset for diagnosing Alzheimer's disease. The specific steps are as follows:

[0047] Step 1.1. Collect 4133 pieces of T1 structural magnetic resonance imaging including 1.5T and 3T from the ADNI database, and divide these data into a training set, a validation set, and a test set according to the subjects, with the proportions of each part being 8:1:1.

[0048] Step 1.2. Collect 335 pieces of T1 structural magnetic resonance imaging and 997 pieces of T1 structural magnetic resonance imaging from the OASIS and AIBL databases respectively. These two datasets are used as test sets to test the generalization and robustness of the trained model.

[0049] Step 2. Data preprocessing: The data preprocessing mainly includes: converting the format; filtering noise; bias field correction; skull stripping; image registration; resampling; intensity normalization; converting the format again; and dividing the processed data into a training set, a validation set, and a test set in sequence according to the subjects (ensuring no data leakage). The specific steps are as follows:

[0050] Step 2.1. Convert the original data from DICOM format to NIfTI format.

[0051] Step 2.2. Use the N4 bias field correction algorithm to process the image to eliminate the influence of the low-frequency bias field caused by magnetic field inhomogeneity.

[0052] Step 2.3: Use the HD-BET neural network to perform skull stripping on the structural magnetic resonance imaging (sMRI) and remove the non-brain tissue parts.

[0053] Step 2.4: Register the processed sMRI images to the MNI152 standard template to ensure spatial consistency.

[0054] Step 2.5: Perform resampling operations on the images to standardize the voxel spacing, adjust the voxel size to (1.75mm × 1.75mm × 1.75mm), and unify the spatial resolution of the volume data to (128 × 128 × 128) voxels.

[0055] Step 2.6: Standardize the image intensity of all voxels using the zero-mean unit-variance normalization method.

[0056]

[0057] where x is the value of the original data; x ′ is the value after normalization; μ is the mean of the original data (Mean), and σ is the standard deviation of the original data (Standard Deviation).

[0058] Step 2.7: Convert the normalized NIfTI images to the.npy format suitable for the model to read.

[0059] Step 3: Build the model: The whole model is divided into two parts, the contrastive learning framework and the MMK model. In the contrastive learning module, first randomly select one dimension of the sMRI image, and then randomly shuffle the order of the slices in this dimension. The data samples generated in this way and the original samples form a positive sample pair, while the other sMRI data in the same batch are used as negative samples; The other part of the MMK model consists of four parts: the feature encoding module, the multi-view multi-plane feature embedding module, the Mamba module, and the classifier module.

[0060] The specific steps are as follows:

[0061] Step 3.1: Use the contrastive learning framework to train on the generated positive and negative sample pairs. The positive samples are generated by randomly selecting one of the three dimensions of the sagittal, axial, and coronal planes of the sMRI image and shuffling the order of the image slices in this dimension; These samples are used to pre-train the model, and finally a set of pre-trained parameters are generated, which are considered a set of parameters that have learned the underlying features of the data. The following tasks will be carried out on this set of pre-trained parameters.

[0062] Step 3.2: The class-consistent contrast loss function used in the contrastive learning is specifically as follows:

[0063]

[0064] where q is the query sample, and k + is the generated positive sample, and k - is the negative sample, and the temperature parameter τ ensures a smoother similarity distribution. The numerator part: exp(q·k + / τ) represents the similarity between the query q and the positive sample k + . The denominator part: represents the sum of similarities between the query q and all samples (including positive and negative samples).

[0065] Step 3.3, Feature Encoding Module of MMK: Adopt a 5x5x5 convolutional kernel with a stride of 1 and set padding = 2; adopt BatchNorm3d; adopt the activation function GELU; a 3D convolutional network, use a 5x5x5 convolutional kernel with a stride of 1 and set padding = 2; adopt BatchNorm3d; adopt the activation function GELU; this module increases the output channels while keeping the input size unchanged to enhance the feature representation of the input image.

[0066] Step 3.4, Multi-view Multi-plane Feature Embedding Module of MMK: This module consists of two parts: slicing operation and sliced feature extraction. In the slicing operation part, the structural magnetic resonance imaging (sMRI) data is sliced by successively swapping the positions of the axial, sagittal, coronal, and channel dimensions, that is Subsequently, the sliced data is fused along the second dimension to generate a sliced sequence containing information from all three dimensions, that is Next, in the feature extraction part, a pre-trained two-dimensional convolutional neural network ResNet-34 is used to extract features from the sliced sequence, and the final output feature shape is batch×sequence_length×512. Non-linear mapping module: Consists of two layers of multi-layer perceptrons (MLPs) and a RELU activation function. The shapes of the two MLPs are (512×256) and (256×128) respectively. RELU is used to increase the non-linearity of the model, and its specific formula is as follows:

[0067]

[0068] This module maps the feature vector generated by the previous module to a desired 128-dimensional vector to further refine the feature information, thereby enhancing the expressive power of the features and the discriminative performance of the model.

[0069] Step 3.5, Mamba module:

[0070] This module consists of a selective state space model, Mamba, which is used to perform long-range modeling on the sequence of feature vectors output by the above module and capture global context information. Its calculation formula is as follows:

[0071]

[0072] Where and The calculation formula is as follows:

[0073]

[0074] Mamba processes these features through its Selective Scan Mechanism (S6), effectively capturing the context relationship between slices and complex long-range dependencies. At the same time, the output of the Mamba module maintains the same shape as the input.

[0075] Step 3.6, Classifier module: The Adaptive Average Pooling layer is used to reduce the dimension of the input features. The value is selected as 32. By adaptively adjusting the size of the pooling window, a feature map with a fixed output size is generated, thereby enhancing the adaptability of the model to inputs of different sizes. Subsequently, the multi-dimensional feature map is transformed into a one-dimensional vector through the Flatten Layer to prepare for the processing of the subsequent fully connected layer. Then, a Dropout layer is added, where the dropout value is selected as 0.8, which is used to randomly mask some neurons during the training process, thereby effectively preventing overfitting and enhancing the generalization ability of the model. Finally, the Kolmogorov-Arnold Networks (KAN) layer is adopted, and the specific formula is as follows:

[0076] KAN(Z) = (Φ K-1 ° Φ K-2 ° … ° Φ1 ° Φ0)Z, where Z is the input feature vector,

[0077] Through its non-linear mapping ability, it models high-dimensional features and completes classification.

[0078] Step 4, Use the AdamW algorithm to optimize the training of the model and set the corresponding model training parameters: Set β1 = 0.9, β2 = 0.999, train for 50 epochs, the learning rate is 1e-5, the batch size is 3, and the weight decay rate is 1e-4. Use a custom learning rate change strategy to keep the learning rate unchanged at the beginning of training, and then gradually decrease the learning rate as training progresses. Improve the stability and efficiency of model training.

[0079] Step 5, Training of the MMK model:

[0080] After all structural magnetic resonance imaging is preprocessed, in the contrast learning part, positive sample pairs required in contrast learning are generated by randomly shuffling the slice order of a certain dimension of sMRI, and other data in a batch are used as negative samples, which are input into the MMK model. The class-consistent contrast loss function is used, and the parameters of each layer in the network are updated through the backpropagation of the loss and the stochastic gradient descent algorithm. This generates a set of pre-trained weights.

[0081] In the MMK training of the second part, the data in the ADNI dataset are divided into a training set, a validation set, and a test set. The training set is input into the MMK model using the pre-trained weights generated in contrast learning for training. The cross-entropy loss function is adopted to calculate the error between the predicted value output by the model and the label, and the parameters in the model are updated through the backpropagation and gradient descent algorithms. At the end of each epoch, the validation set is used for model selection.

[0082] Step 6. Model testing and evaluation:

[0083] The trained MMK model is used to test on the test set, and tests on the ADNI, AIBL, and OASIS datasets are obtained respectively.

[0084] The present invention uses four indicators, namely accuracy (ACC), area under the curve (AUC), sensitivity (Sensitivity, SEN), and specificity (Specificity, SPE), to evaluate the performance of the classification model. These indicators are very commonly used in medical image analysis and binary classification tasks. Their definitions and formulas are as follows:

[0085] Step 6.1. Accuracy (ACC), the calculation formula is as follows:

[0086] Step 6.2. Sensitivity (Sensitivity, SEN): Also called recall, it is the ability of the model to identify positive class samples. The calculation formula is:

[0087] Step 6.3. Specificity (Specificity, SPE), which is the ability of the model to identify negative class samples. The calculation formula is as follows:

[0088] Among them: TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative respectively. Among them, TN: the number of samples correctly classified as negative class, FP: the number of negative class samples misclassified as positive class, TP: the sample data correctly classified as positive class, FN: the number of positive samples misclassified as negative samples.

[0089] Step 6.4, AUC (Area Under the Curve) is an important metric for evaluating the performance of a classification model, which refers to the area under the ROC curve. The abscissa of the ROC curve is the false positive rate, and the ordinate is the true positive rate. The larger this area, the better the model performance.

[0090] The method of the present invention was compared with advanced methods at home and abroad, and the comparison results are shown in Tables 1, 2, and 3 below. It can be seen from this that compared with other methods, the Alzheimer's disease diagnosis method based on contrast learning and Mamba proposed by the present invention has advantages such as high accuracy and high robustness.

[0091] Table 1 Comparison table of diagnostic results of different methods on the structural magnetic resonance imaging collected from the ADNI dataset by the present invention

[0092] Method ACC AUC SEN SPE 3D ResNet152 0.8344 0.8692 0.7504 0.8292 3D ViT 0.8524 0.8135 0.7367 0.8489 MRNet 0.8796 0.9316 0.8204 0.8974 MedicalNet 0.8889 0.8880 0.8162 0.8739 M3T 0.8980 0.9104 0.8367 0.9563 The present invention 0.9402 0.9453 0.8381 0.9856

[0093] Table 2 Comparison table of diagnostic results of different methods on the structural magnetic resonance imaging collected from the AIBL dataset by the present invention

[0094]

[0095]

[0096] Table 3 Comparison table of diagnostic results of different methods on the structural magnetic resonance imaging collected from the OASIS dataset by the present invention

[0097] Method ACC AUC SEN SPE 3D ResNet152 0.7134 0.7287 0.6639 0.8533 3D ViT 0.7569 0.7713 0.7109 0.8441 MRNet 0.7077 0.8197 0.7001 0.8527 MedicalNet 0.7385 0.7272 0.6938 0.8482 M3T 0.8047 0.8167 0.7895 0.9109 The present invention 0.8574 0.8459 0.7914 0.9218

[0098] As shown in the results of Tables 1, 2, and 3, the Alzheimer's disease diagnosis method based on contrast learning and Mamba proposed by the present invention achieved the best performance in all metrics on the three public datasets of ANDI, AIBL, and OASIS, and was superior to the performance shown by 3D ResNet152, 3D ViT, MRNet, MedicalNet, and M3T. These results prove the effectiveness of the method proposed by the present invention.

[0099] Only the preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A method for diagnosing Alzheimer's disease based on contrastive learning and Mamba, characterized in that, It includes the following steps: S1. Selection and establishment of the dataset: Select and collect T1-weighted structural magnetic resonance imaging (sMRI) data from different public datasets to construct a dataset for diagnosing Alzheimer's disease; S2. Data preprocessing: The data preprocessing process includes: format conversion, noise removal, bias field correction, skull stripping, image registration, resampling, intensity normalization, and secondary format conversion; S3. Model construction: This model consists of two parts; the first part is a contrastive learning framework, aiming to help the model learn the underlying features of images; this module generates the positive sample pairs required in contrastive learning by randomly shuffling the order of slices in a certain dimension of the MRI image; the second part is the MMK model; the MMK model consists of four parts, a feature fusion module, a multi-view multi-plane feature extraction module, a Mamba module, and a classifier module, and the dataset is used to train the model to obtain a trained model; S4. Use the AdamW algorithm to optimize the training of the model and set the corresponding model training parameters; S5. Train the model: Use the training set, validation set, and test set to train, validate, and test the model. Use the cross-entropy loss function, save the best model during the validation process according to the evaluation metrics, and use the test set to conduct experimental tests on the effectiveness of the proposed model; all data is divided according to the subjects to ensure that there is no data leakage.

2. The Alzheimer's disease diagnosis method based on contrastive learning and Mamba according to claim 1, wherein, The method for selecting and collecting T1-weighted structural magnetic resonance imaging (sMRI) data from different public datasets in S1 is: Select and collect T1-weighted structural magnetic resonance imaging (sMRI) data from the ADNI, AIBL, and OASIS databases, delete the unavailable data contained therein, collect T1-weighted structural magnetic resonance imaging (sMRI) data from the ADNI database to form the ADNI dataset, and this dataset is split into a training set, a validation set, and a test set. This division is based on the subjects, so as to ensure that the data of the same subject will not appear in the training set and the test set or the validation set at the same time, avoiding data leakage problems; use the T1-weighted structural magnetic resonance imaging (sMRI) data collected from the AIBL dataset and the OASIS dataset as the test set to test and verify the generalization and robustness of the model.

3. The Alzheimer's disease diagnosis method based on contrastive learning and Mamba according to claim 1, wherein, The method for the preliminary processing of the input data in S2 is: First, convert the original data from DICOM format to NIfTI format; then, use the N4 bias field correction algorithm to process the image to eliminate the influence of the low-frequency bias field caused by magnetic field inhomogeneity; subsequently, use the HD-BET neural network to perform skull stripping on the sMRI data to remove the non-brain tissue part; next, register the processed sMRI to the MNI152 standard template to ensure spatial consistency; then, resample the image to adjust the voxel spacing; after that, use the zero-mean unit-variance normalization method to standardize the image intensity of all voxels; finally, convert the normalized NIfTI image to the.npy format suitable for the model to read. After preprocessing, the data size is H = W = D.

4. A method for diagnosing Alzheimer's disease based on contrastive learning and Mamba according to claim 1, characterized in that In S3, a contrastive learning framework is used to pre-train the model, enabling the model to initially capture the underlying features of sMRI, and the selective state space models Mamba and KAN are used to process the three-dimensional information after vectorization.

5. A method for diagnosing Alzheimer's disease based on contrastive learning and Mamba according to claim 4, characterized in that The method for training the model is as follows: S3.1: Use the contrastive learning framework to train on the generated positive and negative sample pairs. The positive samples are generated by randomly selecting one of the three dimensions of sMRI, namely the sagittal plane, the axial plane, and the coronal plane, and shuffling the order of the image slices in one of the dimensions. S3.2: The MMK model is pre-trained using positive and negative samples to obtain the weights of a pre-trained model, and all subsequent work will be carried out on this model. S3.3: Construct an MMK model for processing structural magnetic resonance imaging (sMRI) data using the pre-trained weights.

6. The Alzheimer's disease diagnosis method based on contrastive learning and Mamba according to claim 5, wherein The method for constructing an MMK model for processing sMRI data using the pre-trained weights in S3.3 is as follows: S3.3.1: Feature encoding module: This module is centered around a three-dimensional convolutional neural network 3D-CNN. The initial convolutional layer uses a convolutional kernel of size 5x5x5, a stride of 1, and a padding of 2 is set to ensure that the input size remains unchanged after the convolutional operation. Subsequently, batch normalization is implemented through BatchNorm3d to accelerate model training and improve the stability of the network. The activation function uses GELU to introduce a non-linear transformation to enhance the expressive power of the model. Then, the module contains a second three-dimensional convolutional layer, still using a convolutional kernel of 5x5x5, a stride of 1, and a padding of 2, and is combined with BatchNorm3d for normalization processing. The entire module further enhances the ability to capture and express the features of the input image by gradually increasing the number of output channels while keeping the input size constant. S3.3.2 Multi - perspective and multi - plane feature embedding module: This module includes two stages: slicing each dimension and feature extraction from the slices; in the slicing stage, first, the positions of the axial plane, sagittal plane, coronal plane, and channel dimension are sequentially swapped to achieve multi - dimensional slicing of sMRI data, that is Next, the sliced data is merged in the second dimension, that is A sequence of slices containing three dimensions is formed; After entering the feature extraction stage, the pre-trained ResNet34 model is used to process the slices, and each slice is mapped into a vector representation from different dimensions through this model. Subsequently, through the non-linear mapping module: This module consists of two layers of multi-layer perceptron MLP and a RELU activation function, and is used to perform non-linear mapping on the feature vectors output by the previous module to adjust the shape of the feature vectors to the target dimension to meet the requirements of subsequent processing. S3.3.3: Construction of the Mamba module: This module consists of a selective state space model Mamba, which is used to perform long-range modeling and capture global context information on the sequence of feature vectors output by the above module. Its calculation formula is as follows: Among them and The calculation formula is as follows: The shape of the data does not change after passing through this module, which is convenient for subsequent processing. S3.3.4: Classifier module: This module consists of an adaptive average pooling layer, a flatten layer Flatten Layer, a Dropout layer, and a KAN layer. Among them, the Adaptive Average Pooling layer is used to reduce the dimension of the input features. By adaptively adjusting the size of the pooling window, a feature map with a fixed output size is generated, thereby enhancing the adaptability of the model to inputs of different sizes. Subsequently, the Flatten Layer is used to convert the multi-dimensional feature map into a one-dimensional vector to prepare for the processing of the subsequent fully connected layer. Then, a Dropout layer is added to randomly mask some neurons during the training process, thereby effectively preventing overfitting and enhancing the generalization ability of the model. Finally, the KAN layer is adopted. Through its powerful non-linear mapping ability, it models high-dimensional features and completes classification.

7. A diagnostic method for Alzheimer's disease based on contrastive learning and Mamba according to claim 1, characterized in that: In S4, the AdamW algorithm is used to optimize the training of the model, and β = 0.9, β2 = 0.999 are set. It is trained for 50 epochs, the learning rate is 1e-5, and the batch size is 3.

8. A method for diagnosing Alzheimer's disease based on contrastive learning and Mamba according to claim 1, characterized in that: The model training process in S5 is as follows: After all structural magnetic resonance imaging is preprocessed, in the contrast learning part, the positive sample pairs required in the contrast learning are generated by randomly shuffling the order of the slices in a certain dimension of the sMRI, and the other data in a batch are used as negative samples and input into the MMK model. The class-consistent contrast loss function is used, and the parameters of each layer in the network are updated through the backpropagation of the loss and the stochastic gradient descent algorithm; this will generate a set of pre-trained weights. In the second part of the MMK training, the data in the ADNI dataset are divided into a training set, a validation set, and a test set; the training set is input into the MMK model using the pre-trained weights generated in the contrast learning for training. The cross-entropy loss function is adopted to calculate the error between the predicted value output by the model and the label, and the parameters in the model are updated through the backpropagation and gradient descent algorithms; at the same time, the validation set is used for model selection at the end of each epoch.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method and device based on Kan-Mamba model

    CN119399473A

  • Early Alzheimer's disease classification method based on multi-view comparative learning

    CN119540641A

  • Detection system for alzheimer's disease using brain structural and / or functional magnetic resonance image processing and machine learning techniques

    WO2024252413A1

  • AU2020102569A4