Alzheimer disease image classification method based on Mama model
Through the three-dimensional PET image classification method of multi-stage progressive Lmamba block feature extraction, the subjectivity and computing efficiency of manual interpretation of PET images in the prior art are solved, and efficient and accurate early diagnosis and treatment support for AD are achieved.
Patent Information
- Application Number
- CN202510749601.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-02
AI Technical Summary
The prior art relies on manual interpretation of PET images in the diagnosis of Alzheimer's disease, which has the problem of strong subjectivity, inefficiency and difficulty in early diagnosis of patients with mild cognitive impairment. Deep learning models are expensive to calculate and difficult to capture long-range dependencies when processing 3D medical images.
A three-dimensional positron emission tomography data image classification method based on multi-stage progressive Lmamba block feature extraction is adopted. Through standardized preprocessing, data enhancement, multi-layer feature extraction, cross-scale feature fusion and detail enhancement, a state space hybrid convolution model is constructed to realize automated diagnosis.
Improves the accuracy and generalization of AD classification, reduces computational complexity, and can provide accurate diagnosis at an early stage, supporting clinical intervention and treatment.
Smart Images

Figure CN120580503A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image analysis and artificial intelligence technology, and in particular to a three-dimensional positron emission tomography (PET) data image classification method based on multi-stage progressive Lmamba block feature extraction, which is used for automated diagnosis of Alzheimer's Disease (AD). Technical Background
[0002] Alzheimer's disease (AD) is the most common and devastating neurodegenerative disease among dementias. Patients typically experience memory loss, loss of daily self-care abilities, and ultimately a significant decline in cognitive function. With the intensification of the global aging population, AD has become a significant medical and societal issue.
[0003] The pathological characteristics of AD primarily include the abnormal accumulation of amyloid plaques and tau protein tangles within brain cells. These pathological changes disrupt communication between nerve cells, leading to a gradual loss of neurological function. Positron emission tomography (PET) is an imaging technique that can visualize brain metabolic processes at the molecular level. 18F-FDG-PET scanning, in particular, can effectively measure brain glucose metabolism and is recognized as an important biomarker for identifying neurodegenerative diseases such as AD. However, the interpretation of PET images currently relies primarily on the experience of nuclear medicine and neuroimaging physicians, which is subject to high subjectivity, low efficiency, and difficulty in early diagnosis of whether patients with mild cognitive impairment (MCI) will develop AD.
[0004] In recent years, the application of deep learning technology in medical image analysis has provided new possibilities for solving the above problems. Unlike traditional machine learning methods that rely on experts to manually extract region of interest (ROI) features, deep learning has the characteristics of end-to-end learning and can automatically extract multi-level features in images, significantly improving the generalization ability of the model. Deep learning models represented by convolutional neural networks (CNNs) have achieved remarkable success in 2D image classification tasks, but face challenges such as high computational overhead and difficulty in capturing long-range dependencies when processing 3D medical images. In addition, although the Transformer model performs well in capturing global relationships, the quadratic computational complexity of its self-attention mechanism makes it inefficient when processing large-scale 3D data, making it difficult to promote in practical applications. Summary of the Invention
[0005] This paper provides a 3D positron emission tomography (PET) data classification method based on multi-stage progressive Lmamba block feature extraction for automated AD diagnosis. This method preprocesses data to obtain a standardized PET dataset, automatically learns disease differentiation through a model, and ultimately accurately predicts disease categories, thereby improving the efficiency of early AD detection and providing reliable support for clinical diagnosis.
[0006] To achieve the above object, the technical solution adopted by the present invention comprises the following steps:
[0007] The present invention provides a three-dimensional positron emission tomography image classification method based on a state-space hybrid convolution model, which specifically includes the following steps:
[0008] S1: Standardization of raw PET images. PET raw image data for four disease categories, including AD, normal control (NC), stable mild cognitive impairment (sMCI), and progressive mild cognitive impairment (pMCI), were obtained from the public ADNI database. To eliminate deviations caused by head size, shape, position, and noise interference during scanning, the statistical parametric mapping toolbox (SPM12) in MATLAB platform was used to complete the standard preprocessing process of the raw data, including spatial registration, skull removal, smoothing, and intensity normalization, to eliminate imaging differences and improve data quality.
[0009] S2: Enhanced Dataset Diversity. The preprocessed dataset was randomly divided into training, validation, and test sets according to a scientific ratio, ensuring that the data in each subset was used independently. To enhance the generalization performance of the model, a random data augmentation strategy was designed to improve dataset diversity. When the training set was imported into the model, the optimal random probability was set to trigger the augmentation process for each image.
[0010] S3: Shallow feature extraction. Aiming at the high-resolution characteristics of the input image, a convolutional backbone module is designed based on 3D convolution to extract coarse features and generate high-spatial-resolution feature maps.
[0011] S4: Trunk-branch modeling. In the model's trunk input module, semantic features at different scales are gradually extracted through a multi-stage progressive feature extraction module. Each stage includes multiple Lmamba blocks and downsampling operations, gradually reducing the resolution of the feature map while increasing the number of channels (channel dimension) to construct a global information representation of the whole-brain image. This process effectively captures multi-level features in the image, providing rich input for subsequent feature fusion and classification.
[0012] S5: Local feature modeling. Corresponding to the trunk branch feature flow in S4, a step-by-step cross-scale channel attention fusion branch (CSCAF) is designed to fuse global and local semantic features. CSCAF dynamically adjusts the weights of features at different scales through the attention mechanism, interactively fusing the trunk features (F0 represents the trunk features at each stage) with the cross-scale branch features F1 (F1 represents the cross-scale branch features of the branches) to enhance the multi-scale feature expression capability. This process can capture details and contextual information that cannot be obtained at a single scale, improve the model's ability to recognize complex pathological features, and thus enhance the robustness and accuracy of the classification model.
[0013] S6: Feature integration: The global and local information of the whole-brain image are integrated through the Channel and Spatial Perception Mechanism (CSPM).
[0014] S7: Detail Enhancement: The inverted bottleneck block (IBB) is used to blend long-range spatial and positional information to improve the ability to express detailed features.
[0015] S8: Classification prediction. On the model’s classifier network, class prediction is performed on the fused feature map.
[0016] Preferably, the standard preprocessing process described in S1 includes the following six independently performed preprocessing sub-processes: image structure adaptive non-local mean denoising, voxel interval sampling of (1.50mm×1.50mm×1.50mm), ICBM152 standard brain template alignment, non-uniformity intensity normalization, Gaussian smoothing and skull stripping.
[0017] Preferably, a set of random data enhancement strategies of S2 includes random rotation: generating different viewing angles within a specified range of (-30° to +30°); gamma adjustment: adjusting the contrast of the image; and random masking: masking voxels in random areas with zero settings.
[0018] Preferably, the convolution backbone module described in S3 uses a convolution operation with a convolution kernel size of K=7 to expand the receptive field, and combines batch normalization (BN) and ReLU activation function to ensure stable gradient transfer.
[0019] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0020] (1) Improve diagnostic accuracy: The present invention significantly improves the accuracy and effectiveness of AD classification by fusing global and local information of images. The main branch focuses on modeling the overall structural characteristics of brain images and capturing global contextual information, while the branch focuses on capturing local detail information through a multi-scale feature extraction mechanism. The synergistic effect of the two enhances the perception of complex lesion characteristics. (2) Enhance generalization: The introduction of data enhancement strategies and standardized preprocessing processes further improves the adaptability and generalization performance of the model to diverse data. (3) Improve computational efficiency: By introducing a selective state space mechanism, while maintaining linear time complexity, performance comparable to Transformer is achieved. (4) Early diagnosis advantage: The method of the present invention can provide more accurate diagnosis in the early stages of AD through an improved deep learning network architecture. It is of great significance for early intervention and treatment of the disease, can delay the progression of the disease, and improve the quality of life of patients. It provides an efficient and accurate solution for automated medical imaging diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flowchart of a specific implementation of the three-dimensional positron emission tomography image classification method based on the Mamba model proposed in the present invention;
[0022] Figure 2 This is a model framework diagram of the three-dimensional positron emission tomography image classification method based on the Mamba model proposed in the present invention;
[0023] Figure 3 This is a schematic diagram of the structure of the Lmamba block proposed in the present invention;
[0024] Figure 4 This is a schematic diagram of the spatial feature fusion module structure proposed in the present invention;
[0025] Figure 5 This is a schematic diagram of the structure of the CSCAF proposed in the present invention;
[0026] Figure 6 This is a schematic diagram of the structure of the channel space perception mechanism proposed in the present invention;
[0027] Figure 7 This is a schematic diagram of the structure of the IBB block proposed in the present invention; DETAILED DESCRIPTION
[0028] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting the present invention;
[0029] In order to better illustrate this embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the drawings.
[0030] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of the Mamba model-based AD image classification method proposed in the present invention. Unless otherwise defined, all technical and scientific terms used in this invention have the same meanings as those commonly understood by those skilled in the art to which this invention belongs.
[0031] Please refer to the instruction manual Figure 1 , which is a flow chart of a specific embodiment of the method proposed in the present invention, comprising the following steps:
[0032] Step S1: PET dataset acquisition and preprocessing, the dataset is divided into training set, validation set and test set in a 4:1:1 ratio.
[0033] The PET imaging data used in this paper all come from the ADNI. The resulting dataset includes four disease categories: AD, NC, sMCI, and pMCI. Because the ADNI database's native DCM format is typically not directly suitable for modeling and analysis, the raw images were preprocessed using the Statistical Parametric Mapping Toolbox (SPM12) on the MATLAB platform. This toolbox integrates a variety of standardized image processing tools to suppress the effects of feature space misalignment between raw scans and acquisition noise, thereby obtaining a training dataset for the algorithm.
[0034] Specifically, the raw image preprocessing process in step S1 is described as follows. On the one hand, due to the influence of the subject's head size, shape, and position during scanning, the image feature space is usually in a non-identical coordinate system. On the other hand, PET often contains various types of noise, including magnetic field inhomogeneity, acquisition noise, and motion artifacts. To eliminate these effects, this study adopted a systematic image preprocessing process, including: structure-adaptive non-local mean denoising, voxel interval resampling to (1.5mm×1.5mm×1.5mm), Montreal Neurological Institute (MNI) standard space alignment, intensity normalization, skull stripping, and Gaussian smoothing. In addition, to facilitate model loading and analysis, all PET images were resampled to the same resolution of 112px×112px×112px, and the NIfTI format was unified into the NumPy format.
[0035] Step S2: Data diversity enhancement.
[0036] During the model training process, the present invention designed a data enhancement process before importing the training set images, and randomly applied data enhancement strategies to 50% of the data, including random rotation, gamma adjustment and random masking, to enhance sample diversity and enhance the feature learning robustness of the model.
[0037] Specifically, the data enhancement process designed in step S2 is based on the distribution characteristics of medical image features and performs data enhancement with a random probability of 50%. It includes the following image processing steps: first, the central axis of the coronal plane of the 3D PET is used as the rotation axis, and a random angle in the range of (-30° to +30°) is adopted; second, the voxel intensity value is adjusted by nonlinear transformation to adjust the brightness and contrast of the image; finally, a space of size (10px×10px×10px) is randomly generated in the entire 3D image space for voxel zero masking;
[0038] Step S3: shallow feature extraction is achieved through the convolutional backbone module.
[0039] Specifically, as shown in the appendix to the manual Figure 2 As shown in the figure, to pre-extract features and increase the channel dimensionality of the preprocessed PET images, the image input size after the convolutional backbone module is reduced from B×1×112×112×112 to B×48×112×112×112, where B is the batch size. The convolutional backbone module primarily consists of 3D convolutional layers and instance normalization for 3D data to stabilize gradient propagation. The 3D convolution kernel size is 7×7×7, with a stride of 2 and padding of 1. The number of output channels in the subsequent downsampling is doubled, and the spatial scale of the output feature map is downsampled to half. The output feature map scale of the entire backbone network is B×384×7×7×7.
[0040] Step S4: Long-range dynamic modeling of the imagery is performed using stacked Lmamba blocks.
[0041] The output of the convolutional backbone module serves as the input to the main branch, where Lmamba blocks and downsampling modules are sequentially stacked to dynamically adjust the number of channels and spatial size of the data stream. In four stages, the main branch generates feature maps with spatial scales of 56×56×56, 28×28×28, 14×14×14, and 7×7×7, respectively, with 48, 96, 192, and 384 channels. This multi-scale design effectively captures multi-level information from local details to global context, gradually expanding the receptive field and enhancing the understanding of global structure.
[0042] Specifically, the Lmamba block in step S4 is described in the appendix of the specification. Figure 3 The submodules and processing steps included in the multi-stage progressive feature extraction module are as follows:
[0043] 1. Design Lmamba blocks to achieve efficient modeling of long-range dependencies between features. In order to effectively perceive and retain spatial information before feature serialization, SCB blocks are designed, as shown in the attached manual. Figure 4 As shown. The SCB block uses a convolution block with a dual-branch structure to model features at multiple scales. The input features pass through multiple convolution blocks including normalization, 3D convolution and nonlinear activation. The upper branch uses two layers of 3×3×3 convolution kernels, a step size of S=1, and a padding layer of P=1 to perform large window feature extraction on the three dimensions of the input data while maintaining spatial resolution. The introduction of the activation function ReLU enhances the nonlinear expression ability of the model. The difference between the lower branch is that a single layer of 1×1×1 convolution kernel is used without padding to mix information from different channels while maintaining the spatial dimension. The features of the two branches are fused through element-by-element multiplication to enhance the feature interaction ability and capture more complex patterns. Then a 1×1×1 convolution layer is performed for post-processing, and the fused features are instance normalized. Finally, the processed features are added to the original input to form a residual connection to alleviate the gradient disappearance problem and improve the training stability and performance of the model. This step can be expressed in the formula:
[0044] SCB(f)=f+Conv1((Conv3(Conv3(f)))·Conv1(f)
[0045] Where f represents the three-dimensional features of the input, Conv i Represents a convolution block with a convolution kernel size of i.
[0046] 2. The State Space Hybrid Convolutional Block (SSHCB) is a deep learning network layer.
[0047] By combining depthwise separable convolution, linear attention mechanism, and multi-scale feature interaction, it extracts and transforms features from the input tensor. Its core functions include: (1) spatial-channel decoupling modeling: flattening the high-dimensional features of the input into a sequence form, separating the spatial and channel dimensions, and adapting to lightweight feature transformations; (2) dynamic feature enhancement: injecting position information through convolutional positional encoding (CPE) to enhance the model's sensitivity to local structure; (3) using linear attention instead of traditional self-attention to significantly reduce computational complexity and be suitable for long sequence input; (4) residual multi-path fusion: interacting with multi-branch feature multiplication through residual connections to improve gradient flow and capture cross-scale feature dependencies. Its main operational steps are implemented through Mamba-Type Linear Attention (MTLA) and Multilayer Perceptron (MLP).
[0048] 3. The MTLA module is an improvement on the Mamba block, adding additional depth convolution and residual links, using a linear attention mechanism, replacing the traditional self-attention O(L) with a complexity of O(L) through linear projection. 2 ) computation to address the Transformer's quadratic computational complexity overhead. Dynamic modulation of feature activation and attention weights is achieved through gated interactions. A built-in DropPath randomly drops paths to improve generalization. Finally, the processed sequence features are restored to their original spatial dimensions, maintaining input and output size consistency.
[0049] The specific operation process of the Lmamba block is as follows:
[0050] 1. Effectively perceive and preserve spatial information before feature serialization through the SCB module.
[0051] 2. Spatial-channel decoupling modeling: Flatten the input dimensions [B, C, W, H, D] into a sequence form [B, L, C] (L = D × H × W). Perform layer normalization at the same time.
[0052] 3. Normalization and Position Encoding: Enhance position awareness through layer normalization and CPE.
[0053] 4. Calculate feature correlations through lightweight attention, combine dynamic gating (SiLU activation) to modulate attention output, and strengthen important features.
[0054] 5. Add the attention output to the original input to form a residual fusion, preserving the underlying information.
[0055] 6. Further nonlinear transformation is performed through multi-layer perceptron to enhance feature expression and restore sequence features to their original spatial dimensions.
[0056] Step S5: Global contextual semantic adaptive fusion is achieved through a layer-by-layer cross-scale channel attention fusion module.
[0057] Corresponding to the global branch in step S4, the output feature map of the convolutional trunk module is also used as the input of the branch branch. The feature flow of the branch branch is also progressively fused, and CSCAF is set between the stages. Each stage of the branch branch is mainly performed by CSCAF on the channel dimension to fuse F0 and F1. The spatial scale and number of channels of the feature maps generated in the three stages of the branch branch are consistent with those of the trunk branch.
[0058] Specifically, the CSCAF proposed in step S5 is described in the appendix of the specification. Figure 5 As shown in Figure 2, the module implements cross-scale feature interaction through channel adjustment, residual connection, and ECA attention mechanism. The core functions of the module include:
[0059] 1. Multi-scale feature alignment: The spatial size of high-dimensional feature maps is downsampled through adaptive average pooling, and the number of channels is adjusted using 1×1×1 convolution to match the number of channels of feature maps of different scales.
[0060] 2. Cross-scale Feature Fusion Module: This module implements cross-scale feature interaction through channel adjustment, ECA attention mechanism, and depthwise separable convolution. Specifically, features are first element-wise multiplied and then concatenated with the base. A 3×3×3 grouped convolution is then performed on the concatenated features to achieve lightweight feature extraction. A 1×1×1 convolution is used to adjust the number of channels to a preset value. After enhancing important features using the ECA attention mechanism, the fused result is superimposed on the original input to preserve the underlying information.
[0061] Step S6: Obtain decision-level feature representation by fusing F0 and F1.
[0062] The global branch and branch branch focus on the global representation feature map and local representation feature map of PET image features respectively. By fusing the local and global feature maps through the channel space perception mechanism, a more advanced category decision-level feature representation can be generated.
[0063] Specifically, please refer to the attached manual for the CSPM proposed in step S6. Figure 6 The channel space perception structure shown. CSPM includes channel perception network and space perception network:
[0064] S61: Channel-aware network captures local cross-channel interactions through adaptive average pooling and 1D convolution. The input features of the channel-aware mechanism are the backbone high-dimensional features. The size is B×384×7×7×7. The target size is fixed through adaptive average pooling, which reduces the computational complexity while retaining the main information.
[0065]
[0066] Among them F cp Represents channel-aware output, performs lightweight channel interaction through one-dimensional convolution (Conv1d), σ represents the sigmoid activation function, Avgpool represents average pooling, and * represents element-by-element multiplication.
[0067] S62: The spatial perception network extracts features from the width, height, and depth dimensions, generates a channel spatial attention weight coefficient matrix, and then adds it to the initial input F in Multiplying together, the final output is a high-dimensional feature map with a shape of B×768×7×7×7. The specific operation process is as follows:
[0068] S621: F cp with F1 3 Combined to form complementary information:
[0069] F in =F cp +F1 3 ;
[0070] S622: Split from width, height, and depth dimensions and use average pooling:
[0071]
[0072] S623: Splicing and using convolution blocks for spatial perception to generate feature weights for each dimension:
[0073] (f' W ,f' H ,f' D )=split[σ[BN[Conv[cat(f w ,f h ,f d )]]]]];
[0074] Where cat represents the concatenation operation, Conv represents 1D convolution, BN represents layer normalization, σ represents the GELU activation function, and split represents segmentation along the height, width, and depth dimensions.
[0075] S624: Use the activation function σ2 to generate a weight matrix with a mask in the range [0-1] and F in Features are weighted:
[0076] M s=σ2(f' W ,f' H ,f' D ), F CSPM =M s *F in ;
[0077] Among them, M s represents the weight matrix, σ2 represents the sigmoid activation function, F CSPM Output feature maps for CSPM. Finally, CSPM fuses global and local information and outputs high-dimensional semantic features, providing rich feature representation for subsequent classification tasks.
[0078] Step S7: Mixing long-range spatial and position information via IBB.
[0079] Specifically, the step S7 uses IBB to mix the long-distance space and position information, as shown in the appendix of the specification. Figure 7 The IBB architecture is shown in Figure 2. It uses depthwise separable convolution and pointwise convolution. It first expands the channel dimension through depthwise separable convolution, then extracts features through pointwise convolution. Finally, batch normalization and activation functions enhance the expressiveness of features. This design significantly improves computational efficiency while maintaining the expressiveness of high-dimensional features. This allows high-dimensional features to have richer semantic and spatial information, thereby enhancing the model's ability to capture detailed features.
[0080] Step S8: Predict the disease category probability through the classifier.
[0081] The high-dimensional category decision-level feature map first undergoes spatial pooling compression through a global average pooling layer, outputting a feature map with a shape of B × 768 × 1 × 1 × 1. A spatial flattening operation then produces a high-dimensional feature vector with a dimension of 768. Finally, a fully connected layer with 768 input nodes and 1 output node outputs the patient's probability of having AD.
[0082] Specifically, the classifier in step S8 predicts the disease category probability. The specific operation process is as follows: using a global average pooling layer to compress the feature space and then perform a dimensional flattening operation. Then, a fully connected layer outputs the category prediction probability. The high-dimensional semantic features are mapped to the category space, and the disease classification result corresponding to the patient image is output. This process realizes the automated diagnosis of AD and provides reliable support for clinical decision-making.
[0083] To verify the effectiveness of the Alzheimer's disease image classification method based on the Mamba model proposed in the present invention, this example conducted multiple sets of experiments on a PET dataset and highlighted the advantages of the technical solution of the present invention in multiple evaluation indicators.
[0084] Table 1 Dataset sample category distribution details
[0085]
[0086] Table 1 shows the sample category distribution of the dataset after preprocessing in step S1. Subsequently, the dataset was randomly divided into training, validation, and test sets according to the dataset partitioning ratio in step S2. During the model training phase, the number of iterations was set to 150, and the batch size was set to 6. The Adam optimizer with a weight decay of 1e-5 was selected, and the learning rate was flexibly adjusted using a double cosine annealing strategy with an oscillation period of 75. Before loading the training set into the model, random data augmentation was performed according to step S2.
[0087] The present invention evaluates two classification tasks: AD classification (AD vs. NC) and MCI conversion prediction (sMCI vs. pMCI). The advantages of the present invention are reflected below through the comparative experiment and ablation experiment results of this embodiment.
[0088] Table 2 Ablation test results of main components of the model according to the present invention
[0089]
[0090] Table 2 shows the ablation experiment results of this embodiment on the test set. The evaluation indicators used in the table include accuracy (ACC), sensitivity (SEN), specificity (SPE) and area under the curve (AUC). In order to test the impact of design choices on AD classification using the model proposed in the present invention, a series of ablation experiments were conducted to verify the impact of different modules on model performance. For fair comparison, the Lmamba block without SCB was used as the baseline backbone network, and all models were trained using the experimental parameter configuration above, as shown in Table 2. The analysis of the ablation experiment results shows that each module proposed in this article has a significant contribution to the model performance. In the AD and NC classification tasks, the introduction of the SCB block alone increased the accuracy from 79.56% to 86.13%, and the AUC value from 86.34% to 90.86%, indicating that SCB effectively enhances the spatial feature extraction capability. While the IBB module achieved an 8.16% improvement in accuracy, similar to the 6.57% improvement achieved by SCB, it better balanced sensitivity and specificity, achieving a higher AUC (94.46%). The combined application of CSCAF and CSPM demonstrated even stronger performance (ACC 94.85% / AUC 97.24%), demonstrating the critical role of cross-scale feature fusion in the early diagnosis of Alzheimer's disease.
[0091] Finally, the proposed method achieved the best performance in the AD classification task, with an ACC of 97.03% and an AUC of 98.23%, significantly outperforming other ablation configurations.
[0092] In the more challenging MCI conversion prediction task, each module improved results compared to the baseline. However, the synergistic effect of the combined modules was particularly significant. While maintaining high sensitivity, our method increased its specificity to 89.23% and its AUC to 83.57%, a 12.47% improvement over the baseline. Notably, the combination of CSCAF and CSPM demonstrated performance gains exceeding those of the individual modules in both tasks (ACC 74.59% / AUC 77.56%), validating the universal effectiveness of cross-scale channel attention feature fusion and channel spatial perception. Finally, the complete model integrating SCB, IBB, and CSCAF-CSPM achieved the best performance in both tasks, confirming that the modules jointly optimize the model's ability to represent neurodegenerative diseases. Table 3 shows the performance evaluation of our method compared with existing research results. These comparison methods used the ADNI public dataset. The performance of the proposed model was compared with models from recent state-of-the-art papers. These models are currently leading methods in the field. Experimental results show that our method demonstrates significant advantages in the classification of AD and NC. The method of the present invention achieved an accuracy of 97.03%, a 0.13 percentage point improvement over the existing best-performing method by Rehman et al.; its sensitivity was 97.49%, slightly higher than Rehman's 96.10%; and its specificity was 96.72%, slightly lower than Rehman's 97.50%, but still superior to other compared methods. Furthermore, the AUC value of the method of the present invention reached 98.23%, significantly higher than Pan et al.'s 97.11% and Rehman et al.'s 96.90%, demonstrating that the model has a higher ability to distinguish between AD and NC, confirming the high efficiency of the present invention.
[0093] Table 3 Model performance evaluation of the method of the present invention compared with existing research results
[0094]
[0095] The theoretical analysis and experimental results of the above embodiments reflect the effectiveness and advantages of the designed three-dimensional positron emission tomography (PET) image classification method based on the Mamba model and the constructed system, providing a complete and effective implementation example.
[0096] The same or similar reference numerals correspond to the same or similar components;
[0097] The terms used in the drawings to describe positional relationships are for illustrative purposes only and are not to be construed as limiting the present invention.
[0098] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A three-dimensional positron emission tomography data image classification method based on multi-stage progressive Lmamba block feature extraction, characterized in that: The following steps are involved: S1: A systematic preprocessing process is performed on the raw positron emission tomography imaging data to construct the image data format for modeling; S2: At the beginning of training, data augmentation strategies are randomly applied to the positron emission tomography image dataset at a ratio of 50%, including random rotation, gamma adjustment and random masking; S3: In the backbone input module of the method model, a large kernel convolution layer is used for coarse feature extraction; S4: In the main branch of the method model, a global information representation of the whole brain image is constructed through multi-stage progressive feature extraction modules; S5: On the branch branches of the method model, the cross-scale channel attention fusion module (CSCAF) is used step by step to fuse global and local semantic features. S6: In the method model, the main features and branch features of the whole brain image are integrated through the channel spatial perception mechanism (CSPM) to output high-dimensional semantic features. S7: On the method model, long-range spatial and position information are mixed via an inverted bottleneck block. S8: On the classifier network of the method model, the fused feature map is predicted to obtain the disease classification result corresponding to the patient image.
2. The method for three-dimensional positron emission tomography data image classification based on multi-stage progressive feature extraction according to claim 1, characterized in that: The pre-processing process of S1 includes: S11: Register the original PET images to the Montreal Neurological Institute (MNI) standard space; S12: removal of skull and non-brain tissues; S13: The image was smoothed using an 8mm full-width at half maximum Gaussian smoothing filter; Where G(x,y) represents the image pixel value after Gaussian smoothing, and σ represents the standard deviation of the Gaussian kernel. S14: Intensity normalization is performed on the value of each voxel, that is, the value of each voxel is normalized to the original value minus the minimum value divided by the difference between the maximum and minimum values: Where I' represents the normalized voxel value, I represents the original pixel value, and I max and I min Represent the maximum and minimum values of the voxel respectively.
3. The method according to claim 1, characterized in that The data augmentation strategies in S2 include: random rotation (generating different viewing angles within a specified range of -30° to +30°); gamma adjustment (adjusting image contrast); and random masking (zeroing voxels in random regions). Randomly applying data augmentation strategies includes one or more of random rotation, gamma adjustment, and random masking.
4. The method according to claim 1, wherein The large-kernel convolution layer of S3 implements coarse feature extraction in the following manner: a three-dimensional convolution operation with a convolution kernel size of K=7 is used.
5. The method according to claim 1, wherein The multi-stage progressive Lmamba feature extraction module of S4 is used to extract image features step by step in multiple stages, each stage including: 1) Lmamba Block, used for feature extraction. Each Lmamba Block includes a Space Convergent Block (SCB) and a State Space Hybrid Convolution Block (SSHCB). SSHCB combines a selective state-space model, depthwise separable convolution, a linear attention mechanism, an activation function, and a gating mechanism to design a Mamba-like module (Mamba Tye Linear Attention, MTLA). 2) Use convolutional downsampling layers to perform spatial downsampling between stages.
6. The multi-stage progressive feature extraction module according to claim 5, characterized in that: The core of the Lmamba block consists of an SCB block and a SSHCB block. The specific operation process includes the following: Coarsely extracted features first pass through the SCB module, and the output is divided into two branches. Branch 1 undergoes layer normalization and SSHCB, while branch 2 combines the SCB output with branch 1 to form a residual connection. The output of the residual connection is further divided into two branches: one undergoes layer normalization and a multi-layer perceptron, while the other branch remains unchanged to form a second residual structure. The final output passes through a downsampling layer and enters the next stage.
7. The method according to claim 1, characterized in that The CSCFM of S5 achieves feature fusion in the following way: the output feature of the backbone stage is The output feature of cross-scale branch fusion is F1 i-1 , where subscript 0 represents the output of each stage of the trunk after downsampling, subscript 1 represents the output of the branch, and superscript i represents the i-th stage of the trunk output (i=1,2,3,4). When i=1, the branch F1 0 is the output of the main layer. This process is expressed by the formula: Among them, ECA represents the channel attention mechanism, cat represents the concatenation of feature maps along the channel dimension, and σ represents the GELU activation function.
8. The method according to claim 1, characterized in that The S6 CSPM includes the following steps: S61: Capturing local cross-channel interactions using channel-aware networks: Among them F cp Represents channel-aware output, performs lightweight channel interaction through one-dimensional convolution (Conv1d), σ represents the sigmoid activation function, Avgpool represents average pooling, and * represents element-by-element multiplication. S62: Weighting features using spatially aware networks: S621: F cp with F1 3 Combined to form complementary information: F in =F cp +F1 3 ; S622: Split from width, height, and depth dimensions and use average pooling: S623: Splicing and using convolution blocks for spatial perception to generate feature weights for each dimension: (f' W ,f' H ,f' D )=split[σ[BN[Conv[cat(f w ,f h ,f d )]]]]]; Where cat represents the concatenation operation, Conv represents 1D convolution, BN represents layer normalization, σ represents the GELU activation function, and split represents segmentation along the height, width, and depth dimensions. S624: Use the activation function σ2 to generate a weight matrix with a mask in the range [0-1] and F in Features are weighted: M s =σ2(f' W ,f' H ,f' D ),F CSPM =M s *F in ; Among them, M s represents the weight matrix, σ2 represents the sigmoid activation function, F CSPM Output feature map for CSPM. A spatial perception network is used to extract features from three dimensions: width, height, and depth, and a channel spatial attention weight coefficient matrix is generated through the sigmoid function.
9. The method according to claim 1, characterized in that The inverted bottleneck block (IBB) of S7 includes the following steps: S71: Expand the channel dimension through depthwise separable convolution, use GELU function for nonlinear activation and then perform batch normalization, and finally introduce the input to form the residual. This process can be expressed as: F' out =BN(σ(DW(F CSPM )))+F CSPM ; where F′ out Represents the output after the depthwise separable convolutional layer. S72: Feature extraction through point convolution: The hidden dimension between two point convolutions is set to four times the input dimension; the expressive power of the features is enhanced through batch normalization and activation function. This process is expressed as: F out =BN(σ(PW(F′ out )))。 10. The method according to claim 1, characterized in that The S8 classifier network achieves category prediction by using a global average pooling layer, a flattening operation, and a linear fully connected layer for classification. This process is expressed as: classify =Linear(View(Avgpool(F out ))).
Citation Information
Cited By
Remote sensing image semantic segmentation method and device, equipment and medium
CN121213935A
Oral cavity contour modeling method based on context enhanced mixed vision
CN121259217A
Medical image classification method and system of structure perception state space model
CN121505366A
Histopathological image cancer auxiliary diagnosis method based on pathologist cognitive simulation
CN121812116A