Alzheimer's disease early-stage image detection and classification method based on 3D-ViT and single-mode MRI
By combining 3D-ViT with single-modal MRI, using deep belief networks for integrated learning, the problems of feature extraction limitations and high computational complexity of Alzheimer's disease image detection in the prior art are solved, and efficient and accurate early image detection and classification are achieved.
Patent Information
- Application Number
- CN202510356893.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-22
AI Technical Summary
The existing Alzheimer's disease image detection method based on MRI has problems such as limitations in feature extraction, high data requirements, high computational complexity and insufficient model generalization capabilities.
Using a method based on 3D-ViT and single-modal MRI, an integrated learning model of brain extraction, image registration, deep learning analysis and deep belief networks is integrated to generate the final classification results.
It improves the accuracy and interpretability of image detection classification, reduces the computational complexity, reduces the dependence on large-scale annotation data, and enhances the generalization ability and clinical practicality of the model.
Smart Images

Figure CN120355979A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image detection technology, and in particular, to an early image detection and classification method for Alzheimer's disease based on 3D-ViT and single-modal MRI. Background Art
[0002] Alzheimer's Disease (AD) is a common neurodegenerative disease, mainly manifested as cognitive decline and memory impairment. With the aggravation of the global aging problem, the incidence of Alzheimer's disease has been increasing year by year, and it has become one of the major challenges in the field of global public health. Early detection and diagnosis are of great significance for delaying the progression of the disease, improving the quality of life of patients, and reducing the social medical burden. In recent years, with the rapid development of medical imaging technology, magnetic resonance imaging (MRI) has become an important tool for studying Alzheimer's disease, especially structural magnetic resonance imaging (sMRI), which can provide high-resolution brain structure information for detecting brain atrophy and structural changes. Although significant progress has been made in the research and clinical application of MRI-based Alzheimer's disease detection methods, there are still many technical bottlenecks and practical challenges.
[0003] In the prior art, the image detection methods for Alzheimer's disease based on MRI are mainly divided into two categories: traditional machine learning methods and deep learning methods. Traditional machine learning methods usually rely on manually extracting features, such as gray matter volume, white matter volume, cerebrospinal fluid volume, etc., which can reflect the macroscopic changes of the brain structure. By combining classifiers (such as support vector machine SVM, random forest, etc.), traditional machine learning methods can achieve the classification of Alzheimer's disease. However, the process of manually extracting features is cumbersome and time-consuming, and may miss important non-linear information or local details. Secondly, the quality of feature extraction highly depends on the experience of domain experts, which may lead to subjective biases. In addition, traditional machine learning methods are prone to overfitting problems when dealing with high-dimensional data, especially when the sample size is limited, and the generalization ability of the model is poor. Summary of the Invention
[0004] The embodiments of the present application provide an early image detection and classification method for Alzheimer's disease based on 3D-ViT and single-modal MRI, which solves the problems of limited feature extraction, high data requirements, large computational complexity, and insufficient model generalization ability of the existing methods.
[0005] In a first aspect, the embodiments of the present application provide an early image detection and classification method for Alzheimer's disease based on 3D-ViT and single-modal MRI, and the method includes the following steps:
[0006] Obtain an MRI image, perform brain extraction on the MRI image to isolate the brain tissue, remove the non-brain tissue part, and obtain a ROI image;
[0007] Register a number of ROI images onto a predefined standard template through a set process;
[0008] Perform in-depth analysis on the ROI image based on a deep learning model;
[0009] Adopt a deep belief network as an ensemble learning model, integrate the prediction results of multiple deep learning models, and generate a final classification result.
[0010] Furthermore, the performing brain extraction on the MRI image to isolate the brain tissue, remove the non-brain tissue part, and obtain a ROI image includes:
[0011] Perform non-uniformity correction on the MRI image through a correction algorithm to compensate for and correct the non-uniformity artifacts in the image;
[0012] Adopt the Pincram method to perform brain extraction on the MRI image, identify and extract the boundary of the brain tissue to isolate the brain tissue, remove the non-brain tissue part, and obtain a ROI image.
[0013] Furthermore, the registering a number of ROI images onto a predefined standard template through a set process includes:
[0014] Use the MALPEM4D process for image registration to register a number of ROI images onto a predefined standard template;
[0015] Perform ROI identification. Through MALPEM4D, structural segmentations of a number of brain regions are created, and each brain region is assigned a unique voxel value corresponding to the region number in the mask volume.
[0016] Furthermore, the performing in-depth analysis on the ROI image based on a deep learning model includes:
[0017] Adopt a 3D-ViT model to perform deep learning analysis on the ROI image. Through the way of tubelet embedding, non-overlapping tubelets in the 3D MRI input volume are extracted and linearly embedded;
[0018] The input sequence z ∈ RN is obtained through the following formula:
[0019] z = [Ex1, Ex2, …, ExN] + p;
[0020] where x represents the features corresponding to each brain region obtained after feature extraction; E is a projection operation, obtained through a learned 3D convolutional filter and flattened into a 1D vector;
[0021] $z\in\mathbb{R}^N$ is the position embedding index from 1 to N, which transforms the 3D MRI data into sequence data suitable for processing by the Transformer model;
[0022] The D-ViT model uses L encoder transformer layers to process tokens, and each layer contains layer normalization and a multi-head self-attention block, specifically as follows:
[0023] $y_l = MSA(LN(z)) + z_l$;
[0024] $z_{l + 1} = MLP(GELU(MLP(LN(y_l)))) + y_l$;
[0025] where GELU is a non-linear activation function; both MSA and MLP represent multi-head self-attention blocks; LN represents layer normalization; $y_l$ represents the output feature of the $l$-th layer obtained by combining the multi-head self-attention mechanism and layer normalization with residual connection.
[0026] Furthermore, the use of a deep belief network as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate a final classification result includes:
[0027] A deep belief network DBN is used as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate a final classification result; where DBN consists of multiple restricted Boltzmann machines (RBMs).
[0028] Furthermore, the obtaining of the MRI image includes:
[0029] Obtaining the MRI image through a magnetic resonance imaging machine.
[0030] In a second aspect, an Alzheimer's disease early image detection and classification device based on 3D-ViT and single-modal MRI provided by an embodiment of the present application includes:
[0031] A brain extraction module and an ROI recognition module, used to obtain an MRI image and perform brain extraction on the MRI image to separate brain tissues, remove non-brain tissue parts, and obtain an ROI image;
[0032] An image registration module, used to register several ROI images onto a predefined standard template through a set process;
[0033] A 3D-ViT module, used to perform in-depth analysis on the ROI image based on a deep learning model;
[0034] A DBN integration module, used to use a deep belief network as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate a final classification result.
[0035] In a third aspect, an embodiment of the present application further provides a computer device, including: a memory and one or more processors;
[0036] The memory is used to store one or more programs;
[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement an early Alzheimer's disease image detection and classification method based on 3D-ViT and single-modal MRI as described above.
[0038] In a fourth aspect, an embodiment of the present application further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute an early Alzheimer's disease image detection and classification method based on 3D-ViT and single-modal MRI as described above when executed by a computer processor.
[0039] In a fifth aspect, an embodiment of the present application provides a computer program product, the computer program product includes instructions, and when the instructions are executed by a computer, the computer implements the method described above.
[0040] The embodiment of the present application acquires an MRI image, performs brain extraction on the MRI image to isolate brain tissue, removes non-brain tissue parts, and obtains a ROI image; registers a number of ROI images to a predefined standard template through a set process; uses a deep belief network as an ensemble learning model to perform in-depth analysis on the ROI image based on a deep learning model, integrates the prediction results of multiple deep learning models, and generates a final classification result; utilizes the powerful capabilities of deep learning to extract and analyze fine-grained features from specific regions of the brain, thereby improving the accuracy and interpretability of image detection and classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of an early Alzheimer's disease image detection and classification method based on 3D-ViT and single-modal MRI provided by an embodiment of the present application;
[0042] Figure 2 is a schematic structural diagram of an early Alzheimer's disease image detection and classification device based on 3D-ViT and single-modal MRI provided by an embodiment of the present application;
[0043] Figure 3 is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of this application more clear, the following further describes specific embodiments of this application in conjunction with the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application. Additionally, it should be noted that for the convenience of description, only parts related to this application are shown in the drawings rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0045] Magnetic resonance imaging (MRI) is a non-invasive imaging technique that can provide high-resolution images of the brain structure. MRI has significant advantages in detecting brain structure changes related to AD and MCI. In recent years, deep learning techniques have made remarkable progress in medical image analysis, especially in the automatic diagnosis and classification of AD and MCI. Deep learning models can directly learn complex and high-level features from MRI data to achieve accurate and timely diagnosis.
[0046] An early image detection and classification method for Alzheimer's disease based on 3D-ViT and single-modal MRI is established in the embodiments of this application to solve the problems of existing feature extraction limitations, high data requirements, large computational complexity, and insufficient model generalization ability.
[0047] The early image detection and classification method for Alzheimer's disease based on 3D-ViT and single-modal MRI provided in the embodiments can be executed by an early image detection and classification device for Alzheimer's disease based on 3D-ViT and single-modal MRI. The early image detection and classification device for Alzheimer's disease based on 3D-ViT and single-modal MRI can be implemented in software and / or hardware and integrated in an early image detection and classification device for Alzheimer's disease based on 3D-ViT and single-modal MRI. Among them, the early image detection and classification device for Alzheimer's disease based on 3D-ViT and single-modal MRI can be a device such as a computer.
[0048] Figure 1 This is a flowchart of an early image detection and classification method for Alzheimer's disease based on 3D-ViT and single-modal MRI provided in the embodiments of this application. Refer to Figure 1 , the method includes the following steps:
[0049] 100. Obtain an MRI image, perform brain extraction on the MRI image to isolate the brain tissue, remove the non-brain tissue part, and obtain a ROI image.
[0050] Specifically, obtain the MRI image through a magnetic resonance imaging machine; perform non-uniformity correction on the MRI image through a correction algorithm to compensate for and correct the non-uniformity artifacts in the image; use the Pincram method to perform brain extraction on the MRI image, identify and extract the boundary of the brain tissue to isolate the brain tissue, remove the non-brain tissue part, and obtain a ROI image.
[0051] Exemplarily, first perform brain extraction on the MRI image. Its core objective is to accurately isolate the brain tissue from the MRI scan image, remove the irrelevant non-brain tissue part, so that subsequent analysis can focus on the brain region, improving the analysis efficiency and accuracy. The embodiment of the present application uses the Pincram method for brain extraction. This method is known for its flexibility and high precision and can accurately label the adult brain region. The Pincram method analyzes the 3D T1-weighted magnetic resonance head image through a series of complex algorithms and image processing techniques, identifies and extracts the boundary of the brain tissue, and then realizes the accurate extraction of the brain region. This process not only improves the quality of the brain image but also provides a clear image basis for subsequent image registration and ROI identification.
[0052] After brain extraction, non-uniformity correction becomes another important link in the preprocessing. During the MRI image scanning process, due to the influence of various factors, such as magnetic field non-uniformity, equipment performance differences, etc., non-uniformity artifacts often appear in the MRI image. These artifacts will have a negative impact on the image quality and analysis results. Therefore, the embodiment of the present application performs non-uniformity correction on the MRI image in the preprocessing stage. Through an advanced correction algorithm, it compensates for and corrects the non-uniformity artifacts in the image, making the signal intensity distribution of the image more uniform, thereby improving the image quality and the accuracy of analysis. This correction process is of great significance for subsequent image analysis and diagnosis and can effectively reduce the interference of artifacts on the diagnosis results.
[0053] 200. Register a number of ROI images onto a predefined standard template through a set process.
[0054] Specifically, use the MALPEM4D process for image registration to register a number of ROI images onto a predefined standard template; perform ROI identification. Through MALPEM4D, a structural segmentation of several brain regions is created, and each brain region is assigned a unique voxel value corresponding to the region number in the mask volume.
[0055] Exemplarily, to achieve consistent region of interest (ROI) extraction in the human brain, the embodiments of the present application use the MALPEM4D process for image registration. This process is an efficient image registration method that can accurately register MRI images of various human brains onto a predefined standard template. Through this process, the brain images of different human brains are unified and standardized in space, providing a consistent reference framework for subsequent ROI identification and analysis.
[0056] Based on image registration, the embodiments of the present application further perform ROI identification. The purpose of ROI identification is to accurately locate and identify specific brain regions related to AD diagnosis in the standardized brain images. The embodiments of the present application create a structural segmentation of 138 brain regions through MALPEM4D, and each brain region is assigned a unique voxel value corresponding to the region number in the mask volume. This fine brain segmentation and identification method enables subsequent deep learning analysis to perform feature extraction and analysis for specific brain regions, improving the pertinence and accuracy of diagnosis. Through in-depth analysis of these specific brain regions, the pathological characteristics and development trends of AD in the brain can be better understood, providing an important basis for the early diagnosis and treatment of the disease.
[0057] 300. Perform in-depth analysis on the ROI image based on the deep learning model.
[0058] Specifically, a 3D-ViT model is used to perform deep learning analysis on the ROI image. Through the way of tubelet embedding, non-overlapping tubelets in the 3D MRI input volume are extracted and linearly embedded;
[0059] The input sequence z∈RN is obtained through the following formula:
[0060] z = [Ex1, Ex2, …, ExN] + p;
[0061] Among them, x refers to the feature corresponding to each brain region obtained after feature extraction, which identifies the gray matter volume information from each segmented brain region. E is a projection operation, obtained through a learned 3D convolutional filter and flattened into a 1D vector;
[0062] z∈RN is the position embedding index from 1 to N, which converts the 3D MRI data into sequence data suitable for processing by the Transformer model;
[0063] The D-ViT model uses L = 8 encoder transformer layers to process tokens, and each layer contains layer normalization and a multi-head self-attention block, specifically as follows:
[0064] yl = MSA(LN(z)) + zl;
[0065] zl + 1 = MLP(GELU(MLP(LN(yl)))) + yl;
[0066] Among them, GELU is a non - linear activation function; MSA and MLP both represent multi - head self - attention blocks; LN represents layer normalization; yl represents the output feature of the l - th layer obtained by combining the multi - head self - attention mechanism and layer normalization processing with residual connection.
[0067] Through this multi - layer encoder transformation, the model can perform deep feature extraction and analysis on the input sequence data, and learn complex features and patterns related to the AD category. The encoded output passes through normalization, global average pooling layer and another MLP, and finally calculates the network output through the softmax function, and trains the network using the sparse categorical cross - entropy loss function. This process enables the model to accurately classify and diagnose AD, improving the accuracy and reliability of the diagnosis.
[0068] 400. Use a deep belief network as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate the final classification result.
[0069] Specifically, a deep belief network (DBN) is used as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate the final classification result; among them, DBN is composed of multiple restricted Boltzmann machines (RBMs).
[0070] The strengths are comparable. To further improve the prediction accuracy and robustness, the embodiment of this application uses a deep belief network (DBN) as an ensemble learning model. DBN is composed of multiple restricted Boltzmann machines (RBMs) and can effectively integrate the prediction results of multiple 3D - ViTs. During the training process, to prevent overfitting and reduce the impact of class imbalance, the embodiment of this application introduces a dropout rate of 0.1 in the multi - head attention layer, allowing some features to be randomly "ignored" during training. In addition, to enable the same network architecture to be used in different regions, the embodiment of this application uses spline interpolation to resize the ROI to a size of 28×28×28 voxels and uses a kernel and stride of 8×8×8 in the projection operation EE. The comprehensive application of these technical means enables the embodiment of this application to have higher accuracy and robustness in AD diagnosis, providing strong support for the development of the medical image diagnosis field.
[0071] As described above, the embodiments of the present application acquire MRI images, perform brain extraction on the MRI images to isolate brain tissues, remove non-brain tissue parts, and obtain ROI images; register a number of ROI images onto a predefined standard template through a set process; use a deep belief network as an ensemble learning model for in-depth analysis of ROI images based on a deep learning model, integrate the prediction results of multiple deep learning models, and generate a final classification result; utilize the powerful capabilities of deep learning to extract and analyze fine-grained features from specific regions of the brain, thereby improving the accuracy and interpretability of image detection and classification.
[0072] The embodiments of the present application solve the limitations of feature extraction. In the prior art, traditional machine learning methods rely on manual feature extraction, which is a cumbersome process and may miss important information, while the features extracted by deep learning methods lack interpretability. The embodiments of the present application use a hybrid feature extraction method, combining traditional manual features and high-level features automatically extracted by deep learning, to make up for the limitations of a single method.
[0073] Solve the high data requirements and difficult annotation. Deep learning methods require a large amount of labeled data, while medical image data is costly to obtain and difficult to annotate. The embodiments of the present application are based on single-modal MRI data and combine traditional machine learning methods, reducing the dependence on large-scale labeled data, and at the same time utilizing the advantages of deep learning to improve the classification performance under small-sample data.
[0074] Solve the high computational complexity. In the prior art, deep learning models and multi-modal methods have high computational complexity and are difficult to be popularized in clinical environments with limited resources. The embodiments of the present application optimize the model architecture, reduce the number of convolutional layers and the number of feature maps, and use single-modal data, reducing the computational complexity and making it more suitable for actual clinical environments.
[0075] Solve the insufficient model generalization ability. Existing methods have significant differences in performance on different datasets and have poor generalization ability. The embodiments of the present application improve the generalization ability of the model and avoid the overfitting problem through feature fusion and model optimization, combining traditional features and deep learning features.
[0076] Solve the insufficient interpretability and clinical practicality. Traditional machine learning methods have strong interpretability but limited performance, and the "black box" characteristics of deep learning methods limit their clinical practicality. The embodiments of the present application use a hybrid feature extraction method, combining manual features and deep learning features, improving the interpretability of the model while ensuring performance, and facilitating clinicians to understand and accept.
[0077] The potential of unimodal data is not fully exploited. In the prior art, multimodal methods are complex and costly, while the potential of unimodal MRI data has not been fully explored. The embodiments of the present application focus on unimodal data and fully exploit its potential through hybrid feature extraction and model optimization, avoiding the complexity and high cost of multimodal data.
[0078] Based on the above embodiments, please refer to Figure 2 , an early Alzheimer's disease image detection and classification device based on 3D-ViT and unimodal MRI provided by the embodiments of the present application, the early Alzheimer's disease image detection and classification device based on 3D-ViT and unimodal MRI specifically includes: a brain extraction module 201, an ROI recognition module 202, an image registration module 203, a 3D-ViT module 204, and a DBN integration module 205.
[0079] Among them, the brain extraction module 201 and the ROI recognition module 202 are used to obtain an MRI image and perform brain extraction on the MRI image to separate brain tissues, remove non-brain tissue parts, and obtain an ROI image; the image registration module 203 is used to register a plurality of ROI images onto a predefined standard template through a set process; the 3D-ViT module 204 is used to perform in-depth analysis on the ROI image based on a deep learning model; the DBN integration module 205 is used to use a deep belief network as an ensemble learning model, integrate the prediction results of multiple deep learning models, and generate a final classification result.
[0080] Exemplarily, the brain extraction module is one of the core components of the system, and its main function is to preprocess the MRI image and remove irrelevant non-brain tissue parts. This module uses the Pincram method and analyzes the 3D T1-weighted magnetic resonance head image through a series of complex algorithms and image processing techniques to identify and extract the boundaries of brain tissues, achieving accurate extraction of the brain region. This process not only improves the quality of brain images but also provides a clear image basis for subsequent image registration and ROI recognition. The efficiency and accuracy of the brain extraction module are crucial for the performance of the entire system, and it directly affects the accuracy and reliability of subsequent analysis.
[0081] The image registration module is responsible for unifying the MRI images of different human brains onto a standard template. This module uses the MALPEM4D process for image registration and precisely registers the brain images of various human brains onto a predefined standard template through an efficient image registration algorithm. This process realizes the extraction of consistent regions of interest (ROIs) across human brains and provides a consistent reference framework for subsequent ROI recognition and analysis. The accuracy and stability of the image registration module have an important impact on the overall performance of the system, and it ensures the consistency and comparability of different human brain images.
[0082] The main task of the ROI recognition module is to accurately locate and identify specific brain regions related to AD diagnosis in standardized brain images. This module creates a structural segmentation of 138 brain regions through MALPEM4D, and each brain region is assigned a unique voxel value corresponding to the region number in the mask volume. This fine-grained brain region segmentation and recognition method enable subsequent deep learning analysis to perform feature extraction and analysis for specific brain regions, improving the pertinence and accuracy of diagnosis. The accuracy and reliability of the ROI recognition module are crucial for the diagnostic performance of the system, which is directly related to the accuracy and effectiveness of subsequent analysis.
[0083] The 3D-ViT module is the core analysis component of the system, responsible for performing deep learning analysis on each identified ROI. This module adopts an advanced 3D-ViT model. Through the way of tubelet embedding, non-overlapping tubelets in the 3D MRI input volume are extracted and linearly embedded, transformed into sequence data suitable for processing by the Transformer model. Then, through multiple layers of encoder transformations, the model performs deep feature extraction and analysis on the input sequence data, learning complex features and patterns related to the AD category. Finally, the network output is calculated through the softmax function, and the network is trained using the sparse categorical cross-entropy loss function to achieve accurate classification and diagnosis of AD. The efficiency and accuracy of the 3D-ViT module are the key to the diagnostic performance of the system, which directly affects the accuracy and reliability of the diagnostic results.
[0084] The DBN integration module is the ensemble learning component of the system, responsible for integrating the prediction results of multiple 3D-ViTs to generate the final classification result. This module adopts a deep belief network (DBN), which consists of multiple restricted Boltzmann machines (RBMs), and can effectively integrate the prediction results of multiple 3D-ViTs, improving the accuracy and robustness of prediction. During the training process, to prevent overfitting and mitigate the impact of class imbalance, the DBN integration module introduces a dropout rate of 0.1 in the multi-head attention layer, allowing some features to be randomly "ignored" during training. In addition, to enable the same network architecture to be used for different regions, the DBN integration module also uses spline interpolation to resize the ROI to a size of 28×28×28 voxels and uses a kernel and stride of 8×8×8 in the projection operation. The efficiency and stability of the DBN integration module have an important impact on the overall performance of the system, ensuring the accuracy and reliability of the final diagnostic results.
[0085] The connection relationships between these key components are sequential, that is, the output of the previous module serves as the input of the next module. This sequential connection method makes the working process of the entire system clear and efficient, enabling the modules to cooperate closely and jointly achieve the efficient processing of MRI data and the accurate diagnosis of AD. Starting from the brain extraction module, it passes through the image registration module, ROI recognition module, 3D-ViT module in sequence, and finally the DBN integration module generates the final diagnosis result. The design of the entire system fully considers the functions and performances of each module. Through reasonable module division and connection methods, it realizes the efficient processing of MRI data and the accurate diagnosis of AD, providing strong support for the development of the medical imaging diagnosis field.
[0086] The key innovation of the embodiments of this application lies in the combination of three-dimensional vision transformer (3D-ViT) and region of interest (ROI) analysis, and the use of deep belief network (DBN) for integrated learning, thereby significantly improving the accuracy and interpretability of the model in the early detection of Alzheimer's disease (AD) and mild cognitive impairment (MCI). By finely dividing the brain into 138 predefined regions, the embodiments of this application can not only deeply analyze the contribution of each region to disease classification, but also provide strong data support for the formulation of clinical diagnosis and treatment strategies. This innovative method not only demonstrates excellent performance in multiple classification tasks, but also opens up a new research path for understanding the complex relationship between brain regions and disease progression, and is expected to promote significant progress in AD and MCI diagnosis technologies.
[0087] In summary, the embodiments of this application provide a new solution for the early detection of AD and MCI by cleverly integrating ROI analysis and 3D-ViT technology and leveraging the powerful integrated learning ability of DBN. This solution is not only innovative in technical implementation, but also shows great potential and value in practical applications. It can not only improve the accuracy of diagnosis and reduce the misdiagnosis rate, but also help doctors better understand the pathogenesis of the disease, so as to formulate more personalized and accurate treatment plans and improve the quality of life of patients.
[0088] The Alzheimer's disease early image detection and classification device based on 3D-ViT and single-modal MRI provided by the embodiments of this application can be used to execute the Alzheimer's disease early image detection and classification method based on 3D-ViT and single-modal MRI provided by the above embodiments, and has corresponding functions and beneficial effects.
[0089] The embodiments of this application also provide a computer device, which can integrate the Alzheimer's disease early image detection and classification device based on 3D-ViT and single-modal MRI provided by the embodiments of this application. Figure 3 is a schematic structural diagram of a computer device provided by the embodiments of this application. Refer toFigure 3 , the computer device includes: an input device 33, an output device 34, a memory 32, and one or more processors 31; the memory 32 is used to store one or more programs; when the one or more programs are executed by the one or more processors 31, the one or more processors 31 implement the early Alzheimer's disease image detection and classification method based on 3D-ViT and single-modal MRI provided in the above embodiments. Among them, the input device 33, the output device 34, the memory 32, and the processor 31 can be connected through a bus or other means, Figure 3 taking the connection through the bus as an example.
[0090] The processor 31 executes various functional applications and data processing of the device by running software programs, instructions, and modules stored in the memory 32, that is, implements the above-mentioned early Alzheimer's disease image detection and classification method based on 3D-ViT and single-modal MRI.
[0091] The computer device provided above can be used to execute the early Alzheimer's disease image detection and classification method based on 3D-ViT and single-modal MRI provided in the above embodiments, and has corresponding functions and beneficial effects.
[0092] The embodiment of the present application also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a method for early Alzheimer's disease image detection and classification based on 3D-ViT and single-modal MRI when executed by a computer processor. The method for early Alzheimer's disease image detection and classification based on 3D-ViT and single-modal MRI includes: adding scheduling information through the scheduling system background, where the scheduling information includes point location, path, and action information; automatically generating a running path based on the starting point and target point of the robot and dispatching it to the robot; according to the real-time task status and real-time working status of the elevator, combined with the real-time action information of the robot, adjusting the running path in real time and controlling the robot to run.
[0093] Storage medium - Any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks or tape drives; computer device memories or random access memories such as DRAM, DDRRAM, SRAM, EDORAM, Rambus RAM, etc.; non-volatile memories such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium may also include other types of memory or combinations thereof. Additionally, the storage medium may be located in a first computer device in which the program is executed, or may be located in a different second computer device that is connected to the first computer device via a network (such as the Internet). The second computer device may provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media that may reside in different locations (such as in different computer devices connected via a network). The storage medium may store program instructions executable by one or more processors (e.g., embodied as a computer program).
[0094] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present application, the computer-executable instructions are not limited to the Alzheimer's disease early image detection and classification method based on 3D-ViT and unimodal MRI as described above, and can also perform related operations in the Alzheimer's disease early image detection and classification method provided by any embodiment of the present application.
[0095] The Alzheimer's disease early image detection and classification device, storage medium, and computer device provided in the above embodiments can execute the Alzheimer's disease early image detection and classification method provided by any embodiment of the present application. For technical details not described in detail in the above embodiments, reference can be made to the Alzheimer's disease early image detection and classification method provided by any embodiment of the present application.
[0096] The embodiments of the present application further provide a computer program product. The methods described in the embodiments of the present application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM (Open Application Model), or other programmable devices.
[0097] The above is only the preferred embodiment of the present application and the technical principles applied. The present application is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions that can be made by those skilled in the art will not depart from the protection scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, it can also include more other equivalent embodiments, and the scope of the present application is determined by the scope of the claims.
Claims
1. A method for early image detection and classification of Alzheimer's disease based on 3D-ViT and unimodal MRI, characterized in that, The method includes the following steps: Obtain an MRI image, perform brain extraction on the MRI image to isolate brain tissue, remove non-brain tissue parts, and obtain a ROI image; Register a number of ROI images onto a predefined standard template through a set process; Perform in-depth analysis on the ROI image based on a deep learning model; Adopt a deep belief network as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate a final classification result.
2. The early image detection and classification method for Alzheimer's disease based on 3D-ViT and unimodal MRI according to claim 1, characterized in that, The performing brain extraction on the MRI image to isolate brain tissue, remove non-brain tissue parts, and obtain a ROI image includes: Perform non-uniformity correction on the MRI image through a correction algorithm to compensate for and correct non-uniformity artifacts in the image; Adopt the Pincram method to perform brain extraction on the MRI image, identify and extract the boundaries of brain tissue to isolate brain tissue, remove non-brain tissue parts, and obtain a ROI image.
3. The early image detection and classification method for Alzheimer's disease based on 3D-ViT and unimodal MRI according to claim 1, characterized in that, The registering a number of ROI images onto a predefined standard template through a set process includes: Use the MALPEM4D process for image registration to register a number of ROI images onto a predefined standard template; Perform ROI identification. Through MALPEM4D, structural segmentation of several brain regions is created, and each brain region is assigned a unique voxel value corresponding to the region number in the mask volume.
4. The method for early image detection and classification of Alzheimer's disease based on 3D-ViT and unimodal MRI according to claim 1, characterized in that The performing in-depth analysis on the ROI image based on a deep learning model includes: Adopt a 3D-ViT model to perform deep learning analysis on the ROI image. Through the way of tubelet embedding, extract and linearly embed non-overlapping tubelets in the 3D MRI input volume; The input sequence z ∈ RN is obtained through the following formula: z = [Ex1, Ex2, …, ExN] + p; Where, x represents the features corresponding to each brain region obtained after feature extraction; E is a projection operation, obtained through a learned 3D convolutional filter and flattened into a 1D vector; z ∈ RN is the position embedding index from 1 to N, which converts 3D MRI data into sequence data suitable for processing by the Transformer model; The D-ViT model uses L encoder transformer layers to process tokens, and each layer includes layer normalization and a multi-head self-attention block, specifically as follows: yl = MSA(LN(z)) + zl; zl+1 = MLP(GELU(MLP(LN(yl)))) + yl; Where, GELU is a non-linear activation function; both MSA and MLP represent multi-head self-attention blocks; LN represents layer normalization; yl represents the output feature of the l-th layer obtained after the multi-head self-attention mechanism and layer normalization processing, combined with a residual connection.
5. The early image detection and classification method for Alzheimer's disease based on 3D-ViT and unimodal MRI according to claim 1, characterized in that, The adopting a deep belief network as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate a final classification result includes: Adopt a deep belief network DBN as an ensemble learning model to integrate the prediction results of multiple deep learning models and generate a final classification result; where, DBN is composed of multiple restricted Boltzmann machines (RBM).
6. The early image detection and classification method for Alzheimer's disease based on 3D-ViT and unimodal MRI according to claim 1, characterized in that, The obtaining an MRI image includes: An MRI image is acquired by a magnetic resonance imaging machine.
7. An early image detection and classification device for Alzheimer's disease based on 3D-ViT and unimodal MRI, characterized in that, Including: A brain extraction module and an ROI recognition module, which are used to acquire an MRI image, perform brain extraction on the MRI image to isolate brain tissues, remove non-brain tissue parts, and obtain an ROI image; An image registration module, which is used to register a number of ROI images onto a predefined standard template through a set process; A 3D-ViT module, which is used to perform in-depth analysis on the ROI image based on a deep learning model; A DBN integration module, which is used to adopt a deep belief network as an ensemble learning model, integrate the prediction results of multiple deep learning models, and generate a final classification result.
8. A computer device, characterized in that, Including: A memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement an early Alzheimer's disease image detection and classification method based on 3D-ViT and unimodal MRI as described in any one of claims 1-6.
9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute an early Alzheimer's disease image detection and classification method based on 3D-ViT and unimodal MRI as described in any one of claims 1-6 when executed by a computer processor.
10. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a computer, cause the computer to implement the method according to any one of claims 1-6.