A method and system for automatically predicting Alzheimer's disease based on deep learning
By introducing a coordinated attention model and a motivation and squeeze attention model in the deep learning framework, the problems of artificial feature extraction and channel weight consistency in traditional methods are solved, and high-accuracy prediction and early diagnosis of Alzheimer's disease are achieved.
Patent Information
- Application Number
- CN202111510101.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-10
AI Technical Summary
When processing medical images in the prior art, traditional machine learning methods rely on manual feature extraction, which has poor generalization and high computational cost. The deep learning model has the same weights in the convolution and pooling process, and it is impossible to effectively capture the importance of different features in classification.
The deep learning framework is adopted, combining the coordinated attention model, deep learning model and the excitation and extrusion attention model, and the position perception is extracted through the coordinated attention model, the deep learning model extracts features, and the excitation and extrusion attention model adjusts the channel weight to solve problems with different importance of different channels.
Significantly improves the accuracy of Alzheimer's disease prediction, enables early diagnosis, reduces physician workload, and improves the performance of deep learning networks.
Smart Images

Figure CN114359164B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly, to a method and system for automatically predicting Alzheimer's disease based on deep learning. Background Art
[0002] Alzheimer's disease (AD) is a progressive neurodegenerative disease of the nervous system. Clinically, it is manifested as a decline in cognitive function, mental symptoms and behavioral disorders, and a gradual decline in the ability to perform daily life activities. The onset of this disease is slow, and it is impossible to determine when the disease starts. As people age, the probability of the elderly getting the disease increases, and the number of deaths each year also increases, seriously affecting the health of the elderly and normal family life. Alzheimer's disease is irreversible. After the onset, only drugs can be relied on to relieve symptoms, and the disease process of patients cannot be changed. Since the causative factors involve many aspects and cannot be treated solely with drugs, it is necessary to conduct screening and diagnosis at the earliest possible stage in order to take appropriate intervention measures before further cognitive impairment occurs. Mild cognitive impairment (MCI) is the early stage of dementia. Patients experience a decline in cognitive function in one or more areas, but maintain the ability to live independently and have not yet reached the criteria for dementia. MCI can be divided into stable mild cognitive impairment (sMCI) and progressive mild cognitive impairment (pMCI). The cognitive state of sMCI remains stable, while the cognitive state of pMCI gradually declines and progresses to Alzheimer's disease. Therefore, classifying patients at the MCI stage helps with early intervention and treatment.
[0003] Currently, using deep learning to solve medical image problems has also become a research hotspot. Deep learning can automatically learn the relationship between tasks and features without the need for manual feature extraction, thereby improving the recognition effect of different categories. In recent years, deep learning has demonstrated excellent performance in image classification, and deep learning has also begun to be applied to the field of medical images. Deep learning models include recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep belief networks (DBNs), and stacked autoencoders (SAEs), etc. Kanghan et al. classified AD and NC based on unsupervised learning of a convolutional autoencoder (CAE) and used transfer learning for the classification of pMCI and sMCI. They proposed an end-to-end concept for classification and proposed data augmentation and regularization, and applied visualization techniques based on gradient backpropagation to the learned model. Some research has proposed a fully stacked bidirectional long short-term memory (FSBi-LSTM) that analyzes both MRI and PET data simultaneously. The long short-term memory (LSTM) can solve the problems of gradient explosion or gradient disappearance. The MRI and PET are input into a 3D CNN to extract features, and the FSBi-LSTM is used to replace the traditional fully connected layer to extract high-level semantic and spatial information. The FSBi-LSTM can obtain spatial and semantic information from the feature map.
[0004] After analysis, in the prior art, traditional machine learning methods generally extract features from medical images manually and then use a classifier for classification. This method relies on prior knowledge to extract features and requires in-depth analysis of the dataset, which is time-consuming and laborious. When dealing with 3D data, the voxel-based feature extraction method takes into account global information, but requires a large amount of computing power and computational cost. The region-based extraction method focuses features on a specified region, reducing the computational amount for the entire image, but inevitably ignores the global structural information. In short, when dealing with specific simple tasks, manual feature extraction in traditional machine learning is simple and effective, but has poor generalization and is only applicable to specific tasks. In the current solution of automatically extracting features using deep learning, in the convolution and pooling processes of deep learning, the same weights are assigned to the features extracted from different channels. However, in actual problems, the importance of different features in the classification effect should be different. Summary of the Invention
[0005] The object of the present invention is to overcome the above-mentioned defects of the prior art and provide a method and system for automatically predicting Alzheimer's disease based on deep learning.
[0006] According to the first aspect of the present invention, there is provided a method for automatically predicting Alzheimer's disease based on deep learning. The method includes the following steps:
[0007] Obtain a target image to be detected;
[0008] Input the target image into a trained deep learning framework to obtain a predicted classification result of Alzheimer's disease. The deep learning framework sequentially includes a coordinated attention model, a deep learning model, and an excitation and squeeze attention model. The coordinated attention model separately extracts individual position perceptions from different directions of the input feature map, fuses the obtained spatial information with the input feature map after weighting on the channels, and obtains a first feature map; the deep learning model uses the first feature map as input and extracts a second feature map; the excitation and squeeze attention model performs a squeeze operation on the second feature map to obtain the global feature of the channel, and performs an excitation operation on the global feature to obtain the weights of different channels.
[0009] According to the second aspect of the present invention, there is provided a system for automatically predicting Alzheimer's disease based on deep learning. The system includes:
[0010] An image acquisition unit: used to obtain a target image to be detected;
[0011] Prediction unit: It is used to input the target image into a trained deep learning framework to obtain the prediction classification result of Alzheimer's disease. The deep learning framework sequentially includes a coordinated attention model, a deep learning model, and an excitation and squeeze attention model. The coordinated attention model separately extracts individual position perceptions from different directions of the input feature map, fuses the obtained spatial information after weighting on the channels with the input feature map to obtain a first feature map; the deep learning model uses the first feature map as the input to extract a second feature map; the excitation and squeeze attention model performs a squeeze operation on the second feature map to obtain the global feature of the channels, and performs an excitation operation on the global feature to obtain the weights of different channels.
[0012] Compared with the prior art, the advantages of the present invention are as follows: It automatically predicts Alzheimer's disease, realizes the early diagnosis of Alzheimer's disease, adds an attention mechanism to the deep learning network structure, solves the loss problem caused by the different importance of different channels, significantly improves the prediction accuracy, can assist doctors in diagnosis, and thus reduces the workload of doctors.
[0013] Other features and advantages of the present invention will become clear through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.
[0015] Figure 1 It is a schematic diagram of a deep learning framework for automatically predicting Alzheimer's disease according to an embodiment of the present invention;
[0016] Figure 2 It is a schematic diagram of a dense convolutional network model according to an embodiment of the present invention;
[0017] Figure 3 It is a schematic diagram of an excitation and squeeze attention model according to an embodiment of the present invention;
[0018] Figure 4 It is a schematic diagram of a coordinated attention model according to an embodiment of the present invention;
[0019] In the accompanying drawings, Conv - Convolutional layer; CA Block - Coordinate Attention module; Dense Block - Dense block; Avg - Pool - Average pooling; SE Block - Squeeze - and - Excitation attention module; FC - Fully connected; Global pooling - Global pooling; Fully Connected - Fully connected; Re - weight - Re - adjust weights; non - liner - Non - linear; Transitionlayers - Transition layers. Detailed implementation manners
[0020] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.
[0021] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation to the present invention and its application or use.
[0022] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.
[0023] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.
[0024] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0025] The present invention proposes a deep - learning framework for automatically predicting Alzheimer's disease. Refer to Figure 1 As shown, the framework mainly includes a deep - learning model, a Coordinate Attention model (CA), and a Squeeze - and - Excitation attention model (SE). The following takes MRI data as the target image as an example for illustration, but it should be understood that the present invention is equally applicable to other types of medical images, such as PET and CSF, etc. However, PET examinations are expensive, and CSF is an invasive examination method.
[0026] Hereinafter, the deep - learning model, the Squeeze - and - Excitation attention model (or Squeeze - and - Excitation attention module), and the Coordinate Attention model (or Coordinate Attention module) will be specifically described respectively.
[0027] 1), Deep learning model
[0028] The deep learning model can adopt various types, such as convolutional neural network or deep belief network, etc. Considering that in a deep learning network, as the number of network layers continues to deepen, the problem of gradient disappearance becomes more and more serious when the input information and gradient information are transmitted. In a preferred embodiment, a dense convolutional network (DenseNet) is used to solve this problem to ensure that the maximum amount of information can be transmitted between layers of the network, and directly splice all the previous layers. The structure of the dense convolutional network is as Figure 2 shown.
[0029] The densely connected method is equivalent to directly connecting the input and the loss for each layer. Therefore, it can alleviate the problem of gradient disappearance and enable the network to be constructed deeper and deeper. Through this connection method, the transmission of features and gradients becomes more effective, and each layer can utilize the information input at the beginning, which is helpful for the training of the network. In addition, the dense convolutional network also strengthens the transmission of features, makes more effective use of features, and reduces the number of parameters to a certain extent
[0030] 2), Excitation and squeeze attention model
[0031] Combined with Figure 1 shown, the purpose of convolution is to aggregate the information in space and the information in the feature dimension in the local receptive field. The convolutional neural network consists of a series of convolutional layers, non-linear layers and downsampling layers, enabling the network to capture information from the global receptive field. The operation of convolution is default to fuse all channels of the input feature map. The SE model focuses the attention on the relationship between channels and can automatically learn the importance of different channel features.
[0032] Referring to Figure 3 shown, the SE model first performs a squeeze operation on the obtained feature map to obtain the global features of the channels, then performs an excitation operation on the global features to learn the relationship between each channel, obtains the weights of different channels, and finally assigns weights to the original feature map. This attention mechanism can make the model pay more attention to the channel features with the largest amount of information, enhance the important features, and suppress the unimportant channel features. The SE model can also be easily integrated into the existing network and can improve the performance of the network at a small cost. It should be understood that the excitation operation or activation operation involved can also adopt other non-linear processing functions in addition to the Sigmoid function.
[0033] 3), Coordinate attention model
[0034] Currently, the SE model only considers the internal channel information and ignores the importance of location information. However, the spatial structure of objects is very important in vision. The CA model adds location information to channel attention, enabling the network to obtain larger regional information while avoiding large overhead. To avoid the loss of location information caused by global pooling, the CA model performs average pooling separately in the horizontal and vertical directions to extract features and efficiently integrate spatial coordinate information.
[0035] See Figure 4 As shown, the CA model separately extracts two individual location perceptions from the vertical and horizontal directions, encodes the feature maps with information in specific directions respectively, and fuses the obtained spatial information by weighting on the channels. The CA model considers both channel and location information, can not only capture cross-channel information, but also contains location and direction sensitive information. This model is flexible and lightweight and can be easily inserted into existing networks.
[0036] Correspondingly, the present invention also provides a system for automatically predicting Alzheimer's disease based on deep learning, which is used to implement one or more aspects of the above method. For example, the system includes: an image acquisition unit for acquiring a target image to be detected; a prediction unit for inputting the target image into a trained deep learning framework to obtain a prediction classification result of Alzheimer's disease, where the deep learning framework sequentially includes a coordinated attention model, a deep learning model, and an excitation and squeeze attention model. The coordinated attention model separately extracts individual location perceptions from different directions of the input feature map, fuses the obtained spatial information with the input feature map after weighting on the channels to obtain a first feature map; the deep learning model uses the first feature map as input to extract a second feature map; the excitation and squeeze attention model performs a squeeze operation on the second feature map to obtain the global feature of the channels, and performs an excitation operation on the global feature to obtain the weights of different channels. Each unit involved can be implemented using dedicated hardware, a processor, or an FPGA.
[0037] In summary, the present invention combines two attention mechanisms with DenseNet. The attention mechanism is based on the mechanism of the brain processing visual signals, finds the key points for different pictures, and focuses on the regions related to the target task. The dense connection of DenseNet has a strong regularization effect and can reduce overfitting on a small training set. SE automatically obtains the importance of each channel through learning from the perspective of feature channels, enhancing the ability to extract important information. CA aggregates features from two spatial directions and embeds location information into channel attention, not only obtaining channel information, but also acquiring direction and location information.
[0038] In summary, the present invention uses a dense convolutional network as the main framework, and integrates two attention models in the network. First, it combines channel and position information to learn the importance of features, and then enhances useful features and suppresses features that are less important for the current task according to the importance. For the proposed attention models, the coordinated attention model with both channel and position information is set at the front of the network to obtain more feature information, while the excitation and squeeze attention model with only channel information is set at the end of the network. After continuous downsampling of the network, the dimension of the final output image is smaller and the amount of position information it can provide is also less. At this time, choosing the SE model is more conducive to improving the efficiency and accuracy of classification and recognition. In short, the present invention adds an attention mechanism to the existing deep learning network framework, learns feature weights through the loss function, can obtain the interdependent relationship between feature positions and channels, and recalibrates the features. The attention module is flexible and lightweight, can be easily embedded into the existing network model, and the growth of the model's parameters and computational amount due to attention embedding is negligible, and the performance of the network is also improved.
[0039] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0040] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structure in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0041] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0042] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0043] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0044] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions which implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0045] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0046] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.
[0047] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A method for automatically predicting Alzheimer's disease based on deep learning, comprising the following steps; Obtain a target image to be detected; Input the target image into a trained deep learning framework to obtain a predicted classification result of Alzheimer's disease. The deep learning framework sequentially includes a coordinated attention model, a deep learning model, and an excitation and squeeze attention model. The coordinated attention model separately extracts individual position perceptions from different directions of the input feature map, weights the obtained spatial information on the channels, and then fuses it with the input feature map to obtain a first feature map; The deep learning model uses the first feature map as input to extract a second feature map; the excitation and squeeze attention model performs a squeeze operation on the second feature map to obtain the global features of the channels, and performs an excitation operation on the global features to obtain the weights of different channels.
2. The method according to claim 1, wherein, The deep learning model is a dense convolutional network model, which includes multiple dense blocks.
3. The method according to claim 1, wherein, The coordinated attention model separately extracts two individual position perceptions from the vertical and horizontal directions of the input feature map, separately encodes the information feature maps with specific directions, and then weights the obtained spatial information on the channels.
4. The method according to claim 3, wherein, The coordinated attention model performs average pooling and three-dimensional convolution on the input feature map from the horizontal and vertical directions respectively, then performs fusion and nonlinear processing, and then obtains weights after separate excitation.
5. The method according to claim 1, wherein, The excitation and squeeze attention model sequentially includes a global pooling layer, a first fully connected layer, a nonlinear processing layer, a second fully connected layer, and an activation layer.
6. The method according to claim 5, wherein, The activation layer is processed using the Sigmoid function.
7. The method according to claim 1, wherein, The target image is magnetic resonance imaging.
8. A system for automatically predicting Alzheimer's disease based on deep learning, comprising: An image acquisition unit: used to obtain a target image to be detected; A prediction unit: used to input the target image into a trained deep learning framework to obtain a predicted classification result of Alzheimer's disease. The deep learning framework sequentially includes a coordinated attention model, a deep learning model, and an excitation and squeeze attention model. The coordinated attention model separately extracts individual position perceptions from different directions of the input feature map, weights the obtained spatial information on the channels, and then fuses it with the input feature map to obtain a first feature map; The deep learning model uses the first feature map as input to extract a second feature map; the excitation and squeeze attention model performs a squeeze operation on the second feature map to obtain the global features of the channels, and performs an excitation operation on the global features to obtain the weights of different channels.
9. A computer-readable storage medium, on which a computer program is stored, wherein, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer device, comprising a memory and a processor, and a computer program capable of running on the processor is stored on the memory, wherein, when the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Alzheimer's disease classification and prediction system based on multi-task learning
CN111488914A
Method for classifying images by constructing global embedded attention residual network
CN113111970A