A 3D feature recognition method for Alzheimer's disease in MRI images

By introducing 3D asymmetric convolution blocks and attention feature fusion modules into the 3D MRI image recognition model, the problem of difficulty in feature extraction in 3D MRI images in the prior art is solved, and higher diagnostic accuracy and early recognition effects are achieved.

CN114820524BActive Publication Date: 2025-05-09SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210457193.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-05-09
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

The prior art is difficult to extract discriminant features that distinguish each stage from 3D MRI images, resulting in difficulty in multi-classification diagnosis of Alzheimer's disease, low accuracy, and inability to achieve early recognition.

Method used

Using a 3D feature recognition model including an input layer, multiple ResNet blocks, adaptive average pooling modules and fully connected modules, local and global features in 3D MRI images are extracted and fused by adding 3D asymmetric convolution blocks and attention feature fusion addition modules in each ResNet block.

Benefits of technology

More discriminant features are effectively extracted, which improves the diagnostic accuracy of Alzheimer's disease at all stages, and realizes early recognition, avoiding the loss of feature information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820524B_ABST
    Figure CN114820524B_ABST
Patent Text Reader

Abstract

This invention discloses a 3D feature recognition method for Alzheimer's disease in MRI images. The method uses an Alzheimer's disease 3D feature recognition model to extract 3D features. The model includes an input layer, multiple sequentially arranged ResNet blocks, an adaptive average pooling module, and a fully connected module. The input layer receives a 3D MRI image. Each ResNet block contains a 3D asymmetric convolutional block that extracts discriminative features from the 3D MRI image. An attention feature fusion module is added after each ResNet block to fuse local and global feature contexts, performing multi-scale channel attention feature fusion. After feature extraction from multiple ResNet blocks, the adaptive average pooling module and the fully connected module output the final recognition result. This invention can extract more discriminative features, achieving earlier recognition and higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and in particular relates to a 3D feature recognition method for Alzheimer's disease in MRI images. Background Art

[0002] Alzheimer's disease (AD) is a common neurodegenerative brain disease in the elderly. AD is a typical dementia that is difficult to reverse and difficult to detect early. Clinically, it manifests as memory impairment, aphasia, cognitive impairment, and behavioral abnormalities. Unfortunately, patients are usually diagnosed clinically at an advanced stage, so early diagnosis can control and understand the patient's condition in a timely manner. Doctors usually diagnose patients' conditions through 3D brain magnetic resonance imaging (MRI). However, since the 3D MRI structures of adjacent disease stages are almost similar, multi-classification diagnosis of Alzheimer's disease becomes very difficult. Therefore, it is necessary to improve the feature extraction capability, extract more discriminant features from 3D MRI, and promote more accurate diagnosis. In addition, not only the entire MRI has overall changes, but also local changes in the MRI. Therefore, it is necessary to pay attention to the changes in the entire image and local areas and fuse features of different scales.

[0003] Mild cognitive impairment (MCI) is the transitional stage between cognitive decline in normal elderly people and Alzheimer's disease, and is also the earliest clinically detectable stage in the progression of Alzheimer's disease. Mild cognitive impairment (MCI) is further divided into early mild cognitive impairment (EMCI) and late mild cognitive impairment (LMCI) stages. If we can identify the disease stage more accurately, it will be more conducive to early treatment of patients.

[0004] Since the structures of 3D brain MRI are similar, the disease stages of adjacent MRIs are almost the same. Therefore, the features in 3D MRI are not easy to extract, and the features in 3D MRI need to be extracted step by step. In previous work, it is usually done by using 3D convolutional neural network (CNN) and 2D CNN in sequence, or processing 3D MRI into 2D slices and then using 2D CNN for feature extraction. The loss of feature information accompanies the conversion of 3D to 2D.

[0005] The existing image feature recognition of Alzheimer's disease has poor performance in terms of accuracy and cannot achieve early recognition and high accuracy. The reason is that no discriminative features that can distinguish each stage have been extracted. Some technical methods use more labeled training samples and prior knowledge to improve accuracy and early recognition. Generally, labeled samples require a gold standard marked by very experienced doctors, which makes the research complicated and the labeling of samples prone to errors. Some technical methods use the integration of several models to improve accuracy by using the learning ability of different models. However, the training of several models increases the training time, and at the same time, the learning ability of different models cannot be controlled. Different models may have different recognition effects at the same stage. Other technical methods use deeper and more complex neural network models to improve recognition effects. The neural network model is too deep, and coupled with the lack of hundreds of MRI data, the network model is usually not fully trained. Summary of the invention

[0006] In order to solve the above problems, the present invention proposes a 3D feature recognition method for Alzheimer's disease in MRI images, which can extract more discriminative features and achieve early recognition and higher accuracy.

[0007] To achieve the above object, the technical solution adopted by the present invention is: a 3D feature recognition method of Alzheimer's disease in MRI images, comprising: inputting a 3D MRI image into an Alzheimer's disease 3D feature recognition model to extract the 3D features of Alzheimer's disease;

[0008] The Alzheimer's disease 3D feature recognition model includes an input layer, a plurality of sequentially arranged ResNet blocks, an adaptive average pooling module and a fully connected module; a 3D MRI image is received by the input layer; a 3D asymmetric convolution block is added to each ResNet block, and the 3D asymmetric convolution block extracts discriminative features in the 3D MRI image; an attention feature fusion addition module is set after each ResNet block to fuse local and global feature contexts and perform multi-scale channel attention feature fusion; after feature extraction of multiple ResNet blocks, an adaptive average pooling module and a fully connected module are used to output the final recognition result.

[0009] Furthermore, the ResNet block includes a first convolutional layer of 1*1*1, a second convolutional layer of 3*3*3, a third convolutional layer of 1*1*1, and a fourth convolutional layer of 1*1*1. The first convolutional layer is connected to the second convolutional layer, the second convolutional layer is connected to the 3D asymmetric convolutional block, the 3D asymmetric convolutional block is connected to the third convolutional layer, and the third convolutional layer is connected to the fourth convolutional layer.

[0010] Furthermore, the 3D asymmetric convolution block includes four parallel branches, and the four parallel branches are respectively a 3×3×3 convolution kernel, a 3×1×1 convolution kernel, a 1×3×1 convolution kernel and a 1×1×3 convolution kernel; normalization is performed after each of the four branches, and then the outputs of the four branches are added as the output of the 3D asymmetric convolution block.

[0011] Furthermore, an attention feature fusion addition module is set after each ResNet block to fuse local and global feature contexts, including: selecting point-by-point convolution to use point-by-point channel interaction for each spatial position as a local channel context aggregator.

[0012] Furthermore, the local feature context is calculated: L(X)∈R C×H×W×L , C, H, W, L represent the number of channels, height, width and length of L(X);

[0013] Obtained through the bottleneck structure: L(X) = B(PW Conv2(δ(B(PWConv1(X)))));

[0014] Among them, the kernel sizes of PW Conv1 and PW Conv2 are respectively and r is the channel reduction rate, δ represents the rectified linear unit, and B represents batch normalization.

[0015] Furthermore, the global feature context g(X)∈R is calculated by the global average pooling operation C×1×1×1 , calculation formula:

[0016]

[0017] Among them, X[i,j,k] represents the input image features;

[0018] For the global feature context g(X) and the local feature context L(X), a multi-scale channel attention mechanism is used for further processing. By changing the size of the spatial pool, the multi-scale channel attention mechanism fuses multi-scale context information along the channel dimension.

[0019] Furthermore, the output after processing by the multi-scale channel attention mechanism is M(X)=R C×W×H×L , that is, M(X) represents the attention weight generated by the multi-scale channel attention mechanism, and the formula is: Among them, σ is the Sigmoid function, Indicates broadcast addition;

[0020] Multi-scale channel attention feature fusion of two features X,Y∈R C×H×W×L , assuming that Y is a feature map generated by a larger receptive field;

[0021] In the short-hop connection scenario: X is the feature obtained by identity mapping, and Y is the residual feature learned in the ResNet block. Based on the multi-scale channel attention mechanism, after being processed by the multi-scale channel attention mechanism, it is expressed as:

[0022]

[0023] Among them, ⊥ represents the initial feature integration, and the element-by-element summation is selected as the initial integral;

[0024] Define Z∈R C×H×W×L To add the fused features of the module output for attention feature fusion, the formula is:

[0025]

[0026] The beneficial effects of adopting this technical solution are:

[0027] The present invention not only prevents the loss of feature information, but also improves the extraction of more discriminative features, which makes the diagnosis of each stage of Alzheimer's disease more accurate and the early identification of Alzheimer's disease more timely. In addition, combined with the characteristics of Alzheimer's disease image MRI, it better integrates features with various semantics and scales in the attention mechanism, improves the fusion of 3D features, and further improves the accuracy of feature recognition of Alzheimer's disease.

[0028] The 3D asymmetric convolution structure introduced in the present invention can improve the neural network's ability to extract 3D features and avoid feature information loss, because 3D asymmetric convolution can help extract more discriminative features from three spatial dimensions, improving the neural network's ability to extract 3D features. At the same time, the multi-scale channel attention mechanism is combined to realize the feature fusion in the attention mechanism, and the fusion of global features and local features is realized, which can detect both the overall changes of MRI and the subtle local changes, thereby improving the recognition accuracy of each stage of Alzheimer's disease. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A schematic flow chart of a method for identifying 3D features of Alzheimer's disease in MRI images according to the present invention;

[0030] Figure 2 Schematic diagram of the principle of the multi-scale channel attention mechanism in an embodiment of the present invention;

[0031] Figure 3 This is a schematic diagram of the principle of adding a module for feature fusion according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings.

[0033] In this embodiment, see Figure 1 As shown, the present invention proposes a 3D feature recognition method for Alzheimer's disease in MRI images, comprising: inputting a 3D MRI image into an Alzheimer's disease 3D feature recognition model to extract the Alzheimer's disease 3D features;

[0034] The Alzheimer's disease 3D feature recognition model includes an input layer, a plurality of sequentially arranged ResNet blocks, an adaptive average pooling module and a fully connected module; a 3D MRI image is received by the input layer; a 3D asymmetric convolution block is added to each ResNet block, and the 3D asymmetric convolution block extracts discriminative features in the 3D MRI image; an attention feature fusion addition module is set after each ResNet block to fuse local and global feature contexts and perform multi-scale channel attention feature fusion; after feature extraction of multiple ResNet blocks, an adaptive average pooling module and a fully connected module are used to output the final recognition result.

[0035] As an optimization scheme of the above embodiment, the ResNet block includes a first convolution layer of 1*1*1, a second convolution layer of 3*3*3, a third convolution layer of 1*1*1, and a fourth convolution layer of 1*1*1. The first convolution layer is connected to the second convolution layer, the second convolution layer is connected to the 3D asymmetric convolution block, the 3D asymmetric convolution block is connected to the third convolution layer, and the third convolution layer is connected to the fourth convolution layer.

[0036] Among them, the 3D asymmetric convolution block includes four parallel branches, and the four parallel branches are respectively a 3×3×3 convolution kernel, a 3×1×1 convolution kernel, a 1×3×1 convolution kernel and a 1×1×3 convolution kernel; normalization is performed after each of the four branches, and then the outputs of the four branches are added as the output of the 3D asymmetric convolution block.

[0037] By adding asymmetric convolution kernels, 3D asymmetric convolution kernels can enhance feature extraction in different directions, and more discriminative features can be extracted from 3D MRI.

[0038] As an optimization scheme for the above embodiment, due to the different target sizes in 3D MRI, for example, the entire MRI is variable, and the whole brain atrophy is different at different stages of the disease. Some local areas of MRI have also changed, such as local hippocampal shrinkage and brain ventricular dilation. Therefore, to capture the changes in brain MRI, it is necessary to pay attention to both the entire image and the local areas of the image, and fuse features of different scales. The present invention proposes multi-scale channel attention feature fusion (MS-CAFF). The core idea is to change the size of the spatial pool.

[0039] After each ResNet block, an attention feature fusion addition module is set to fuse local and global feature contexts, including: selecting point-wise convolution to use point-wise channel interaction for each spatial location as a local channel context aggregator.

[0040] Among them, the local feature context is calculated: L(X)∈R C×H×W×L , C, H, W, L represent the number of channels, height, width and length of L(X);

[0041] Obtained through the bottleneck structure: L(X) = B(PW Conv2(δ(B(PWConv1(X)))));

[0042] Among them, the kernel sizes of PW Conv1 and PW Conv2 are respectively and r is the channel reduction rate, δ represents the rectified linear unit, and B represents batch normalization (BN).

[0043] After the above processing, since L(X) has the same shape as the input feature, the low-level subtle detail features can be retained and focused.

[0044] Among them, the global feature context g(X)∈R is calculated by the global average pooling operation C×1×1×1 , calculation formula:

[0045]

[0046] Among them, X[i,j,k] represents the input image features;

[0047] For the global feature context g(X) and the local feature context L(X), a multi-scale channel attention mechanism is used for further processing. By changing the size of the spatial pool, the multi-scale channel attention mechanism fuses multi-scale context information along the channel dimension. It can highlight both global and local changes, making it easier for the network to identify and detect targets under scale changes.

[0048] Among them, the output after processing by the multi-scale channel attention mechanism is M(X)=R C×W×H×L , that is, M(X) represents the attention weight generated by the multi-scale channel attention mechanism, and the formula is:

[0049]

[0050] Among them, σ is the Sigmoid function, Indicates a broadcast addition.

[0051] Detailed structure such as Figure 2 As shown. In the ResNet neural network, the input feature is defined as: X∈RC×H×W×L , the output feature is: X'∈R C×H×W×L Then the output features are as follows:

[0052] in, Represents element-wise multiplication.

[0053] Multi-scale channel attention feature fusion of two features X,Y∈R C×H×W×L , assuming that Y is a feature map generated by a larger receptive field;

[0054] In the short-hop connection scenario: X is the feature obtained by identity mapping, and Y is the residual feature learned in the ResNet block. Based on the multi-scale channel attention mechanism, after being processed by the multi-scale channel attention mechanism, it is expressed as:

[0055]

[0056] Among them, ⊥ represents the initial feature integration, and the element-by-element summation is selected as the initial integral;

[0057] Define Z∈R C×H×W×L To add the fused features of the module output for attention feature fusion, the formula is:

[0058]

[0059] Figure 3 In the figure, the dotted arrows represent 1-M(X⊥Y), in particular, the fusion weights M(X⊥Y) and 1-M(X⊥Y) consist of real numbers between 0 and 1, so that the network performs a soft selection or weighted average between X and Y.

[0060] In order to extract more discriminative features, the present invention will use asymmetric convolution and multi-scale channel attention mechanism to respectively improve feature extraction and fusion. The usual practice is to use 3D convolutional neural network (CNN) and 2D CNN to process in sequence, or to process 3D MRI into 2D slices, and then use 2D CNN for feature extraction. As MRI is converted from 3D to 2D, it is also accompanied by the loss of feature information. On the contrary, directly extracting 3D features can avoid the loss of feature information. In this regard, 3D asymmetric convolution can help extract more discriminative features from three spatial dimensions. Asymmetric convolution is used to improve the ability of neural networks to extract 3D features. By adding asymmetric convolution, features can be extracted from three-dimensional MRI, which can enhance feature extraction in different directions and extract more discriminative features. In addition, the present invention combines the 3D MRI characteristics of Alzheimer's disease. The overall structure of MRI at various stages of the disease is similar, and there are differences in local areas in the image. This is also the fundamental reason why the accuracy of various other technical methods is not high. The recognition of different stages requires attention to the changes in the entire MRI and the local changes within the image. Therefore, the present invention proposes a multi-scale channel attention feature fusion (MS-CAFF), which can simultaneously emphasize the more global distribution of the entire MRI and highlight smaller changes in local distribution. More importantly, in order to better fuse features with various semantics and scales, the present invention fuses features of different scales in the attention mechanism. This technology can start from extracting more discriminative features, and can also combine the imaging characteristics of the disease to improve the feature recognition of Alzheimer's disease.

[0061] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A 3D feature recognition method for Alzheimer's disease in MRI images, characterized in that: include: Input 3D MRI images into the Alzheimer's disease 3D feature recognition model to extract Alzheimer's disease 3D features; The Alzheimer's disease 3D feature recognition model includes an input layer, a plurality of sequentially arranged ResNet blocks, an adaptive average pooling module and a fully connected module; a 3D MRI image is received by the input layer; a 3D asymmetric convolution block is added to each ResNet block, and the 3D asymmetric convolution block extracts discriminative features in the 3D MRI image; an attention feature fusion addition module is set after each ResNet block to fuse local and global feature contexts and perform multi-scale channel attention feature fusion; After extracting features from multiple ResNet blocks, the adaptive average pooling module and the fully connected module output are used to finally obtain the recognition result; The 3D asymmetric convolution block includes four parallel branches, and the four parallel branches are respectively a 3×3×3 convolution kernel, a 3×1×1 convolution kernel, a 1×3×1 convolution kernel, and a 1×1×3 convolution kernel; normalization is performed after each of the four branches, and then the outputs of the four branches are added as the output of the 3D asymmetric convolution block; After each ResNet block, an attention feature fusion addition module is set to fuse local and global feature contexts, including: selecting point-wise convolution to use point-wise channel interaction for each spatial location as a local channel context aggregator.

2. The 3D feature recognition method of Alzheimer's disease in MRI images according to claim 1, characterized in that: The ResNet block includes a first convolutional layer of 1*1*1, a second convolutional layer of 3*3*3, a third convolutional layer of 1*1*1, and a fourth convolutional layer of 1*1*1. The first convolutional layer is connected to the second convolutional layer, the second convolutional layer is connected to the 3D asymmetric convolutional block, the 3D asymmetric convolutional block is connected to the third convolutional layer, and the third convolutional layer is connected to the fourth convolutional layer.

3. The 3D feature recognition method of Alzheimer's disease in MRI images according to claim 1, characterized in that: Calculate local feature context: L(X)∈R C×H×W×L , C, H, W, L represent the number of channels, height, width and length of L(X); Obtained through the bottleneck structure: L(X) = B(PW Conv2(δ(B(PWConv1(X))))); Among them, the kernel sizes of PW Conv1 and PW Conv2 are respectively and r is the channel reduction rate, δ represents the rectified linear unit, and B represents batch normalization.

4. The method for 3D feature recognition of Alzheimer's disease in MRI images according to claim 3, characterized in that: The global feature context g(X)∈R is calculated by global average pooling operation C×1×1×1 , calculation formula: Among them, X[i,j,k] represents the input image features; For the global feature context g(X) and the local feature context L(X), a multi-scale channel attention mechanism is used for further processing. By changing the size of the spatial pool, the multi-scale channel attention mechanism fuses multi-scale context information along the channel dimension.

5. The method for 3D feature recognition of Alzheimer's disease in MRI images according to claim 4, characterized in that: The output after processing by the multi-scale channel attention mechanism is M(X)=R C×W×H×L , that is, M(X) represents the attention weight generated by the multi-scale channel attention mechanism, and the formula is: Among them, σ is the Sigmoid function, Indicates broadcast addition; Multi-scale channel attention feature fusion of two features X,Y∈R C×H×W×L , assuming that Y is a feature map generated by a larger receptive field; In the short-hop connection scenario: X is the feature obtained by identity mapping, and Y is the residual feature learned in the ResNet block. Based on the multi-scale channel attention mechanism, after being processed by the multi-scale channel attention mechanism, it is expressed as: Among them, ⊥ represents the initial feature integration, and the element-by-element summation is selected as the initial integral; Define Z∈R C×H×W×L To add the fused features of the module output for attention feature fusion, the formula is: