Alzheimer's disease picture classification method and system based on bimodal iteration cross attention fusion

By adopting the bimodal iterative cross attention fusion method in Alzheimer's picture classification, a model including residual network and bimodal iterative cross attention module was constructed, and the problem of single feature sources and lack of complementary information in the traditional singlemodal method was solved, and a more efficient and accurate Alzheimer's picture classification was achieved.

CN119992175AActive Publication Date: 2025-05-13GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510050942.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The traditional single-modal image data classification method has a single feature source, lacks complementary information, weak generalization ability and high risk of overfitting, making it difficult to effectively classify Alzheimer's related MRI and PET image data.

Method used

Using a method based on dual-modal iterative cross attention fusion, a dual-modal iterative cross attention fusion model including residual network module, spatial feature shrinkage module and dual-modal iterative cross attention module is constructed, and iterative learning and feature fusion are carried out to enhance the complementarity and generalization ability of image features.

Benefits of technology

It effectively solves the problem of single feature sources and lack of complementary information in traditional methods, improves the accuracy of Alzheimer's picture classification, reduces the risk of overfitting, and achieves more efficient and accurate inter-modal and in-modal feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992175A_ABST
    Figure CN119992175A_ABST
Patent Text Reader

Abstract

The invention discloses an Alzheimer's disease image classification method based on bimodal iteration cross attention fusion, and relates to the technical field of Alzheimer's disease medical image classification, and the method comprises the steps: obtaining sMRI and PET image data of AD, and carrying out the preprocessing of the sMRI and PET image data; cutting the preprocessed data into blocks and carrying out data augmentation; constructing a bimodal iteration cross attention fusion model and taking the bimodal iteration cross attention fusion model as a base classifier, and training the base classifier by using the block image; in the training process, a block image is input into a residual network module to extract features, redundant information of the features is reduced by using a spatial feature shrinkage module, and complementary information is extracted by using a bimodal iterative cross attention module in combination with an iterative learning strategy to enhance the features and fuse the features; inputting the verification set into the trained base classifiers for classification, and selecting the base classifiers of which the classification accuracy is greater than a preset threshold; and constructing a meta classifier, classifying the to-be-classified block image by using the selected base classifier, and inputting the fusion features into the meta classifier for classification to obtain a classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Alzheimer's disease medical image classification, and more specifically, to an Alzheimer's disease image classification method and system based on bimodal iterative cross-attention fusion. Background Art

[0002] Alzheimer's Disease (AD) is an irreversible neurodegenerative disease of the brain that destroys memory and causes cognitive thinking ability to deteriorate. Its prevalence rate increases significantly with age. The classification of AD-related images is of great significance for AD drug development, AD prevention and treatment, and delaying the progression of the disease.

[0003] At present, most methods for classifying AD-related images not only analyze MRI single-modality data, but also usually require manual feature extraction based on expert knowledge, which has certain limitations. Although MRI single-modality data can contain a large amount of AD-related feature information, the traditional single-modality image data classification method has a single feature source, lacks complementary information, has weak generalization ability, and has a high risk of overfitting. Multimodal images based on MRI and PET can help understand the anatomical and neural changes related to AD. Studies have shown that the fusion of multimodal features can improve the classification performance of AD. In addition, the research on deep learning methods has made great progress in recent years, especially in extracting information features for computer vision and medical image analysis. Automatically extracting image features in a data-driven way can reduce the subjective factors and labor costs brought by manual feature extraction. Therefore, there is an urgent need for an Alzheimer's disease image classification method based on bimodal iterative cross-attention fusion. Summary of the invention

[0004] The purpose of the present invention is to overcome the defects of traditional unimodal image data classification methods in the prior art, such as single feature source, lack of complementary information, weak generalization ability and high risk of overfitting, and to provide an Alzheimer's disease image classification method and system based on bimodal iterative cross-attention fusion.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] In a first aspect, the present invention provides an Alzheimer's disease image classification method based on bimodal iterative cross attention fusion, comprising the following steps:

[0007] Acquire structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) imaging data of Alzheimer's disease and perform preprocessing;

[0008] Cut the preprocessed sMRI and PET image data into blocks respectively to obtain equal numbers of sMRI and PET block images and randomly select a preset percentage of the block images for data augmentation;

[0009] Constructing a bimodal iterative cross-attention fusion model including a residual network module, a spatial feature shrinkage module and a bimodal iterative cross-attention module, taking the bimodal iterative cross-attention fusion model as a base classifier, and using a number of the sliced ​​images after data augmentation to train a corresponding number of the base classifiers;

[0010] In the training process of each base classifier, the cut-up images of the sMRI and PET are input into the residual network module to extract sMRI features and PET features, a spatial feature shrinkage module is used to reduce redundant information in the sMRI features and PET features, and then a bimodal iterative cross-attention module is used in combination with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then the enhanced sMRI features and PET features are fused;

[0011] A validation set is preset and input into the trained base classifier for classification, and a base classifier whose classification result accuracy in the validation set is greater than a preset threshold is selected;

[0012] A meta-classifier is constructed, and the selected base classifier is used to classify the sMRI and PET slice images to be classified. The fusion features obtained in the classification process are input into the meta-classifier for classification to obtain the final classification result.

[0013] Preferably, the pretreatment comprises:

[0014] The sMRI images were de-skulled, registered to the Montreal Neurological Institute standard space, and smoothed, followed by grayscale normalization so that the pixel value of each sMRI image data was between 0 and 1;

[0015] The PET images were registered to the PET template, affine-registered to the corresponding sMRI template, and the voxels outside the brain in the image data were deleted, intensity normalized and converted to images with uniform isotropic resolution;

[0016] After preprocessing, the sMRI and PET images have the same size and voxel size.

[0017] Preferably, the step of slicing the preprocessed sMRI and PET image data to obtain an equal number of sMRI and PET sliced ​​images and randomly taking a preset percentage of the sliced ​​images for data augmentation comprises:

[0018] The sMRI and PET image data are padded with 0 on both sides of each dimension, and then a triple loop is used to traverse the padded image data, taking a cubic block each time, and finally obtaining an equal number of sMRI and PET block images. 50% of the block images are randomly flipped along the X-axis, Y-axis or Z-axis for data augmentation.

[0019] Preferably, the step of inputting the cut-up images of the sMRI and PET into a residual network module to extract sMRI features and PET features, using a spatial feature shrinkage module to reduce redundant information in the sMRI features and PET features, and then using a bimodal iterative cross-attention module in combination with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then fusing the enhanced sMRI features and PET features comprises the following steps:

[0020] The preprocessed sMRI and PET block images are used as the input of the residual network module for feature extraction to obtain the three-dimensional feature vectors of sMRI and PET, and then the spatial feature shrinkage module is used to perform feature shrinkage to obtain the two-dimensional feature vectors of sMRI and PET respectively. The two-dimensional feature vectors of sMRI and PET are input into the bimodal iterative cross attention module for feature enhancement to obtain enhanced two-dimensional feature vectors of sMRI and PET respectively. The enhanced two-dimensional feature vector of sMRI is connected element by element with the three-dimensional feature vector of sMRI to obtain the enhanced three-dimensional feature vector of sMRI. At the same time, the enhanced two-dimensional feature vector of PET is connected element by element with the three-dimensional feature vector of PET to obtain the enhanced three-dimensional feature vector of PET. The three-dimensional feature vector of sMRI, the enhanced three-dimensional feature vector of sMRI, the three-dimensional feature vector of PET, and the enhanced three-dimensional feature vector of PET are connected channel by channel and output to the fully connected layer for classification, and the classification result is output.

[0021] Preferably, the residual network module includes an sMRI branch and a PET branch, and both branches include a three-dimensional convolution layer, a layer normalization layer, a batch normalization layer, a ReLU activation function layer, a maximum pooling operation layer and three residual blocks.

[0022] Preferably, the output of a previous residual block among the three residual blocks is used as the input of a subsequent residual block, and the numbers of feature channels of the three residual blocks are 64, 128, and 256, respectively.

[0023] Preferably, the spatial feature shrinkage module includes a pooling operation and a convolution operation;

[0024] For the convolution operation, a method for dimensionality reduction based on convolution operation is used, and the expression of the method is as follows:

[0025] Fconv =conv 1×1 (Reshape(F))

[0026] Among them, F represents the feature map formed by the three-dimensional feature vector of sMRI or PET output by the residual network module, and conv represents the compression of the feature map by 1×1 convolution. Specifically, the spatial information of the feature is converted into the channel dimension by reshaping the dimension of the feature map, and then the channel dimension is compressed by 1×1 convolution operation;

[0027] For the pooling operation, a method of adaptively aggregating average pooling and maximum pooling is adopted. By adaptively aggregating average pooling and maximum pooling, significant features and global background information are retained at the same time to enhance the feature expression ability of the network. The formula of the adaptive aggregation method is as follows:

[0028] F a =AvgPooling(F,S)

[0029] F m =MaxPooling(F,S)

[0030] F o =α·F a +(1-α)·F m

[0031] Among them, S represents the scale factor of the feature map, F a and F m They represent the feature maps compressed by average pooling AvgPooling(·) and maximum pooling MaxPooling(·), respectively. α is the weight used to adaptively aggregate the feature maps after average pooling and maximum pooling. The value of α is between 0 and 1, which is a learnable parameter.

[0032] Preferably, the bimodal iterative cross-attention module includes an sMRI branch and a PET branch, each branch applies a learnable coefficient and each branch includes two cross-attention modules, each cross-attention module includes a multi-head attention module, a layer normalization layer and a feedforward neural network, and cross-modal complementary information is extracted through the two cross-attention modules of the sMRI branch and the PET branch respectively to enhance the features of sMRI and PET.

[0033] Preferably, the bimodal iterative cross-attention module combines the iterative learning strategy with the cross-attention module, gradually deepens the network depth by multiple iterations and sharing parameters, and gradually optimizes the cross-modal complementary information without increasing the number of parameters. The expression of the iterative learning strategy is as follows:

[0034]

[0035] in, represents the output of the bimodal iterative cross attention module after n iterations, {V M , V P} represents the input of the bimodal iterative cross attention module, f ICA (·) represents a bimodal iterative criss-cross attention module. The output of each iterative operation is used as the input of the next iterative operation, and parameters are shared between each iterative operation. In addition, at the output, the bimodal iterative criss-cross attention module converts the output sequence and into a feature map, which is then recalibrated to the original size of the feature map through bilinear interpolation.

[0036] In a second aspect, the present invention provides an Alzheimer's disease picture classification system based on bimodal iterative cross-attention fusion, applying the Alzheimer's disease picture classification method based on bimodal iterative cross-attention fusion described in the above technical solution, the system comprises:

[0037] Data acquisition and preprocessing module, used to acquire structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) imaging data of Alzheimer's disease and perform preprocessing;

[0038] An image slicing and data augmentation module, used to slice the preprocessed sMRI and PET image data respectively to obtain an equal number of sMRI and PET sliced ​​images and randomly select a preset percentage of the sliced ​​images for data augmentation;

[0039] A model construction and training module, used to construct a bimodal iterative cross-attention fusion model including a residual network module, a spatial feature contraction module and a bimodal iterative cross-attention module, taking the bimodal iterative cross-attention fusion model as a base classifier, and using a number of the sliced ​​images after data augmentation to train a corresponding number of the base classifiers;

[0040] A feature processing module, used for inputting the cut-up images of the sMRI and PET into a residual network module to extract sMRI features and PET features during the training process of each base classifier, using a spatial feature shrinkage module to reduce redundant information in the sMRI features and PET features, and then using a bimodal iterative cross-attention module combined with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then fusing the enhanced sMRI features and PET features;

[0041] A base classifier selection module is used to preset a verification set and input it into the trained base classifier for classification, and select a base classifier whose classification result accuracy in the verification set is greater than a preset threshold;

[0042] The classification result acquisition module is used to construct a meta-classifier, use the selected base classifier to classify the sMRI and PET slice images to be classified, input the fusion features obtained in the classification process into the meta-classifier for classification, and obtain the final classification result.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The present invention constructs and trains a bimodal iterative cross-attention fusion model as a base classifier, combines an iterative learning strategy to avoid overfitting, and more efficiently and accurately learns to extract complementary information between and within modalities and enhances the features of sMRI and PET images. A meta-classifier is used to classify the enhanced fusion features of sMRI and PET images obtained by the selected base classifier to obtain a classification result, which effectively solves the defects of traditional single-modal image data classification methods, such as a single feature source, lack of complementary information, weak generalization ability and high risk of overfitting, and improves the accuracy of Alzheimer's disease image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a schematic diagram of the steps of the Alzheimer's disease image classification method based on bimodal iterative cross-attention fusion according to Example 1 of the present application;

[0046] Figure 2 A schematic diagram of the steps for preprocessing the sMRI image data of Example 1 of the present application;

[0047] Figure 3 A schematic diagram of the bimodal iterative cross-attention fusion integrated framework (BICAFE) provided in Example 1 of the present application;

[0048] Figure 4 A schematic diagram of a bimodal iterative cross attention fusion model (BICAF) provided in Example 1 of the present application;

[0049] Figure 5 A schematic diagram of a residual block provided in Example 1 of the present application;

[0050] Figure 6 A schematic diagram of a bimodal iterative cross-attention module provided in Example 1 of the present application;

[0051] Figure 7 This is a result diagram of the visual analysis of the experiment using the method of the present invention provided in Example 3 of the present application. DETAILED DESCRIPTION

[0052] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0053] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0054] Example 1

[0055] See also Figure 1 Embodiment 1 of the present invention provides an Alzheimer's disease image classification method based on bimodal iterative cross attention fusion, comprising the following steps S1-S6:

[0056] S1: Acquire structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) imaging data of Alzheimer's disease and perform preprocessing.

[0057] In a preferred embodiment, the structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) image data obtained in step S1 are derived from the public Alzheimer's Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu), and the data set used in the embodiment of the present invention can be obtained from the website (www.loni.ucla.edu / ADNI). The main purpose of ADNI is to detect whether a series of magnetic resonance imaging (MRI), positron emission tomography (PET), other biomarkers, and clinical and neuropsychological assessments can be used in combination to measure the progression of mild cognitive impairment (MCI) and early Alzheimer's disease (AD).

[0058] Since the sMRI and PET raw image data downloaded directly from the ADNI database cannot be used directly, they need to be preprocessed to make the data from different modalities more similar.

[0059] In this embodiment 1, two preprocessing methods are used to preprocess sMRI and PET data respectively. For the preprocessing of sMRI image data, please refer to Figure 2, the sMRI image is de-skulled, registered to the Montreal Neurological Institute (MNI) standard space and image smoothing is performed, and then grayscale normalization is performed so that the pixel value of each sMRI image data is between 0 and 1. In a specific embodiment, the original image data in DICOM format can be converted into NII format using mricron software, and the toolkit CAT12 is used in MATLAB software to perform a series of processing such as de-skulling, registration to the MNI standard space and image smoothing on the MRI in NII format, and then grayscale normalization is performed on each sMRI data so that the pixel value of each sMRI image data is between 0 and 1.

[0060] Regarding the preprocessing of PET image data, after converting the PET raw image data into HDR format, the PET image in HDR format was registered to the PET template using the toolkit spm12 in the MATLAB software, and then affine registered to the corresponding sMRI. At the same time, the voxels outside the brain in the image data were deleted, and the intensity was normalized and converted to a uniform isotropic resolution image with 8mm FWHM (Full Widthat Half Maximum). The image was smoothed to the same extent in all directions. After preprocessing, the size of the sMRI and PET images was 121mm×145mm×121mm, and the voxel size was 1.5.

[0061] S2: Cut the preprocessed sMRI and PET image data into blocks respectively to obtain an equal number of sMRI and PET block images and randomly select a preset percentage of the block images for data augmentation.

[0062] For the preprocessed sMRI and PET image data, in order to reduce the difficulty of training the neural network model, reduce the consumption of computing resources, and train multiple base classifiers at the same time, the embodiment 1 of the present invention cuts the three-dimensional images of the two modalities into 150 small blocks respectively. In order to obtain cubic blocks and reduce the loss of edge information caused by the blocks, the image with a size of 121mm×145mm×121mm is first filled with 0 on both sides of each dimension to obtain an image with a size of 125mm×150mm×125mm, and then a triple loop is used to traverse the filled image data, and a cubic block with a size of 25mm×25mm×25mm is taken each time, and finally 150 blocks are obtained for sMRI and PET respectively. Since the present invention subsequently uses a base classifier with a dual-branch structure, the data of the two modalities need to be stored separately, and each block also needs to be stored separately. The storage method is as follows: For the modalities sMRI and PET, 150 folders are created and named block0 to block149 respectively. Block0 corresponds to the first block, and the first blocks of other samples of the same modality are also stored in block0. Similarly, block149 corresponds to the 150th block of all samples in the same modality. This storage not only facilitates the data reading of the base classifier, but also ensures that the input of each sample is the same block, avoiding training problems caused by data confusion.

[0063] In addition, in order to increase data diversity, Example 1 of the present invention performs data augmentation on the sliced ​​images, and the data augmentation method is to randomly select 50% of the slices and randomly flip them along the X-axis, Y-axis or Z-axis.

[0064] S3: Construct a bimodal iterative cross-attention fusion model including a residual network module, a spatial feature shrinkage module and a bimodal iterative cross-attention module, use the bimodal iterative cross-attention fusion model as a base classifier, and use a number of the sliced ​​images after data augmentation to train a corresponding number of the base classifiers.

[0065] See also Figure 4 , Figure 4 Schematic diagram of the bimodal iterative cross-attention fusion model (BICAF) provided in Example 1 of the present application. In this Example 1, the bimodal iterative cross-attention fusion model (BICAF) is used as a base classifier, including a residual network module, a spatial feature shrinkage module and a bimodal iterative cross-attention module, and also includes a fully connected layer FC. BICAF uses 150 pre-processed sMRI and PET slice images 150 base classifiers are trained as input. First, the residual network module is used to extract features to obtain a three-dimensional feature vector Then the feature is shrunk by the spatial feature shrinkage module SFS to obtain a two-dimensional feature vector The data are then input into the bimodal iterative cross attention module ICA to extract the complementarity within and between the two modalities for feature enhancement. The enhanced two-dimensional feature vectors of sMRI and PET are 3D feature vectors from sMRI and PET Element-by-element connection to obtain the enhanced three-dimensional feature vector Finally, it is combined with the three-dimensional feature vector Concatenate channel by channel, and then classify through the fully connected layer FC to get the classification result

[0066] See also Figure 4 In this embodiment 1, the residual network module (ResNet) of the dual-modal iterative cross-attention fusion model for feature extraction is a 3DCNN based on the ResNet-18 architecture. Compared with the standard 3D ResNet-18, the embodiment 1 of the present invention has cropped and optimized the network to make it more suitable for small-scale block data sets, and effectively reduces the computational overhead of the model and prevents overfitting while maintaining sufficient feature capabilities. The residual network module ResNet in this embodiment 1 includes an sMRI branch and a PET branch. Both branches include a three-dimensional convolution layer (Conv 3D), a layer normalization layer (LN), a ReLU activation function layer, a maximum pooling operation layer (Max Pooling) and three residual blocks (ResidualBlock) connected in sequence. Taking the sMRI branch as an example, the input of the three-dimensional convolution layer is Among them, 25×25×25 is the spatial dimension H×W×D of sMRI, 1 represents the number of channels, and first passes through a three-dimensional convolution operation with 64 filters, a kernel size of 7×7×7, a stride of 2, and the same padding, which can be expressed as:

[0067] H0=ReLU(LN(Conv3D(P M ,64,7,2,same)))

[0068] Among them, LN represents layer normalization; then the maximum pooling operation is performed, the pooling window is 3×3×3, the stride is 2, and the maximum pooling operation can be expressed as:

[0069] H I =MaxPool3D(H0,3,2,same)

[0070] Then, after three residual blocks, each residual block can be expressed as frcs (H, F, s). Among them, H is the input feature vector, F is the number of output feature channels, and s is the stride of the convolution.

[0071] See also Figure 5 , Figure 5 Schematic diagram of the residual block provided in Example 1 of the present application. In the residual block, H first passes through a three-dimensional convolution layer, whose output feature channel number is F, kernel size is 3×3×3, stride is s, then passes through batch normalization (BN) and ReLU activation function, then passes through a three-dimensional convolution layer whose output feature channel number is F, kernel size is 3×3×3, stride is 1, and then passes through a BN to obtain the feature vector Y. Finally, the residual connection is performed to obtain the final output:

[0072] H out =ReLU(Y+S)

[0073] Where S is defined as:

[0074]

[0075] In order to make ResNet gradually increase its feature expression ability from the initial stage of feature extraction to the aggregation of high-order features, the three Residual Blocks use different F and s, which are expressed as follows:

[0076] H2=f res (H1,64,1)

[0077] H3=f res (H2,128,2)

[0078] F M =f res (H3,256,2)

[0079] Among them, the output of the previous residual block in the three residual blocks is used as the input of the next residual block, and the number of feature channels of the three residual blocks is 64, 128, and 256 respectively. H2 is the output of the first Residual Block, H3 is the output of the second Residual Block, and F M is the final output of the sMRI branch in the residual network module ResNet. Similarly, after passing through the PET branch in the residual network module ResNet, the PET feature vector F P .

[0080] In order to further reduce the redundant information in the features, highlight the key features, and reduce the subsequent computational cost, the spatial feature shrinkage module SFS is applied before the ICA module. SFS consists of two parts: pooling operation and convolution operation.

[0081] For the convolution operation, a dimensionality reduction method based on convolution operation is first used, as shown in the following formula.

[0082] F conv =conv 1×1 (Reshape(F))

[0083] Among them, F represents the feature map formed by the three-dimensional feature vector of sMRI or PET output by the residual network module, and conv represents the compression of the feature map by 1×1 convolution. Specifically, the spatial information of the feature is converted into the channel dimension by reshaping the dimension of the feature map, and then the channel dimension is compressed by 1×1 convolution operation.

[0084] For pooling operations, traditional average pooling and maximum pooling methods have their own advantages and disadvantages. Maximum pooling pays more attention to local significant features, such as edges and textures, but may lose background information; average pooling can retain overall information more smoothly, but it is easy to lose details. Therefore, the SFS module of Example 1 of the present invention adopts a method of adaptively aggregating average pooling and maximum pooling. By adaptively aggregating average pooling and maximum pooling, significant features and global background information can be retained at the same time, thereby enhancing the feature expression ability of the network. The formula of the adaptive aggregation method is as follows:

[0085] F a =AvgPooling(F,S)

[0086] F m =MaxPooling(F,S)

[0087] F o =α·F a +(1-α)·F m

[0088] Among them, S represents the scale factor of the feature map, F a and F m They represent the feature maps compressed by average pooling AvgPooling(·) and maximum pooling MaxPooling(·), respectively. α is the weight used to adaptively aggregate the feature maps after average pooling and maximum pooling. The value of α is between 0 and 1, which is a learnable parameter.

[0089] The ICA module combines a novel iterative learning strategy with the Cross-Attention (CA) module to extract complementary information from inter-modality and intra-modality features to enhance the features of sMRI and PET, further improving the performance of the model. Unlike previous studies that capture local features of different modalities, the CA proposed in Example 1 of the present invention enables a single modality to learn more complementary information from the auxiliary modality (i.e., another modality) from a global perspective. Please refer to Figure 6 , Figure 6 This is a schematic diagram of the bimodal iterative cross-attention module CA provided in Example 1 of the present application. CA consists of a multi-head attention module, a layer normalization layer (Layer Norma, LN) and a feed-forward neural network (Feed-Forward Network, FFN).

[0090] S4: During the training process of each base classifier, the cut images of the sMRI and PET are input into the residual network module to extract sMRI features and PET features, and the spatial feature shrinkage module is used to reduce redundant information in the sMRI features and PET features. Then, a bimodal iterative cross-attention module is used in combination with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then the enhanced sMRI features and PET features are fused.

[0091] S5: Preset a verification set and input it into the trained base classifier for classification, and select a base classifier whose classification result accuracy in the verification set is greater than a preset threshold.

[0092] In Example 1 of the present invention, 150 sMRI and PET cut-up images are used to train 150 base classifiers respectively, a validation set is preset and input into the trained base classifiers for classification, and a base classifier with an accuracy rate greater than 75% on the validation set is selected.

[0093] S6: construct a meta-classifier, use the selected base classifier to classify the sMRI and PET slice images to be classified, input the fusion features obtained in the classification process into the meta-classifier for classification, and obtain the final classification result.

[0094] In Example 1 of the present invention, the meta-classifier consists of a convolutional layer and a fully connected layer, and its input is the fusion features obtained in the process of classifying the sMRI and PET block images to be classified by a base classifier with an accuracy rate greater than 75%. Classification is performed in the meta-classifier under the guidance of the labels to obtain the final classification result.

[0095] See also Figure 3 , Figure 3 This is a schematic diagram of the bimodal iterative cross-attention fusion integrated framework (BICAFE) formed by the above steps of Example 1 of the present application, wherein the Selection module is used to select the base classifier in step S5, and the Meta-classifier module represents the meta-classifier.

[0096] The bimodal iterative cross attention module ICA and iterative learning strategy are further explained as follows:

[0097] The bimodal iterative cross-attention module includes an sMRI branch and a PET branch. Each branch applies a learnable coefficient and each branch includes two cross-attention modules. Each cross-attention module includes a multi-head attention module, a layer normalization layer and a feedforward neural network. The two cross-attention modules of the sMRI branch and the PET branch respectively extract cross-modal complementary information to enhance the features of sMRI and PET.

[0098] Given an input feature map F M and The present invention first flattens each feature map into a set of tokens and adds a learnable position embedding, which is a trainable parameter of dimension HWD×C, used to encode the spatial information between different tokens. Then, a set of sMRI and PET tokens with position embeddings are obtained: V M and V P , and use it as the input of the CA module. Since sMRI and PET features usually have differences in spatial expression, both the sMRI branch and the PET branch in the ICA module use dual CA to extract complementary information within and between modalities, respectively, to enhance sMRI and PET features. The following takes the CA module of the sMRI branch as an example:

[0099] For the CA module of the sMRI branch, first convert the sMRI token: V M Projection to a separate matrix PET’s token V P Projection into two separate matrices Perform multi-head attention calculation in and get the feature vector C M , as shown in the following formula:

[0100] C M =Concat(head1,...,head h )W O

[0101] Among them, the attention calculation formula of the i-th head is:

[0102]

[0103] Then, the feature vector C M It is reprojected back to the original space through a nonlinear transformation and added to the input sequence through a residual connection. The formula is as follows:

[0104] V′ M =α·V M +β·C M W O

[0105] in, represents the output weight matrix before LN. Then it passes through LN. Finally, the same two-layer fully connected FFN as in the standard Transformer is used to further refine the global information to improve the robustness and accuracy of the model and output the enhanced features The formula is as follows:

[0106]

[0107] In order to adaptively learn data from different branches to achieve performance improvement, the present invention applies a learnable coefficient to each branch of the residual connection, where α, β, γ and δ are learnable parameters and are initialized to 1 during the training process.

[0108] Similarly, another CA module is used to enhance the features of the PET branch, but in contrast to the sMRI branch, V P Q, V projected to the multi-head attention module M Then project it to K, V to more comprehensively capture the correlation between the modes, thereby improving the model performance.

[0109] Traditional methods usually improve performance by stacking multiple modules, but this strategy of increasing model depth will not only significantly increase the number of parameters, but may also cause overfitting. In order to overcome the shortcomings of traditional methods, the iterative learning strategy used in the present invention gradually deepens the network depth by multiple iterations and sharing parameters. At the same time, without increasing the number of parameters, it gradually optimizes the complementary information across modalities, and dynamically allocates weights to sMRI and PET features to achieve efficient fusion. The formula of the adaptive aggregation method is as follows:

[0110]

[0111] in, represents the output of the ICA module after n iterations, {V M ,V P} represents the input of ICA. ICA (·) represents the ICA module proposed in the present invention, which integrates two CA modules of the sMRI branch and the PET branch respectively. The output of each iterative operation is used as the input of the next iterative operation, and the parameters are shared between each iterative operation. In addition, the output sequence and of the ICA module are converted into feature maps and then recalibrated to the original size of the feature map by bilinear interpolation.

[0112] Example 2

[0113] Based on Example 1, this Example 2 provides an Alzheimer's disease picture classification system based on bimodal iterative cross-attention fusion, and applies the Alzheimer's disease picture classification method based on bimodal iterative cross-attention fusion described in Example 1. The system includes:

[0114] Data acquisition and preprocessing module, used to acquire structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) imaging data of Alzheimer's disease and perform preprocessing;

[0115] An image slicing and data augmentation module, used to slice the preprocessed sMRI and PET image data respectively to obtain an equal number of sMRI and PET sliced ​​images and randomly select a preset percentage of the sliced ​​images for data augmentation;

[0116] A model construction and training module, used to construct a bimodal iterative cross-attention fusion model including a residual network module, a spatial feature contraction module and a bimodal iterative cross-attention module, taking the bimodal iterative cross-attention fusion model as a base classifier, and using a number of the sliced ​​images after data augmentation to train a corresponding number of the base classifiers;

[0117] A feature processing module, used for inputting the cut-up images of the sMRI and PET into a residual network module to extract sMRI features and PET features during the training process of each base classifier, using a spatial feature shrinkage module to reduce redundant information in the sMRI features and PET features, and then using a bimodal iterative cross-attention module combined with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then fusing the enhanced sMRI features and PET features;

[0118] A base classifier selection module is used to preset a verification set and input it into the trained base classifier for classification, and select a base classifier whose classification result accuracy in the verification set is greater than a preset threshold;

[0119] The classification result acquisition module is used to construct a meta-classifier, use the selected base classifier to classify the sMRI and PET slice images to be classified, input the fusion features obtained in the classification process into the meta-classifier for classification, and obtain the final classification result.

[0120] The other steps or technical details of this embodiment 2 are the same as those of embodiment 1 and will not be repeated here.

[0121] Example 3

[0122] Example 3 of the present invention is based on Example 1 and Example 2. In order to verify the effectiveness and advancement of the present invention, the imaging data of subjects with both sMRI and PET image data are randomly selected for experiments and visual analysis is performed, as follows:

[0123] In Example 3 of the present invention, a total of 418 subjects with both sMRI and PET image data were selected from the ADNI database. The subjects were divided into two groups: AD and HC (192 AD and 226 HC). The AD group was a subject with clinically diagnosed AD symptoms, and the HC was a subject without any cognitive impairment. The experiment used five classification indicators: classification accuracy (ACC), area under the ROC curve (AUC), sensitivity (SEN), specificity (SPE), and Matthews correlation coefficient (MCC) to measure model performance.

[0124] A 5-fold cross-validation strategy was used to evaluate the classification performance.

[0125]

[0126] Table 1 Comparative experimental results

[0127] The comparative experimental results obtained by comparing the method of the present invention with the current mainstream multimodal fusion method are shown in Table 1. The comparative methods include the 2DCNN+BGRU method proposed by Huang et al., the dual-branch three-dimensional convolutional neural network model proposed by Sharma et al., the MCAD framework proposed by Zhang et al., the MADDi method proposed by Golovanevsky et al., and the DMFF method proposed by Shen et al. As shown in Table 1, the classification performance of the method of the present invention is better than that of other methods. It is worth noting that when using the same data set, the ACC of the method of the present invention is 4.9%, 4.9%, 2.4%, 0.8%, 1.5% higher than that of 2DCNN+BGRU, Sharma, MCAD, MADDi and DMFF, respectively, and the AUC is 2.2%, 1%, 5.7%, 0.7%, 0.7%, and MCC is 9.9%, 9.5%, 4.7%, 0.8%, 3% higher than that of DMFF. This shows that the method can better capture the important features in the data, thereby improving the accuracy of classification. In addition, the improvement of AUC further reflects the classification ability of the model at different thresholds, proving its reliability in practical applications. The improvement of MCC, especially the significant increase of 9.9%, shows that the proposed method is more robust in dealing with class imbalance problems. It also shows the effectiveness of the method in dealing with complex data sets, further demonstrating its value in practical application scenarios.

[0128]

[0129] Table 2 Experimental comparison results of single-modal and dual-modal fusion

[0130] The experimental comparison of single-modality and dual-modality fusion of Example 3 of the present invention is shown in Table 2. In the ADvs.HC experiment, the classification effect of the PET modality is better than that of the sMRI modality as a whole, and the classification performance of the dual-modality fusion is better than that of the single sMRI modality or the single PET modality. The ACC is improved by 6.6% and 1% compared with single sMRI and single PET, respectively, the AUC is improved by 2.9% and 0.6%, respectively, and the MCC is improved by 13% and 2%, respectively, and the dual-modality fusion has a lower standard deviation. This shows that the dual-modality fusion method proposed in the present invention can effectively combine the structural information provided by sMRI and the functional and metabolic information provided by PET, thereby achieving information complementarity and improving the accuracy of AD image classification.

[0131] In order to better illustrate how different modalities affect network performance and make the model more interpretable, Example 3 of the present invention provides a visual result analysis, using Grad-CAM to display the areas of interest of the single-modality and bi-modal fusion models.

[0132] See also Figure 7 , Figure 7The result diagram of the visual analysis of the experiment using the method of the present invention provided in Example 3 of the present invention. The visual results are as follows Figure 7 As shown, the results show that the single-modality model has a large difference in the areas of interest for sMRI and PET, which indicates that the model extracts different brain area information for different modalities. The areas of interest for the dual-modality fusion model overlap to a certain extent with the areas of interest for the single-modality, and also present more comprehensive information. In order to further display the brain areas of interest for the single-modality model and the dual-modality fusion model, the present invention uses Brainnetome Atlas to annotate the brain areas of Patch106, and compares the specific brain areas of interest for the models under different modalities, as shown in Table 3 below.

[0133]

[0134] Table 3 Specific brain regions that the model focuses on under different modalities

[0135] Table 3 shows that the main brain regions that sMRI focuses on include those related to brain morphology and structural changes, such as the superior temporal gyrus (73), prefrontal cortex (77, 79), and areas related to memory and emotional processing, such as the hippocampus (215). PET, on the other hand, tends to focus on those brain regions related to metabolic activity, blood flow, and neurotransmitter distribution, such as the ventral insular cortex (165) and dorsal insular cortex (167), which are related to cognitive and emotional regulation of the brain. The brain regions that the dual-modality fusion model focuses on not only include the significant regions of sMRI and PET, but also cover the intersection of multiple brain regions and have a wider range of brain functional connections. This shows that by fusing sMRI and PET data, the dual-modality fusion model can capture more comprehensive brain region characteristics and improve the ability to identify and classify complex brain lesions. Compared with a single modality, the dual-modality fusion model can better reflect the interaction between different brain regions and may reveal subtle differences that cannot be identified by a single modality.

[0136] The other steps of this embodiment 3 are the same as those of embodiment 1 or embodiment 2 and will not be repeated here.

[0137] In summary, the present invention constructs and trains a bimodal iterative cross-attention fusion model as a base classifier, combines iterative learning strategies, avoids overfitting, learns to extract complementary information between and within modalities more efficiently and accurately, and enhances the features of sMRI and PET images. The fusion features of the enhanced sMRI and PET images obtained by the selected base classifier are classified using a meta-classifier to obtain classification results, which effectively solves the defects of the traditional single-modal image data classification method, such as single feature source, lack of complementary information, weak generalization ability, and high risk of overfitting, and improves the accuracy of Alzheimer's disease image classification. A new iterative cross-attention fusion strategy proposed by the present invention is combined with a bimodal cross-attention fusion module to strengthen the memory of complementary information from inter-modal and intra-modal features, so as to learn complementary information of different modalities more efficiently and accurately, and perform AD image classification. Experimental results prove the effectiveness of this method, with an accuracy of 94.3% in AD and HC. Compared with other methods, there is a significant improvement. Visual analysis confirmed that the method of the present invention extracted different brain area information for different modalities, and the dual-modal fusion model complemented the brain area information of different modalities, thereby improving the accuracy of AD image classification and recognition.

Claims

1. Alzheimer's disease image classification method based on bimodal iterative cross attention fusion, characterized by: The following steps are involved: Acquire structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) imaging data of Alzheimer's disease and perform preprocessing; Cut the preprocessed sMRI and PET image data into blocks respectively to obtain equal numbers of sMRI and PET block images and randomly select a preset percentage of the block images for data augmentation; Constructing a bimodal iterative cross-attention fusion model including a residual network module, a spatial feature shrinkage module and a bimodal iterative cross-attention module, taking the bimodal iterative cross-attention fusion model as a base classifier, and using a number of the sliced ​​images after data augmentation to train a corresponding number of the base classifiers; In the training process of each base classifier, the cut-up images of the sMRI and PET are input into the residual network module to extract sMRI features and PET features, a spatial feature shrinkage module is used to reduce redundant information in the sMRI features and PET features, and then a bimodal iterative cross-attention module is used in combination with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then the enhanced sMRI features and PET features are fused; A validation set is preset and input into the trained base classifier for classification, and a base classifier whose classification result accuracy in the validation set is greater than a preset threshold is selected; A meta-classifier is constructed, and the selected base classifier is used to classify the sMRI and PET slice images to be classified. The fusion features obtained in the classification process are input into the meta-classifier for classification to obtain the final classification result.

2. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 1 is characterized in that: The pre-processing comprises: The sMRI images were de-skulled, registered to the Montreal Neurological Institute standard space, and smoothed, followed by grayscale normalization so that the pixel value of each sMRI image data was between 0 and 1; The PET images were registered to the PET template, affine-registered to the corresponding sMRI template, and the voxels outside the brain in the image data were deleted, intensity normalized and converted to images with uniform isotropic resolution; After preprocessing, the sMRI and PET images have the same size and voxel size.

3. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 2 is characterized in that: The steps of slicing the pre-processed sMRI and PET image data to obtain equal numbers of sMRI and PET sliced ​​images and randomly taking a preset percentage of the sliced ​​images for data augmentation include: The sMRI and PET image data are padded with 0 on both sides of each dimension, and then a triple loop is used to traverse the padded image data, taking a cubic block each time, and finally obtaining an equal number of sMRI and PET block images. 50% of the block images are randomly flipped along the X-axis, Y-axis or Z-axis for data augmentation.

4. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 1 is characterized in that: The step of inputting the cut-up images of the sMRI and PET into a residual network module to extract sMRI features and PET features, using a spatial feature shrinkage module to reduce redundant information in the sMRI features and PET features, and then using a bimodal iterative cross-attention module in combination with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then fusing the enhanced sMRI features and PET features comprises the following steps: The preprocessed sMRI and PET block images are used as the input of the residual network module for feature extraction to obtain the three-dimensional feature vectors of sMRI and PET, and then the spatial feature shrinkage module is used to perform feature shrinkage to obtain the two-dimensional feature vectors of sMRI and PET respectively. The two-dimensional feature vectors of sMRI and PET are input into the bimodal iterative cross attention module for feature enhancement to obtain enhanced two-dimensional feature vectors of sMRI and PET respectively. The enhanced two-dimensional feature vector of sMRI is connected element by element with the three-dimensional feature vector of sMRI to obtain the enhanced three-dimensional feature vector of sMRI. At the same time, the enhanced two-dimensional feature vector of PET is connected element by element with the three-dimensional feature vector of PET to obtain the enhanced three-dimensional feature vector of PET. The three-dimensional feature vector of sMRI, the enhanced three-dimensional feature vector of sMRI, the three-dimensional feature vector of PET, and the enhanced three-dimensional feature vector of PET are connected channel by channel and output to the fully connected layer for classification, and the classification result is output.

5. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 4 is characterized in that: The residual network module includes an sMRI branch and a PET branch, and both branches include a three-dimensional convolution layer, a layer normalization layer, a ReLU activation function layer, a maximum pooling operation layer and three residual blocks.

6. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 5 is characterized in that: The output of the previous residual block among the three residual blocks is used as the input of the next residual block, and the numbers of feature channels of the three residual blocks are 64, 128, and 256 respectively.

7. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 4 is characterized in that: The spatial feature shrinkage module includes a pooling operation and a convolution operation; For the convolution operation, a method for dimensionality reduction based on convolution operation is used, and the expression of the method is as follows: F conv =conv 1×1 (Reshape(F)) Among them, F represents the feature map formed by the three-dimensional feature vector of sMRI or PET output by the residual network module, and conv represents the compression of the feature map by 1×1 convolution. Specifically, the spatial information of the feature is converted into the channel dimension by reshaping the dimension of the feature map, and then the channel dimension is compressed by 1×1 convolution operation; For the pooling operation, a method of adaptively aggregating average pooling and maximum pooling is adopted. By adaptively aggregating average pooling and maximum pooling, significant features and global background information are retained at the same time to enhance the feature expression ability of the network. The formula of the adaptive aggregation method is as follows: F a =AvgPooling(F,S) F m =MaxPooling(F,S) F o =α·F a +(1-α)·F m Among them, S represents the scale factor of the feature map, F a and F m They represent the feature maps compressed by average pooling AvgPooling(·) and maximum pooling MaxPooling(·), respectively. α is the weight used to adaptively aggregate the feature maps after average pooling and maximum pooling. The value of α is between 0 and 1, which is a learnable parameter.

8. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 7 is characterized in that: The bimodal iterative cross-attention module includes an sMRI branch and a PET branch, each branch applies a learnable coefficient and each branch includes two cross-attention modules, each cross-attention module includes a multi-head attention module, a layer normalization layer and a feedforward neural network, and cross-modal complementary information is extracted through the two cross-attention modules of the sMRI branch and the PET branch to enhance the features of sMRI and PET.

9. The Alzheimer's disease image classification method based on bimodal iterative cross attention fusion according to claim 8, characterized in that: The bimodal iterative cross attention module combines the iterative learning strategy with the cross attention module, gradually deepens the network depth by multiple iterations and sharing parameters, and gradually optimizes the cross-modal complementary information without increasing the number of parameters. The expression of the iterative learning strategy is as follows: in, represents the output of the bimodal iterative cross attention module after n iterations, {V M ,V P } represents the input of the bimodal iterative cross attention module, f ICA (·) represents a bimodal iterative criss-cross attention module. The output of each iterative operation is used as the input of the next iterative operation, and parameters are shared between each iterative operation. In addition, at the output, the bimodal iterative criss-cross attention module converts the output sequence and into a feature map, which is then recalibrated to the original size of the feature map through bilinear interpolation.

10. An Alzheimer's disease image classification system based on bimodal iterative cross-attention fusion, applying the Alzheimer's disease image classification method based on bimodal iterative cross-attention fusion described in claims 1-9, characterized in that: The system comprises: Data acquisition and preprocessing module, used to acquire structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) imaging data of Alzheimer's disease and perform preprocessing; An image slicing and data augmentation module, used to slice the preprocessed sMRI and PET image data respectively to obtain an equal number of sMRI and PET sliced ​​images and randomly select a preset percentage of the sliced ​​images for data augmentation; A model construction and training module, used to construct a bimodal iterative cross-attention fusion model including a residual network module, a spatial feature contraction module and a bimodal iterative cross-attention module, taking the bimodal iterative cross-attention fusion model as a base classifier, and using a number of the sliced ​​images after data augmentation to train a corresponding number of the base classifiers; A feature processing module, used for inputting the cut-up images of the sMRI and PET into a residual network module to extract sMRI features and PET features during the training process of each base classifier, using a spatial feature shrinkage module to reduce redundant information in the sMRI features and PET features, and then using a bimodal iterative cross-attention module combined with an iterative learning strategy to extract complementary information to enhance the sMRI features and PET features, and then fusing the enhanced sMRI features and PET features; A base classifier selection module is used to preset a verification set and input it into the trained base classifier for classification, and select a base classifier whose classification result accuracy in the verification set is greater than a preset threshold; The classification result acquisition module is used to construct a meta-classifier, use the selected base classifier to classify the sMRI and PET slice images to be classified, input the fusion features obtained in the classification process into the meta-classifier for classification, and obtain the final classification result.

Citation Information

Patent Citations

  • Target segmentation method based on multi-source cross attention fusion

    CN119131396A

  • Low-light environment multispectral pedestrian detection method based on cross-modal attention fusion network

    CN119206791A

  • Computer aided diagnosis system for mild cognitive impairment

    US20200126221A1

Cited By

  • Non-destructive detection method for content of chlorophyll and carotenoid in tobacco leaves

    CN120510153A

  • Classification method and device based on graph neural network, computer equipment and medium

    CN120707937A