Alzheimer's disease diagnosis method based on missing mode generation

By using the DAMSA-GAN network with a multi-scale generator and dual discriminator structure and the CM-DFNet network, the problem of incomplete MRI and PET modal information is solved, generating high-quality PET images and improving the multimodal diagnostic effect of Alzheimer's disease.

CN120977537APending Publication Date: 2025-11-18CHONGQING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510977659.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing MRI and PET modalities suffer from incomplete information in the diagnosis of Alzheimer's disease. Traditional GAN ​​methods struggle to generate high-quality images of missing modalities, and multimodal fusion methods fail to fully utilize multi-level information, resulting in poor diagnostic outcomes.

Method used

We employ a generator based on multi-scale and RWKV self-attention, combined with a dual discriminator structure of point discriminator and disease perception discriminator, and design DAMSA-GAN network and CM-DFNet network. Through a multimodal feature fusion module, we generate high-quality PET images and improve the effect of multimodal fusion classification.

Benefits of technology

The generator can effectively capture multi-scale features and long-range dependencies in MRI images, the point discriminator accurately captures brain structures, and the disease perception discriminator enhances the ability to model pathological features, thereby achieving high-quality PET image reconstruction and improved multimodal diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977537A_ABST
    Figure CN120977537A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical diagnosis, in particular to an Alzheimer's disease diagnosis method based on deletion mode generation. The method comprises the following steps: S1, constructing a self-attention generator based on multiple scales and RWKV; s2, constructing a double-discriminator structure based on a point discriminator and a disease perception discriminator; and S3, constructing a feature fusion module based on different levels. According to the Alzheimer's disease diagnosis method based on missing modal generation, features of different scales are extracted through convolution kernels of different sizes in a coding path of a generator, limitation of single-scale feature extraction is avoided, and multi-scale features effectively extracted through an encoder can be used for diagnosis of the Alzheimer's disease. According to the method, the MRI data can be captured, nuisance and important features in the MRI data can be captured, so that the brain structure of the MRI image can be effectively processed and analyzed, meanwhile, an RWKV self-attention module is introduced into a decoder, the long-distance dependency relationship in the brain structure is captured, and a generator can generate a high-quality PET image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical diagnostic technology, and in particular to a diagnostic method for Alzheimer's disease based on missing modality generation. Background Technology

[0002] Alzheimer's disease (AD), also known as senile dementia, is a neurodegenerative disease of the central nervous system. It has an insidious onset and a chronic, progressive course, and is the most common type of dementia in the elderly. Its main manifestations include progressive memory impairment, cognitive dysfunction, personality changes, and language disorders, severely affecting social, occupational, and daily life functions. When making a diagnosis, doctors should first determine if it is dementia, then identify the cause, and finally assess the severity. Currently, there is no specific method for a definitive diagnosis of Alzheimer's disease; it primarily relies on pathological examination.

[0003] Existing diagnostic methods have the following drawbacks

[0004] Magnetic resonance imaging (MRI) and positron emission tomography (PET) provide complementary information about brain structure and function, respectively, and are therefore widely used in AD diagnostic research. However, in clinical practice, the lack of PET modalities leads to incomplete diagnostic information, making it challenging to utilize this missing modality information to improve diagnostic accuracy.

[0005] In the process of generating missing modalities from existing modalities, the images generated by classic GAN methods have limitations. Using the same convolutional kernel in the generator makes it difficult to capture feature information at different levels, hindering the generation of high-quality missing images from multi-layered information. Secondly, long-range correlations in 3D brain structures in medical images enhance image detail generation, a point not considered by basic GANs. Furthermore, traditional voxel discriminators, due to their fixed mesh division, suffer from blurred peripheral brain regions.

[0006] In the traditional Generative Adversarial Network (GAN) framework, the core objective of the model is to reconstruct images through adversarial training, typically focusing on the generation of global content and the optimization of overall visual effects. However, this global generation strategy has significant limitations in the field of medical imaging, especially when it comes to the reconstruction of disease-specific pathological features. Traditional GANs often struggle to accurately capture and reproduce subtle abnormal structures that are crucial for diagnosis.

[0007] Existing multimodal fusion methods utilize feature concatenation and simple attention to achieve feature fusion, without considering multi-stage and multi-level fusion, or the redundant information caused by modal fusion. This results in insufficient realization of deep feature interaction between the two modalities, leading to poor diagnostic performance.

[0008] Compared with existing technologies:

[0009] Compared to 2D CNN architectures, 3D CNN architectures are more suitable for medical image reconstruction. Compared to the limitations of 2D networks in processing single slices, 3D convolutional kernels use voxels to model the spatial relationships in the axial, coronal, and sagittal planes. This 3D modeling maintains the spatial consistency of brain structure and metabolic intensity, avoiding the blurring of reconstructed images caused by information fragmentation between slices in traditional methods.

[0010] Adversarial learning trains generators and discriminators. The core goal of the model is to achieve high-quality image reconstruction through adversarial training. It usually focuses on the generation of global content and the optimization of the overall visual effect. The generator reconstructs the image through an encoding and decoding structure, while the discriminator distinguishes between the generated image and the real image. By introducing adversarial loss and training the discriminator, the quality of the reconstructed image is improved.

[0011] Traditional discriminators typically make distinctions based on voxels or image patches, while point discriminators convert 3D images into point set representations and distinguish between real and fake images by analyzing structural information and contextual relationships within the point set. This design fully leverages the geometric advantages of point representations, enabling it to more sensitively capture structural differences (such as organ edges and the morphology of small-sized tissues).

[0012] Multimodal feature fusion employs various fusion strategies to extract features from different modalities. These include input fusion strategies that perform pixel-level fusion of images from different modalities, hierarchical strategies that utilize different fusion blocks across multiple branches for multi-stage, multi-level fusion, and output fusion methods that extract features from independent branches and then perform weighted and concatenated summaries. These different fusion strategies are used for multimodal disease diagnosis.

[0013] To address this issue, an Alzheimer's disease diagnostic method based on missing mode generation is designed to provide a technical solution for the aforementioned technical problems. Summary of the Invention

[0014] Therefore, it is necessary to provide an Alzheimer's disease diagnostic method based on missing mode generation to address the aforementioned technical problems.

[0015] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0016] The Alzheimer's disease diagnostic method based on missing modalities consists of the following steps:

[0017] S1: Construct a self-attention generator based on multi-scale and RWKV;

[0018] S2: Construct a dual discriminator structure based on a point discriminator and a disease perception discriminator;

[0019] S3: Construct feature fusion modules based on different levels;

[0020] S4: Construct a DAMSA-GAN network and a multimodal fusion classification CM-DFNet network to generate missing PET modalities. This allows the DAMSA-GAN network to generate missing PET images using existing MRI images, and utilizes additional information from PET to improve the disease diagnosis performance of the multimodal fusion classification CM-DFNet network.

[0021] In a preferred embodiment of the Alzheimer's disease diagnosis method based on missing mode generation provided by the present invention, step S1 includes the following steps:

[0022] The generator part of DAMSA-GAN consists of a contracting path composed of multiple downsampling operations and an expanding path composed of multiple upsampling operations.

[0023] In the contracting path, key feature information of MRI images is extracted through multi-layer convolution operations, and a multi-scale convolution module is introduced to enhance feature representation.

[0024] In the expanding path, high-resolution image details are reconstructed from low-resolution feature maps through upsampling operations.

[0025] As a preferred embodiment of the Alzheimer's disease diagnosis method based on missing mode generation provided by the present invention, the upsampling path includes the following steps:

[0026] Introducing the RWKV self-attention mechanism into 3D medical image generation;

[0027] By combining multi-scale deep convolutions with omnidirectional displacement layers, local context awareness is enhanced, and subtle structures at the edge of lesions are captured.

[0028] In the channel mixing module, a squared ReLU activation function is introduced to enhance the nonlinear interaction between channels.

[0029] In a preferred embodiment of the Alzheimer's disease diagnosis method based on missing mode generation provided by the present invention, step S2 includes the following steps:

[0030] Replace the standard discriminator with a point discriminator.

[0031] The point discriminator enhances local-global relationship modeling through a context clustering module;

[0032] The point discriminator maps PET voxels into a point set representation that fuses three-dimensional coordinates with brain texture features.

[0033] In a preferred embodiment of the Alzheimer's disease diagnosis method based on missing mode generation provided by the present invention, step S2 includes the following steps:

[0034] We construct a disease perception discriminator, combining the disease classification task with the GAN framework. By introducing disease-related discriminative constraints during the generation process, we enhance the model's ability to model pathological features.

[0035] A disease perception discriminator is introduced into a multi-task learning mechanism to simultaneously optimize image realism and disease classification accuracy with a point discriminator. At the same time, the loss constraint of the disease perception discriminator is added to the loss during the training process.

[0036] As a preferred embodiment of the Alzheimer's disease diagnosis method based on missing mode generation provided by the present invention, the disease perception discriminator is composed of three DM Block modules connected in series. The DM Block module is composed of dense connection blocks, Mamba modules, and convolutions.

[0037] The enhanced feature X with channel number C in the densely connected block is first processed by the densely connected block, then normalized using a LayerNorm layer, and subsequently uniformly divided into four sub-features, namely...

[0038] The number of channels for each sub-feature is C / 4;

[0039] Each sub-feature is input into a Mamba module for processing, and the output of the Mamba module is optimized through residual connections and adjustment factors;

[0040] The four processed sub-features are then merged back into a single feature X with C channels using the Concat operation. out ;

[0041] X out The DM Block is output by sequentially passing through LayerNorm, Projection, and convolution layers.

[0042] In a preferred embodiment of the Alzheimer's disease diagnosis method based on missing mode generation provided by the present invention, step S3 includes the following steps:

[0043] Design two independent symmetrical path structures, each branch consisting of three dense blocks, a convolutional network, and a pooling operation;

[0044] The fused features output by each fusion module are concatenated with the original features of the previous layer and passed as input to subsequent dense blocks.

[0045] It is clear without a doubt that the technical solution described above in this application can solve the technical problem that this application aims to address.

[0046] Meanwhile, through the above technical solutions, the present invention has at least the following beneficial effects:

[0047] 1. The Alzheimer's disease diagnosis method based on missing modality generation provided by this invention extracts features of different scales through convolutional kernels of different sizes in the encoding path of the generator, avoiding the limitations of extracting single-scale features. The multi-scale features effectively extracted by the encoder can capture subtle differences and important features in MRI data, enabling the brain structure of MRI images to be effectively processed and analyzed. At the same time, the RWKV self-attention module is introduced into the decoder to capture long-distance dependencies in brain structures, enabling the generator to generate high-quality PET images.

[0048] 2. In order to make the generated image as close as possible to the real image, this invention introduces a point discriminator instead of the traditional standard discriminator. It maps voxel values ​​and three-dimensional coordinates into a point set. This point set representation can explicitly preserve the spatial distribution information of brain structure. It uses a context clustering module to build models of local and global brain regions, which can accurately capture subtle regional morphological features and avoid the problem of blurred brain edge regions in the reconstruction image due to fixed grid division of traditional voxel discriminators.

[0049] 3. In order to reconstruct subtle abnormal structures that are crucial for diagnosis, this invention designs a disease discriminator. Disease-related discriminative constraints are introduced during the generation process to enhance the model's ability to model pathological features. Global features are extracted through the Mamba module and local features are extracted through Denseblock. The combination of the two effectively preserves pathological information that is highly relevant to disease diagnosis and improves the diagnostic effect of downstream classification tasks.

[0050] 4. To effectively extract unique features from MRI and PET modalities and avoid feature confusion during the feature extraction process, this invention designs two independent feature extraction branches. This design can effectively extract unique features from each modality, and through a dual-path dense connection and a multi-stage fusion block structure, it achieves full interaction and deep fusion of features in both spatial and channel dimensions.

[0051] 5. In order to improve the fusion effect of the two modalities, this invention designs shallow and deep feature fusion modules. By using the DSC module in DFFB to reduce redundant information in feature extraction, the cross-attention mechanism effectively captures complementary information between MRI and PET modalities, making up for the limitations of a single modality, realizing global feature interaction across modalities, and improving the disease discrimination ability of fused features. Attached Figure Description

[0052] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a general framework diagram of the model of the present invention;

[0054] Figure 2 This is a detailed structural diagram of the DAMSA-GAN of the present invention;

[0055] Figure 3 This is a detailed structural diagram of the DM module of the present invention;

[0056] Figure 4 This is a detailed structural diagram of the CM-DFNet of the present invention;

[0057] Figure 5 This is a comparative experimental effect diagram of the present invention;

[0058] Figure 6 This is a diagram illustrating the ablation experiment results of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0060] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0061] It should be noted that, unless otherwise specified, the embodiments and features and technical solutions in the present invention can be combined with each other.

[0062] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0063] Example 1

[0064] Reference Figures 1-4 A diagnostic method for Alzheimer's disease based on missing modalities.

[0065] 1. Overall framework based on two phases

[0066] Based on real-world medical data, modality loss and multimodal fusion methods are major challenges in multimodal classification of medical images. To address these challenges, this invention proposes a deep learning framework consisting of two deep learning networks: a DAMSA-GAN network for generating missing PET modalities and a CM-DFNet network for multimodal fusion classification. Figure 1 As shown, the DAMSA-GAN network is used to generate missing PET images from existing MRI images, and the additional information from PET images is used to improve the disease diagnosis performance of the multimodal fusion classification CM-DFNet network.

[0067] 2. Multi-scale and RWKV-based self-attention generator

[0068] The generator of the model is based on the 3D U-net framework, whose 3D convolutional kernels are naturally adapted to the stereoscopic spatial characteristics of brain images. The generator part of DAMSA-GAN consists of a contracting path composed of multiple downsampling operations and an expanding path composed of multiple upsampling operations, such as... Figure 2 As shown.

[0069] In the downsampling path, the network extracts key feature information from MRI images through multi-layer convolution operations. However, relying on the progressive expansion of the receptive field by stacking single-scale convolution kernels makes it difficult to effectively capture the multi-scale features of anatomical structures and lesion areas in MRI images. This invention introduces a multi-scale convolution module in the contracting path to capture multi-level detailed features of MRI images, thereby enhancing feature representation capabilities. Specifically, convolution kernels of different scales focus on multi-level anatomical structures and pathological information through differentiated receptive fields: small-sized convolution kernels (3×3×3) accurately capture local fine structures, such as subtle structural changes in lesion areas in MRI, while large-scale convolution kernels (7×7×7) can capture the dependence between brain lesion areas and surrounding tissues through wide-range perception.

[0070] In the upsampling path, the network reconstructs high-resolution image details from low-resolution feature maps through upsampling operations. The expanding path contains three deconvolutional modules that concatenate the downsampled features with the upsampled features via skip connections. Anatomical structures and pathological features in medical images often exhibit cross-regional correlations; therefore, long-range dependency modeling is a core requirement for accurate diagnosis and image reconstruction. To capture long-range dependencies in images, the RWKV self-attention mechanism is introduced into 3D medical image generation, such as... Figure 2As shown, in the spatial blending module, multi-scale deep convolutional combinations using omni-shift layers enhance local context awareness, capturing subtle structures at lesion edges. In the channel blending module, a squared ReLU activation function is introduced to enhance nonlinear interactions between channels. Traditional self-attention mechanisms suffer from computational complexity that increases quadratically with sequence length. RWKV addresses this by employing a linearly complex recurrent attention mechanism (Re-WKV) to model global dependencies in medical images with linear complexity, generating high-quality images. This structure overcomes the shortcomings of traditional 3D UNET structures, such as insufficient feature extraction capabilities and difficulty in modeling global dependencies.

[0071] The advantages of generators include:

[0072] Firstly, spatial consistency is maintained through 3D modeling, avoiding the blurry reconstructed image caused by information fragmentation between slices in traditional methods.

[0073] Secondly, multi-scale feature extraction is achieved by using downsampling with different convolution kernels, and the RWKV attention mechanism is used to realize the global dependency relationship of brain image modeling, ensuring the rationality of the image structure generated in three-dimensional space, and providing a high-quality image foundation for downstream multimodal diagnostic tasks.

[0074] 3. A dual discriminator structure based on a point discriminator and a disease perception discriminator

[0075] Inspired by the point set representation and contextual clustering strategy of PC-GAN models, this invention introduces a point discriminator to replace the traditional standard discriminator. When generating PET images, traditional methods often suffer from local structural distortions that affect the diagnostic accuracy of the generated images. The point discriminator, however, strengthens the local-global relationship modeling through contextual clustering modules (CoC Blocks), effectively constraining the rationality of the generated brain image structure. The point discriminator maps PET voxels to a point set representation that fuses three-dimensional coordinates with brain texture features, explicitly preserving the spatial distribution information of brain structures. This allows for precise capture of the morphological features of small organs or minute lesions, avoiding the problem of blurred brain peripheral regions caused by fixed grid division in traditional voxel discriminators. This characteristic is particularly crucial in medical image generation.

[0076] The loss function of the point discriminator is defined as:

[0077]

[0078] Among them, D p Representative point discriminator, Real data representing the distribution of PET images, The PET image generated by the generator from the MRI image represents the image generated by the generator.

[0079] In traditional Generative Adversarial Network (GAN) frameworks, the core objective of the model is to achieve high-quality image reconstruction through adversarial training, typically focusing on the generation of global content and optimization of overall visual effects. However, this global generation strategy has significant limitations in the field of medical imaging, especially when it comes to reconstructing disease-specific pathological features. Traditional GANs often struggle to accurately capture and reproduce subtle abnormal structures that are crucial for diagnosis. To overcome this limitation, this invention designs a disease-perceiving discriminator, combining disease classification tasks with the GAN framework. By introducing disease-related discriminative constraints during the generation process, the model's ability to model pathological features is enhanced. A multi-task learning mechanism is introduced through the disease-perceiving discriminator, simultaneously optimizing image realism and disease classification accuracy along with the point discriminator. Furthermore, by adding loss constraints to the disease-perceiving discriminator during training, the generative model is guided to focus on details related to lesion regions, enabling the generated images to produce more pathological markers crucial for diagnosis, thus providing more reliable disease diagnostic results for downstream tasks.

[0080] The core of the disease perception and discrimination module consists of three DM Block modules connected in series, such as... Figure 2 As shown, the DM Block module consists of densely connected blocks, Mamba modules, and convolutions, as follows. Figure 3 As shown.

[0081] In the DM Block, densely connected blocks connect the output features of each layer with the features of previous layers through a feature reuse mechanism. This mechanism enhances the expressive power of local features and improves the discriminator's ability to represent local features. The Mamba module, on the other hand, captures the dependencies between the long-range spatial locations of features in each layer, effectively compensating for the limitation of traditional convolutions in capturing only local features. Densely connected blocks focus on enhancing the expressive power of local features, while the Mamba module focuses on capturing long-range spatial dependencies. The combination of these two approaches achieves complementarity between local and global features, effectively improving the performance of the disease perception discriminator and enabling the inclusion of more pathological features in the image generation process.

[0082] The enhanced feature X with channel number C in the densely connected block is first processed by the densely connected block, then normalized using a LayerNorm layer, and subsequently uniformly divided into four sub-features, namely... Each sub-feature has C / 4 channels. Next, each sub-feature is processed separately into a Mamba module to enhance its expressive power. The output of the Mamba module is optimized through residual connections and adjustment factors, thereby improving the model's performance in capturing long-range spatial information. The four processed sub-features are then merged back into a single feature X with C channels using a Concat operation. outFinally, X out The DM Block is output by sequentially passing through LayerNorm, Projection, and convolution layers.

[0083]

[0084] Where Dense represents dense connections, LN represents layer normalization, and Sp represents channel splitting operation. This represents the sub-features processed by the Mamba operation after segmentation, where Mamba refers to the Mamba operation. This represents the sub-features processed by the Mamba operation after segmentation, θ is the adjustment factor in the residual join, Concat represents the concatenation operation, and X out express.

[0085] To further refine task-related features, a 3×3×3 max-pooling layer (stride 2) is used to downsample the feature map. Finally, a fully connected layer and a SoftMax activation function are used to output the disease classification probability. This module jointly optimizes the adversarial loss and classification loss, allowing the generator to retain pathological information highly relevant to disease diagnosis during reconstruction, thereby improving the downstream multimodal classification model's ability to distinguish between diseased and healthy subjects. The loss function of the disease perception discriminator is defined as:

[0086]

[0087] Among them, D disease This represents a disease perception and discriminator, which outputs a predicted label for the disease. i The label represents the actual input image.

[0088] The loss of the entire generative model combines L1 loss, multi-scale structural similarity loss, point discriminator loss, and disease perception discriminator loss, and is defined as:

[0089]

[0090] Where, γ, λ, These are the weighting coefficients for each individual loss.

[0091] 4. Feature fusion modules based on different levels

[0092] In the network architecture design, to facilitate the extraction and effective fusion of unique features from two modalities, this invention simultaneously designs two independent symmetrical path structures. Each branch consists of three dense blocks, a convolutional network, and pooling operations. The fused features output by each fusion module are concatenated with the original features from the previous layer and passed as input to subsequent dense blocks. Through this dual-path dense connection structure, full interaction and deep fusion of features in both spatial and channel dimensions are achieved.

[0093] A deep feature fusion module (DFFB) based on a cross-attention mechanism is used to fuse deep features from two modalities, such as... Figure 4 As shown, DFFB first utilizes the DSC Block to enhance the features of both branches. The DSC Block extracts multi-level features through the DenseBlock, preserving rich local and global information. Combined with the SEBlock's channel attention mechanism, it dynamically adjusts feature weights, enhancing task-related features and suppressing redundant information. Furthermore, it extracts local spatial context information through 3×3×3 convolution. This design effectively enhances the discriminative power of the features. The enhanced features are then projected, transforming them into keys and values. The formula is as follows:

[0094] K i V i =Reshape(DSC(F) i ));

[0095] Among them, F i K represents two modal features of the input. i V i This represents the key and value generated by the feature transformation.

[0096] To fully exploit the complementary information between MRI and PET multimodal features, this invention generates a query by concatenating these two features, as shown in the following expression:

[0097] Q = Reshape(DSC(Concat(F)) MRI ,F PET ));

[0098] Where Q represents the generated query generated by feature transformation.

[0099] After calculating the key, value, and query, this invention utilizes a cross-modal attention mechanism to compute a modality-specific attention graph for each modality.

[0100] A = softmax(QK) T );

[0101] Where A represents the generated modality-specific attention map.

[0102] Subsequently, this value is multiplied by the attention weight to generate features containing global context information. This invention adds the global context features to the original features of another branch and concatenates the resulting features along the channel dimension. Finally, the concatenated features are input into a convolutional layer to generate the final fused features.

[0103] F fusion =Conv(Concat(F MRI ⊕Proj(A PET V PET ),F PET ⊕

[0104] Proj(A MRI V MRI ));

[0105] Among them, F MRI F represents the input MRI modal features. PET A represents the input PET modal features. PET Showing the generated PET modality-specific attention map, A MRI The generated PET modality-specific attention map, V PET V represents the value generated by the PET modal feature transformation. MRI This represents the value generated by MRI modality feature transformation.

[0106] DFFB achieves cross-modal global feature interaction through a cross-attention mechanism, thereby improving the disease discrimination capability of fused features.

[0107] Example 2

[0108] refer to Figures 5-6 Based on the above embodiment one, the effects and steps are disclosed.

[0109] 1. Image generation results

[0110] In the DAMSA-GAN model, the effectiveness of the proposed method in image generation tasks was evaluated and compared with other state-of-the-art methods. Among the image evaluation metrics, higher values ​​for the Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR) indicate better image quality, while lower values ​​for Mean Squared Error (MSE) and Mean Absolute Error (MAE) are better.

[0111] Table 1: Comparison of performance indicators of different generation methods in PET image generation

[0112]

[0113] The proposed generative model is compared with four image generation models: 3D Uet, Attention Uet, BrainStatTrans-GAN, and TPA GAN. Table 1 shows the metrics for each model's generation task. Figure 5 The comparison of horizontal slice error maps for images generated by various models is presented. 3DUet, as the baseline model, suffers from variations in detail and structure due to the limited receptive field of its convolutional kernels and the lack of a discriminator to effectively constrain the generated images. Attention Uet, by introducing an attention module at skip connections, optimizes feature selection in these connections. Attention Uet achieves a 2.4% improvement in SSIM compared to 3DUnet, and a 0.6dB increase in PSNR, reducing the voxel value error between generated and real images. BrainStatTrans-GAN, by incorporating a discriminator and attention mechanism, effectively models the extracted deep brain features globally. The introduction of an adversarial discriminator constraint optimizes the quality of generated images during training. BrainStatTrans-GAN achieves a 2.1% improvement in SSIM compared to 3D Unet, and a 0.7dB increase in PSNR, but variations in detail and structure still exist in the cerebellum and brainstem regions. The TPA GAN self-attention module is introduced into the last upsampling block, enabling the modeling of structural correlations in distant regions of the image, thereby generating clearer details. The addition of a discriminator forces the generated PET images to retain key features relevant to disease classification. TPA GAN achieves a 0.5% improvement in SSIM compared to BrainStatTrans GAN, and a simultaneous 0.5 dB improvement in PSNR, further enhancing the detail quality of the generated images and reducing voxel-level errors. The model of this invention achieves optimal image quality metrics and minimizes voxel-level errors. The generated images retain complete details and structures, and significantly outperform other models in key AD regions such as the hippocampus and cerebellum, validating the effectiveness and advancement of the proposed method.

[0114] Table 2: Ablation Experiment Results

[0115]

[0116]

[0117] To verify the effectiveness of the DM Block in the disease perception discriminator, this invention first designed ablation experiments with different feature extraction modules and compared the performance differences of different feature extraction modules, as shown in Table 2. Specifically, this invention replaced the DM Block with a convolutional module, a Dense Block, and a Conv+Mamba Block, respectively, while keeping other structures unchanged. In Conv, this invention only uses the convolutional module to extract local features of brain structures. Since the convolutional module can only capture local features and fails to consider the dependencies of global features, the generated image has the worst performance in all metrics, indicating that local features are difficult to model the complex anatomical relationships of the brain. In the Dense Block, the SSIM of the generated image is improved by 0.1% compared to Conv, and the PSNR is also improved by 0.1dB. The feature reuse mechanism of the Dense Block improves local details. In the Conv+Mamba Block, a Mamba Block is added after each convolutional module. The convolutional module captures local features, while the Mamba Block captures global features. By combining local and global feature extraction, the quality metrics of the generated image are improved by 0.4% in SSIM and 0.3dB in PSNR compared to using only the Conv feature extraction module. In the DM Block, this module combines the advantages of the Dense Block and the Mamba Block. The Dense Block excels in local feature extraction, while the Mamba Block effectively captures the dependencies of global features. The DM Block performs best in the generation task, which verifies the effectiveness of the DM Block in combining local and global features.

[0118] Meanwhile, ablation experiments were conducted on each module, including

[0119] A. Basic 3D Unet as a baseline model

[0120] B. Based on (1), add a point discriminator to construct a Point GAN.

[0121] C. Based on (2), multi-scale convolution and RWKV self-attention module are further introduced.

[0122] D. Add a disease perception discriminator to (3) to form the final improved model.

[0123] Table 2 shows the metrics for generating images. Figure 6This is a demonstration of horizontal slices of the generated image. The voxel values ​​of the brain image generated by the baseline 3D Unet differ significantly from the original image, showing some differences in detail and structure. After introducing a point discriminator on top of 3D Unet, Point GAN's SSIM is improved by 1.1% compared to 3D Unet's SSIM, and PSNR is simultaneously improved by 0.2dB. As can be seen from the error map, the edge errors of brain regions in the generated image are improved. The point discriminator explicitly preserves brain structural information through point sets, accurately capturing the morphological features of small lesions, which effectively improves the detail representation of the image. Further adding multi-scale convolution and RWKV self-attention module improves SSIM by 0.7% compared to Point GAN, and PSNR is simultaneously improved by 0.6dB, further enhancing the quality of the generated image. Among them, multi-scale convolution enhances the model's ability to capture features at different scales, while the RWKV self-attention module effectively models the dependencies of global features. After introducing a disease-aware discriminator into the final model, the SSIM of the generated images improved by 1.1% compared to the model without the discriminator, and the PSNR also improved by 0.5 dB. The disease-aware discriminator, through the constraints of the disease task, further optimized the structure and detail of the images, especially demonstrating outstanding performance in generating disease-related regions. With the gradual introduction of each module, the details and structure of the generated images became increasingly clear, particularly in disease-related regions, where the performance more closely resembled that of realistic images, validating the effectiveness of the proposed method.

[0124] 2. Disease Classification Results and Analysis

[0125] In the Cross-Modality Dense Fusion Network (CM-DFNet) model, in order to evaluate the improvement effect of multimodal fusion on disease classification performance, the classification metrics are based on four core metrics: accuracy (ACC), sensitivity (SEN), specificity (SPE), and area under the curve (AUC).

[0126] Table 3: Ablation Experiment Results

[0127]

[0128] After generating the missing data, the generated paired data is added to the training set of CM-DFNet to fully utilize the missing modality data to improve diagnostic performance. This invention first constructs a single-modality classification network (Single CM-DFNet) for T1-weighted magnetic resonance imaging (T1w-MRI) and positron emission tomography (PET), using a single-branch structure with a single modality as input. Table 3 shows that the diagnostic accuracy of single-modality T1w-MRI and PET is 86.94% and 88.98%, respectively. Both modalities diagnose diseases from the perspectives of metabolic activity and changes in brain structure. In the table, "Concat CM-DFNet" indicates that features extracted from the two branches are concatenated in a fully connected layer. The improved accuracy achieved by concatenating the features of the two branches verifies that multimodal data provides more diagnostic information compared to single-modality data. Multimodal methods are significantly superior to single-modality methods in disease diagnosis, exhibiting an inherent advantage.

[0129] Meanwhile, to verify the effectiveness of the shallow and deep fusion modules of this invention, a deep feature fusion module (DFFB) is added to the simple concatenated two-branch network. Our modal (w / o SFFB) indicates the addition of the deep feature fusion module (DFFB) to the two-branch network. DFFB can effectively fuse features from two different modalities through a cross-modal attention mechanism, achieving global feature interaction across modalities. Compared with before adding the deep feature fusion module, all indicators are improved. In the table, Our modal indicates the addition of deep and shallow feature fusion modules to the network. The shallow fusion module fuses edge and texture features of brain regions in the early stages of the network, while the channel attention in SFFB suppresses redundant features from both modalities. This allows the deep feature fusion model of this invention to make full use of the key complementary information of different modalities, thereby achieving the best classification performance. Our modal (w / o GAN) indicates that the generated PET data was not added to the training set, and it can be seen that the classification indicators did not reach the best results. By adding the missing generated data to the training set, CM-DFNet can effectively utilize the information of the missing modalities during training, improving the diagnostic accuracy of the model.

[0130] Table 4: Comparison of Indicators of Different Classification Methods in Multimodal Classification

[0131]

[0132] To further verify the effectiveness of the proposed method, it was compared with existing multimodal classification methods based on PET and MRI data from the ADNI database. The experimental results are shown in Table 5. The data from the comparative experiments in the table are all from the ADNI database. The VGGNet, ResNet, DenseNet, and Google Net methods utilize VGG16, ResNet18, DenseNet121, and Google Net as backbone networks for feature extraction, respectively, and then fuse the two modalities by concatenating the extracted deep features. The fusion mechanism is limited to feature concatenation, the fusion strategy is too simple, and the classification effect depends on the feature extraction capability of the backbone network, failing to consider other fusion strategies. The RMFN classification method achieves fusion by superimposing cross-branch features through residual connections, but this method fails to achieve sufficient fusion. The PT-DCN model fuses the two images in a fusion block, which consists of multiple convolutions. This method's fusion strategy struggles to establish a global relationship between the two modalities, and the improvement in classification effect using convolutional fusion is not significant. The RegBN method extracts independent features for each modality in a dual-branch path using a cross-attention mechanism, and then concatenates and fuses these redundant features. However, this approach fails to consider the multi-stage fusion of features at different scales. The proposed AD diagnosis method utilizes features from different stages for different fusion blocks, outperforming existing methods and further demonstrating its superiority in multimodal data fusion and disease classification.

[0133] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. An Alzheimer's disease diagnostic method based on missing mode generation, characterized in that, The steps are as follows: S1: Construct a self-attention generator based on multi-scale and RWKV; S2: Construct a dual discriminator structure based on a point discriminator and a disease perception discriminator; S3: Construct feature fusion modules based on different levels; S4: Construct a DAMSA-GAN network and a multimodal fusion classification CM-DFNet network to generate missing PET modalities. This allows the DAMSA-GAN network to generate missing PET images using existing MRI images, and utilizes additional information from PET to improve the disease diagnosis performance of the multimodal fusion classification CM-DFNet network.

2. The Alzheimer's disease diagnostic method based on missing mode generation according to claim 1, characterized in that, In step S1, the steps are as follows: The generator part of DAMSA-GAN consists of a contracting path composed of multiple downsampling operations and an expanding path composed of multiple upsampling operations. In the downsampling path, key feature information of MRI images is extracted through multi-layer convolution operations. A multi-scale convolution module is introduced in the contracting path to enhance feature representation. In the upsampling path, high-resolution image details are reconstructed from low-resolution feature maps through upsampling operations.

3. The Alzheimer's disease diagnostic method based on missing mode generation according to claim 2, characterized in that, The upsampling path involves the following steps: Introducing the RWKV self-attention mechanism into 3D medical image generation; By combining multi-scale deep convolutions with omnidirectional displacement layers, local context awareness is enhanced, and subtle structures at the edge of lesions are captured. In the channel mixing module, a squared ReLU activation function is introduced to enhance the nonlinear interaction between channels.

4. The Alzheimer's disease diagnostic method based on missing mode generation according to claim 1, characterized in that, In step S2, the steps are as follows: Replace the standard discriminator with a point discriminator. The point discriminator enhances local-global relationship modeling through a context clustering module; The point discriminator maps PET voxels into a point set representation that fuses three-dimensional coordinates with brain texture features.

5. The Alzheimer's disease diagnostic method based on missing mode generation according to claim 4, characterized in that, In step S2, the steps are as follows: We construct a disease perception discriminator, combining the disease classification task with the GAN framework. By introducing disease-related discriminative constraints during the generation process, we enhance the model's ability to model pathological features. A disease perception discriminator is introduced into a multi-task learning mechanism to simultaneously optimize image realism and disease classification accuracy with a point discriminator. At the same time, the loss constraint of the disease perception discriminator is added to the loss during the training process.

6. The Alzheimer's disease diagnostic method based on missing mode generation according to claim 5, characterized in that, The disease perception discriminator is composed of three DM Block modules connected in series. Each DM Block module consists of a dense connection block, a Mamba module, and a convolution. The enhanced feature X with channel number C in the densely connected block is first processed by the densely connected block, then normalized using a LayerNorm layer, and subsequently uniformly divided into four sub-features, namely... The number of channels for each sub-feature is C / 4; Each sub-feature is input into a Mamba module for processing, and the output of the Mamba module is optimized through residual connections and adjustment factors; The four processed sub-features are then merged back into a single feature Xout with C channels using the Concat operation. Xout passes through LayerNorm, Projection, and convolution layers in sequence to complete the final output of the DM Block.

7. The Alzheimer's disease diagnostic method based on missing mode generation according to claim 1, characterized in that, In step S3, the steps are as follows: Design two independent symmetrical path structures, each branch consisting of three dense blocks, a convolutional network, and a pooling operation; The fused features output by each fusion module are concatenated with the original features of the previous layer and passed as input to subsequent dense blocks.

Citation Information

Cited By

  • Alzheimer's disease data completion method based on multi-modal generation and fusion

    CN121439269A