Medical image analysis method and device based on multi-mode fusion network
Patent Information
- Application Number
- CN202610393285.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-27
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-03-27
AI Technical Summary
[0003]然而,现有多模态融合方法普遍采用跨模态统一架构,并依赖静态融合策略,在临床实践中,不同患者的sMRI和PET影像可能呈现不同程度的病理变化,某些患者可能在一种模态中表现出更明显的异常
本申请通过分别对结构磁共振成像的第一图像和正电子发射断层扫描的第二图像进行特征提取,并将特征图分别沿空间维度展平为序列,保留了各模态的特征完整性,其次,对第一序列和第二序列分别进行注意力加权处理得到判别特征,实现了对疾病相关区域的有效定位。然后,通过对第一判别特征和第二判别特征进行跨模态融合,得到综合了两种模态互补信息的融合特征,该过程能够根据待分析对象的个体特性动态调整不同模态的贡献权重,有效适应患者个体间的异质性。最后,将融合特征输入分类器输出医学图像分析结果,提升了医学图像分析的准确率,实现了对阿尔茨海默病尤其是早期阶段的精确识别。
Smart Images

Figure CN122049543B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing technology, specifically to a medical image analysis method and device based on a multi-mode fusion network. Background Technology
[0002] With the accelerating aging of the global population, Alzheimer's disease (AD) has become one of the leading neurodegenerative diseases affecting the health and quality of life of the elderly, impacting more than 50 million people worldwide. Alzheimer's disease often presents with subtle symptoms in its early stages (such as mild cognitive impairment, MCI), yet this stage is a critical window for intervention and treatment. Therefore, developing accurate early diagnostic methods is crucial for clinical intervention and disease management. Medical imaging technologies, particularly structural magnetic resonance imaging (sMRI) and positron emission tomography (PET), have become important tools for the early diagnosis of Alzheimer's disease, providing complementary anatomical and metabolic functional information, respectively.
[0003] However, existing multimodal fusion methods generally employ a unified cross-modal architecture and rely on static fusion strategies. In clinical practice, different patients' sMRI and PET images may exhibit varying degrees of pathological changes, and some patients may show more pronounced abnormalities in one modality. Static fusion methods struggle to adapt to this heterogeneity among individual patients, resulting in low accuracy in the analysis of medical images of the analyzed subjects. Summary of the Invention
[0004] In view of this, this application provides a medical image analysis method and device based on a multi-mode fusion network.
[0005] In a first aspect, this application provides a medical image analysis method based on a multi-modal fusion network, the method comprising: Acquire a first image of the object to be analyzed in structural magnetic resonance imaging and a second image in positron emission tomography (PET); Feature extraction is performed on the first image and the second image respectively to obtain a first feature map and a second feature map, and the first feature map and the second feature map are flattened along the spatial dimension into a first sequence and a second sequence respectively; The first sequence and the second sequence are subjected to attention weighting processing respectively to obtain the first discriminant feature and the second discriminant feature; Cross-modal fusion is performed on the first discriminant feature and the second discriminant feature to obtain the fused feature; The fused features are input into the classifier, which outputs the medical image analysis results of the object to be analyzed.
[0006] By employing the aforementioned technical solution, features are extracted from the first image of structural magnetic resonance imaging (SMRI) and the second image of positron emission tomography (PET), and the feature maps are flattened into sequences along the spatial dimension, preserving the feature integrity of each modality. Next, attention-weighted processing is applied to the first and second sequences to obtain discriminative features, achieving effective localization of disease-related regions. Then, cross-modal fusion of the first and second discriminative features yields fused features that integrate complementary information from both modalities. This process dynamically adjusts the contribution weights of different modalities based on the individual characteristics of the analyzed object, effectively adapting to the heterogeneity among individual patients. Finally, the fused features are input into a classifier to output medical image analysis results, improving the accuracy of medical image analysis and achieving precise identification of Alzheimer's disease, especially in its early stages.
[0007] Optionally, feature extraction is performed on the first image and the second image respectively to obtain a first feature map and a second feature map, including: The first image and the second image are respectively input into the backbone of the three-dimensional residual network for convolution processing to obtain the first initial feature map and the second initial feature map. Calculate the attention weights of the first initial feature map in the spatial dimension, and perform weighted processing on the first initial feature map to obtain the first feature map; Calculate the attention weights of the second initial feature map along the channel dimension, and then perform weighted processing on the second initial feature map to obtain the second feature map. Optionally, attention-weighted processing is performed on the first sequence and the second sequence respectively to obtain a first discriminant feature and a second discriminant feature, including: The basic attention signal and the classifier-guided attention signal are calculated for the first sequence; a first selection ratio is adaptively determined based on the global features of the first sequence; the classifier-guided attention signal is selectively retained based on the first selection ratio; the basic attention signal and the retained classifier-guided attention signal are fused to obtain a first attention weight; the first attention weight is used to weight the first sequence to obtain the first discriminative feature. The basic attention signal and the classifier-guided attention signal are calculated for the second sequence; a second selection ratio is adaptively determined based on the global features of the second sequence; the classifier-guided attention signal is selectively retained based on the second selection ratio; the basic attention signal and the retained classifier-guided attention signal are fused to obtain the second attention weight; the second attention weight is used to weight the second sequence to obtain the second discriminative feature.
[0008] Optionally, cross-modal fusion is performed on the first discriminative feature and the second discriminative feature to obtain fused features, including: The first discriminant feature and the second discriminant feature are respectively input into a multilayer perceptron for projection to obtain the first projection feature and the second projection feature; Multi-head attention calculation is performed on the first projection feature and the second projection feature to obtain the first enhanced feature and the second enhanced feature; Global average pooling is performed on the first enhanced feature and the second enhanced feature respectively to obtain the first global statistical information and the second global statistical information; Calculate similarity weights based on the first global statistical information and the second global statistical information; The first enhanced feature and the second enhanced feature are weighted and summed according to the similarity weight to obtain the fused feature.
[0009] Optionally, multi-head attention computation is performed on the first projection feature and the second projection feature to obtain a first enhanced feature and a second enhanced feature, including: Using the first projection feature as the query vector and the second projection feature as the key vector and value vector, perform the first multi-head attention calculation to obtain the first intermediate feature; Using the second projection feature as the query vector and the first projection feature as the key vector and value vector, a second multi-head attention calculation is performed to obtain the second intermediate feature. The first intermediate feature and the first projected feature are residually connected to obtain the first enhanced feature, and the second intermediate feature and the second projected feature are residually connected to obtain the second enhanced feature.
[0010] Optionally, the fused features are input into a classifier, and the medical image analysis results of the object to be analyzed are output, including: The fused features are subjected to global average pooling to obtain a global feature vector; The global feature vector is input into a fully connected layer for classification mapping; Calculate the probability distribution of each category using the softmax activation function; The medical image analysis results of the object to be analyzed are determined based on the probability distribution.
[0011] Optionally, after outputting the medical image analysis results of the object to be analyzed, the method further includes: A first verification result is output based on the first discriminant feature, and a second verification result is output based on the second discriminant feature; Calculate the consistency score between the medical image analysis results and the first verification result and the second verification result; When the consistency score is lower than a preset threshold, the medical image analysis result is marked as low confidence and a prompt message indicating that manual review is required is output.
[0012] A second aspect of this application provides an electronic device for medical image analysis based on a multi-mode fusion network, the electronic device comprising: one or more processors and a memory; the memory being coupled to the one or more processors, the memory being used to store computer program code including computer instructions, the one or more processors calling the computer instructions to cause the electronic device for medical image analysis based on the multi-mode fusion network to perform the method described in the first aspect and any possible implementation thereof.
[0013] A third aspect of this application provides a computer program product containing instructions that, when run on an electronic device for medical image analysis based on a multi-modal fusion network, causes the electronic device to perform the method described in the first aspect and any possible implementation thereof.
[0014] A fourth aspect of this application provides a computer-readable storage medium including instructions that, when executed on an electronic device for medical image analysis based on a multi-modal fusion network, cause the electronic device to perform the method described in the first aspect and any possible implementation thereof.
[0015] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: This application extracts features from a first image obtained from structural magnetic resonance imaging (MRI) and a second image obtained from positron emission tomography (PET), and flattens the feature maps into sequences along the spatial dimension, preserving the feature integrity of each modality. Next, attention-weighted processing is applied to the first and second sequences to obtain discriminative features, achieving effective localization of disease-related regions. Then, cross-modal fusion of the first and second discriminative features yields fused features that integrate complementary information from both modalities. This process dynamically adjusts the contribution weights of different modalities based on the individual characteristics of the analyzed object, effectively adapting to the heterogeneity among individual patients. Finally, the fused features are input into a classifier to output medical image analysis results, improving the accuracy of medical image analysis and achieving precise identification of Alzheimer's disease, especially in its early stages. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a medical image analysis method based on a multi-modal fusion network provided in an embodiment of this application. Figure 2 This is a flowchart of a logic-guided attention mechanism provided in an embodiment of this application; Figure 3 This is a schematic diagram of a multimodal medical image feature fusion architecture based on similarity-guided gating fusion provided in an embodiment of this application; Figure 4 This is an overall architecture diagram of medical image analysis based on a multi-modal fusion network provided in an embodiment of this application; Figure 5 This is a schematic diagram of an exemplary hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0018] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0019] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0020] Please refer to Figure 1 This paper presents a flowchart illustrating a medical image analysis method based on a multi-mode fusion network. This method can be implemented using a computer program, a microcontroller, or run on a device based on a multi-mode fusion network for medical image analysis. The computer program can be integrated into the computer device or run as a standalone application. Specifically, the method includes steps 10 to 50, as follows: Step 10: Acquire the first image of the object to be analyzed in structural magnetic resonance imaging and the second image in positron emission tomography.
[0021] In this embodiment of the application, the subject to be analyzed refers to a patient or subject diagnosed with Alzheimer's disease. The first image refers to an image of the brain anatomy of the subject to be analyzed, obtained through structural magnetic resonance imaging (sMRI) technology. This image mainly reflects the anatomical features of the brain, such as structural information like gray matter volume and white matter integrity. The second image refers to an image of the brain's metabolic function, obtained through positron emission tomography (PET) technology. This image mainly reflects the brain's metabolic activity, such as glucose metabolism and amyloid deposition.
[0022] Specifically, the system can retrieve structural magnetic resonance imaging (sMRI) and positron emission tomography (PET) image data of the subject to be analyzed from the hospital's authorized Picture Archiving and Communication System (PACS). Using the standard DICOM (Digital Imaging and Communications in Medicine) protocol, the system can access and read medical image files stored on the PACS server. Operators only need to enter the unique identifier of the subject to be analyzed (such as medical record number, ID number, etc.) into the system interface, select the corresponding examination type (sMRI and PET) and examination date, and the system will automatically retrieve and download the corresponding first and second images from the PACS server.
[0023] In addition to acquiring images from the PACS system, the system in this embodiment also supports direct import of DICOM format image files from local storage devices (such as hospital workstations, dedicated servers, or external storage media). Operators can use the system's file browser function to locate and select the sMRI and PET image files of the object to be analyzed for import. The system supports batch import functionality, allowing simultaneous import of multimodal image data from multiple patients for batch processing and analysis. After acquiring the first and second images, the system automatically checks the integrity and quality of the image data to ensure that the data meets the requirements for subsequent processing. If missing or poor-quality data is found, the system will prompt the operator to re-acquire or select alternative data.
[0024] Step 20: Extract features from the first image and the second image respectively to obtain the first feature map and the second feature map, and flatten the first feature map and the second feature map along the spatial dimension to obtain the first sequence and the second sequence respectively.
[0025] In this embodiment of the application, the first feature map refers to the feature representation obtained by extracting features from the first image of structural magnetic resonance imaging. It contains anatomical structural information in the original first image and exists in the form of a three-dimensional feature map, in which the feature vector of each location reflects the structural characteristics of the corresponding brain region.
[0026] The second feature map refers to the feature representation obtained by extracting features from the second image of the positron emission tomography scan. It contains metabolic function information from the original second image and also exists in the form of a three-dimensional feature map, where the feature vector at each location reflects the metabolic activity characteristics of the corresponding brain region.
[0027] The first sequence refers to the feature sequence formed by flattening the first feature map along the spatial dimension. It converts the three-dimensional structural feature map into a two-dimensional sequence representation, which is convenient for subsequent sequence processing models (such as attention mechanisms) to process it. Specifically, if the dimension of the first feature map is [C1×H×W×D] (where C1 is the number of channels, and H, W, and D are the height, width, and depth, respectively), then the dimension of the flattened first sequence is [N1×C1], where N1=H×W×D, representing the number of spatial positions.
[0028] The second sequence refers to the feature sequence formed by flattening the second feature map along the spatial dimension, converting the three-dimensional functional feature map into a two-dimensional sequence representation, which also facilitates subsequent sequence processing. Specifically, if the dimension of the second feature map is [C2×H×W×D] (where C2 is the number of channels, and H, W, and D are the height, width, and depth, respectively), then the dimension of the flattened second sequence is [N2×C2], where N2=H×W×D, representing the number of spatial locations.
[0029] As an optional embodiment, the step of extracting features from the first image and the second image to obtain a first feature map and a second feature map may further include the following steps: Step 101: Input the first image and the second image into the backbone of the three-dimensional residual network for convolution processing to obtain the first initial feature map and the second initial feature map.
[0030] Specifically, in order to fully extract feature information from different modalities of medical images, this application embodiment designs two unique 3D ResNet encoders to extract relevant features from sMRI (Ig) and PET (Ip) modalities, respectively. The first image (sMRI image) and the second image (PET image) of the object to be analyzed are respectively input into the corresponding three-dimensional residual network (3DResNet) backbone for processing.
[0031] The input dimensions of both the first and second images are Im∈R. (B×100×120×100) Where B represents the batch size, and 100×120×100 represents the size of the input 3D shape. The images of the two modalities are processed by convolution through their respective 3D ResNet backbone networks, where each 3D ResNet contains an initial convolutional layer and multiple residual blocks, which can effectively process the spatial information of 3D medical image data.
[0032] Through processing by the 3D ResNet backbone network, the system converts the first image into a first initial feature map Fg and the second image into a second initial feature map Fp. Both initial feature maps have a dimension of R. (B×256×5×6×5 ), where B is the batch size, 256 represents the number of feature channels, and 5×6×5 corresponds to the spatial dimension.
[0033] Step 102: Calculate the attention weights of the first initial feature map in the spatial dimension, and perform weighted processing on the first initial feature map to obtain the first feature map.
[0034] Specifically, a spatial self-attention (SA) mechanism is applied to the first initial feature map Fg to enhance the spatial correlation of the anatomical structure. First, the first initial feature map Fg is flattened along the spatial dimension into a patch sequence Xg∈R^(B×150×256), where 150=5×6×5 represents the number of spatial locations.
[0035] Next, the spatial self-attention mechanism Φg is applied. att The feature sequence Xg is processed to calculate the correlations between different spatial locations, enhancing the spatial representation of anatomical structures. The spatial self-attention mechanism establishes long-distance dependencies between different spatial locations in the feature map by calculating the relationships between queries, keys, and values, making it particularly suitable for handling complex spatial relationships of anatomical structures in sMRI images. After spatial self-attention processing, the result is reshaped back to the original feature map shape, resulting in the first feature map, represented as: , m=g.
[0036] Step 103: Calculate the attention weights of the second initial feature map in the channel dimension, and perform weighted processing on the second initial feature map to obtain the second feature map.
[0037] Specifically, a channel attention (CA) mechanism is applied to the second initial feature map Fp to highlight the channel specificity of metabolic patterns. First, the second initial feature map Fp is flattened along the spatial dimension into a patch sequence Xp∈R. (B×150×256) , where 150 = 5 × 6 × 5 represents the number of spatial locations.
[0038] Next, the channel attention mechanism Φp is applied. att The feature sequence Xp is processed to focus on the relationships between different channels, which is particularly important for capturing metabolic patterns in PET images. The channel attention mechanism automatically identifies and enhances the most discriminative channels by calculating the importance weights of different channels; these channels often correspond to specific types of metabolic patterns or disease-related functional activities. After channel attention processing, the result is reshaped back to the original feature map shape, resulting in a second feature map, represented as follows: , m=p.
[0039] For example, the process of extracting features from the first image and the second image to obtain the first feature map and the second feature map can be mathematically expressed as follows: ; in, These represent spatial attention and channel attention operations, respectively. The two backbone networks generate feature maps respectively. and Where B is the batch size, 256 represents the number of feature channels, and 5×6×5 corresponds to the spatial dimension. Each feature map is flattened along the spatial dimension to form a patch sequence. Then, attention mechanisms are used to further capture cross-modal correlations.
[0040] Step 30: Perform attention weighting on the first sequence and the second sequence respectively to obtain the first discriminant feature and the second discriminant feature.
[0041] Among them, the first discriminative feature refers to the discriminative feature representation extracted from the first feature map (i.e., the sMRI feature map after processing by the spatial self-attention mechanism). It mainly contains the key features of the brain anatomical structure information of the object to be analyzed, and can effectively distinguish different disease states or classifications.
[0042] The second discriminative feature refers to the discriminative feature representation extracted from the second feature map (i.e., the PET feature map after processing by the channel attention mechanism). It mainly contains key features of the brain metabolic function information of the subject to be analyzed, and can also effectively distinguish different disease states or classifications.
[0043] As an optional embodiment, the step of performing attention-weighted processing on the first sequence and the second sequence to obtain the first discriminant feature and the second discriminant feature may further include the following steps: Step 201: Calculate the basic attention signal and the classifier-guided attention signal for the first sequence; adaptively determine the first selection ratio based on the global features of the first sequence; selectively retain the classifier-guided attention signal based on the first selection ratio; fuse the basic attention signal and the retained classifier-guided attention signal to obtain the first attention weight; use the first attention weight to weight the first sequence to obtain the first discriminative feature.
[0044] Step 202: Calculate the basic attention signal and the classifier-guided attention signal for the second sequence; adaptively determine the second selection ratio based on the global features of the second sequence; selectively retain the classifier-guided attention signal based on the second selection ratio; fuse the basic attention signal and the retained classifier-guided attention signal to obtain the second attention weight; use the second attention weight to weight the second sequence to obtain the second discriminative feature.
[0045] Specifically, to enhance the ability to identify Alzheimer's disease characteristics, this application proposes a logic-guided attention module for processing the first sequence (sMRI modality, m=g) and the second sequence (PET modality, m=p). The specific implementation process is as follows: Given a sequence of input image patches Xm∈R from each modality (B×N×C) (Where N=150, C=256), the system first calculates two attention signals for each sequence. The basic attention signal is generated through a fully connected layer and an activation function, denoted as σ(W1·δ(W2·(Xm))), where W1∈R (64×256) W2∈R (256×64) These are learnable parameters, where δ(·) represents the ReLU activation function and σ(·) represents the sigmoid function. The classifier-guided attention signal is calculated by multiplying the classifier weights by the feature sequence, denoted as W. cls ·Xm, where W cls ∈R (2×256) This is the classifier weight matrix. Next, the system adaptively determines the selection ratio τm based on the global features of each sequence, calculated using the following formula: Where τmin=0.1 and τmax=0.9 are the preset minimum and maximum thresholds, W t These are learnable parameters, and GAP(·) represents the global average pooling operation. This adaptive mechanism ensures that the system can dynamically adjust the preferences of the attention mechanism according to the overall characteristics of different modal features. Then, the system selectively retains the classifier-guided attention signal based on its respective selection ratio τm, which is implemented using the TopKτ operation, denoted as TopKτ(W cls ·Xm). This operation retains the elements with the highest scores (τm) in the classifier-guided attention signal, while setting the remaining elements to zero, thus focusing on the most discriminative feature locations. Subsequently, the system fuses the base attention signal with the retained classifier-guided attention signal to obtain the attention weights α for each modality. m The calculation formula is: ; in, , These represent ReLU, sigmoid, and global average pooling, respectively. Used to balance the weights of the two attention signals, and These represent the attention weights and features of the nth image patch, respectively. These are discriminative brain pattern features that are used for subsequent fusion.
[0046] Through the above processing, the system obtains the first discriminative feature. Second discriminant features They captured anatomical features in the sMRI modality and metabolic function features in the PET modality, respectively, all of which focused on locations important for the diagnosis of Alzheimer's disease.
[0047] Please refer to Figure 2 This document presents a flowchart of a logic-guided attention mechanism provided in an embodiment of this application. It receives a feature map Fm as input and first reshapes it into a sequence Xm using a reshape operation. The processing then proceeds in three parallel paths: the upper path generates logits signals using a fully connected layer (FC) and takes the maximum value; the middle path calculates an adaptive selection ratio using global average pooling (GAP) and a multilayer perceptron (MLP) structure, which is then processed by a sigmoid function to control the TopK operation; the lower path generates basic attention weights α using another MLP structure and a sigmoid function. The TopK output of the middle path is added to and fused with the attention weights α from the lower path (⊕) to obtain a comprehensive attention weight. This weight is then multiplied by the input sequence Xm (⊗) to achieve feature weighting. Finally, the weighted features undergo weighted processing and L2 normalization to output a feature representation F'm with high discriminative power. This design introduces the classifier's supervisory information into the attention mechanism, enabling the system to adaptively focus on the feature regions most valuable for disease diagnosis, effectively improving the ability to identify Alzheimer's disease features.
[0048] Step 40: Perform cross-modal fusion on the first discriminant feature and the second discriminant feature to obtain the fused feature.
[0049] In this context, fusion features refer to the comprehensive feature representation obtained by fusing first discriminant features from the sMRI modality and second discriminant features from the PET modality. Specifically, fusion features are obtained by complementary integration of first discriminant features capturing information about brain anatomy and second discriminant features representing information about brain metabolic function through a specific multimodal fusion algorithm, thereby forming a comprehensive feature that can simultaneously represent both brain structure and function. This fusion feature fully utilizes the complementary advantages of different modalities of medical imaging, providing a more comprehensive and accurate disease representation than a single modality, and offering more reliable feature support for subsequent Alzheimer's disease diagnosis.
[0050] As an optional embodiment, the step of performing cross-modal fusion of the first and second discriminant features to obtain the fused features may further include the following steps: Step 301: Input the first discriminant feature and the second discriminant feature into the multilayer perceptron for projection to obtain the first projection feature and the second projection feature.
[0051] Specifically, the feature space is first transformed using a multilayer perceptron. Specifically, the first discriminative features of the sMRI modality are input into a dedicated multilayer perceptron (MLPg) for nonlinear transformation to obtain the first projected features. Similarly, the second discriminative feature of the PET modality is input into another multilayer perceptron (MLPp) to obtain the second projected feature. This projection process can be represented as: .
[0052] Step 302: Perform multi-head attention calculation on the first projection feature and the second projection feature to obtain the first enhanced feature and the second enhanced feature.
[0053] Specifically, after obtaining the projected features, a multi-head attention (MHA) mechanism is used for cross-modal feature enhancement. The MHA mechanism effectively allows each modality to focus on complementary information from another modality. The first projected features are then used. Second projection features As input to the multi-head attention computation, the enhanced features are calculated separately: .
[0054] As an optional embodiment, the step of performing multi-head attention calculation on the first projection feature and the second projection feature to obtain the first enhanced feature and the second enhanced feature may further include the following steps: Step 401: Using the first projected feature as the query vector and the second projected feature as the key vector and value vector, perform the first multi-head attention calculation to obtain the first intermediate feature; Step 402: Using the second projected feature as the query vector and the first projected feature as the key vector and value vector, perform a second multi-head attention calculation to obtain the second intermediate feature; Step 403: Perform a residual connection between the first intermediate feature and the first projected feature to obtain the first enhanced feature, and perform a residual connection between the second intermediate feature and the second projected feature to obtain the second enhanced feature.
[0055] Specifically, to fully realize the deep complementary fusion of sMRI and PET modal features, this application employs a cross-modal feature enhancement method based on a multi-head attention mechanism. This method achieves effective information interaction and complementary enhancement between modalities by enabling different modal features to pay attention to each other's information.
[0056] The first projection feature of the sMRI modality The second projection feature of the PET modality as the query vector (Q) The first multi-head attention calculation is performed using the key vector (K) and value vector (V). This calculation process can be detailed as follows: First, the query vector is... Key vector Sum value vector The sMRI features are mapped to h different subspaces (i.e., h attention heads) via linear projection. Then, in each subspace, the dot product similarity between the query vector and the key vector is calculated and normalized using a softmax function to obtain attention weights. Next, these attention weights are used to perform a weighted summation of the value vectors. Finally, the outputs of the h heads are concatenated and subjected to a linear transformation to obtain the first intermediate feature F'g_mid. Through this step, sMRI features can selectively focus on key information in the PET modality, thereby capturing metabolic functional features valuable for disease diagnosis in the PET modality. Since the two modalities capture different aspects of brain information (anatomical structure and metabolic function), this cross-modal attention mechanism allows sMRI features to acquire supplementary information that is difficult to observe in a single modality.
[0057] Similarly, the second projection features of the PET modality The first projection feature of the sMRI modality as the query vector (Q) As the key vector (K) and value vector (V), a second multi-head attention computation is performed to obtain the second intermediate feature. Through this computation, PET features can selectively focus on key information in the sMRI modality, capturing anatomical structural features valuable for disease diagnosis in the sMRI modality. This bidirectional attention computation ensures that the two modalities can fully exchange information, learn from each other, and enhance each other.
[0058] After performing cross-modal attention calculation, in order to retain the original feature information and incorporate the newly acquired complementary information, a residual connection method is used to add the intermediate features to the original projected features. That is, the first intermediate feature and the first projected feature are residually connected to obtain the first enhanced feature. And by performing a residual connection between the second intermediate feature and the second projected feature, a second enhanced feature is obtained. .
[0059] Step 303: Perform global average pooling on the first enhanced feature and the second enhanced feature respectively to obtain the first global statistical information and the second global statistical information.
[0060] Specifically, in order to obtain global statistical information on the enhanced features, the first enhanced feature... Second Enhancement Features Global average pooling (GAP) is applied to extract statistical information that characterizes the overall feature distribution: First global statistical information Second global statistics These global statistical information provide an important basis for subsequent calculation of similarity weights between modes, enabling the system to adaptively determine the fusion strategy.
[0061] Step 304: Calculate the similarity weight based on the first global statistical information and the second global statistical information.
[0062] Specifically, after obtaining global statistical information, this information is weighted using learnable parameters, and the similarity weight λ between modalities is calculated using the sigmoid function: ,in, It is a learnable parameter matrix. Let σ be the bias vector, σ(·) denote the sigmoid function, and [;] denote the vector concatenation operation. Through this calculation, the system can adaptively evaluate the relative importance between two modal features and generate a similarity weight λ in the range [0, 1] to balance the contributions of different modalities.
[0063] Step 305: The first enhanced feature and the second enhanced feature are weighted and summed according to the similarity weight to obtain the fused feature.
[0064] Specifically, the enhanced features are weighted and fused using the calculated similarity weight λ to obtain the final fused features: Through this adaptive weighted fusion method, the system can dynamically adjust the importance of each modality based on the characteristics of the current input, emphasizing more discriminative modalities for different cases, thus achieving more intelligent and effective modality fusion. This selective gating fusion method fully integrates the complementary information from sMRI and PET modalities, providing a more comprehensive and accurate feature representation for subsequent Alzheimer's disease diagnosis.
[0065] Please see Figure 3 This is a schematic diagram of a multimodal medical image feature fusion architecture based on similarity-guided gating fusion provided in an embodiment of this application. The left side shows two modal features input. First, the features are converted into projective features using a shared Linear-Layer Norm-ReLU network. The SGGF module in the middle implements a cross-modal attention mechanism by calculating inter-modal similarity, generating enhanced features. This effectively captures complementary information. The right-side fusion stage employs a dual-path design: the upper path captures higher-order interactions through element-wise multiplication, the lower path extracts abstract representations through MLP, and finally, element-wise addition is used to integrate them into fused features. This "enhancement first, then fusion" design fully leverages modal complementarity.
[0066] Step 50: Input the fused features into the classifier and output the medical image analysis results of the object to be analyzed.
[0067] The classifier refers to the neural network classification component used for Alzheimer's disease diagnosis. This component receives fused features as input and outputs disease diagnosis probabilities through fully connected layers and a softmax function. Specifically, the classifier contains one or more fully connected layers that map high-dimensional features to a disease category space, ultimately outputting a probability distribution of whether a sample belongs to different categories such as normal cognition (NC), mild cognitive impairment (MCI), or Alzheimer's disease (AD). During training, the classifier learns the decision boundaries that distinguish different disease states using a cross-entropy loss function, and accurately classifies new patient data during the inference phase.
[0068] Medical image analysis results refer to diagnostic conclusions obtained after processing and analyzing the patient's sMRI and PET brain images using the multimodal fusion method of this application. These conclusions include disease classification results (such as normal cognition, mild cognitive impairment, or Alzheimer's disease) and their corresponding confidence scores. These analysis results reflect the patient's brain health status and possible pathological changes, providing clinicians with objective diagnostic references, facilitating early detection and intervention of Alzheimer's disease, developing personalized treatment plans, and tracking disease progression.
[0069] As an optional embodiment, the step of inputting fused features into the classifier and outputting the medical image analysis results of the object to be analyzed may further include the following steps: Step 501: Perform global average pooling on the fused features to obtain the global feature vector.
[0070] Specifically, the fused features obtained in the preceding steps are first subjected to global average pooling to obtain a compact global feature representation. Global average pooling is a dimensionality reduction technique that reduces the spatial dimension of the feature map while preserving information in the channel dimension, thereby generating a fixed-length feature vector. By averaging each channel of the fused features, we obtain a global feature vector containing key discriminative information from the entire brain image. This step effectively reduces the number of parameters, avoids the risk of overfitting, and preserves the key discriminative information in the fused features, providing a more concise and effective feature representation for subsequent classification tasks. Global average pooling also has a certain regularization effect, which helps improve the model's generalization ability. Through this operation, we compress the fused features into a global feature vector that can comprehensively represent the patient's brain health status.
[0071] Step 502: Input the global feature vector into the fully connected layer for classification mapping, and calculate the probability distribution of each category using the softmax activation function.
[0072] Specifically, after obtaining the global feature vector, it is input into a multi-layer fully connected network for nonlinear transformation and classification mapping. Specifically, the global feature vector first undergoes feature transformation through a hidden layer, which typically contains 256 or 512 neurons and uses the ReLU activation function to introduce nonlinear characteristics. This hidden layer can further extract high-level abstract features, enhancing the model's expressive power. Subsequently, the output of the hidden layer is mapped to the category space through an output layer. In this embodiment, the output layer contains three neurons, corresponding to the three categories of normal cognition (NC), mild cognitive impairment (MCI), and Alzheimer's disease (AD), respectively. The raw scores generated by the output layer are then converted into a normalized probability distribution through a softmax activation function. The softmax function converts the raw scores into probability values ranging from 0 to 1, with the sum of the probabilities of all categories being 1, forming a valid probability distribution. Through this calculation, the system can generate a three-dimensional probability vector for each sample, representing the probability of it belonging to each category.
[0073] Step 503: Determine the medical image analysis results of the object to be analyzed based on the probability distribution.
[0074] Specifically, when finalizing the medical image analysis results, the system uses the "maximum probability principle" to determine the sample's category label based on the probability distribution calculated in step 502. Specifically, the system selects the category with the highest probability as the final diagnosis. For example, for a sample with the probability distribution [0.05, 0.35, 0.60], the system would classify it as Alzheimer's disease (AD) because the third category has the highest probability. In addition to the category label, the system also outputs the complete probability distribution as a confidence index, providing more granular information for clinical decision-making. This is particularly important for cases at the disease boundary or with uncertainty. When the probabilities of two categories are close (e.g., the difference is less than 0.1), the system marks it as "requiring further investigation" to remind clinicians to conduct a more detailed evaluation.
[0075] Please see Figure 4 This document presents an overall architecture diagram for medical image analysis based on a multi-modal fusion network, as provided in this application embodiment. First, it processes the input sMRI and PET images through two parallel feature extraction paths. The top path uses a spatial attention module composed of 3D convolutional blocks to process the sMRI image, generating feature Fg; the bottom path applies a channel attention module to process the PET image, generating feature Fp. This dual-path design effectively captures the unique information characteristics of different modalities. The core of the architecture lies in the two key fusion modules in the middle: Logical Guided Attention (LGA) module (light blue area at the top): This module receives sMRI features and enhances feature representation through multi-branch processing. The upper branch uses a fully connected layer (FC), logits operation, and max function to generate attention weights; the middle branch generates a Top-K attention map through global average pooling (GAP), MLP, and sigmoid activation function; the lower branch generates a gating signal α through MLP and sigmoid function. The outputs of these branches are processed by addition (⊕) and multiplication (⊗) operations, and after weighted pooling, the enhanced feature F'_g is generated.
[0076] Similarity-Guided Gated Fusion (SGGF) module (pink area below): This module integrates features F'_g and F'_p from two modalities. First, it generates projected features through a Linear-LayerNorm-ReLU network. and Then, based on the similarity matrix, cross-modal attention interaction is achieved to generate enhanced features and Finally, a comprehensive feature F is generated through dual-path fusion (element-wise multiplication and element-wise addition after MLP transformation). fused .
[0077] The rightmost module is the classifier module, which receives fused features and outputs a prediction of whether a sample belongs to Alzheimer's disease (AD), mild cognitive impairment (MCI), or normal cognition (CN). This hierarchical multimodal fusion architecture fully utilizes the complementary information of different image modalities and effectively enhances the feature representation of disease-related regions through attention mechanisms and similarity guidance.
[0078] As an optional embodiment, after the step of outputting the medical image analysis results of the object to be analyzed, the following steps may also be included: Step 601: Output the first verification result based on the first discriminant feature, and output the second verification result based on the second discriminant feature.
[0079] Specifically, to enhance the reliability of diagnostic results, a multimodal cross-validation mechanism is introduced. After generating the main medical image analysis results (based on fusion features), the system also performs independent validation using the discriminant features of each single modality. Specifically, the first discriminant feature from the sMRI modality is first input into a dedicated validation classifier. This classifier has a similar structure to the main classifier, including global average pooling, fully connected layers, and a softmax activation function. Through this validation classifier, the system outputs a diagnostic result based on the sMRI single modality, i.e., the first validation result. For example, a patient's first validation result might be "mild cognitive impairment (MCI)" with a confidence level of 65%. Similarly, the second discriminant feature from the PET modality is input into another validation classifier, outputting a diagnostic result based on the PET single modality, i.e., the second validation result. For example, the same patient's second validation result might be "Alzheimer's disease (AD)" with a confidence level of 58%. These validation classifiers are trained synchronously with the main classifier during the training phase, but only use their respective single-modality features, thus ensuring that they can fully utilize the unique diagnostic value of each modality. This step yielded three independent but related diagnostic results: the primary diagnosis (based on fusion features), the first validation result (based on sMRI features), and the second validation result (based on PET features).
[0080] Step 602: Calculate the consistency score between the medical image analysis results and the first and second verification results.
[0081] Specifically, after obtaining three diagnostic results, the system assesses their degree of consistency and generates a quantitative consistency score. The consistency score reflects the degree of synergy between different modalities and fusion models in diagnostic decision-making and is an important indicator of diagnostic reliability. There are various methods for calculating the consistency score; in this embodiment, a weighted voting mechanism is used. First, the system checks whether the category labels of the three diagnostic results are consistent. If all three results are completely consistent (e.g., all are AD), the consistency score is set to 1.0 (highest). If two results are consistent but the third is different, a weighted consistency score is calculated based on the confidence levels of the consistent and inconsistent results. If all three results are inconsistent, the consistency score will be low.
[0082] For example, if the primary diagnosis is AD (70% confidence), the first validation result is MCI (65% confidence), and the second validation result is AD (58% confidence), the system will calculate a moderate level of consistency score, approximately 0.65.
[0083] Step 603: When the consistency score is lower than the preset threshold, mark the medical image analysis result as low confidence and output a prompt message that manual review is required.
[0084] Specifically, based on the calculated consistency score, the system evaluates the reliability of the diagnostic results. In this embodiment, a preset threshold (usually 0.6 or 0.7) is set to distinguish between high-reliability and low-reliability diagnostic results. When the consistency score is higher than the preset threshold, it indicates that different modalities and fusion models have reached a high degree of consensus on diagnosis, and the system will mark the diagnostic result as "high-reliability".
[0085] Conversely, when the consistency score is below a preset threshold, it indicates a diagnostic discrepancy between different modalities or models. The system will mark the diagnosis as "low confidence" and generate a clear prompt message to remind the doctor to conduct a manual review. For example, the system may output: "Diagnosis: Mild cognitive impairment (MCI), Confidence: Low, Consistency score: 0.45, Recommendation: Further evaluation by a professional physician is required."
[0086] This application also provides a computer storage medium that can store multiple instructions. The instructions are adapted to be loaded and executed by a processor. The above-described medical image analysis method based on a multi-mode fusion network is described in detail below. The specific execution process can be found in the detailed description of the above-described embodiments, which will not be repeated here.
[0087] The following describes an electronic device for medical image analysis based on a multi-mode fusion network, provided by an embodiment of this application. Figure 5 This is a schematic diagram of an exemplary hardware structure of an electronic device provided in an embodiment of this application.
[0088] In some embodiments, the electronic device for medical image analysis based on a multi-mode fusion network is a computer device or includes a computer device. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data. The network interface of the computer device is used to communicate with other external terminals or servers via a network connection. In some embodiments, the network interface can be a wired network interface; in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, it implements the methods in the embodiments of this application.
[0089] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0090] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0091] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0092] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A medical image analysis method based on a multi-modal fusion network, characterized in that, The method includes: Acquire a first image of the object to be analyzed in structural magnetic resonance imaging and a second image in positron emission tomography (PET); Feature extraction is performed on the first image and the second image respectively to obtain a first feature map and a second feature map, and the first feature map and the second feature map are flattened along the spatial dimension into a first sequence and a second sequence respectively; The first sequence and the second sequence are subjected to attention weighting processing to obtain a first discriminant feature and a second discriminant feature, including: calculating a basic attention signal and a classifier-guided attention signal for the first sequence; adaptively determining a first selection ratio based on the global features of the first sequence; selectively retaining the classifier-guided attention signal based on the first selection ratio; fusing the basic attention signal and the retained classifier-guided attention signal to obtain a first attention weight; weighting the first sequence using the first attention weight to obtain the first discriminant feature; calculating a basic attention signal and a classifier-guided attention signal for the second sequence; adaptively determining a second selection ratio based on the global features of the second sequence; selectively retaining the classifier-guided attention signal based on the second selection ratio; fusing the basic attention signal and the retained classifier-guided attention signal to obtain a second attention weight; and weighting the second sequence using the second attention weight to obtain the second discriminant feature. Cross-modal fusion is performed on the first discriminant feature and the second discriminant feature to obtain the fused feature; The fused features are input into the classifier, which outputs the medical image analysis results of the object to be analyzed.
2. The medical image analysis method based on a multi-mode fusion network according to claim 1, characterized in that, Feature extraction is performed on the first image and the second image respectively to obtain a first feature map and a second feature map, including: The first image and the second image are respectively input into the backbone of the three-dimensional residual network for convolution processing to obtain the first initial feature map and the second initial feature map. Calculate the attention weights of the first initial feature map in the spatial dimension, and perform weighted processing on the first initial feature map to obtain the first feature map; Calculate the attention weights of the second initial feature map in the channel dimension, and perform weighted processing on the second initial feature map to obtain the second feature map.
3. The medical image analysis method based on a multi-mode fusion network according to claim 1, characterized in that, Cross-modal fusion is performed on the first discriminant feature and the second discriminant feature to obtain fused features, including: The first discriminant feature and the second discriminant feature are respectively input into a multilayer perceptron for projection to obtain a first projection feature and a second projection feature; Multi-head attention calculation is performed on the first projection feature and the second projection feature to obtain the first enhanced feature and the second enhanced feature; Global average pooling is performed on the first enhanced feature and the second enhanced feature respectively to obtain the first global statistical information and the second global statistical information; Calculate similarity weights based on the first global statistical information and the second global statistical information; The first enhanced feature and the second enhanced feature are weighted and summed according to the similarity weight to obtain the fused feature.
4. The medical image analysis method based on a multi-mode fusion network according to claim 3, characterized in that, Multi-head attention computation is performed on the first projection feature and the second projection feature to obtain the first enhanced feature and the second enhanced feature, including: Using the first projection feature as the query vector and the second projection feature as the key vector and value vector, perform the first multi-head attention calculation to obtain the first intermediate feature; Using the second projection feature as the query vector and the first projection feature as the key vector and value vector, a second multi-head attention calculation is performed to obtain the second intermediate feature. The first intermediate feature and the first projected feature are residually connected to obtain the first enhanced feature, and the second intermediate feature and the second projected feature are residually connected to obtain the second enhanced feature.
5. The medical image analysis method based on a multi-mode fusion network according to claim 1, characterized in that, The fused features are input into a classifier, which outputs the medical image analysis results of the object to be analyzed, including: The fused features are subjected to global average pooling to obtain a global feature vector; The global feature vector is input into a fully connected layer for classification mapping, and the probability distribution of each category is calculated using the softmax activation function. The medical image analysis results of the object to be analyzed are determined based on the probability distribution.
6. The medical image analysis method based on a multi-mode fusion network according to claim 1, characterized in that, After outputting the medical image analysis results of the object to be analyzed, the method further includes: A first verification result is output based on the first discriminant feature, and a second verification result is output based on the second discriminant feature; Calculate the consistency score between the medical image analysis results and the first verification result and the second verification result; When the consistency score is lower than a preset threshold, the medical image analysis result is marked as low confidence and a prompt message indicating that manual review is required is output.
7. An electronic device for medical image analysis based on a multi-mode fusion network, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-6.
8. A computer program product containing instructions, characterized in that, When the computer program product is run on an electronic device for medical image analysis based on a multi-modal fusion network, the electronic device performs the method as described in any one of claims 1-6.
9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device for medical image analysis based on a multimodal fusion network, the electronic device causes the electronic device to perform the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Image classification method and device, equipment and storage medium
CN115423754A
Multi-modal medical image fusion and disease prediction method, computer program and terminal
CN118967480A