Incomplete MRI image segmentation method based on feature completion and feature fusion

The multimodal segmentation model with feature optimization solves the problem of MRI image mode loss, achieving a more efficient and robust image segmentation effect, which is suitable for clinical diagnosis of incomplete MRI images.

CN120339611APending Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510398434.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When processing incomplete MRI images, the prior art has problems such as degradation in segmentation performance, high model complexity and poor robustness caused by modal loss. In particular, the acquisition of Flair and T1ce modal images is difficult to obtain and the generated image feature information is unreal, making it difficult to use for clinical diagnosis.

Method used

The pre-trained multimodal segmentation model based on feature optimization is adopted, including encoder, multi-scale feature optimization module, modal Transformer module and multi-modal feature adaptive fusion module. Through feature completion and fusion, the single-modal effect is enhanced, the model complexity and calculation cost are reduced, and the robustness is improved.

Benefits of technology

Through feature completion and fusion, the model's adaptability to modal loss is enhanced, the accuracy and robustness of MRI image segmentation are improved, the complexity and calculation cost of the model are reduced, and the segmentation effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339611A_ABST
    Figure CN120339611A_ABST
Patent Text Reader

Abstract

The invention relates to the field of medical image segmentation, in particular to an incomplete MRI image segmentation method based on feature completion and feature fusion, and the method comprises the steps: obtaining a to-be-processed incomplete medical image, processing the to-be-processed incomplete medical image through a pre-trained feature optimization-based multi-modal segmentation model, and obtaining a segmentation result; the multi-modal segmentation model based on feature optimization comprises an encoder, a multi-scale feature optimization module, an intra-modal Transform module, a multi-modal feature adaptive fusion module and a decoder; according to the method, the redundancy is reduced and the single-mode effect is enhanced by retaining the key information of each mode and fusing the high-quality features, and meanwhile, the mode missing is simulated by means of the mode mask, so that the segmentation precision is improved, the model cost and complexity are reduced, and the segmentation effect and robustness are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image segmentation, and particularly relates to an incomplete MRI image segmentation method based on feature completion and feature fusion. Background Art

[0002] In actual clinical practice, due to limitations such as equipment and manpower, modality images are often missing in brain tumor medical image MRI (Magnetc Resonance Imagng). Existing incomplete medical image segmentation technologies mainly rely on exhaustive methods or generative networks based on GAN and GNN.

[0003] In the existing technical solutions in the field of multi-modal feature research, there is a problem that the context connection is not tight. The methods of generating corresponding networks for specific modalities and based on feature complementarity and feature generation both have certain defects. The former has high costs and time consumption and the model does not have robustness. The latter does not consider that different modalities have different focuses on features and the generated image feature information is not real, making it difficult to be used in real clinical diagnosis. Summary of the Invention

[0004] In view of this, the present invention discloses an incomplete MRI image segmentation method based on feature completion and feature fusion to solve the above problems, including: obtaining an incomplete medical image to be processed, and processing the incomplete medical image to be processed by a pre-trained multi-modal segmentation model based on feature optimization to obtain a segmentation result;

[0005] Further, the multi-modal segmentation model based on feature optimization includes: an encoder, configured to obtain an incomplete medical image to be processed, extract features of the incomplete medical image through downsampling to obtain shallow features; a multi-size feature optimization module, configured to process the shallow features to obtain an enhanced modality; a Transformer module within the modality, configured to process the enhanced modality to obtain a single-modal enhanced feature; a multi-modal feature adaptive fusion module, configured to fuse the single-modal enhanced features to obtain a multi-modal fusion feature; and a decoder, configured to gradually restore the multi-modal fusion feature to the size of the incomplete medical image to be processed to generate a segmentation result.

[0006] The beneficial effects of the present invention include:

[0007] By strengthening the feature fusion between multiple modalities, the effect of a single modality is enhanced, thereby weakening the impact of missing modality information on the model segmentation performance, and further reducing the model complexity and computational cost; by constructing a multi-size feature optimization module and designing a feature library to solve the problem of missing Flair and T1ce modality images; by constructing a multi-modal feature adaptive fusion module to simulate various modality missing situations for the input incomplete MRI images, so that the model proposed by the present invention can handle all modality incomplete situations, and improve the segmentation effect and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 It is a schematic structural diagram of a multi-modal segmentation model based on feature optimization in an embodiment of the present invention;

[0009] Figure 2 It is the segmentation result of four-modal brain cross-sections of the same case in an embodiment of the present invention;

[0010] Figure 3 It is a schematic diagram of the segmentation effect of different MRI sequence combinations on the tumor area in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] In order to make the objectives, technical solutions, features and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0012] This embodiment includes a method for segmenting incomplete MRI images based on feature completion and feature fusion, including: obtaining the incomplete medical images to be processed, and processing the incomplete medical images to be processed by using a pre-trained multi-modal segmentation model based on feature optimization to obtain a segmentation result.

[0013] Further, as Figure 1 shown, the multi-modal segmentation model based on feature optimization includes: an encoder, a multi-size feature optimization module, an intra-modal Transformer module, a multi-modal feature adaptive fusion module, and a decoder.

[0014] The processing of the encoder for data includes: obtaining the incomplete medical images to be processed, performing deep encoding on the incomplete medical images, and extracting the features of the incomplete medical images through downsampling operations to obtain shallow features.

[0015] Specifically, the incomplete medical images include T1, T2, flair, and T1ce modalities, which are MRI images of size (128, 128, 128) with 1 channel. Downsampling is performed through 3D convolution, InstanceNorm3d normalization, and LeakyReLU activation function operations to sequentially generate features of sizes (128, 128, 128), (64, 64, 64), (32, 32, 32), (16, 16, 16), and (8, 8, 8), with the number of channels being 8, 16, 32, 64, and 128 respectively. This is equivalent to reducing the size of the features by half and doubling the number of channels with each downsampling. The 3D convolution has a size of (3, 3, 3) and a stride of (2, 2, 2).

[0016] The multi-scale feature optimization module (MSFO, Multi-Scale Feature Optimization) is used to process shallow features to obtain enhanced modalities.

[0017] Specifically, MSFO contains four branches, corresponding to the encoder T1, T2, flair, and T1ce modalities respectively. In each branch, atrous spatial pyramid pooling (ASPP) operations are performed on the shallow features of different modalities to capture context information of different sizes in the image, obtaining context features of multiple sizes. Convolution and upsampling operations are performed on the context features of multiple sizes to further enrich the information of the features. Finally, a 1×1×1 convolution is used to fuse the features of different sizes into one feature, obtaining the fused features after multi-scale convolution of different modalities.

[0018] Among them, the ASPP operation includes: passing the shallow features through 4 convolutions and a global average pooling to obtain multi-scale features. The 4 convolutions are a regular convolution of (1, 1, 1) and three dilated convolutions of (3, 3, 3) with dilation rates of 6, 12, and 18 respectively. Specifically, in the face of MRI image segmentation tasks, the receptive field of ordinary convolution operations is too small. To solve this problem, the present invention introduces dilated convolution, also known as atrous convolution. Dilated convolution has a certain interval between internal elements, that is, holes, and the dilation rate r is usually used to represent the size of the interval. For example, when the convolution kernel is (3, 3, 3) and r is 2, it is equivalent to a regular convolution kernel with an actual coverage area of (5, 5, 5), which can be understood as having a blank space between each position and filling a value of 0. Dilated convolution solves the problem of the small receptive field of ordinary convolution without increasing the number of parameters. The formula for dilated convolution is as follows:

[0019]

[0020] Among them, x[i+r·m,j+r·n,k+r·p]·w[m,n,p] represents the input feature map, w[m,n,p] represents the convolution kernel, r represents the dilation rate, and y[i,j,k] represents the output feature map of the dilated convolution.

[0021] Furthermore, in actual clinical applications, due to limitations of factors such as equipment and manpower, it is more difficult to obtain the Flair and T1ce modalities, and it is easier to have the situation of missing modality images. The present invention solves this problem by designing a feature bank Feature Bank.

[0022] Specifically, after obtaining the fused features, the fused features are stored in the feature bank, and more excellent features are selected through the feature bank, and the Flair and T1ce modalities are complemented through the feature bank to obtain enhanced modalities.

[0023] The complementing includes: performing interpolate interpolation on the target image to make its size the same as the obtained T1 and T2 modality features, calculating the cross entropy loss (CrossEntropyLoss) as the selection criterion for excellent features, and converting the current features into a matrix and recording it for subsequent replacement of the feature bank.

[0024] Among them, the feature bank stores the first 100 minimum loss features in each round and stores them in the feature bank. The size of the feature bank is set to 16000. When it is not full, new features are default directly added. When it exceeds the set value, priority replacement is performed. In order to clean the storage pool, a cleaning operation is set to be executed every 50 rounds to optimize the storage quality. Every 5 rounds, the clustering centers are calculated and updated according to the features stored in the feature bank. Since the features corresponding to each clustering center are always similar, this step is equivalent to concentrating similar features, so as to use these screened and updated T1 modality and T2 modality features to complement the Flair modality features and T1ce modality features.

[0025] The processing of data by the Transformer module within the modality includes: establishing global associations based on the enhanced modalities of different branches, and dynamically allocating weights through the self-attention mechanism of the Transformer framework to capture important information in the enhanced modalities and obtain single-modal enhanced features.

[0026] The multi-modal feature adaptive fusion module (MFAF, Multi-modal Feature Adaptive Fusion) is used to fuse the single-modal enhanced features to obtain multi-modal fusion features.

[0027] Specifically, in order to deal with the negative impact of modality loss, the present invention proposes MFAF to establish connections between different modalities and integrate key information of each modality; MFAF processes data including:

[0028] Step 1: Transform the feature size of the single-modal enhanced features to obtain a merged tensor.

[0029] Specifically, through the three-dimensional convolution operation, the channel dimension of the single-modal enhanced feature is converted from 512 to 128, and features of different modes with a size of (1, 128, 8, 8, 8) are obtained; the number dimension of the modality is added to the four modes respectively, and the features are merged to obtain the merged tensor x.

[0030] Step 2: Add a mask to the merged tensor to obtain the effective modal features.

[0031] Specifically, a modality mask is added to the merged tensor of each modality to simulate whether it is missing. The modality mask [λ1,λ2,λ3,λ4] is used to simulate 15 different modality input situations of the four modalities of MRI data, where λ is 1 or 0, 1 indicates the modality exists, and 0 indicates the modality is missing; Boolean judgment and dimension expansion operations are performed on the modality mask to obtain a valid mask of size and shape (1,4,128,8,8,8), which is convenient for subsequent modality-by-modality feature processing. Multiply the merged tensor x by the valid mask mask to obtain the effective modality feature F orig .

[0032] Step 3: Perform multimodal feature fusion operation on the effective modal features to obtain enhanced features.

[0033] Sum the effective modal features in the modal dimension to obtain the total effective modal features F of size (1,128,8,8,8) modal The modal average fusion is used to reduce the impact of the missing modality on the segmentation effect; the average fusion modality of all valid modal features is calculated. The formula is:

[0034]

[0035] Among them, N modal Indicates the number of valid modes.

[0036] The average fusion mode is used to enhance all effective modal features, and all effective modal features are spliced in the channel dimension. A 1×1×1 three-dimensional convolution is further performed channel by channel without changing the spatial size to obtain the enhanced feature F enh .

[0037] Step 4: Generate attention weights based on the enhanced features, and weight the effective modal features with the attention weights to obtain weighted features.

[0038] Perform fully connected and activation function processing on F enh to obtain the attention map for each type of effective modality. The activation function used is ReLU. Due to modality missing, the model has poor generalization ability when facing different modality input situations. Normalize the attention map of the effective modality through the softmax operation. Considering the influence of the invalid modality on the finally generated weights, further add a mask to set the invalid weights to 0 to obtain the effective feature weights A. The formula is:

[0039] A = σ(W2 * ReLU(W1 * F enh ))

[0040] where W1 and W2 represent two different 1×1×1 three-dimensional convolutions, r represents the reduction ratio, and σ(·) represents Softmax normalization. Generate attention weights according to the effective feature weights. The formula is:

[0041]

[0042] where M represents the modality mask. When M is 1, it indicates that the marked modality is effective; when M is 0, it indicates that the marked modality is invalid. ∈ is a sufficiently small positive number, and the ∈ term is used to prevent the denominator from being 0.

[0043] Furthermore, weight the effective modal features with the attention weights to obtain the weighted feature F attn . The weighting formula is:

[0044]

[0045] Step 5: Concatenate the weighted features and the enhanced features, and perform feature fusion through the conv_interaction convolution operation to obtain the output feature F out of the multi-modal feature adaptive fusion module. The formula is:

[0046] F out = W3 * [F attn || F enh

[0047] where || represents the concatenation operation in the channel dimension, W3 represents a 1×1×1 three-dimensional convolution for the final feature fusion, and F out represents the final output feature of MFAF, that is, the multi-modal fusion feature.

[0048] The decoder is used to gradually restore the multi-modal fusion feature to the size of the incomplete medical image to be processed and generate the segmentation result. ​

[0049] Furthermore, the training process of the present invention is based on deep supervision learning. Deep supervision learning includes: recording the features of each downsampling of the modality in the encoder, providing an additional supervision signal for the intermediate levels of the multi-modal segmentation model optimized based on features, and calculating the loss function of the intermediate layer. Deep supervision learning and the entire segmentation model are combined to form a multi-task training mode, and each modality passes through the decoder separately for multi-task learning.

[0050] In the deep supervision learning stage, Softmax weighted cross-entropy loss and Dice loss are adopted to improve the overall training effect of the network. The loss function formula in the deep supervision stage is:

[0051]

[0052] Among them, represents the Softmax weighted cross-entropy loss in the deep supervision process, represents the Dice loss in the deep supervision process. The formula for the Softmax weighted cross-entropy loss is:

[0053]

[0054] Among them, C represents the number of categories of the image; w c represents the weight of category c, which is used to balance the class imbalance problem, y c represents the true label, in the form of one-hot encoding, p c represents the predicted softmax probability, and log(p c ) represents the log probability of the predicted category. The formula for the Dice loss is:

[0055]

[0056] Among them, DSC represents the Dice coefficient, p represents the predicted segmentation result of the model, y represents the true target region, in the form of one-hot encoding; ∑(py) represents the intersection of the predicted value and the true value, which is the sum of the intersecting pixel points; ∑p 2 and ∑y 2 respectively represent the region sizes of the predicted value and the true value; ∈ is a decimal used to prevent the denominator from being zero, preferably 10 -6 .

[0057] Calculate the loss between the features optimized by the encoder and the MSFO module and the transformed target. The loss function of the feature library is as follows:

[0058]

[0059] The loss function in the image segmentation loss stage is:

[0060]

[0061] Among them, represents the loss of the final input result, that is, the final output of the decoder; represents the segmentation situation of each layer in the final decoder stage.

[0062] Furthermore, the overall loss function L Total has the following formula:

[0063] L Total = L Deep + L bank + L Seg

[0064] Furthermore, in this embodiment, the BraTS2020 dataset is used to test the multi-modal segmentation model based on feature optimization. Figure 2 shows the visualization results of the four input modalities of this model under the same case, showing the visualization of the brain tissues in the cross-sectional slices of the brain of four different MRI sequences of the same patient, Figure 3 shows the segmentation effect of different MRI sequence combinations on the tumor region when using the BraTS2020 dataset for testing, and compares it with the Ground Truth (true label). Among them, different colors represent different tissue regions. Red (Enhancing Tumor, ET) represents the tumor enhancement region; blue (Tum or Core, TC) represents the tumor core region; green (Whole Tumor, WT) represents the whole tumor region. The black region represents the background and is not the object of concern of the segmentation model.

[0065] Table 1. Quantitative comparison results of DSC on the BraTS2020 dataset

[0066]

[0067]

[0068] Table 1 shows the comparison test results of the multi-modal segmentation model based on feature optimization and the commonly used MMformer model (baseline) in the prior art under fifteen different multi-modal missing combinations of the BraTS2020 dataset. ● indicates the presence of the modality, and ○ indicates the absence of the modality. Using the Dice coefficient as the evaluation index, the segmentation situations of different regions of the brain tumor: the whole tumor WT, the tumor core TC, and the enhanced tumor ET are represented numerically. It can be seen that under various different modality missing situations, the multi-modal segmentation model based on feature optimization proposed by the present invention has better performance.

[0069] Finally, it should be noted that the above only describes some embodiments of the present invention. For those skilled in the art, various changes, modifications, substitutions, and deformations can be conceived without departing from the principle and spirit of the present invention. The protection scope of the present invention is defined by the appended claims and their equivalents, and the above actions should all be covered within the protection scope of the present invention.

Claims

1. An incomplete MRI image segmentation method based on feature completion and feature fusion, characterized in that, Including: Obtain an incomplete medical image to be processed; Process the incomplete medical image to be processed using a pre-trained multi-modal segmentation model based on feature optimization; Obtain a segmentation result; The multi-modal segmentation model based on feature optimization includes: an encoder, which is used to obtain the incomplete medical image to be processed, extract the features of the incomplete medical image through downsampling, and obtain shallow features; a multi-scale feature optimization module, which is used to process the shallow features to obtain an enhanced modality; a Transformer module within the modality, which is used to process the enhanced modality to obtain a single-modal enhanced feature; a multi-modal feature adaptive fusion module, which is used to fuse the single-modal enhanced features to obtain a multi-modal fusion feature; a decoder, which is used to gradually restore the multi-modal fusion feature to the size of the incomplete medical image to be processed and generate a segmentation result.

2. The incomplete MRI image segmentation method based on feature completion and feature fusion according to claim 1, wherein The downsampling includes: successively performing three-dimensional convolution, InstanceNorm3d normalization, and LeakyReLU activation processing on the incomplete medical image to obtain shallow features of different modalities; among them, the incomplete medical image includes one or several of MRI images of T1, T2, flair, and T1ce modalities, the size of the three-dimensional convolution is (3, 3, 3), and the stride is (2, 2, 2).

3. The incomplete MRI image segmentation method based on feature completion and feature fusion according to claim 1, wherein The processing of the multi-scale feature optimization module for data includes: performing atrous spatial pyramid pooling operation on the shallow features of different modalities to obtain context features of multiple sizes; performing convolution and upsampling on the context features of multiple sizes, and using 1×1×1 convolution to fuse the features of different sizes to obtain a multi-modal fusion feature; storing the fusion feature in a feature library, and complementing the missing modality to obtain an enhanced modality.

4. The incomplete MRI image segmentation method based on feature completion and feature fusion according to claim 3, characterized in that, The complementing includes: performing interpolate interpolation on the image of the modality to be complemented to make its size the same as the modality feature that does not need to be complemented, and converting the current feature into a matrix.

5. The incomplete MRI image segmentation method based on feature completion and feature fusion according to claim 3, characterized in that The feature library stores the first 100 minimum cross-entropy loss features in each round of propagation. The size of the feature library is 16000. When it is not full, new features are directly added. When a feature with a smaller cross-entropy loss appears, priority replacement is performed; a cleaning operation is performed on the feature library every 50 propagation rounds, and the clustering center is updated every 5 propagation rounds.

6. The incomplete MRI image segmentation method based on feature completion and feature fusion according to claim 1, wherein The processing of the multi-modal feature adaptive fusion module for data includes: Step 1: Perform feature size transformation on the single-modal enhanced feature to obtain a merged tensor; Step 2: Add a mask to the merged tensor to obtain an effective modality feature; Step 3: Perform multi-modal feature fusion operation on the effective modality feature to obtain a strengthened feature; Step 4: Generate attention weights according to the strengthened feature, and use the attention weights to weight the effective modality feature to obtain a weighted feature; Step 5: Concatenate the weighted feature and the strengthened feature, and perform feature fusion through convolution operation to obtain a multi-modal fusion feature.

7. The method for incomplete MRI image segmentation based on feature completion and feature fusion according to claim 6, wherein Adding a mask for the merged tensor includes: using a modality mask [λ1, λ2, λ3, λ4] to simulate 15 different modality input cases existing in the four modalities of MRI data, where λ is 1 or 0, 1 indicates the presence of a modality, and 0 indicates the absence of a modality; performing a boolean judgment and a dimension expansion operation on the modality mask to obtain an effective mask; multiplying the merged tensor by the effective mask to obtain effective modality features.

8. The incomplete MRI image segmentation method based on feature completion and feature fusion according to claim 1, characterized in that The pre-training process uses deep supervised learning.

Citation Information

Cited By

  • Asset information intelligent completion method and system fused with multi-modal large model

    CN120747981A