Brain tumor segmentation method based on multi-modal fusion

By adding a multi-scale progressive fusion module to the Unet network to fuse the cross-modal correlation features in multimodal data, the problem of insufficient utilization of multimodal features in the prior art is solved, and more efficient brain tumor image segmentation performance is achieved.

CN120013954APending Publication Date: 2025-05-16HENAN UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411987915.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing brain tumor segmentation methods lack the utilization of multimodal features and lack the modeling of complex relationships between multimodal information, which limits the improvement of segmentation performance.

Method used

Add a multi-scale progressive fusion module to the Unet network structure to fuse complex cross-modal correlation features and complementary clues in multimodal data to enhance the segmentation performance of the model.

Benefits of technology

It improves the utilization rate of multimodal feature information, enhances the model's understanding of complex structures, and improves the performance of brain tumor image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013954A_ABST
    Figure CN120013954A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image segmentation, and provides a brain tumor segmentation method based on multi-modal fusion. The method comprises the following steps: 1, obtaining a multi-modal brain tumor image, and preprocessing the multi-modal brain tumor image to construct a multi-modal brain tumor data set; step 2, constructing a brain tumor segmentation network structure, including: selecting a Unet network as a reference model, and adding a multi-scale progressive fusion MFCM module at a jump joint of each coding layer and each decoding layer in the Unet network; 3, training the brain tumor segmentation network structure by using the multi-modal brain tumor data set to obtain a brain tumor segmentation model; and 4, preprocessing a to-be-detected multi-modal brain tumor image, and inputting the preprocessed to-be-detected multi-modal brain tumor image into the brain tumor segmentation model to obtain a segmentation result. According to the method, the segmentation performance of the model can be improved while the multi-modal features are fully utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular to a brain tumor segmentation method based on multimodal fusion. Background Art

[0002] Brain tumors are diseases that seriously threaten human health, and early and accurate diagnosis and treatment are crucial to improving survival rates. Magnetic resonance imaging (MRI), as a non-invasive technology, is widely used in the detection and diagnosis of brain tumors. MRI images of different modalities, such as T1-weighted, T2-weighted, FLAIR, and enhanced T1-weighted images, can provide diverse tumor information and help to fully describe its morphology. However, a single modality is often difficult to accurately segment the tumor area, especially in cases of complex morphology and fuzzy boundaries. To overcome this challenge, multimodal fusion technology has emerged, which can effectively improve the accuracy and robustness of segmentation.

[0003] In recent years, brain tumor segmentation methods based on deep convolutional neural networks have developed rapidly. The U-Net architecture and its variants are widely used in this field. U-Net realizes feature extraction and image reconstruction through the symmetrical structure of the encoder and decoder, and transmits high-resolution local features in the jump connection to improve the fine-grained accuracy of segmentation. U-Net++ uses dense short connection design to make the information exchange between the encoder and decoder layers closer, alleviating the semantic gap problem. With the rise of Transformer, some studies introduced it into the tumor segmentation model, using its global modeling ability to capture the deep spatial dependencies in the image, solving the shortcomings of the U-Net structure in capturing long-range dependencies. Existing brain tumor segmentation methods have the problem of insufficient utilization of multimodal features and lack of modeling of complex relationships between multimodal information, which limits the improvement of segmentation performance. Summary of the invention

[0004] In order to solve the problems of insufficient utilization of multimodal features and lack of modeling of complex relationships between multimodal information in existing brain tumor segmentation methods, the present invention provides a brain tumor segmentation method based on multimodal fusion. By adding a multi-scale progressive fusion module to the Unet network structure, the complex cross-modal correlation features and complementary clues in multimodal data can be fused, and the multimodal features can be fully utilized while improving the segmentation performance of the model.

[0005] The present invention provides a brain tumor segmentation method based on multimodal fusion, comprising:

[0006] Step 1: obtaining a multimodal brain tumor image, and preprocessing the multimodal brain tumor image to construct a multimodal brain tumor dataset;

[0007] Step 2: constructing a brain tumor segmentation network structure, including: selecting a Unet network as a benchmark model, and adding a multi-scale progressive fusion MFCM module at the jump connection between each encoding layer and decoding layer in the Unet network; wherein the Unet network includes four independent encoders and one decoder, and the multi-scale progressive fusion MFCM module is used to fuse the multi-modal features extracted by the independent encoders;

[0008] Step 3: Using the multimodal brain tumor dataset to train the brain tumor segmentation network structure to obtain a brain tumor segmentation model;

[0009] Step 4: After preprocessing, the multimodal brain tumor image to be detected is input into the brain tumor segmentation model to obtain a segmentation result.

[0010] Furthermore, the multimodal brain tumor image includes T1, T1 ce , Flair and T2 modes.

[0011] Furthermore, the preprocessing includes converting the three-dimensional MRI image into two-dimensional slice data, converting the format of the annotation file, and enhancing normalization.

[0012] Furthermore, the multi-scale progressive fusion MFCM module includes T1-T1 ce Fusion branch, Flair-T2 fusion branch and spatial channel fusion module;

[0013] Correspondingly, the multi-scale progressive fusion MFCM module is used to fuse the multimodal features extracted by the independent encoders, including:

[0014] Determine that the features of T1, T1ce, Flair and T2 output by independent encoders are M1, M2, M3 and M4;

[0015] The T1-T1ce fusion branch fuses the features of M1 and M2 output by the encoder to obtain M 12 ;

[0016] The Flair-T2 fusion branch fuses the features of M3 and M4 output by the encoder to obtain M 21 ;

[0017] The spatial channel fusion module is used to combine M 12 and M 21 Perform cross-group feature fusion to obtain M fuse .

[0018] Furthermore, the T1-T1 ceThe fusion branch and the Flair-T2 fusion branch have the same structure, including: a first branch and a second branch; wherein the first branch includes two 3×3 convolution blocks, and the second branch includes two 5×5 convolution blocks, and each convolution block is connected to a splicing layer;

[0019] Correspondingly, the T1-T1 ce The specific operations of the fusion branch are as follows:

[0020] M1 and M2 are respectively sent to 1×1 convolution for dimensionality reduction to obtain f1 and f2;

[0021] After f1 and f2 are input into the first concatenation layer of the first branch for concatenation, they are sent into the first 3×3 convolutional block of the first branch to obtain f 12 At the same time, f1 and f2 are input into the first concatenation layer of the second branch for concatenation, and then sent into the first 5×5 convolution block of the second branch to obtain f 21 ;

[0022] f 12 and f 21 After the second concatenation layer of the first branch is input for concatenation, it is sent to the second 3×3 convolution block of the first branch to obtain f′ 12 At the same time, f 12 and f 21 After concatenation with the first concatenation layer of the second branch, it is fed into the second 5×5 convolutional block of the second branch to obtain f′ 21 ;

[0023] Use element-by-element multiplication to convert f′ 12 and f′ 21 After the combination, it is added to f1 and f2 through the residual connection, and the fused features are further integrated through 3×3 convolution to obtain M 12 ;

[0024] M3 and M4 are fused into M through the Flair-T2 branch 21 .

[0025] Furthermore, the spatial channel fusion module is used to combine M 12 and M 21 Perform cross-group feature fusion to obtain M fuse , specifically including: first, M 12 and M 21 After the splicing operation, it is sent to the 3×3 convolution block to achieve cross-group feature fusion, and then sent to the SCConv layer to generate M fuse .

[0026] Furthermore, the construction of the brain tumor segmentation network structure also includes: adding a first context semantics-guided CIBM I module after the multi-scale progressive fusion modules of the third and fourth layers of the brain tumor segmentation network structure, and adding a second context semantics-guided CIBMII module after the multi-scale progressive fusion modules of the first and second layers, respectively; wherein the CIBM I module and the CIBMII module are used to fuse the output of the multi-scale progressive fusion module and the previous decoding layer.

[0027] Further, the first contextual semantics guided CIBM I module includes a first spatial attention branch, a second spatial attention branch, a 3×3 convolutional block and a channel attention module;

[0028] Correspondingly, the specific operations of the CIBM I module are as follows:

[0029] Determine the output M of the multi-scale progressive fusion module fuse and the output H of the previous decoding layer as the input of the CIBM I module;

[0030] The first spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by multiplication to obtain f mul , then f mul Through the spatial attention module, we can obtain f′ mul ;

[0031] The second spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by addition to obtain f sum , then f sum Through the spatial attention module, we can obtain f′ sum ;

[0032] f′ mul and f′ sum After concatenation in the channel dimension, features are automatically selected through a 3×3 convolutional block to obtain f lh ;

[0033] f lh The feature weights are adaptively adjusted through the channel attention module to obtain F lh .

[0034] Furthermore, the second context semantics guided CIBMII module includes two spatial attention branches, a large core selection attention module, a 3×3 convolution block and a channel attention module;

[0035] Correspondingly, the specific operations of the CIBMII module are as follows:

[0036] Determine the output M of the multi-scale progressive fusion module fuse and the output H of the previous decoding layer as the input of the CIBMII module;

[0037] The first spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by multiplication to obtain f mul , then f mul Through the spatial attention module, we can obtain f′ mul ;

[0038] The second spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by addition to obtain f sum , then f sum Through the spatial attention module, we can obtain f′ sum ;

[0039] After the 1×1 convolution block, M fuse After being concatenated with H, it is input into the large core selection attention module to obtain f′ g ;

[0040] f′ mul , f′ g and f′ sum After concatenation in the channel dimension, features are automatically selected through a 3×3 convolutional block to obtain f lh ;

[0041] f lh The feature weights are adaptively adjusted through the channel attention module to obtain F lh .

[0042] Furthermore, the processing process of the large core selection attention module is expressed by the following formula:

[0043] f k1 =S conv5×5 (f g )

[0044]

[0045] f′ k1 =S conv1×1 (f k1 ), f′ k2 =S conv1×1 (f k2 )

[0046] f k = Cat[f′ k1 , f′k2 ]

[0047] f sig =σ(S conv7×7 (Cat[AvgPool(f k ), MaxPool(f k )]))

[0048]

[0049] Among them, f g Indicates M fuse and H are the concatenated outputs after passing through 1×1 convolution blocks, S conv5×5 represents a 5×5 convolutional block, represents 2D dilated convolution, with a kernel size of 7×7 and a dilation rate of 3; S conv1×1 represents a 1×1 convolutional block, Cat represents a concatenation operation, and S conv7×7 represents a 7×7 convolutional block, AvgPool represents average pooling, MaxPool represents maximum pooling, σ represents the sigmoid activation function, represents multiplication, Represents addition.

[0050] Beneficial effects of the present invention:

[0051] The present invention uses an improved multi-coding Unet network to segment lesions in brain tumor images. On the basis of the original network structure, a multi-scale progressive fusion module is constructed to extract multi-scale information while fusing cross-modal correlation features and complementary clues in multimodal data, thereby enhancing the model's feature representation capabilities when processing complex tumor structures. The utilization rate of multimodal feature information is improved, the purpose of enhancing the integrity of target information is achieved, and the model's ability to understand complex structures is enhanced. According to the comparison of experimental results, compared with other mainstream models, the present invention has better performance in brain tumor image segmentation and can effectively complete the task of clinical auxiliary diagnosis.

[0052] The present invention adds a contextual semantic guidance module, which enhances the model's ability to recover fine-grained details by fusing the characteristics of deep features with shallow features, effectively integrates rich spatial detail features with global semantic features, performs adaptive recalibration of features, reduces semantic conflicts and information redundancy that occur during the fusion process of features at different levels, and helps the decoder better recover the fine-grained features of the image.

[0053] The present invention adds a large kernel selective attention module to the partial context semantic guidance module. In the process of brain tumor image segmentation, if only local background information is considered, some normal brain tissues or complex vascular structures with similar density and morphology to tumors are often misclassified as tumors. By introducing large kernel selective attention and modeling multi-scale context information, this problem can be effectively solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A flowchart of a brain tumor segmentation method based on multimodal fusion provided by an embodiment of the present invention;

[0055] Figure 2 A schematic diagram of a brain tumor segmentation network structure provided by an embodiment of the present invention;

[0056] Figure 3 A schematic diagram of the structure of a multi-scale progressive fusion module provided in an embodiment of the present invention;

[0057] Figure 4 A schematic diagram of the structure of a first context semantic guidance module provided in an embodiment of the present invention;

[0058] Figure 5 A schematic diagram of the structure of a second context semantic guidance module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] In the field of brain tumor segmentation, although the existing deep learning-based segmentation methods have made significant progress in accuracy and generalization performance, there are still the following technical difficulties that affect the overall performance and practical application of the model.

[0061] (1) The first is the insufficient utilization of multimodal features. Current methods usually simply integrate features from different MRI modalities by splicing or adding them together, attempting to make up for the shortcomings of single modality information through the combination of multimodal features. However, this processing method lacks modeling of the complex relationship between multimodal information, cannot fully utilize the advantages of different modality images in describing brain tumor structures, and limits the improvement of segmentation performance.

[0062] (2) There may be semantic conflicts and information redundancy when fusing shallow and deep features. The skip connection mechanism in the U-shaped network architecture supplements the deep features with detailed information by directly passing the shallow features in the encoder to the decoder. However, shallow features often contain more noise or redundant information that is irrelevant to the segmentation task, and direct use may introduce errors. In addition, the significant semantic differences between shallow and deep features may cause conflicts during fusion, weakening the segmentation effect. Therefore, how to effectively handle the semantic differences between shallow and deep features while maintaining detailed information has become a major challenge in current research.

[0063] The existence of these problems indicates that the existing brain tumor segmentation methods still need to be further optimized, especially in the fusion of multimodal information and the processing of feature levels. More sophisticated feature integration and fusion strategies need to be designed to improve the accuracy and robustness of segmentation.

[0064] like Figure 1 As shown, an embodiment of the present invention provides a brain tumor segmentation method based on multimodal fusion, comprising:

[0065] Step 1: Obtain multimodal brain tumor images and preprocess the multimodal brain tumor images to construct a multimodal brain tumor dataset; the preprocessing includes converting the three-dimensional MRI images into two-dimensional slice data, converting the format of the annotation file, and enhancing normalization. In this embodiment, the multimodal brain tumor images include T1, T1 ce , Flair and T2 four different modalities, and resized the images in the dataset to 160×160.

[0066] Step 2: Construct a brain tumor segmentation network structure, such as Figure 2 As shown, it includes: selecting the Unet network as the benchmark model, adding a multi-scale progressive fusion MFCM module at the jump connection of each encoding layer and decoding layer in the Unet network; wherein the Unet network includes four independent encoders and one decoder, and the multi-scale progressive fusion MFCM module is used to fuse the multi-modal features extracted by the independent encoders.

[0067] Specifically, Figure 3 As shown, the multi-scale progressive fusion MFCM module includes T1-T1 ce Fusion branch, Flair-T2 fusion branch and spatial channel fusion module;

[0068] Correspondingly, the multi-scale progressive fusion MFCM module is used to fuse the multimodal features extracted by the independent encoders, including:

[0069] Determine that the features of T1, T1ce, Flair and T2 output by independent encoders are M1, M2, M3 and M4, multimodal features Corresponding to T1, T1ce, T2 and Flair sequences respectively, they are divided into two groups {M1, M2} and {M3, M4} for multi-scale feature extraction and cross-fusion:

[0070] The T1-T1ce fusion branch fuses the features of M1 and M2 output by the encoder to obtain M 12 ;

[0071] Specifically, the T1-T1ce fusion branch and the Flair-T2 fusion branch have the same structure, including: a first branch and a second branch; wherein the first branch includes two 3×3 convolution blocks, the second branch includes two 5×5 convolution blocks, each convolution block is connected with a first convolution layer, a first 3×3 convolution block, a second convolution layer and a second 3×3 convolution block, and the second branch includes a third convolution layer, a first 5×5 convolution block, a fourth convolution layer and a second 5×5 convolution block; the specific operations are as follows:

[0072] M1 and M2 are respectively sent to 1×1 convolution for dimensionality reduction to obtain f1 and f2:

[0073]

[0074] where f i ∈R H×W×C / 2 , i∈{1,2}.

[0075] After f1 and f2 are concatenated in the first concatenation layer of the first branch, they are fed into the first 3×3 convolutional block of the first branch to obtain f 12 At the same time, f1 and f2 are input into the first concatenation layer of the second branch for concatenation, and then sent to the first 5×5 convolution block of the second branch to obtain f 21 :

[0076]

[0077] f 12 and f 21 After concatenation with the second concatenation layer of the first branch, it is fed into the second 3×3 convolutional block of the first branch to obtain f′ 12 At the same time, f 12 and f 21 After concatenation, the first concatenation layer of the second branch is input and then sent to the second 5×5 convolution block of the second branch to obtain f′ 21 :

[0078]

[0079] It is understandable that f i Convolution processing at different scales can establish pixel-level information association between modalities.

[0080] Use element-by-element multiplication to convert f′ 12 and f′ 21 After the combination, it is added to f1 and f2 through the residual connection, and the fused features are further integrated through 3×3 convolution to obtain M 12 :

[0081]

[0082] The Flair-T2 fusion branch fuses the features of M3 and M4 output by the encoder to obtain M 21 .

[0083] Specifically, the structure of the Flair-T2 fusion branch is the same as that of the T1-T1ce fusion branch. M3 and M4 are obtained by the same operation of the Flair-T2 fusion branch. 21 , which will not be elaborated in detail here.

[0084] The spatial channel fusion module is used to combine M 12 and M 21 Perform cross-group feature fusion to obtain M fuse .

[0085] Specifically, the spatial channel fusion module is used to transform M 12 and M 21 Perform cross-group feature fusion to obtain M fuse , including: first, M 12 and M 21 After the splicing operation, it is sent to the 3×3 convolution block to achieve cross-group feature fusion:

[0086] M′=S conv3×3 (Cat[M 12 , M 21 ])

[0087] The features after cross-group fusion Then send it to the spatial and channel convolution SCConv layer to generate M fuse :

[0088] M fuse =SCConv(M′)

[0089] Step 3: Use the multimodal brain tumor dataset to train the brain tumor segmentation network structure to obtain a brain tumor segmentation model;

[0090] Step 4: After preprocessing, the multimodal brain tumor image to be detected is input into the brain tumor segmentation model to obtain the segmentation result.

[0091] In summary, the improved multi-encoding Unet network is used to segment the lesions of brain tumor images. Based on the original network structure, a multi-scale progressive fusion module is constructed to extract multi-scale information while fusing cross-modal correlation features and complementary clues in multi-modal data, enhancing the model's feature representation ability when processing complex tumor structures. The utilization rate of multi-modal feature information is improved, the purpose of enhancing the integrity of target information is achieved, and the model's ability to understand complex structures is strengthened.

[0092] On the basis of the above embodiment, further, constructing a brain tumor segmentation network structure in step 3 also includes: adding a first context semantics-guided CIBM I module after the multi-scale progressive fusion modules of the third and fourth layers of the brain tumor segmentation network structure, and adding a second context semantics-guided CIBMII module after the multi-scale progressive fusion modules of the first and second layers, respectively; wherein the CIBM I module and the CIBMII module are used to fuse the output of the multi-scale progressive fusion module with the output of the previous decoding layer.

[0093] It can be understood that adding contextual semantic guidance modules CIBM I and CIBMII to perform adaptive recalibration of features can reduce semantic conflicts and information redundancy in the fusion process of features at different levels, and help the decoder to better restore the fine-grained features of the image.

[0094] Specifically, Figure 4 As shown in FIG. 1 , the first context semantics guided CIBM I module includes two spatial attention branches, a 3×3 convolution block and a channel attention module. The specific operations of the CIBM I module are as follows:

[0095] Determine the output M of the multi-scale progressive fusion module fuse The output H of the previous decoding layer is used as the input of the CIBM I module, where the output of the multi-scale progressive fusion module is a shallow feature, the output of the previous decoding layer A deep feature.

[0096] The first spatial attention branch takes M fuse After passing through 1×1 convolution blocks, H and H are combined by multiplication to obtain f mul :

[0097]

[0098] Then f mul Through the spatial attention module, we can obtain f′ mul :

[0099]

[0100] The second spatial attention branch takes M fuse After passing through 1×1 convolution blocks, H and H are combined by addition to obtain f sum :

[0101]

[0102] In order to enhance the feature expression, f sum Through the spatial attention module, we can obtain f′ sum :

[0103]

[0104] Among them, σ represents the sigmoid activation function.

[0105] In order to further optimize the fusion effect of deep and shallow features, f′ mul and f′ sum After concatenation in the channel dimension, features are automatically selected through a 3×3 convolutional block to obtain f lh :

[0106] f lh =S conv3×3 (Cat[f′ mul , f′ sum ])

[0107] f lh The feature weights are adaptively adjusted through the channel attention module to obtain F lh .

[0108]

[0109] Specifically, the second contextual semantics guided CIBMII module includes two spatial attention branches, a large kernel selection attention module, a 3×3 convolutional block, and a channel attention module;

[0110] Specifically, Figure 5 As shown in Figure 1, the second context semantics guided CIBMII module includes two spatial attention branches, a large core selection attention module, a 3×3 convolution block and a channel attention module. Correspondingly, the specific operations of the CIBMII module are as follows:

[0111] Determine the output M of the multi-scale progressive fusion module fuse And the output H of the previous decoding layer is used as the input of the CIBMII module;

[0112] The first spatial attention branch takes M fuse After passing through 1×1 convolution blocks, H and H are combined by multiplication to obtain f mul , then f mulThrough the spatial attention module, we can obtain f′ mul ;

[0113] The second spatial attention branch takes M fuse After passing through 1×1 convolution blocks, H and H are combined by addition to obtain f sum , then f sum Through the spatial attention module, we can obtain f′ sum ;

[0114] After the 1×1 convolution block, M fuse After concatenating with H, it is input into the large core selection attention module to obtain f′ g ,like Figure 5 As shown in Figure 2, the processing of the large core selection attention module is expressed by the following formula:

[0115] f k1 =S conv5×5 (f g )

[0116]

[0117] f′ k1 =S conv1×1 (f k1 ), f′ k2 =Sc onv1×1 (f k2 )

[0118] f k = Cat[f′ k1 , f′ k2 ]

[0119] f sig =σ(S conv7×7 (Cat[AvgPool(f k ), MaxPool(f k )]))

[0120]

[0121] Among them, f g Indicates M fuse and H are the concatenated outputs after passing through 1×1 convolution blocks, S conv5×5 represents a 5×5 convolutional block, represents 2D dilated convolution, with a kernel size of 7×7 and a dilation rate of 3; S conv1×1 represents an I×1 convolutional block, Cat represents a concatenation operation, and S conv7×7 represents a 7×7 convolutional block, AvgPool represents average pooling, MaxPool represents maximum pooling, σ represents the sigmoid activation function, represents multiplication, Represents addition.

[0122] f′ mul , f′ g and f′ sum After concatenation in the channel dimension, features are automatically selected through a 3×3 convolutional block to obtain f lh ;

[0123] f lh The feature weights are adaptively adjusted through the channel attention module to obtain F lh .

[0124] The embodiment of the present invention adds a contextual semantic guidance module, which enhances the model's ability to recover fine-grained details by fusing the characteristics of deep features and shallow features, effectively integrates rich spatial detail features with global semantic features, and adaptively recalibrates features, thereby reducing semantic conflicts and information redundancy in the fusion process of features at different levels, and helping the decoder to better restore the fine-grained features of the image. A large kernel selective attention module is added to some contextual semantic guidance modules. During the brain tumor image segmentation process, if only local background information is considered, some normal brain tissues or complex vascular structures with density and morphology similar to tumors are often misjudged as tumors. This problem can be effectively solved by introducing large kernel selective attention and modeling multi-scale context information.

[0125] Table 1 Comparison with other networks on the Brats2018 validation dataset

[0126]

[0127]

[0128] As shown in Table 1, the method proposed in the present invention outperforms existing advanced methods in terms of performance indicators. Enhanced tumors are usually located between edema and necrosis areas, and are therefore the most challenging of the three subregions of brain tumor segmentation. The Dice Score of the method of the present invention on enhanced tumors is 79.2%, and the result on Hausdorff Distance is 3.24 mm, indicating that the maximum error distance between the automatic segmentation result and the manual annotation result is 3.24 mm. This error distance shows that the method of the present invention is very accurate in locating the boundary of the ET area, almost close to the accuracy of manual annotation, showing its great potential in actual clinical applications.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A brain tumor segmentation method based on multimodal fusion, characterized in that: include: Step 1: obtaining a multimodal brain tumor image, and preprocessing the multimodal brain tumor image to construct a multimodal brain tumor dataset; Step 2: constructing a brain tumor segmentation network structure, including: selecting a Unet network as a benchmark model, and adding a multi-scale progressive fusion MFCM module at the jump connection between each encoding layer and decoding layer in the Unet network; wherein the Unet network includes four independent encoders and one decoder, and the multi-scale progressive fusion MFCM module is used to fuse the multi-modal features extracted by the independent encoders; Step 3: Using the multimodal brain tumor dataset to train the brain tumor segmentation network structure to obtain a brain tumor segmentation model; Step 4: After preprocessing, the multimodal brain tumor image to be detected is input into the brain tumor segmentation model to obtain a segmentation result.

2. The brain tumor segmentation method based on multimodal fusion according to claim 1, characterized in that: The multimodal brain tumor image includes T1, T1ce, Flair and T2 modalities.

3. The brain tumor segmentation method based on multimodal fusion according to claim 1, characterized in that: The preprocessing includes converting the three-dimensional MRI image into two-dimensional slice data, converting the format of the annotation file, and strengthening normalization.

4. The brain tumor segmentation method based on multimodal fusion according to claim 2, characterized in that: The multi-scale progressive fusion MFCM module includes a T1-T1ce fusion branch, a Flair-T2 fusion branch and a spatial channel fusion module; Correspondingly, the multi-scale progressive fusion MFCM module is used to fuse the multi-modal features extracted by the independent encoders, including: Determine that the features of T1, T1ce, Flair and T2 output by independent encoders are M1, M2, M3 and M4; The T1-T1ce fusion branch fuses the features of M1 and M2 output by the encoder to obtain M 12 ; The Flair-T2 fusion branch fuses the features of M3 and M4 output by the encoder to obtain M 21 ; The spatial channel fusion module is used to combine M 12 and M 21 Perform cross-group feature fusion to obtain M fuse .

5. The brain tumor segmentation method based on multimodal fusion according to claim 4, characterized in that: The T1-T1ce fusion branch and the Flair-T2 fusion branch have the same structure, including: a first branch and a second branch; wherein the first branch includes two 3×3 convolution blocks, and the second branch includes two 5×5 convolution blocks, and each convolution block is connected to a splicing layer; Correspondingly, the specific operations of the T1-T1ce fusion branch are as follows: M1 and M2 are respectively sent to 1×1 convolution for dimensionality reduction to obtain f1 and f2; After f1 and f2 are input into the first concatenation layer of the first branch for concatenation, they are sent into the first 3×3 convolutional block of the first branch to obtain f 12 At the same time, f1 and f2 are input into the first concatenation layer of the second branch for concatenation, and then sent into the first 5×5 convolution block of the second branch to obtain f 21 ; f 12 and f 21 After the second concatenation layer of the first branch is input for concatenation, it is sent to the second 3×3 convolution block of the first branch to obtain f1 ′ 2; At the same time, f 12 and f 21 After concatenation with the first concatenated layer of the second branch, it is fed into the second 5×5 convolutional block of the second branch to obtain f2 ′ 1; Use element-by-element multiplication to convert f1 ′ 2 and f2 ′ 1 is combined with f1 and f2 through residual connection, and the fused features are further integrated through 3×3 convolution to obtain M 12 ; M3 and M4 are fused into M through the Flair-T2 branch 21 .

6. The brain tumor segmentation method based on multimodal fusion according to claim 4, characterized in that: The spatial channel fusion module is used to combine M 12 and M 21 Perform cross-group feature fusion to obtain M fuse , specifically including: first, M 12 and M 21 After the splicing operation, it is sent to the 3×3 convolution block to achieve cross-group feature fusion, and then sent to the SCConv layer to generate M fuse .

7. The brain tumor segmentation method based on multimodal fusion according to claim 1, characterized in that: The method for constructing a brain tumor segmentation network structure also includes: adding a first context semantics-guided CIBMⅠ module after the multi-scale progressive fusion modules of the third and fourth layers of the brain tumor segmentation network structure, and adding a second context semantics-guided CIBMⅡ module after the multi-scale progressive fusion modules of the first and second layers, respectively; wherein the CIBMⅠ module and the CIBMⅡ module are used to fuse the output of the multi-scale progressive fusion module and the previous decoding layer.

8. The brain tumor segmentation method based on multimodal fusion according to claim 7, characterized in that: The first context semantics guided CIBMⅠ module includes a first spatial attention branch, a second spatial attention branch, a 3×3 convolution block and a channel attention module; Correspondingly, the specific operations of the CIBMⅠ module are as follows: Determine the output M of the multi-scale progressive fusion module fuse And the output H of the previous decoding layer is used as the input of the CIBMⅠ module; The first spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by multiplication to obtain f mul , then f mul Through the spatial attention module, we can obtain f′ mul ; The second spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by addition to obtain f sum , then f sum Through the spatial attention module, we can obtain f′ sum ; f′ mul and f′ sum After concatenation in the channel dimension, features are automatically selected through a 3×3 convolutional block to obtain f lh ; f lh The feature weights are adaptively adjusted through the channel attention module to obtain F lh .

9. The brain tumor segmentation method based on multimodal fusion according to claim 7, characterized in that: The second context semantics guided CIBMⅡ module includes two spatial attention branches, a large core selection attention module, a 3×3 convolutional block and a channel attention module; Correspondingly, the specific operations of the CIBMⅡ module are as follows: Determine the output M of the multi-scale progressive fusion module fuse and the output H of the previous decoding layer as the input of the CIBMⅡ module; The first spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by multiplication to obtain f mul , then f mul Through the spatial attention module, we can obtain f′ mul ; The second spatial attention branch M fuse After passing through 1×1 convolution blocks, H and H are combined by addition to obtain f sum , then f sum Through the spatial attention module, we can obtain f′ sum ; After the 1×1 convolution block, M fuse After being concatenated with H, it is input into the large core selection attention module to obtain f′ g ; f′ mul , f′ g and f′ sum After concatenation in the channel dimension, features are automatically selected through a 3×3 convolutional block to obtain f lh ; f lh The feature weights are adaptively adjusted through the channel attention module to obtain F lh .

10. The brain tumor segmentation method based on multimodal fusion according to claim 9, characterized in that: The processing process of the large core selection attention module is expressed by the following formula: f k1 =S conv5×5 (f g ) f′ k1 =S conv1×1 (f k1 ),f′ k2 =S conv1×1 (f k2 ) f k =Cat[f′ k1 ,f′ k2 ] f sig =σ(S conv7×7 (Cat[AvgPool(f k ),MaxPool(f k )])) Among them, f g Indicates M fuse and H are the concatenated outputs after passing through 1×1 convolution blocks, S conv5×5 represents a 5×5 convolutional block, represents 2D dilated convolution, with a kernel size of 7×7 and a dilation rate of 3; S conv1×1 represents a 1×1 convolutional block, Cat represents a concatenation operation, and S conv7×7 represents a 7×7 convolutional block, AvgPool represents average pooling, MaxPool represents maximum pooling, σ represents the sigmoid activation function, represents multiplication, Represents addition.

Citation Information

Cited By

  • Multi-modal brain tumor segmentation method based on frequency domain channel attention

    CN120580433A

  • Construction of brain glioma subregion segmentation model based on multi-modal edge feature fusion

    CN122223030A