Brain tumor image segmentation method based on lightweight multimodality

Through a multimodal data fusion feature extraction network based on a binary tree structure and a brain tumor image segmentation network, the problem of high computing power requirements in existing technologies is solved, and efficient brain tumor segmentation is achieved on low-computing power devices. The graphics card requirement is only 8G, and the segmentation results are accurate.

CN120689351AActive Publication Date: 2025-09-23BEIJING ANZHEN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIVERSITY NANCHONG HOSPITAL·NANCHONG CENTRAL HOSPITAL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510687279.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-23
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing brain tumor segmentation networks require high computing power resources during training and inference, have high deployment costs, and have practical deployment difficulties during high-precision training, making it difficult to achieve accurate brain tumor segmentation on low-computing power devices.

Method used

A multimodal data fusion feature extraction network based on a binary tree structure is adopted. By constructing a local feature extractor and a multimodal fusion unit, and combining jump connection parameters for feature fusion, a lightweight brain tumor image segmentation network is constructed, which is suitable for multimodal brain magnetic resonance image datasets.

Benefits of technology

Efficient brain tumor image segmentation was achieved on low-computing-power devices. The segmentation results were not significantly different from existing models, meeting the needs of doctors. In addition, the graphics card requirement was only 8G, far lower than the existing network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689351A_ABST
    Figure CN120689351A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and provides a brain tumor image segmentation method based on lightweight multi-modality, and the main scheme is as follows: obtaining a multi-modality brain nuclear magnetic resonance image data set, and constructing a multi-modality data fusion feature extraction network based on a binary tree structure; the feature extraction network comprises tree-shaped coding layers with the same number as the modals; constructing a local feature extractor for the brain nuclear magnetic resonance image of each modal, and performing feature extraction; constructing a multi-modal fusion unit for feature extraction results of two adjacent modals, and performing inter-modal feature fusion; generating jump connection parameters for each tree coding layer, performing up-sampling on outputs of all the tree coding layers, and then performing connection fusion operation by using the jump connection parameters; constructing a brain tumor image segmentation network, and obtaining a segmentation result of the brain tumor image; and training, verifying and testing the two networks by using a multi-modal brain nuclear magnetic resonance image data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a brain tumor image segmentation method based on lightweight multimodality. Background Art

[0002] Existing brain tumor segmentation networks, such as those based on the transformer architecture and the nnU-Net structure, require significant computing power (typically over 8GB) during training and inference, resulting in high deployment and hardware costs. Training with high-precision parameters requires precision tailoring, and existing networks present certain difficulties in practical application. For example, the computing power requirements complicate deployment.

[0003] Taking nnU-Net (No New-Network U-Net) as an example, its core structure is based on the classic U-Net architecture. U-Net is a symmetric encoder-decoder network consisting of a downsampling path (encoder) and an upsampling path (decoder). The encoder uses a series of convolution and pooling operations to gradually reduce the spatial size of the feature map, extracting high-level semantic features from the image and capturing information at different scales. For example, in medical image segmentation tasks, the encoder uses convolution to continuously extract features such as the outline and texture of organs or lesions in the image. Pooling allows the model to focus on the distribution of features over a wider range. A 3x3 convolution kernel is typically used to effectively extract features while keeping computational complexity relatively low. Decoder: Utilizing upsampling operations (such as deconvolution and bilinear interpolation), the decoder gradually restores the spatial size of the feature map. It also performs jump connections on the feature maps of the corresponding layers in the encoder, fusing low-level spatial information with high-level semantic information, ultimately outputting a segmentation result of the same size as the input image. Deconvolution, for example, maps low-resolution feature maps back to high resolution. Adding or concatenating these with the feature maps of the corresponding layers in the encoder allows the model to restore spatial details while incorporating previously learned semantics to better segment objects in the image. nnU-Net automatically selects a 2D U-Net, a 3D U-Net, or a combination of the two based on the specific task. 2D U-Net is suitable for processing single slice images and offers high computational efficiency, excelling in scenarios requiring real-time performance or with smaller datasets. 3D U-Net can directly process 3D volumetric data, fully leveraging the spatial context of the image. However, this approach is computationally more complex. For brain tumor segmentation, 3D U-Net can more accurately determine the tumor's location and shape in three dimensions because it accounts for the tumor's continuity across slices. Specifically, for the segmentation network of the 3D U-Net structure, in order to achieve accurate and reliable segmentation results, it generally requires the processing of the following parts, namely: Automated preprocessing: Automatically resample, normalize, and crop the input data. Resample medical images of different modalities and resolutions to a uniform voxel spacing to ensure data consistency. For MRI images, the image resolutions acquired by different devices may vary greatly. Resampling enables the model to be at the same scale benchmark when processing data from different sources. Use normalization methods (such as Z-score normalization) to standardize the image pixel values ​​and adjust the mean and variance of the image to a fixed range. This can accelerate the convergence of the model and improve training efficiency. The cropping operation removes parts of the image that are not related to the target segmentation area, reducing the amount of calculation and avoiding interference of irrelevant background information on model training. For example, in the liver segmentation task, a large number of non-liver areas in the image are cropped out. Automatic architecture selection: Automatically selects the appropriate U-Net architecture based on dataset characteristics (such as data dimensionality and sample size). For small datasets or situations requiring high computational resources, a 2D U-Net may be chosen. This is because small datasets struggle to train complex 3D models, and 2D U-Net offers low computational cost, enabling rapid training and prediction with limited resources. For tasks that require full utilization of 3D contextual information, a 3D U-Net is preferred. For example, in cardiovascular segmentation, a 3D U-Net can more accurately segment vascular structures by analyzing the orientation and connectivity of blood vessels in three-dimensional space. Data augmentation: Utilizing various data augmentation techniques, such as random rotation, flipping, scaling, and elastic deformation, we increase the diversity of training data and improve the model's generalization capabilities. Random rotation can simulate images of organs or lesions from different angles, enabling the model to learn the characteristics of the target at different angles and enhancing robustness to angular variations. Flipping increases sample diversity in the horizontal or vertical directions, similar to observing the target from different perspectives. Scaling allows the model to adapt to changes in target size. Elastic deformation simulates tissue deformation that may occur in actual medical images, allowing the model to better cope with complex deformations in images. Ensemble learning: This approach uses multi-model integration to train multiple different models (such as different U-Net architectures and training parameters) and fuse their predictions to improve segmentation accuracy and stability. For example, one model may perform better at segmenting large object areas, while another model may be better at segmenting small object details. Combining the predictions of multiple models through weighted averaging or voting can fully leverage the strengths of each model, reduce the error of a single model, and make the final segmentation results more accurate and reliable. When using 3D U-Net for medical image segmentation, it is generally achieved through the following steps: Step 1: Data Preparation: Organize the medical image data into the format required by nnU-Net, including the images and corresponding labels, and store them in designated folders. Also, define the dataset information, such as the number of categories and modality information. For example, in the liver tumor segmentation task, liver images and labeled images of the tumor regions must be placed in the specified format, with the number of categories clearly defined as liver and tumor, and the modality information clearly defined as MRI or CT. Accurate and clear data preparation is the foundation of all subsequent steps and directly impacts the quality of model training. Step 2: Data Preprocessing: Automatically resample, normalize, and crop the data in preparation for subsequent training. Resampling is adjusted based on the optimal voxel spacing obtained from dataset analysis to ensure consistency across spatial scales across different sample data. Normalization calculates the mean and variance of the image, mapping pixel values ​​to a standard range. Cropping removes image edges unrelated to the segmentation target based on the image's bounding box or other predefined rules. Preprocessed data allows for faster model convergence, reduces training time, and improves the model's adaptability to diverse data.

[0004] Step 3: Network Training: Based on the automatically selected network architecture, training is performed using preprocessed data. Using cross-validation, the dataset is divided into training and validation sets. The model is trained multiple times, performance is evaluated, and optimal hyperparameters are selected. During training, commonly used optimizers such as Adam continuously adjust network parameters based on the model's loss function to minimize loss. Loss functions typically use Dice Loss or Cross-Entropy Loss, taking into account the characteristics of the medical image segmentation task, to measure the difference between the model's predictions and the true labels. Cross-validation allows for a more comprehensive evaluation of the model's performance on different data subsets, avoiding overfitting and selecting optimal hyperparameters such as the learning rate and number of network layers to achieve optimal model performance. Step 4: Model Inference: Use the trained model to perform segmentation predictions on new medical images. During inference, you can choose whether to use test-time augmentation (TTA) to further improve prediction accuracy. TTA performs transformations such as rotation and flipping on the input image, feeds the model multiple times for prediction, and then fuses the predictions. For example, for a medical image to be segmented, predictions are performed for three different scenarios: horizontal flipping, vertical flipping, and non-flipping. The three predictions are then averaged, effectively improving the reliability of the segmentation results. Step 5, Post-processing: Post-process the model's predictions, such as removing small connected regions and filling holes, to obtain more reasonable segmentation results. In medical image segmentation, model predictions may produce isolated small regions or internal holes that do not conform to actual medical structures. Morphological operations (such as opening to remove small regions and closing to fill holes) or connected domain analysis can be used to correct the predictions, making them more consistent with medical reality and providing more valuable information for clinical diagnosis and treatment.

[0005] In summary, existing traditional brain tumor image segmentation methods need to rely on cumbersome parameter adjustments and high computing power deployment costs to obtain accurate and reliable segmentation results. Summary of the Invention

[0006] The purpose of the present invention is to provide a brain tumor image segmentation method based on lightweight multimodality, which can complete the network deployment under the premise of low computing power and can achieve more accurate brain tumor image segmentation tasks.

[0007] The present invention solves the technical problem and adopts the following technical solution: A brain tumor image segmentation method based on lightweight multimodality includes the following steps: Acquire a multimodal brain magnetic resonance imaging dataset and construct a multimodal data fusion feature extraction network based on a binary tree structure. The feature extraction network includes a tree encoding layer that matches the number of modalities. Constructing a local feature extractor for each modality of brain MRI image, and using the local feature extractor to extract features from the current modality of brain MRI image; A multimodal fusion unit is constructed for the feature extraction results of two adjacent modalities, and the feature extraction results of the two adjacent modalities are fused using the multimodal fusion unit. After all inter-modal features are fused, the output of the current tree coding layer is formed; Generate skip connection parameters for each tree coding layer, upsample the outputs of all tree coding layers, and then use the skip connection parameters to perform concatenation and fusion operations; Construct a brain tumor image segmentation network, and use the segmentation network to obtain the brain tumor image segmentation results after completing the connection and fusion operation; The multimodal brain magnetic resonance image dataset was used to train, verify and test the multimodal data fusion feature extraction network based on binary tree structure and the brain tumor image segmentation network.

[0008] As a further optimization, when the multimodal brain MRI image dataset includes brain MRI images of four modalities, the constructed multimodal data fusion feature extraction network based on the binary tree structure includes four tree encoding layers; The fourth tree-structured coding layer inputs four modal brain MRI images. After being processed by the local feature extractor and the multimodal fusion unit, it outputs three images after feature extraction and inter-modal feature fusion to the third tree-structured coding layer. After being processed by the local feature extractor and the multimodal fusion unit, the third tree coding layer outputs two images after feature extraction and inter-modal feature fusion to the second tree coding layer; After being processed by the local feature extractor and the multimodal fusion unit, the second tree coding layer outputs an image after feature extraction and inter-modal feature fusion to the first tree coding layer.

[0009] As a further optimization, the number of channels of each modality in the tree coding layer is consistent, and the number of modal channels of the upper tree coding layer is 1 / 2 times the number of modal channels in the lower tree coding layer.

[0010] As a further optimization, the local feature extractor includes a first convolutional layer, a second convolutional layer and a newly added convolutional layer; The current modality brain MRI image passes through the first convolutional layer, the newly added convolutional layer, and the second convolutional layer in sequence to complete feature extraction of the current modality brain MRI image; The input size and output size of the newly added convolutional layer are consistent with the input size and output size of the first convolutional layer. The input size of the third convolutional layer is consistent with the output size of the newly added convolutional layer, and the output size of the third convolutional layer is 1 / 2 times the input size.

[0011] As a further optimization, the multimodal fusion unit is a bimodal fusion unit; The bimodal fusion unit includes a first channel attention mechanism, a second channel attention mechanism, a low-rank interaction channel attention mechanism and a spatial attention mechanism; The first-channel attention mechanism is used to input the feature extraction results of one modality, and the second-channel attention mechanism is used to input the feature extraction results of the adjacent modalities.

[0012] As a further optimization, after the feature extraction results of the two adjacent modalities are subjected to inter-modal feature fusion by the dual-modal fusion unit, a feature fusion result of a single modality is generated, and the size of the feature fusion result of the single modality is consistent with the size of the feature extraction results of the two adjacent modalities; The feature fusion results of all single modalities in the current tree coding layer constitute the output of the tree coding layer.

[0013] As a further optimization, the jump connection parameters are used to perform the fusion operation, and the formula is: , Among them, A(.) represents the feature fusion function, its input dimension is n*m, n represents the nth mode, m represents the channel dimension of each single mode, and the output dimension of the feature fusion function is m.

[0014] As a further optimization, when the multimodal fusion unit is used to perform inter-modal feature fusion on the feature extraction results of two adjacent modalities, the method includes: Utilize the first-channel attention mechanism and the second-channel attention mechanism to learn the channel attention of two adjacent modalities; The low-rank interactive channel attention mechanism is used to further learn the channel attention of two adjacent modalities using two low-rank matrices; The spatial attention mechanism is used to perform spatial attention learning on the learning results of the low-rank interaction channel attention mechanism.

[0015] As a further optimization, the calculation formula of the first channel attention mechanism is: , The calculation formula of the second channel attention mechanism is: , The calculation formula of the low-rank interactive channel attention mechanism is: , The calculation formula of the spatial attention mechanism is: , Where a represents the ath mode, a+1 represents the mode adjacent to mode a, and a≤n-1.

[0016] As a further optimization, after obtaining the brain tumor image segmentation results and before using the multimodal brain MRI image dataset to train, verify, and test the binary tree-based multimodal data fusion feature extraction network and the brain tumor image segmentation network, the following is also included: The multimodal brain magnetic resonance imaging dataset is divided into training set, validation set and test set; The multimodal brain MRI images in the training, validation, and test sets were cropped to a uniform size and filtered to remove all images with annotated regions. Data augmentation is performed on images of labeled regions in the training and test sets after cropping and filtering operations.

[0017] The beneficial effects of the present invention are as follows: the present invention provides a new multimodal data fusion feature extraction network and a brain tumor image segmentation network based on a binary tree structure. Due to the characteristics of the network structure, the network structure can be rapidly scaled to facilitate users to use multimodal data for training. Moreover, the segmentation results of the tumor area in the final image are not much different from those of the existing model, which can meet the needs of doctors. At the same time, the present invention only needs to be trained on an 8G graphics card, and the requirements for video memory are far lower than those of the existing model network. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flowchart of a brain tumor image segmentation method based on lightweight multimodality in Example 1 of the present invention; Figure 2 Schematic diagram of the network structure of a multimodal data fusion feature extraction network based on a binary tree structure in Example 1 of the present invention; Figure 3 Schematic diagram of the composition structure of feature convolution blocks when sampling an existing 3D Unet network in Example 1 of the present invention; Figure 4 Schematic diagram of the structure of the feature convolution block in Example 1 of the present invention; Figure 5 This is a schematic diagram of the structure of a multimodal fusion unit in accordance with an embodiment of the present invention; Figure 6 Schematic diagram of the network structure of the brain tumor image segmentation network in Example 1 of the present invention; Figure 7 This is a schematic diagram of the graphics card occupancy during non-training inference in Example 3 of the present invention; Figure 8 This is a schematic diagram of the graphics card occupancy during training in Example 3 of the present invention; Figure 9 This is a schematic diagram of the graphics card occupancy during inference in the third embodiment of the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0020] Example 1 This embodiment provides a brain tumor image segmentation method based on lightweight multimodality. The flowchart is shown in Figure 1 , wherein the method comprises the following steps: S1. Acquire a multimodal brain magnetic resonance imaging dataset and construct a multimodal data fusion feature extraction network based on a binary tree structure. The feature extraction network includes a tree-shaped encoding layer that matches the number of modalities. S2. constructing a local feature extractor for each modality of brain MRI image, and using the local feature extractor to extract features from the current modality of brain MRI image; S3. Construct a multimodal fusion unit for the feature extraction results of two adjacent modalities, and use the multimodal fusion unit to perform inter-modal feature fusion on the feature extraction results of the two adjacent modalities. After all inter-modal features are fused, the output of the current tree coding layer is formed; S4. Generate skip connection parameters for each tree coding layer, upsample the outputs of all tree coding layers, and then perform a concatenation and fusion operation using the skip connection parameters. S5. constructing a brain tumor image segmentation network, and using the segmentation network to obtain a brain tumor image segmentation result after completing a connection and fusion operation; S6. Use the multimodal brain magnetic resonance image dataset to train, verify and test the multimodal data fusion feature extraction network based on the binary tree structure and the brain tumor image segmentation network.

[0021] In this embodiment, the brats21 dataset is taken as an example for detailed description.

[0022] To obtain the BRATs21 dataset, first log in to Kaggle (https: / / www.kaggle.com / ) and click the "Sign Up" button to register. Once registered, log in, search for "BRaTS 2021 Task 1 Dataset" in the search box, find the dataset download page, and click the download button to download. The dataset includes a training set, a validation set, and a test set.

[0023] In the training set folder rsna_asnr_miccai_brats2021_training_data, subfolders named by patient number contain the relevant data for each patient. Each patient subfolder contains MRI image files of four modalities: brats2021_xxxx_flair.nii.gz (T2 FLAIR sequence images), brats2021_xxxx_t1.nii.gz (T1 sequence images), brats2021_xxxx_t1ce.nii.gz (T1 enhancement sequence images), brats2021_xxxx_t2.nii.gz (T2 sequence images), and the corresponding segmentation label file brats2021_xxxx_seg.nii.gz.

[0024] In the validation set, each subfolder contains MRI image files of four modalities, but no segmentation label files. There are only four image files: brats2021_xxxx_flair.nii.gz, brats2021_xxxx_t1.nii.gz, brats2021_xxxx_t1ce.nii.gz, and brats2021_xxxx_t2.nii.gz.

[0025] The test set data is not publicly available and requires model submission for testing. Its directory structure is similar to the validation set, containing MRI images of patients across all four modalities, but without segmentation labels. All dataset files are stored in the NIfTI format (.nii.gz). Each of the four MRI images for each case is 240×240×155 pixels in size and shares segmentation labels. Label information primarily includes enhancing tumor (ET), peritumoral edema, invasive tissue (ED), and necrotic tumor core (NCR). The dataset consists of 1251 cases, split approximately 1:4 into a training set and a validation set (test set) with 1000 cases and 251 cases in total.

[0026] For the brats21 dataset, since it contains image data of four modalities, a multimodal data fusion feature extraction network based on a binary tree structure is constructed in this embodiment. Figure 2 , the number of layers of the tree encoding layer of the feature extraction network is also four layers, that is: when the multimodal brain magnetic resonance image dataset contains brain magnetic resonance images of four modalities, the constructed multimodal data fusion feature extraction network based on the binary tree structure includes four tree encoding layers; The fourth tree-structured coding layer inputs four modal brain MRI images. After being processed by the local feature extractor and the multimodal fusion unit, it outputs three images after feature extraction and inter-modal feature fusion to the third tree-structured coding layer. After being processed by the local feature extractor and the multimodal fusion unit, the third tree coding layer outputs two images after feature extraction and inter-modal feature fusion to the second tree coding layer; After being processed by the local feature extractor and the multimodal fusion unit, the second tree coding layer outputs an image after feature extraction and inter-modal feature fusion to the first tree coding layer.

[0027] Here, n is the number of tree coding layers. In the binary tree-structured multimodal data fusion feature extraction network, the input to the tree coding layer decreases from 4 to 1 from top to bottom. Furthermore, the number of channels for each modality in the tree coding layer is consistent, and the number of modal channels in the previous tree coding layer is 1 / 2 times the number of modal channels in the next tree coding layer.

[0028] In actual application, see Figure 3 In the original 3D Unet network, while extracting features from the input, the size of its output is half of the original. However, the original 3D Unet network will cause the model to overfit when trained on the brats21 dataset. Therefore, this embodiment alleviates this problem by expanding the parameters to adapt to the scale of the brats21 dataset. Figure 4 In this embodiment, while expanding to a tree structure, in order to ensure that the network has sufficient parameters for learning, a new layer is added to the feature extraction part, that is, a new convolution. The input size of this layer remains unchanged from the previous layer. Its main purpose is to increase the generalization ability of the model by expanding the network parameters.

[0029] Taking modality 4 as an example, assume that the size of modality 4 is (B, C, H, W, Z), where B is the batch size, C is the channel size, and (H, W, Z) is the 3D size of the data input. The input of the first convolutional layer is (B, C, H, W, Z), and the output is (B, C, H, W, Z). The input of the newly added convolutional layer is (B, C, H, W, Z), and the output is (B, C, H, W, Z). The input of the second convolutional layer is (B, C, H, W, Z), and the output is (B, 2C, H / 2, W / 2, Z / 2). The entire feature convolution module acts as a local feature extractor for learning model features, enhancing the network's learning and generalization capabilities.

[0030] Therefore, in this embodiment, for each modality of brain MRI image, when the multimodal data fusion feature extraction network based on the binary tree structure of this embodiment is used to extract features, it is completed through the feature convolution block. The entire feature convolution module can be used as a local feature extractor for learning model features, thereby enhancing the learning and generalization capabilities of the network. Therefore, the local feature extractor in this embodiment includes a first convolution layer, a second convolution layer, and a newly added convolution layer; The current modality brain MRI image passes through the first convolutional layer, the newly added convolutional layer, and the second convolutional layer in sequence to complete feature extraction of the current modality brain MRI image; The input size and output size of the newly added convolutional layer are consistent with the input size and output size of the first convolutional layer. The input size of the third convolutional layer is consistent with the output size of the newly added convolutional layer, and the output size of the third convolutional layer is 1 / 2 times the input size.

[0031] In this embodiment, in order to perform multimodal fusion, a multimodal fusion unit is designed. Since the number of modes is 4 and it is necessary to perform inter-modal feature fusion on mode 1 and mode 2, mode 2 and mode 3, and mode 3 and mode 4, the multimodal fusion unit is a dual-modal fusion unit. The unit has two modal inputs and the final output is one modality. The dual-modal fusion unit in this embodiment can be Figure 4 The output results of the structure are further fused.

[0032] See also Figure 5 For the dual-modal fusion unit, it includes a first-channel attention mechanism, a second-channel attention mechanism, a low-rank interaction channel attention mechanism and a spatial attention mechanism; the first-channel attention mechanism is used to input the feature extraction results of one modality, and the second-channel attention mechanism is used to input the feature extraction results of the adjacent modalities.

[0033] Moreover, when the feature extraction results of the two adjacent modalities are subjected to inter-modal feature fusion by the dual-modal fusion unit, a feature fusion result of a single modality is generated, and the size of the feature fusion result of the single modality is consistent with the size of the feature extraction results of the two adjacent modalities; the feature fusion results of all single modalities in the current tree coding layer constitute the output of the tree coding layer.

[0034] Specifically, in the dual-modal fusion unit, the two-channel attention mechanism focuses on the relationship between the high-order feature channels of each modality, focusing on learning the channels with lesion areas, while the low-rank interaction channel attention mechanism adopts the principle of LORA's fine-tuning large model method to reduce the model parameters of the network, while further integrating the dual-modal results. Finally, the spatial attention mechanism will focus on the relationship between high-order feature pixels, generalizing the learning of feature information such as the edge and structure of the lesion. Figure 4 The output of the structure is (B, 2C, H / 2, W / 2, Z / 2), and Figure 5 Then there are bimodal features M1=(B,2C,H / 2,W / 2,Z / 2), M2=(B,2C,H / 2,W / 2,Z / 2).

[0035] Here, the channel attention mechanism is a weight list of size 2C, storing the relevant weight information for each channel. The output of the channel attention mechanism is of size (B, 2C, H / 2, W / 2, Z / 2). Subsequently, the input size of the low-rank interaction is {(B, 2C, H / 2, W / 2, Z / 2), (B, 2C, H / 2, W / 2, Z / 2)}, and the output is (B, 2C, H / 2, W / 2, Z / 2). It is worth noting that the feature fusion result of the spatial attention mechanism is also of size (B, 2C, H / 2, W / 2, Z / 2). The entire fusion structure does not change the input size, that is, the size of modal 1 and modal 2, as well as the result of the fused modal, remains unchanged.

[0036] See also Figure 6 In the middle mode, the number of channels is increased to twice that of the previous layer to achieve better generalization learning. The entire network follows the design pattern of the 3D UNet network, introducing skip connections and upsampling. In the skip connection, this embodiment performs a concatenation and fusion operation on the output of each tree coding layer, as shown in the following formula: , Among them, A(.) represents the feature fusion function, its input dimension is n*m, n represents the nth mode, m represents the channel dimension of each single mode, and the output dimension of the feature fusion function is m.

[0037] In multimodal fusion, if the fusion method of formula (1) is directly adopted, it will directly lead to a gap in network learning, which will directly affect the output of the model. For example, the network only learns part of the labeled area. The reason is that the multimodal data comes from different devices and their data distribution is inconsistent. To solve this problem, this embodiment first learns the channel attention of each modality, and then uses two low-rank matrices to form a high-rank matrix to further learn the multimodal channel attention. This model adopts a progressive learning method. Finally, spatial attention is learned on the channel attention learning results. The purpose is to learn the contribution of the corresponding positions of different modalities to segmentation and enhance the generalization of the network.

[0038] Therefore, in this embodiment, when the multimodal fusion unit is used to perform inter-modal feature fusion on the feature extraction results of two adjacent modalities, the following steps may be included: Utilize the first-channel attention mechanism and the second-channel attention mechanism to learn the channel attention of two adjacent modalities; The low-rank interactive channel attention mechanism is used to further learn the channel attention of two adjacent modalities using two low-rank matrices; The spatial attention mechanism is used to perform spatial attention learning on the learning results of the low-rank interaction channel attention mechanism.

[0039] Specifically, the calculation formula of the first channel attention mechanism is: , The calculation formula of the second channel attention mechanism is: , The calculation formula of the low-rank interactive channel attention mechanism is: , The calculation formula of the spatial attention mechanism is: , Here, a represents the ath mode, a+1 represents the mode adjacent to mode a, and a≤n-1. The ChannelAttention(.) function adopts the traditional channel attention mechanism, while CrossChannelAttention(.) implements operations similar to Lora fine-tuning large model training.

[0040] It should be noted that after obtaining the brain tumor image segmentation results and before using the multimodal brain MRI image dataset to train, verify, and test the binary tree-based multimodal data fusion feature extraction network and the brain tumor image segmentation network, the following steps may also be included: The multimodal brain magnetic resonance imaging dataset is divided into training set, validation set and test set; The multimodal brain MRI images in the training, validation, and test sets were cropped to a uniform size and filtered to remove all images with annotated regions. Data augmentation is performed on images of labeled regions in the training and test sets after cropping and filtering operations.

[0041] So far, this embodiment can be Figure 6 The segmentation network structure outputs the final segmentation target, and the last layer of the segmentation network will output it in the form of probability distribution.

[0042] Example 2 Based on the first embodiment, after obtaining the brats21 dataset, this embodiment needs to preprocess the dataset to complete the subsequent training process. In this embodiment, the preprocessing may include data cropping, data enhancement, and data filtering operations.

[0043] Because the original data size of brain MRI images of various modalities in the BRATs21 dataset is large, with most pixel values ​​being zero, the data is cropped to 128*128*128 pixels for computational convenience and reduced computing power requirements. In all annotated datasets, there are four major categories: 0, 1, 2, and 4. 0 represents background, 1 represents necrotic tumor core, 2 represents edema, and 4 represents enhancing tumor.

[0044] When enhancing data, you can use the torchio toolkit for data enhancement, including using tio.ZNormalization for data standardization, using tio.RandomElasticDeformation for elastic deformation transformation of 3D medical data sets to simulate the deformation of biological tissue (parameter num_control_points controls the number of 3D grid points, and parameter max_displacement controls the maximum displacement of each control point in the elastic deformation, affecting the severity of the deformation), using tio.RandomBlur for image blurring (mean and variance parameters are 0 and 1, respectively), using tio.RandomNoise to add noise and improve the generalization ability of the model (mean is 0, variance is 0.001), using tio.RandomGamma to adjust the brightness of the data (log_gamma parameter is (-0.3, 0.3)), using tio.RandomFlip for 3D flipping (left and right), using tio.RandomAffine for affine changes, etc., including translation, rotation, distortion, etc., using tio.RescaleIntensity The pixel intensity of the 3D data is normalized to between 0 and 1, and finally the training data of the labeled area is enhanced.

[0045] For the training set, since some data labels lacked labels for certain categories, we removed these non-compliant data. After preprocessing, the training set was resized to 946 columns. For the test set, in addition to cropping to the input size (128*128*128), we simply used tio.ZNormalization to normalize the data. However, since some data labels lacked labels for certain categories, we removed these non-compliant data. After preprocessing, the validation and evaluation sets were resized to 240 columns.

[0046] After the above-mentioned preprocessing operation on the brain magnetic resonance images of the four modalities, the model can be trained. In this embodiment, the sample size of the training data set is 946 cases, and the test set (i.e., evaluation set) and validation set are 240 cases. In the process of training samples, the training batch size is 1. The MultiStepLR learning rate scheduling plan [60,80] is adopted, and the first 60 batches are trained with 0.001, 60-80 are trained with a learning rate of 0.0001, and greater than 80 are trained with a learning rate of 0.00001. In the process of optimizing the model weights, the Adam optimizer is adopted, and its weight decay parameter is 0.00001, and the momentum decay and square strategy decay parameters are (0.95, 0.995) to control the first-order momentum estimation and second-order momentum estimation of the matrix respectively. In order to facilitate the reproduction of the model effect, the random factor is 42; the total training batch size is 200; cross entropy and dice are used in gradient update. The combined loss function, the cross entropy loss function is to learn the classification of labels, while the dice loss function calculates the outline of the segmentation label and the predicted label. The calculation formula is as follows: , Among them, L CE is the cross entropy loss function. Since there are more zero pixel values ​​and fewer pixel values ​​in the tumor area, the weight parameters of the cross entropy loss function are [1, 3, 3]. The parameter a represents a balance parameter, which is used to control the balance between pixel classification and contour calculation. Different optimization directions should be adopted at the beginning, middle, and end of model training. Therefore, this embodiment introduces the epoch (training iteration number) for further improvement, as shown below: , , Among them, current_epoch is the current iteration number, and epochs is the total number of iterations.

[0047] After training is completed, intermediate data such as optimizer parameters, model parameters, loss function, and current iteration number are saved to facilitate model recovery and continued training.

[0048] Next, we can evaluate the model. This example uses the Dice metric for verification, where X represents the predicted target area, Y represents the actual target area, |XnY| is the number of elements in the intersection of the predicted and actual areas, and |X| and |Y| are the number of elements in the predicted and actual areas, respectively. The Dice value ranges from [0 to 1; the closer it is to 1, the better the model segmentation performance.

[0049] Finally, model inference is performed. During inference, no other data preprocessing is performed, other than normalization and cropping to 128*128*128 size. Post-processing of the model involves filling holes in the segmentation model results and blurring edges to ensure smooth segmentation. Finally, the argmax function is used to convert the multi-channel segmentation predictions to the same size as the annotation labels. This effectively involves applying the model's output probability distribution map to each pixel in the channel.

[0050] Example 3 Based on Examples 1 and 2, after preprocessing the dataset, model training, model evaluation, and model inference, this example deploys and trains the binary tree-based multimodal data fusion feature extraction network and brain tumor image segmentation network involved in the lightweight multimodal brain tumor image segmentation method in Example 1 on a GPU. The model performance of this example is compared with that of existing models. Specific comparison results are shown in Table 1:

[0051] This example can be trained on only an 8GB graphics card, achieving an average dice (excluding edema areas) of over 80% within 30 iterations, while also allowing for continued optimization and learning. As shown in Table 1, the network in this example requires significantly less graphics memory than nnU-net v2, Swin UNETR, DiffSegNet, and LKA-Unet.

[0052] In this example, the DICE results of the network model were statistically analyzed. Experiments were conducted on the WT region (the entire lesion region, including the background region of category 0, the tumor necrosis region of category 1, the edema region of category 2, and the tumor enhancement region of category 4) and the TC region (the background region of category 0, the tumor necrosis region of category 1, and the tumor enhancement region of category 4) of the brats21 dataset, and the relevant DICE results were calculated.

[0053] Experimental results show that the segmentation performance of this embodiment on TC is significantly superior to that of existing segmentation networks, with improvements of 0.078, 0.061, 0.054, and 0.072, respectively, representing corresponding improvements of approximately 7.8%, 6.1%, 5.4%, and 7.2%. Furthermore, the network in this embodiment was trained using two batches and 300 iterations, with an initial training learning rate of 0.01.

[0054] Based on Table 1, see Figure 7 、 Figure 8 and Figure 9During graphics card training and inference, enter the nvidia-smi command through the cmd interface to obtain the graphics card runtime occupancy result. Therefore, the graphics card occupancy during non-training inference, training, and inference of this embodiment can be obtained. It can be seen that the network training video memory of this embodiment occupies 11.425G, which is much lower than the existing multimodal segmentation network, and the video memory used for inference is only 2.083G. It can still run successfully on general consumer-grade graphics cards or NVIDIA 1660-level graphics cards, and has wide application and promotion value. The video memory used for existing network inference is generally estimated and can be described by empirical values. The greater the graphics card occupancy during training, the greater its occupancy during inference. The relationship between the two shows a linear relationship, but both will be greater than 2.083G in this embodiment.

[0055] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A brain tumor image segmentation method based on lightweight multimodality, characterized by: The steps include: Acquire a multimodal brain magnetic resonance imaging dataset and construct a multimodal data fusion feature extraction network based on a binary tree structure. The feature extraction network includes a tree encoding layer that matches the number of modalities. Constructing a local feature extractor for each modality of brain MRI image, and using the local feature extractor to extract features from the current modality of brain MRI image; A multimodal fusion unit is constructed for the feature extraction results of two adjacent modalities, and the feature extraction results of the two adjacent modalities are fused using the multimodal fusion unit. After all inter-modal features are fused, the output of the current tree coding layer is formed; Generate skip connection parameters for each tree coding layer, upsample the outputs of all tree coding layers, and then use the skip connection parameters to perform concatenation and fusion operations; Construct a brain tumor image segmentation network, and use the segmentation network to obtain the brain tumor image segmentation results after completing the connection and fusion operation; The multimodal brain magnetic resonance image dataset was used to train, verify and test the multimodal data fusion feature extraction network based on binary tree structure and the brain tumor image segmentation network.

2. The brain tumor image segmentation method based on lightweight multimodality according to claim 1, characterized in that: When the multimodal brain magnetic resonance image data set includes brain magnetic resonance images of four modalities, the constructed multimodal data fusion feature extraction network based on the binary tree structure includes four tree-shaped encoding layers; The fourth tree-structured coding layer inputs four modal brain MRI images. After being processed by the local feature extractor and the multimodal fusion unit, it outputs three images after feature extraction and inter-modal feature fusion to the third tree-structured coding layer. After being processed by the local feature extractor and the multimodal fusion unit, the third tree coding layer outputs two images after feature extraction and inter-modal feature fusion to the second tree coding layer; After being processed by the local feature extractor and the multimodal fusion unit, the second tree coding layer outputs an image after feature extraction and inter-modal feature fusion to the first tree coding layer.

3. The brain tumor image segmentation method based on lightweight multimodality according to claim 1, characterized in that: The number of channels of each modality in the tree coding layer is consistent, and the number of modal channels of the upper tree coding layer is 1 / 2 times the number of modal channels in the lower tree coding layer.

4. The brain tumor image segmentation method based on lightweight multimodality according to claim 1, characterized in that: The local feature extractor includes a first convolutional layer, a second convolutional layer and a newly added convolutional layer; The current modality brain MRI image passes through the first convolutional layer, the newly added convolutional layer, and the second convolutional layer in sequence to complete feature extraction of the current modality brain MRI image; The input size and output size of the newly added convolutional layer are consistent with the input size and output size of the first convolutional layer. The input size of the third convolutional layer is consistent with the output size of the newly added convolutional layer, and the output size of the third convolutional layer is 1 / 2 times the input size.

5. The brain tumor image segmentation method based on lightweight multimodality according to claim 1, characterized in that: The multimodal fusion unit is a dual-modal fusion unit; The bimodal fusion unit includes a first channel attention mechanism, a second channel attention mechanism, a low-rank interaction channel attention mechanism and a spatial attention mechanism; The first-channel attention mechanism is used to input the feature extraction results of one modality, and the second-channel attention mechanism is used to input the feature extraction results of the adjacent modalities.

6. The brain tumor image segmentation method based on lightweight multimodality according to claim 5, characterized in that: After the feature extraction results of the two adjacent modalities are subjected to inter-modal feature fusion by the dual-modal fusion unit, a feature fusion result of a single modality is generated, and the size of the feature fusion result of the single modality is consistent with the size of the feature extraction results of the two adjacent modalities; The feature fusion results of all single modalities in the current tree coding layer constitute the output of the tree coding layer.

7. The brain tumor image segmentation method based on lightweight multimodality according to claim 5, characterized in that: The jump connection parameters are used to perform the fusion operation, and the formula is: , Among them, A(.) represents the feature fusion function, its input dimension is n*m, n represents the nth mode, m represents the channel dimension of each single mode, and the output dimension of the feature fusion function is m.

8. The brain tumor image segmentation method based on lightweight multimodality according to claim 7, characterized in that: When the multimodal fusion unit is used to perform inter-modal feature fusion on the feature extraction results of two adjacent modalities, the method includes: Utilize the first-channel attention mechanism and the second-channel attention mechanism to learn the channel attention of two adjacent modalities; The low-rank interactive channel attention mechanism is used to further learn the channel attention of two adjacent modalities using two low-rank matrices; The spatial attention mechanism is used to perform spatial attention learning on the learning results of the low-rank interaction channel attention mechanism.

9. The brain tumor image segmentation method based on lightweight multimodality according to claim 8, characterized in that: The calculation formula of the first channel attention mechanism is: , The calculation formula of the second channel attention mechanism is: , The calculation formula of the low-rank interactive channel attention mechanism is: , The calculation formula of the spatial attention mechanism is: , Where a represents the ath mode, a+1 represents the mode adjacent to mode a, and a≤n-1.

10. The brain tumor image segmentation method based on lightweight multimodality according to claim 1, characterized in that: After obtaining the segmentation results of the brain tumor image, and before using the multimodal brain magnetic resonance image dataset to train, verify, and test the multimodal data fusion feature extraction network based on the binary tree structure and the brain tumor image segmentation network, the following is also included: The multimodal brain magnetic resonance imaging dataset is divided into training set, validation set and test set; The multimodal brain MRI images in the training, validation, and test sets were cropped to a uniform size and filtered to remove all images with annotated regions. Data augmentation is performed on images of labeled regions in the training and test sets after cropping and filtering operations.

Citation Information

Patent Citations

  • Improved U-Net brain tumor segmentation method based on attention mechanism and multi-scale feature fusion

    CN115424103A

  • Method, device and equipment for predicting dyskinesia based on multi-modal image data

    CN117079093A

  • MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net

    CN117876399A

  • Method for automatically segmenting focus of brain injury of premature infant

    CN119991699A

  • Methods and systems for digital pathology assessment of cancer via deep learning

    US20250054624A1