A lightweight brain tumor segmentation network for multimodal magnetic resonance imaging and its segmentation method

By using technical means of cross-layer connection and causal convolution modules in brain tumor segmentation networks, the problem of learning confounding factors in the network during training is solved, segmentation accuracy and inference speed are improved, making the model more suitable for practical applications.

CN118334044BActive Publication Date: 2025-05-30NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410698732.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-05-30
Estimated Expiration
2044-05-31

AI Technical Summary

Technical Problem

The existing brain tumor segmentation network is prone to learn confounding factors during training, resulting in low segmentation accuracy and slow model inference speed, making it difficult to meet the practical application needs.

Method used

A lightweight brain tumor segmentation network for multimodal magnetic resonance imaging is designed, using cross-layer connection and causal convolution modules to reduce false correlations and improve segmentation accuracy and inference speed through information fusion and feature cross-processing between the decoder and the encoder.

Benefits of technology

The segmentation accuracy and inference speed of brain tumor segmentation network are improved, making the model more suitable for practical applications, especially in small medical equipment scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118334044B_ABST
    Figure CN118334044B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging and a segmentation method thereof, belonging to the technical field of brain tumor image segmentation; it includes an encoder and a decoder. The encoder includes a first feature processing layer, a second causal cross layer, a third causal cross layer, and a fourth feature fusion layer; the decoder includes a first output layer, a second feature processing layer, a third feature decoding layer, and a fourth output layer; the encoder adopts a cross-layer connection method to realize the information fusion of feature maps of the same scale between different levels and transmits it to the decoder, and the decoder splices the features flowing out of the encoder to perform cross-processing of shared features. By adopting a cross-layer connection method between the decoder and the encoder, the present invention realizes the information fusion between different levels of feature maps of the same scale; the causal cross-connection module shares parameters to promote the information exchange and fusion between features, better captures the common features between modalities, and thus improves the segmentation accuracy of the brain tumor segmentation network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of brain tumor image segmentation, and particularly relates to a lightweight brain tumor segmentation network for multi-modal magnetic resonance imaging and a segmentation method thereof. Background Art

[0002] Brain tumors are abnormal tissue growths that form in a person's brain or other parts of the body and then metastasize to the brain. Among them, gliomas are the most common malignant tumors in the central nervous system. It originates from glial cells in the brain tissue and usually forms in the deep or hemispheric regions of the brain. Since gliomas are highly malignant and invasive tumors, they can cause an increase in intracranial pressure, which compresses the brain tissue and causes brain dysfunction, and in severe cases, even threatens the patient's life. Therefore, early diagnosis and treatment of brain tumors are very important for improving the prognosis. There are many modern medical imaging techniques, especially multi-modal imaging methods. The information between different modalities complements each other, and compared with single modality, it can provide more abundant diagnostic information, greatly improving the accuracy and correctness of doctors' tumor diagnosis. Magnetic resonance imaging is a commonly used method for diagnosing gliomas. It can provide better resolution and contrast, and there is no ionizing radiation damage to the human body, which is a non-invasive imaging method. At the same time, it helps doctors observe and evaluate the patient's brain structure, including the location, size, shape and other information of the tumor. It is these advantages that make multi-modal MRI medical images the main imaging method for brain tumor diagnosis and monitoring.

[0003] Brain tumor segmentation refers to separating the tumor area from the normal area in the brain tissue in medical images. In brain tumor MRI segmentation, the four most commonly used MRI image modalities are FLAIR sequence, T1 sequence, T1c sequence and T2 sequence. At present, brain tumor segmentation mainly still relies on manual segmentation by doctors or experts. Reliable brain tumor segmentation provides an important basis for treatment decisions such as the surgical design, radiotherapy plan and chemotherapy plan of patients. However, manual segmentation of brain tumors is a very time-consuming, expensive and subjective task. Therefore, practical automatic segmentation methods are highly needed. In recent years, methods based on deep learning have made significant progress in the field of medical image segmentation. Compared with traditional segmentation algorithms, these methods can achieve better accuracy and robustness. Deep learning models can automatically learn complex modalities and feature information in medical images and perform end-to-end image segmentation tasks. Thanks to the encoder-decoder structure, convolutional neural networks are widely used in medical image segmentation tasks due to their strong feature representation ability and have achieved good results.

[0004] A method for brain tumor segmentation of MRI images based on an improved 3D-UNet is disclosed in a Chinese patent (application number: 2022116836412). The main steps of the method are as follows: preprocess the brain MRI images in the dataset to obtain a training set; construct an improved 3D-UNet framework by combining dilated convolution, channel attention mechanism, and residual convolution, and use the training set data to train the improved 3D-UNet, which is a brain tumor segmentation network; import the brain MRI image test data to be segmented into the brain tumor segmentation network to obtain the segmented result.

[0005] The following deficiencies exist in the above method: The above patent considered preprocessing the medical dataset, but did not notice the unique structural properties of medical data samples. The number of medical image datasets is very small, and at the same time, the boundaries between the foreground and background of the segmentation are not clear, and the co-occurrence phenomenon of organs and tissues is very serious, which causes the model to learn confounding factors during the training process. At the same time, there are also requirements for the size and inference speed of the model in practical applications.

[0006] Therefore, how to solve the problem of learning confounding factors during the training process of the brain tumor segmentation network, improve the inference speed of the brain tumor segmentation network in practical applications, and thus improve the accuracy and inference speed of the brain tumor segmentation network is the technical problem that the present invention wants to solve. Summary of the Invention

[0007] The purpose of the present invention is to provide a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging to solve the problems raised in the above background technology and the deficiencies of the prior art.

[0008] The purpose of the present invention is achieved as follows: A lightweight brain tumor segmentation network for multimodal magnetic resonance imaging uses convolutional blocks to build an end-to-end convolutional neural network to complete the segmentation task. The brain tumor segmentation network includes an encoder and a decoder. The encoder includes four layers of networks, namely the first feature processing layer, the second causal cross layer, the third causal cross layer, and the fourth feature fusion layer.

[0009] The decoder includes four layers of networks, namely the first output layer, the second feature processing layer, the third feature decoding layer, and the fourth output layer.

[0010] The encoder uses a cross-layer connection method to achieve information fusion between different levels of the same-scale feature maps in the same stage, and transmits the fused information to the decoder. While the decoder splices the features flowing out of the encoder, it performs cross-processing between the shared features.

[0011] Preferably, the first feature processing layer includes four 5×5×5 convolutions with a stride of 1. The four convolutions process the four original modality images input in parallel and share parameters.

[0012] The second causal cross layer consists of two causal convolution modules, corresponding to four input branches. The causal convolution module uses a 3×3×3 convolution with a stride of 1, and the parameters are shared between the two causal convolution modules.

[0013] The third causal cross layer consists of two parallel max pooling cascaded with a single causal convolution block. Both the second causal cross layer and the third causal cross layer adopt a parameter sharing strategy. The two feature maps pass through a 3×3×3 convolution with a stride of 1 in parallel respectively, and the output channels are concatenated with the feature maps of each other. Then, normalization operations are performed respectively, and the output after inputting the activation function passes through a 3×3×3 convolution with a stride of 1 respectively, and then channel dimension concatenation is performed, followed by normalization operations and activation functions to obtain a feature map with the same resolution as the original image.

[0014] Preferably, the causal convolution module is proposed based on the semantic causal chain. The semantic causal chain uses the Do operator for backdoor adjustment through the structural causal model SCM to intervene in the conventional segmentation semantic chain to deconfound.

[0015] Preferably, the fourth feature fusion layer includes max pooling cascaded with a conventional convolution block, and the max pooling cascaded with a conventional convolution block uses two 3×3×3 convolutions with a stride of 1.

[0016] The feature map first passes through a 3×3×3 convolution with a stride of 1, and then normalization operations are performed and the activation function is input for output. Then, it passes through a 3×3×3 convolution with a stride of 1 again, and normalization operations are performed and the activation function is input to obtain a feature map with the same resolution as the original image.

[0017] Preferably, the first output layer includes a transposed convolution cascaded with a conventional convolution block. The first output layer receives the output from the fourth feature fusion layer, and after transposed convolution, it concatenates the features of the same dimension flowing out of the third causal cross layer, and then is processed by the conventional convolution block.

[0018] The second feature processing layer includes a transposed convolution cascaded with a causal convolution block. It receives the output of the first output layer, and after transposed convolution, it concatenates the two feature information flowing out of the second causal cross layer of the encoder respectively. Then, cross processing of the features with parameter sharing is performed through the causal convolution module, and the processed features are input to the third feature encoding layer.

[0019] The transposed convolution cascaded with a conventional convolution block uses a 2×2×2 transposed convolution with a stride of 2.

[0020] Preferably, the third feature decoding layer includes a single 1×1×1 convolution block with a stride of 1 to further extract the deconfounded features.

[0021] The fourth output layer includes a single 1×1×1 convolutional block with a stride of 1 to output the final segmentation map.

[0022] Preferably, a cross-layer connection method is adopted between the decoder and the encoder to achieve information fusion between different levels of the same-scale feature maps at the same stage, and a hierarchical feature fusion strategy is also adopted.

[0023] Both the decoder and the encoder introduce causal convolutional blocks for feature unmixing, aiming to remove the remaining mixtures in the features and help the model better extract stable features.

[0024] A segmentation method for a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging, the method comprising the following steps:

[0025] Step S1: Four input original image sequences are respectively sent to the encoder branch for convolution with a kernel size of 5×5×5 and a stride of 1 to generate four first feature maps with the same resolution as the original images and a channel number of 4.

[0026] Step S2: The four first feature maps are divided into two groups and respectively sent to the causal convolution module to generate two second feature maps with the same resolution as the original images and a channel number of 16, and two branches s11 and s12 are led out for use by the decoding end.

[0027] Step S3: Max-pooling with a stride of 2 is performed on the two second feature maps respectively, and then they are sent into the causal convolution block to generate a third feature map with a resolution of 1 / 2 of the original image and a channel number of 64, and a branch s2 is led out for use by the decoding end.

[0028] Step S4: Max-pooling with a stride of 2 is first performed on the third feature map, and then it is sent into a conventional convolution block to generate a fourth feature map with a resolution of 1 / 4 of the original image and a channel number of 256, and this output is denoted as s3 for use by the decoding end.

[0029] Step S5: The fourth feature map s3 is first passed through a transposed convolution with a kernel size of 2x2x2 and a stride of 2. The output of the transposed convolution is then concatenated with the branch output s2 led out from the third feature map in channels, and then they are sent into a conventional convolution block together to generate a fifth feature map with a resolution of 1 / 2 of the original image and a channel number of 32.

[0030] Step S6: The fifth feature map is first passed through a transposed convolution with a kernel size of 2x2x2 and a stride of 2. The output of the transposed convolution is divided into two branches and then concatenated with the two branch outputs s11 and s12 led out from the second feature map in channels respectively, and then they are sent into the causal convolution block together to generate a sixth feature map with the same resolution as the original image and a channel number of 64.

[0031] Step S7: Feed the sixth feature map into a conventional convolution block to generate a seventh feature map with the same resolution as the original image and 16 channels.

[0032] Step S8: Perform a convolution with a stride of 1×1×1 on the seventh feature map to generate a segmentation map with the same resolution as the original image and 4 channels.

[0033] Compared with the prior art, the present invention has the following improvements and advantages:

[0034] 1. By adopting a cross-layer connection method between the decoder and the encoder, information fusion between different levels of the same-scale feature maps is achieved at the same stage; at the same time, a causal convolution module is used to share parameters to obtain high-level semantic features, promoting the exchange and fusion of information between features, better capturing the common features between modalities, and thus improving the segmentation accuracy of the lightweight brain tumor segmentation network.

[0035] 2. By considering the correlation between different image modalities and using a hierarchical feature fusion strategy to eliminate the spurious correlations mixed during model training, reducing the number of model parameters and alleviating feature bias, the lightweight brain tumor segmentation network designed by the present invention has a faster inference speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is the overall framework diagram of the lightweight brain tumor segmentation network.

[0037] Figure 2 It is the structural schematic diagram of the semantic chain and the semantic causal chain designed by the present invention.

[0038] Figure 3 It is the structural diagram of the conventional convolution block and the causal convolution block designed by the present invention.

[0039] Figure 4 It is the overall flow chart of the method of the present invention.

[0040] Figure 5 It is the effect comparison diagram of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0041] The following further outlines the present invention in conjunction with the accompanying drawings.

[0042] As Figure 1 shown, a lightweight brain tumor segmentation network for multi-modal magnetic resonance imaging, the brain tumor segmentation network includes an encoder and a decoder, the encoder includes four layers of networks, namely the first feature processing layer, the second causal cross layer, the third causal cross layer, and the fourth feature fusion layer;

[0043] The first feature processing layer includes four 5×5×5 convolutions with a stride of 1, and the four convolutions process the input feature maps in parallel and share parameters; the second causal cross layer consists of two causal convolution modules, corresponding to four input branches. The causal convolution module uses a 3×3×3 convolution with a stride of 1, and the two causal convolution modules share parameters; the third causal cross layer consists of two parallel max pooling cascaded with a single causal convolution block. Both the second causal cross layer and the third causal cross layer adopt a parameter sharing strategy. The two feature maps each pass through a 3×3×3 convolution with a stride of 1 in parallel, and the output channels obtained are concatenated with the feature maps of the other party, and then each performs a normalization operation and inputs an activation function to obtain an output. Then each passes through a 3×3×3 convolution with a stride of 1 and then performs channel dimension concatenation, performs a normalization operation and an activation function to obtain a feature map with the same resolution as the original image; the fourth feature fusion layer includes max pooling cascaded with a conventional convolution block, and the max pooling cascaded with a conventional convolution block uses two 3×3×3 convolutions with a stride of 1;

[0044] The feature map first passes through a 3×3×3 convolution with a stride of 1, and then performs a normalization operation and inputs an activation function for output. Then it passes through a 3×3×3 convolution with a stride of 1 and then performs a normalization operation and an activation function to obtain a feature map with the same resolution as the original image.

[0045] After being processed by the first feature processing layer and the second causal cross layer, the four-branch features have been fused into one branch and will be sent to the next layer to further extract the fused features to obtain high-level semantic features; by considering the correlation between different image modalities and using a hierarchical feature fusion strategy to eliminate the spurious correlations that may be mixed in the feature maps during model training.

[0046] A parameter sharing strategy is adopted between the encoder branches. Cross connections are used between the branches to allow the information in the branches to flow and complement each other.

[0047] The decoder includes four layers of networks, namely the first output layer, the second feature processing layer, the third feature decoding layer, and the fourth output layer;

[0048] The first output layer includes a transposed convolution cascaded with a conventional convolution block. The first output layer receives the output from the fourth feature fusion layer, and after being convolved by the transposed convolution cascaded with a conventional convolution block, it is concatenated with the features of the same dimension flowing out of the third causal cross layer;

[0049] The second feature processing layer includes a transposed convolution cascaded with a conventional convolution block. It receives the output of the first output layer, performs convolution through the transposed convolution cascaded with the conventional convolution block and then conducts branch processing, concatenates the feature information flowing out of the second causal cross layer of the encoder, performs cross-processing between features with parameter sharing using a causal convolution block, and inputs the processed features into the third feature decoding layer; the transposed convolution cascaded with the conventional convolution block uses a transposed convolution of 2×2×2 with a stride of 2.

[0050] The third feature decoding layer includes a single 1×1×1 convolution block with a stride of 1 to further extract the features after removing the confounding; the fourth output layer includes a single 1×1×1 convolution block with a stride of 1 to output the final feature map.

[0051] The first output layer not only receives the output from the last layer of the encoder, but also concatenates the features of the same dimension in the decoder process. It is known that the separate processing of features in the encoder part is often more targeted; while the second feature processing layer cleverly concatenates the features flowing out of the encoder, it also performs a cross-processing between features with parameter sharing once. The processed features are input into the third layer of the encoder for feature encoding, and finally input into the last layer for the complete output of the segmentation model;

[0052] Furthermore, a cross-layer connection method is adopted between the decoder and the encoder to achieve information fusion between different levels of the same-scale feature maps in the same stage, and a hierarchical feature fusion strategy is also adopted; both the decoder and the encoder introduce causal convolution blocks for feature de-mixing, aiming to remove the remaining confounding in the features and help the model better extract stable features. A causal intervention feature double-branch cross-connection module is also used in the encoder for modality fusion and feature extraction; a parameter sharing strategy is adopted to reduce the number of model parameters and at the same time reduce feature bias; the lightweight brain tumor segmentation network can efficiently extract the modality information and detail information of the nuclear magnetic data and can be trained end-to-end. Compared with the recent mainstream segmentation networks, the designed lightweight brain tumor segmentation network architecture achieves higher detection accuracy and a lightweight model structure.

[0053] Furthermore, the causal cross-connection module is proposed based on the semantic causal chain. The semantic causal chain uses the Do operator for backdoor adjustment through the structural causal model SCM to intervene in the conventional segmentation semantic chain to remove confounding;

[0054] Such as Figure 2As shown, X represents data samples, Y represents true labels, M represents a model or module, I is the representation of the input of data at the model end, which is semantically equivalent to X; O is the output soft label of the model module, which is semantically equivalent to Y; F represents the feature map of data in the model, which is considered a specific representation of X or I under the context mixture C; C is the mixture factor existing in the feature map, which can be regarded as a context prior at the high-level semantics.

[0055] The mixture factor existing in the feature map is reflected in the fuzziness between feature maps during the model propagation of modalities in the multi-modal dataset, resulting in the false association of irrelevant pixels between feature maps during the classification of the supervised semantic segmentation model; there is a backdoor path between C, I, and O. By using the Do operator for backdoor adjustment through the structural causal model SCM and intervening in I, the causal connection between C and I is cut off.

[0056] As Figure 3 As shown, it is the structure diagram of the conventional convolution block and the causal convolution block designed by the present invention. The conventional convolution block is composed of two 3×3×3 convolution layers with a stride of 1, a normalization layer, and an activation layer in cascade. For the causal convolution block, after two parallel inputs respectively pass through a 3×3×3 convolution layer with a stride of 1, the output of the other branch is concatenated in channels, and then further passes through a normalization layer and an activation layer, and then respectively passes through a 3×3×3 convolution with a stride of 1. Finally, the outputs of the two branches are concatenated in channels, and then a normalization operation and an activation function are performed. Cross-connections are adopted between the branches of the causal convolution block, and parameter sharing is used.

[0057] In the above conventional semantic chain, the feature data are linearly concatenated and then sent into the module for feature extraction. Since multiple feature blocks are simply linearly stacked into different input channels of the convolution model, each input modality will be treated equally by the model at this time, and the semantic chain reflects a general cascaded convolution module.

[0058] The module, as a black-box system, is very likely to try to take shortcuts from the concatenated data to process the task. As a result, only the features that can most attract the attention of the module are utilized, and most features are discarded without being processed. By means of the mandatory branches of the convolution block, different attentions are aimed to be given to each feature block. And by adopting the method of cross-connection of cross-features, the problem of lack of information flow between branches is overcome to a certain extent, and the information communication between different features is simply and efficiently realized.

[0059] Causal cross - connection feature processing module designed based on the idea of semantic causal chain: Considering the different sensitivities of different image features to the segmented regions, the proposed causal network adaptively aggregates various features from different channel dimensions. Combining the above - mentioned semantic chain and semantic causal chain, by establishing cross - connections between different features, the exchange and fusion of information between features can be promoted. Introducing the information flow from the characteristics of other branches can provide the model with richer context and constraints. Parameter sharing between branches can enable some feature representations to be shared between different modalities, thus better capturing the common features between modalities, which includes the stability of causal relationships; at the same time, it also achieves the purpose of a lightweight causal feature extraction module.

[0060] As Figure 4 shown, a segmentation method for a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging, the method comprising the following steps:

[0061] Step S1: The four input original image sequences are respectively fed into the encoder branch for convolution with a kernel size of 5×5×5 and a stride of 1 to generate four first feature maps with the same resolution as the original images and a channel number of 4.

[0062] Step S2: The four first feature maps are divided into two groups and respectively fed into the causal convolution module to generate two second feature maps with the same resolution as the original images and a channel number of 16, and two branch outputs s11 and s12 are led out for use by the decoding end.

[0063] Step S3: The two second feature maps are respectively subjected to max - pooling with a stride of 2, and then fed into the causal convolution block to generate a third feature map with a resolution of 1 / 2 of the original image and a channel number of 64, and a branch output s2 is led out for use by the decoding end.

[0064] Step S4: The third feature map is first subjected to max - pooling with a stride of 2, and then fed into the conventional convolution block to generate a fourth feature map with a resolution of 1 / 4 of the original image and a channel number of 256, and this output is denoted as s3 for use by the decoding end.

[0065] Step S5: The fourth feature map s3 first passes through a transposed convolution with a kernel size of 2x2x2 and a stride of 2. The output of the transposed convolution is then concatenated with the branch output s2 led out from the third feature map, and then they are together fed into the conventional convolution block to generate a fifth feature map with a resolution of 1 / 2 of the original image and a channel number of 32.

[0066] Step S6: First, perform a transposed convolution with a kernel size of 2x2x2 and a stride of 2 on the fifth feature map. The output of the transposed convolution is divided into two branches, and then concatenated with the two branch outputs s11 and s12 derived from the second feature map along the channel dimension. Then, they are fed into the causal convolution block together to generate a sixth feature map with the same resolution as the original image and 64 channels.

[0067] Step S7: Feed the sixth feature map into a conventional convolution block to generate a seventh feature map with the same resolution as the original image and 16 channels.

[0068] Step S8: Perform a convolution with a kernel size of 1×1×1 and a stride of 1 on the seventh feature map to generate a segmentation map with the same resolution as the original image and 4 channels.

[0069] To verify the accuracy and implementation efficiency of the network designed in this invention, the model was trained, evaluated, and predicted on the widely used Brats2021 dataset:

[0070] The Brats2021 dataset includes 1251 patients. Each case in BraTS contains four modalities of magnetic resonance imaging, and the dimension of each modality is 240×240×155.

[0071] As Figure 5 shown, as Figure 5 shown, by improving the encoder-decoder structure under the guidance of causal theory, the segmentation network designed in this invention is very lightweight, with only 1.65M parameters; and it has a faster inference speed. After calculation, the inference time of the network designed in this invention is only 0.09S; the metrics are also much better than other current segmentation models. Benefiting from smaller model parameters and faster inference speed, it is more suitable for use in medical small device scenarios; experimental results show that the semantic causal chain proposed in this invention provides theoretical guidance for the subsequent design of the causal relationship module. The causal convolution module designed based on the causal semantic chain is applicable to both the encoder and decoder of the segmentation model, with high flexibility and plug-and-play; the causal convolution module can promote the exchange and fusion of information between features by establishing cross-links between different features. Introducing the information flow from the characteristics of other branches can provide richer context and constraints for the model; at the same time, due to the existence of causal relationships, parameter sharing between branches can enable different modalities to share some feature representations, thus better capturing the common features between modalities, which includes the stability of causal relationships. At the same time, the purpose of a lightweight causal feature extraction module is also achieved, and the network architecture designed in this invention has achieved higher segmentation accuracy than existing mainstream models.

[0072] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A method for constructing a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging, using convolutional blocks to build an end-to-end convolutional neural network to complete the segmentation task, the brain tumor segmentation network includes an encoder and a decoder, characterized in that: The encoder includes a four-layer network, namely a first feature processing layer, a second causal cross layer, a third causal cross layer and a fourth feature fusion layer; The decoder includes a four-layer network, namely a first output layer, a second feature processing layer, a third feature decoding layer and a fourth output layer; The encoder adopts a cross-layer connection method to achieve information fusion between different layers of the same scale feature map at the same stage, and transmits the fused information to the decoder. The decoder splices the encoder outflow features and performs cross processing between shared features; The first feature processing layer includes four 5×5×5 convolutions with a step size of 1, and the four convolutions process the four input modal original images in parallel; The second causal cross layer consists of two causal convolution modules, corresponding to four input branches and two output branches. The causal convolution module uses a 3×3×3 convolution with a step size of 1, and the two causal convolution modules share parameters. The third causal cross layer is composed of two parallel maximum pooling cascaded single causal convolution modules, and the second causal cross layer and the third causal cross layer both adopt a parameter sharing strategy; the two feature maps are each parallelized through a maximum pooling with a step size of 2, and then sent to the causal convolution module to obtain a feature map; The causal convolution module is proposed based on the semantic causal chain. The semantic causal chain uses the Do operator to perform backdoor adjustment through the structural causal model SCM to intervene in the conventional segmentation semantic chain to remove confusion; The first output layer includes a transposed convolution cascaded conventional convolution block, the first output layer receives the output from the fourth feature fusion layer, concatenates the features of the same dimension flowing out of the third causal cross layer after transposed convolution, and then processes it through a conventional convolution block; The second feature processing layer includes a transposed convolution cascade causal convolution module, which receives the output of the first output layer, and after transposed convolution, respectively concatenates the two feature information flowing out of the second causal cross layer of the encoder, and then performs cross processing between parameter sharing features through the causal convolution module, and inputs the processed features into the third feature encoding layer; The transposed convolution cascade conventional convolution block adopts a 2×2×2 transposed convolution with a step size of 2; The fourth feature fusion layer includes a maximum pooling cascade conventional convolution module, which uses maximum pooling and two 3×3×3 convolutions with a step size of 1; The feature map is first subjected to maximum pooling with a step size of 2, then passed through a 3×3×3 convolution with a step size of 1, and then normalized and input to the activation function output, and then passed through a 3×3×3 convolution with a step size of 1, and then normalized and activated. The processed feature map is obtained; The third feature decoding layer includes a single 1×1×1 convolution block with a step size of 1 to further extract the features after removing the contamination; The fourth output layer includes a single 1×1×1 convolution block with a stride of 1, which outputs the final segmentation map.

2. The method for constructing a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging according to claim 1, characterized in that: The decoder and the encoder use a cross-layer connection method to achieve information fusion between different levels of the same scale feature map at the same stage, and also use a hierarchical feature fusion strategy; Both the decoder and the encoder introduce causal convolution modules for feature unmixing, aiming to remove residual confusion in the features and help the model better extract stable features.

3. A segmentation method for a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging, characterized in that: The method is implemented based on a method for constructing a lightweight brain tumor segmentation network for multimodal magnetic resonance imaging according to any one of claims 1 to 2; the method comprises the following steps: Step S1: The four original image sequences are respectively input into the encoder branch for 5×5×5 convolution with a step size of 1 to generate four first feature maps with the same resolution as the original image and a channel number of 4; Step S2: The four first feature maps are divided into two groups and sent to the first causal convolution module respectively to generate two second feature maps with the same resolution as the original image and 16 channels, and two branch outputs s11 and s12 are derived for use by the decoding end; Step S3: Perform maximum pooling with a step size of 2 on the two second feature maps respectively, and then send them to the second causal convolution module to generate a third feature map with a resolution of 1 / 2 of the original image and 64 channels, and lead to a branch output s2 for use by the decoding end; Step S4: Perform maximum pooling with a step size of 2 on the third feature map, and then send it to the first regular convolution block to generate a fourth feature map with a resolution of 1 / 4 of the original image and a channel number of 256. This output is recorded as s3 for use by the decoding end; Step S5: The fourth feature map s3 is first subjected to a 2x2x2 transposed convolution with a step size of 2. The output of the transposed convolution is then channel-concatenated with the branch output s2 derived from the third feature map, and then sent together to the second regular convolution block to generate a fifth feature map with a resolution of 1 / 2 of the original image and a channel number of 32. Step S6: The fifth feature map is first subjected to a 2x2x2 transposed convolution with a step size of 2. The output of the transposed convolution is divided into two branches, and then the two branch outputs s11 and s12 derived from the second feature map are spliced ​​by channels respectively, and then sent together to the third causal convolution module to generate a sixth feature map with a resolution equal to that of the original image and a channel number of 64; Step S7: sending the sixth feature map to the third conventional convolution block to generate a seventh feature map having a resolution equal to that of the original image and 16 channels; Step S8: Perform a 1×1×1 convolution on the seventh feature map with a step size of 1 to generate a segmentation map with a resolution equal to that of the original image and a channel number of 4.

Citation Information

Patent Citations

  • Multi-modal fusion lightweight segmentation network and segmentation method for brain MRI image

    CN116740513A