Multi-modal segmentation method and system for brain tumor imaging
Through the multimodal segmentation method, multimodal fusion and dynamic convolution technology are used to automatically identify the location of brain tumors, solving the problem of manual recognition errors and improving the accuracy of brain tumor diagnosis.
Patent Information
- Application Number
- CN202510546969.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
During the diagnosis of brain tumors, manual identification is prone to errors and affects the treatment effect, especially when the doctor is not experienced enough.
The multimodal segmentation method is used to obtain brain magnetic resonance imaging of multiple modalities, preprocessing, dynamic convolution feature extraction, multimodal fusion and attention weighting, and position information of brain tumors is obtained.
It improves the accuracy of brain tumor recognition, reduces the error of manual recognition, and realizes automatic identification and accurate diagnosis of brain tumor location.
Smart Images

Figure CN120070480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical imaging technology, and particularly to a multi-modal segmentation method and system for brain tumor imaging. Background Art
[0002] Currently, when identifying brain tumors, generally clinicians will combine various types of medical imaging images for diagnosis to further determine the possible location and signs of brain tumors. For example, during the clinical process, magnetic resonance imaging (MRI) of the patient's brain is collected for the judgment and identification of brain tumors.
[0003] However, in the actual diagnosis process, since it is necessary to accurately identify the boundary between the lesion area and the rest of the image to identify the size and location of the brain tumor. When identifying manually, since the diagnosis result depends on the doctor's experience, when the doctor has insufficient experience, errors often occur, affecting subsequent treatment. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a multi-modal segmentation method and system for brain tumor imaging to solve the problem of errors in manual identification during the brain tumor diagnosis process. The specific technical solutions are as follows: In the first aspect of the embodiments of this application, a multi-modal segmentation method for brain tumor imaging is first provided. The method includes: Obtain magnetic resonance imaging of the brain in multiple modalities; Preprocess the magnetic resonance imaging of the brain in multiple modalities to obtain the preprocessing result of each modality; Through the attention layer, respectively identify the key regions of the preprocessing results of each modality to obtain the key regions corresponding to each modality; through multiple parallel dynamic convolution layers, extract features from the key regions corresponding to each modality to obtain the extracted features corresponding to each modality; respectively perform feature fusion on the extracted features corresponding to each modality to obtain combined features; Perform convolution and pooling processing on the combined features to obtain a first processing result; perform convolution and mapping processing on the combined features to obtain a second processing result; Perform multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features; Perform attention weighting on the multi-modal fusion features to obtain weighted features; through the weighted features, perform receptive field extraction to obtain the location information of the brain tumor.
[0005] In a possible implementation, the preprocessing result includes the region of interest (ROI) of each modality imaging. The preprocessing of the multi-modal brain magnetic resonance imaging to obtain the preprocessing result of each modality includes: Perform bias field correction on the multi-modal brain magnetic resonance imaging; Normalize the corrected imaging; Perform rigid registration on the normalized imaging; Identify the region of interest of the rigidly registered imaging to obtain the region of interest of each modality imaging.
[0006] In a possible implementation, the feature extraction of the key region corresponding to each modality through multiple parallel dynamic convolutional layers to obtain the extracted features corresponding to each modality includes: Randomly discard some neurons through a dropout layer, and use the remaining neurons to extract features from the key region corresponding to each modality to obtain the target extracted features corresponding to each modality; The feature fusion of the extracted features corresponding to each modality respectively to obtain the combined features includes: Perform feature fusion on the target extracted features corresponding to each modality respectively to obtain the combined features.
[0007] In a possible implementation, after the extraction of the receptive field through the weighted features to obtain the location information of the brain tumor, the method further includes: Judge the reliability of the location information of the brain tumor through voxel-level majority voting; When the reliability is greater than a preset threshold, determine the location information of the brain tumor as the calculation result and output it.
[0008] In a possible implementation, after the location information of the brain tumor is determined as the calculation result and output, the method further includes: Calculate the current loss according to the output result and the preset validation set information; Adjust the dynamic convolutional layer according to the calculated current loss.
[0009] In a possible implementation, the multi-modal brain magnetic resonance imaging includes: T1-weighted image, T1-enhanced image, T2-weighted image, fluid-attenuated inversion recovery (FLAIR) image.
[0010] In the second aspect of the embodiments of the present application, a multi-modal segmentation system for brain tumor imaging is provided. The system includes: A multi-modal input module for acquiring multi-modal brain magnetic resonance imaging; A preprocessing module for preprocessing brain magnetic resonance imaging of multiple modalities to obtain preprocessing results for each modality; A dynamic convolutional encoding module for respectively identifying key regions of the preprocessing results of each modality through an attention layer to obtain key regions corresponding to each modality; extracting features of the key regions corresponding to each modality through a plurality of parallel dynamic convolutional layers to obtain extracted features corresponding to each modality; respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features; A dual-channel decoupling module for performing convolutional and pooling processing on the combined features to obtain a first processing result; performing convolutional and mapping processing on the combined features to obtain a second processing result; A multi-modal fusion module for performing multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features; A dynamic convolutional decoding module for performing attention weighting on the multi-modal fusion features to obtain weighted features; extracting a receptive field through the weighted features to obtain the location information of the brain tumor.
[0011] In a possible implementation manner, the preprocessing module is specifically configured to perform bias field correction on the brain magnetic resonance imaging of the multiple modalities; perform normalization on the corrected imaging; perform rigid registration on the normalized imaging; perform identification of a region of interest on the rigidly registered imaging to obtain the region of interest of each modality imaging.
[0012] In a possible implementation manner, the dynamic convolutional encoding module is configured to randomly discard some neurons through a dropout layer, and extract features of the key regions corresponding to each modality through the remaining neurons to obtain target extracted features corresponding to each modality; the respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features includes: respectively performing feature fusion on the target extracted features corresponding to each modality to obtain the combined features.
[0013] In a possible implementation manner, the system further includes: A three-branch output module for judging the reliability of the location information of the brain tumor through voxel-level majority voting; when the reliability is greater than a preset threshold, determining the location information of the brain tumor as a calculation result and outputting it.
[0014] In a possible implementation manner, the system further includes: A loss function calculation module for calculating a current loss according to the output result and preset validation set information; Adjusting the dynamic convolutional layer according to the calculated current loss.
[0015] In a possible implementation, the multi-modal brain magnetic resonance imaging includes: T1-weighted image, T1-enhanced image, T2-weighted image, fluid-attenuated inversion recovery image FLAIR.
[0016] On the other hand, an embodiment of the present application further provides an electronic device, including: A memory for storing a computer program; A processor, when executing the program stored on the memory, implements any of the above multi-modal segmentation methods for brain tumor imaging.
[0017] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any of the above multi-modal segmentation methods for brain tumor imaging is implemented.
[0018] On the other hand, an embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the above multi-modal segmentation methods for brain tumor imaging.
[0019] Beneficial effects of the embodiments of the present application: The embodiments of the present application provide a multi-modal segmentation method and system for brain tumor imaging. The method includes: acquiring multi-modal brain magnetic resonance imaging; preprocessing the multi-modal brain magnetic resonance imaging to obtain a preprocessing result for each modality; identifying key regions of the preprocessing result of each modality through an attention layer to obtain key regions corresponding to each modality; extracting features of the key regions corresponding to each modality through a plurality of parallel dynamic convolution layers to obtain extracted features corresponding to each modality; respectively performing feature fusion on the extracted features corresponding to each modality to obtain a combined feature; performing convolution and pooling processing on the combined feature to obtain a first processing result; performing convolution and mapping processing on the combined feature to obtain a second processing result; performing multi-modal fusion on the first processing result and the second processing result to obtain a multi-modal fusion feature; performing attention weighting on the multi-modal fusion feature to obtain a weighted feature; and extracting a receptive field through the weighted feature to obtain the position information of the brain tumor. Through the method of the embodiments of the present application, automatic identification of the position of the brain tumor in brain imaging can be achieved, thereby solving the problem of errors caused by manual identification and improving the accuracy of brain tumor identification.
[0020] Of course, implementing any product or method of the present application does not necessarily require achieving all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.
[0022] Figure 1a It is a schematic flowchart of a multi-modal segmentation method for brain tumor imaging provided by an embodiment of the present application; Figure 1b It is a schematic flowchart of the dual-channel decoupling provided by an embodiment of the present application; Figure 1c It is a schematic structural diagram of a multi-modal fusion module provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a preprocessing provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of feature extraction for key regions provided by an embodiment of the present application; Figure 4 It is a schematic flowchart of voxel-level majority voting provided by an embodiment of the present application; Figure 5 It is a schematic flowchart of dynamic convolution layer adjustment provided by an embodiment of the present application; Figure 6 It is a schematic flowchart of dynamic convolution kernel generation provided by an embodiment of the present application; Figure 7 It is a schematic diagram of the system architecture corresponding to the multi-modal segmentation method for brain tumor imaging provided by an embodiment of the present application; Figure 8 It is a schematic structural diagram of a multi-modal segmentation system for brain tumor imaging provided by an embodiment of the present application; Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0024] To solve the problem in the prior art that when manually identifying, since the diagnostic result depends on the doctor's experience, when the doctor has insufficient experience, errors often occur, affecting subsequent treatment In the first aspect of the embodiments of the present application, first, a multi-modal segmentation method for brain tumor imaging is provided. Refer to Figure 1a , Figure 1a which is a schematic flowchart of a multi-modal segmentation method for brain tumor imaging provided by the embodiments of the present application. The method includes: Step S11: Obtain magnetic resonance imaging of the brain in multiple modalities; Step S12: Preprocess the magnetic resonance imaging of the brain in multiple modalities to obtain the preprocessing result of each modality; Step S13: Through the attention layer, respectively identify the key regions of the preprocessing results of each modality to obtain the key regions corresponding to each modality; through multiple parallel dynamic convolution layers, extract features from the key regions corresponding to each modality to obtain the extracted features corresponding to each modality; respectively perform feature fusion on the extracted features corresponding to each modality to obtain combined features; Step S14: Perform convolution and pooling processing on the combined features to obtain a first processing result; perform convolution and mapping processing on the combined features to obtain a second processing result; Step S15: Perform multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features; Step S16: Perform attention weighting on the multi-modal fusion features to obtain weighted features; through the weighted features, extract the receptive field to obtain the position information of the brain tumor.
[0025] Corresponding to the above step S11, when obtaining multi-modal brain magnetic resonance imaging, the brain MRI (Magnetic Resonance Imaging) of the patient can be obtained. Specifically, the brain magnetic resonance imaging in the embodiments of the present application is multi-modal imaging. In a possible implementation manner, the multi-modal brain magnetic resonance imaging includes: T1-weighted imaging, Contrast-enhanced T1-weighted imaging, T2-weighted imaging, and Fluid Attenuated Inversion Recovery (FLAIR). Among them, the T1-weighted image is based on the longitudinal relaxation time (T1 time) of the tissue and reflects the speed at which the tissue recovers longitudinal magnetization; the Contrast-enhanced T1-weighted image is T1-weighted imaging performed after intravenous injection of a gadolinium contrast agent, and the contrast agent enters the lesion area through the damaged blood-brain barrier; the T2-weighted image is based on the transverse relaxation time (T2 time) of the tissue and reflects the attenuation speed of transverse magnetization; FLAIR suppresses the signal of free water (such as cerebrospinal fluid) to prominently display the signal of bound water (such as edema or lesion). Among them, the Contrast-enhanced T1-weighted image can highlight the tumor boundary, the T2-weighted image can show the edema area, and FLAIR can show the suppression of cerebrospinal fluid signal. In the embodiments of the present application, the recognition rate of low-contrast regions can be improved through multi-modal information complementarity. In the actual use process, the resolution of the obtained imaging can also be set for subsequent processing. In a possible implementation manner, the resolution of the imaging can be 0.5mm×0.5mm×0.5mm, 1mm×1mm×1mm, 0.5mm×0.5mm×1mm, etc.
[0026] Corresponding to the above step S12, preprocessing of multi-modal brain magnetic resonance imaging can be performed by various methods, such as registration, normalization, etc. Since in the actual use process, the brain magnetic resonance imaging obtained from different hospitals, devices, or patients may have deviations, in the embodiments of the present application, before identifying brain tumors, preprocessing the obtained multi-modal brain magnetic resonance imaging can standardize the obtained imaging, thereby facilitating subsequent processing and reducing errors in the recognition process. Specifically, when preprocessing multi-modal brain magnetic resonance imaging, operations such as resolution adjustment, image normalization, registration, and extraction of regions of interest can be performed to obtain the preprocessing results of each modality.
[0027] Corresponding to the above step S13, this step is mainly used for dynamic convolution feature extraction. Specifically, through the attention layer, that is, the Attention layer, the key regions of the preprocessing results of each modality are respectively identified to obtain the key regions corresponding to each modality. Among them, the key regions identified through the Attention layer can be regions of interest. When identifying the key regions, threshold segmentation can be used, such as single / multi-band thresholding method; and feature classification, histogram optimization and other methods for identification. Through this attention layer, limited attention can be concentrated on key information, that is, the key regions corresponding to each modality, so as to save resources and quickly obtain the most effective information. Then, for the obtained key regions, through multiple parallel dynamic convolution layers, feature extraction is performed on the key regions corresponding to each modality. Then, through multiple parallel dynamic convolution layers, feature extraction is performed on the key regions corresponding to each modality to obtain the extracted features corresponding to each modality. It should be noted that when performing feature extraction on the key regions corresponding to each modality in this application, a dynamic convolution layer is used. The weights of the convolution kernels in this dynamic convolution layer are not fixed, but are dynamically adjusted according to the input features of each time step, which means that the convolution kernels can change according to the input changes, thus providing more flexible and adaptive feature extraction. Compared with the traditional fixed convolution kernel that cannot adapt to the dynamic changes of tumor morphology, the dynamic convolution kernel can adjust the receptive field according to local features. For example, the convolution kernel size is increased at the tumor boundary to capture richer context information. The fitting ability for complex morphologies is enhanced through dynamic parameter adaptation. Finally, by respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features, the fusion of features in multimodal imaging can be realized. Due to the different characteristics corresponding to each modality, through multimodal fusion, the recognition ability of tumor boundaries and heterogeneous features can be strengthened, and the accuracy of brain tumor recognition can be improved.
[0028] Corresponding to the above step S14, this step mainly extracts features through the COS-SOS (dual-channel decoupling) attention mechanism. Specifically, it can be processed in parallel by a channel attention branch and a spatial attention branch. Among them, the channel branch can perform convolution and pooling operations on the combined features to obtain a first processing result; the spatial attention branch can perform convolution and mapping operations on the combined features to obtain a second processing result. In one example, the channel branch uses 1×1×1 convolution and global pooling, and the spatial branch uses 1×1×1 convolution and the Sigmoid function (activation function) for mapping processing. In the embodiments of the present application, through the COS-SOS attention mechanism, the attention calculation in the channel and spatial dimensions can be separated, reducing redundancy and optimizing the feature weight distribution. The inventor's research found that through dual-channel decoupling, the computational complexity is reduced from O(C²HWD) to O(CHWD), where O represents the algorithm complexity, C represents the number of channels, H represents the height, W represents the width, and D represents the depth; real-time segmentation can be achieved while maintaining the accuracy, which is applicable to the clinical rapid diagnosis scenario. Refer to Figure 1b , Figure 1b is a schematic flowchart of the dual-channel decoupling provided by the embodiments of the present application. The input parameter X is processed in two paths. One path is subjected to convolution processing through a (1×1×1) convolutional layer, and after global pooling, matrix transformation, and normalization, it passes through calculation, that is, matrix dot product operation. Finally, after matrix transformation and activation function calculation, it passes through ⊙ calculation, that is, the multiplication operator in the spatial dimension, to obtain the calculation result of the first path; the other path is subjected to convolution processing through a (1×1×1) convolutional layer, and after global pooling, matrix transformation, and normalization, it passes through calculation, that is, matrix dot product operation. Finally, after convolution, normalization, and activation function calculation, it passes through ⊙ calculation, that is, the multiplication operator in the channel dimension, to obtain the calculation result of the second path; finally, through the calculation of ⊕ calculation, that is, calculating the direct sum, the fusion of the two calculation results is performed.
[0029] Corresponding to the above step S15, this step can achieve multimodal fusion. Specifically, the first processing result and the second processing result can be subjected to multimodal fusion to obtain multimodal fusion features. Among them, when performing multimodal fusion, fusion can be carried out by various methods, such as feature splicing, superposition, weighted summation and other methods for feature fusion. Specifically, it can be set according to the actual situation. In the embodiments of the present application, the method of feature splicing is preferably used for fusion. Among them, in the embodiments of the present application, through multimodal fusion, the recognition rate of low-contrast regions can be improved through multimodal information complementarity. The inventor's research found that through the fusion of multiple modalities, it is convenient to correct errors in the subsequent voting process when a certain modality misjudges the necrotic region. Among them, in the cross-modal test, the fluctuation of DSC (Dice similarity coefficient) is <2%, which is 5.7% higher than that of the single-modal model; the segmentation sensitivity of low-contrast regions reaches 86.9%, effectively coping with the challenge of tumor heterogeneity. See Figure 1c , Figure 1c FIG. Figure 1c is a schematic structural diagram of the multimodal fusion module provided by the embodiment of the present application. When the brain magnetic resonance imaging of multiple modalities in the embodiment of the present application is four modalities, the four-modal input can be performed, and the images of the four modalities can be input into a three-level fusion architecture. First, feature concatenation is performed to obtain a cascaded feature architecture of 4C×H×W×D, where C represents the number of channels, H represents the height, W represents the width, and D represents the depth; then attention weighting is performed to obtain weighted features; then dynamic convolution decoding is performed to obtain the segmentation maps of each modality, and then single-modal segmentation is performed according to the segmentation maps of various modalities to obtain a segmentation result structure of 4×H×W×D; finally, through decision fusion, after obtaining the fusion result, the final segmentation result is obtained.
[0030] Corresponding to the above step S16, by performing attention weighting on the multimodal fusion features, through attention weighting, the relevance of each part of the input data to the current task can be calculated, such as vector dot product or similarity function, to assign weights to different positions, and a more focused context representation can be generated after weighted summation. By performing receptive field extraction on the weighted features, the location information of the brain tumor can be obtained. Specifically, in the actual use process, the location information can be the point cloud information corresponding to the location of the brain tumor, and the location, size, and boundary of the brain tumor can be determined through the point cloud information.
[0031] It can be seen that through the method of the embodiments of the present application, multi-modal brain magnetic resonance imaging can be obtained, and then the obtained imaging is preprocessed. Then, dynamic convolution coding and feature extraction are performed on the preprocessed imaging, and then multi-modal fusion is performed on the extracted features. Thus, attention weighting and receptive field extraction are performed through the fused features to obtain the location information of the brain tumor. Therefore, through the parameter self-adaptation of dynamic convolution, the decoupled design of dual-channel attention, and the information complementarity of multi-modal fusion, the bottlenecks of traditional methods in terms of accuracy, efficiency, and robustness are systematically solved. Moreover, by combining the lightweight dynamic network with multi-modal information integration, while improving the performance of medical image segmentation, computational efficiency and clinical practicability are taken into account, providing a new technical solution for the precise diagnosis and treatment of brain tumors.
[0032] In a possible implementation manner, referring to Figure 2 , the preprocessing result in step S12 includes the region of interest of each modal imaging. The preprocessing of the multi-modal brain magnetic resonance imaging to obtain the preprocessing result of each modal includes: Step S121, perform bias field correction on the multi-modal brain magnetic resonance imaging; Step S122, perform normalization on the corrected imaging; Step S123, perform rigid registration on the normalized imaging; Step S124, perform identification of the region of interest on the rigidly registered imaging to obtain the region of interest of each modal imaging.
[0033] Among them, for the multi-modal brain magnetic resonance imaging, bias field correction can be performed, and N4 bias field (N4BiasFieldCorrection) correction can be carried out to correct the signal intensity change caused by magnetic field inhomogeneity. Then, normalization can be performed on the corrected imaging, and Z-score normalization can be carried out. Through Z-score normalization, the corrected imaging can be unified to the same standard scale. Specifically, all pixels in the imaging can be unified to the interval [-1, 1], and the pixel values in the interval [-1, 1] corresponding to the corrected imaging can be obtained. Then, rigid registration is performed on the normalized imaging. Specifically, rigid registration can be carried out through ANTS (Advanced Normalization Tools). During the registration process, the mutual information threshold can be set to >0.8. Finally, ROI (region of interest) identification is performed on the rigidly registered imaging. Specifically, the Otsu threshold method (Otsu algorithm) can be used. This algorithm is an adaptive binary segmentation algorithm based on the gray distribution of the image. The optimal threshold can be determined by maximizing the between-class variance, and through three-dimensional connected component analysis, the regions composed of pixels or voxels with similar properties in the image can be identified and segmented, so as to obtain the regions of interest of each modality imaging. In a possible implementation manner, a fixed output size can be set to facilitate subsequent processing, such as 128×128×128 pixels, 224×224×128 pixels, etc.
[0034] It can be seen that through the method of the embodiment of the present application, before identifying brain tumors, preprocessing can be performed on the obtained multi-modal brain magnetic resonance imaging, and standardization of the obtained imaging can be achieved, thereby facilitating subsequent processing and reducing errors in the identification process.
[0035] In a possible implementation manner, referring to Figure 3 , the feature extraction of the key area corresponding to each modality is performed through multiple parallel dynamic convolution layers to obtain the extraction features corresponding to each modality, including: Step S31, randomly discard some neurons through the dropout layer, and perform feature extraction on the key area corresponding to each modality through the remaining neurons to obtain the target extraction features corresponding to each modality; The feature fusion of the extraction features corresponding to each modality is respectively performed to obtain the combined features, including: Step S32, respectively perform feature fusion on the target extraction features corresponding to each modality to obtain the combined features.
[0036] In deep learning, when there are too many parameters and the training samples are relatively few, the model is prone to overfitting. Overfitting is a common problem in many deep learning and even machine learning algorithms, specifically manifested as high prediction accuracy on the training set and a significant drop in accuracy on the test set. Therefore, in the embodiments of this application, by using a Dropout layer to randomly discard some neurons, and then using the remaining neurons to extract features from the key regions corresponding to each modality, the target extraction features corresponding to each modality are obtained. Then, feature fusion is performed on the target extraction features corresponding to each modality respectively to obtain the combined features. During training, each neuron is retained with probability p, that is, it stops working with probability 1 - p. The neurons retained during each forward propagation are different, which can make the model less dependent on certain local features and have stronger generalization performance. In the actual use process, through the method of the embodiments of this application, the convolutional weights are adjusted according to the feature weights obtained by the attention mechanism, and Dropout is added to improve the generalization ability and convergence speed of the model. On the BraTS2021 (Brain Tumor Segmentation 2021) dataset, the Dice (a set similarity metric function) coefficient of the tumor core (TC) reaches 91.2%, which is 4.1% higher than that of nnUNetv2 (the second-generation automated medical image segmentation neural network framework).
[0037] In a possible implementation manner, after obtaining the position information of the brain tumor by performing receptive field extraction through the weighted features, refer to Figure 4 , the method further includes: Step S41, determining the reliability of the position information of the brain tumor through voxel-level majority voting; Step S42, when the reliability is greater than a preset threshold, determining the position information of the brain tumor as the calculation result and outputting it.
[0038] In the embodiments of this application, after obtaining the position information of the brain tumor by performing receptive field extraction through the weighted features, determining the reliability of the position information of the brain tumor through voxel-level majority voting can implement a three-level fusion architecture including feature concatenation, channel attention weighting (dynamically allocating modality weights), and voxel-level majority voting (three-dimensional spatial consistency constraint). The multi-modal MRI information is integrated. Among them, through voxel-level majority voting, the results of multiple predictions can be integrated to improve the segmentation accuracy and robustness. At each voxel (three-dimensional pixel) position, different prediction results are statistically counted, and the category with the highest frequency of occurrence is selected as the final output. Thus, combined with the multi-modal imaging in this application, when a certain modality misjudges the necrotic area, the voting results of other modalities can correct the error and improve the recognition accuracy.
[0039] In a possible implementation manner, after determining the position information of the brain tumor as a calculation result and outputting it, refer to Figure 5 , the method further includes: Step S51, calculating a current loss according to the output result and preset verification set information; Step S52, adjusting the dynamic convolution layer according to the calculated current loss.
[0040] Refer to Figure 6 , the working principle of the dynamic convolution in the embodiment of the present application is that the input features complete basic feature extraction and normalization through a triple operation, and then enter the parallel convolution layer. By adaptively using different numbers of convolution kernels, differential feature maps are generated. In this structure, the input parameter X is weighted by attention through the attention layer, then convolution calculations are performed in parallel through multiple convolution layers, then convolution calculations are performed through the convolution layer, then partial neurons are discarded through the dropout layer, and then through normalization processing and activation function calculation, the output parameter Y is obtained. Refer to Figure 7 , after multi-modal input and input of multi-modal imaging; preprocessing the imaging through the preprocessing module; then performing convolution calculation on the preprocessed imaging through dynamic convolution encoding; then performing calculation extraction of multi-modal features through dual-channel decoupling; then performing feature fusion through multi-modal fusion; then determining the receptive field through dynamic convolution decoding; finally, through three-branch output, performing three-dimensional voxel-level majority voting and output; finally, after calculating through the loss function and determining the position information of the brain tumor as a calculation result and outputting it, the current loss can be calculated according to the output result and preset verification set information, so as to adjust the dynamic convolution layer according to the calculated current loss. Specifically, the loss can be calculated through a variety of loss functions, such as mean square error loss function, cross-entropy loss function, logarithmic loss function, etc. By adjusting the dynamic convolution layer according to the calculated current loss, the adjustment of the dynamic convolution layer can be realized, so that the receptive field can be adjusted according to local features. For example, the convolution kernel size is increased at the tumor boundary to capture richer context information, and the fitting ability for complex shapes is enhanced through dynamic parameter adaptation.
[0041] In the second aspect of the embodiment of the present application, a multi-modal segmentation system for brain tumor imaging is provided. Refer to Figure 8 , the system includes: A multi-modal input module 801, configured to obtain multi-modal brain magnetic resonance imaging; A preprocessing module 802, configured to preprocess the multi-modal brain magnetic resonance imaging to obtain a preprocessing result for each modality; The dynamic convolutional encoding module 803 is used to identify key regions of the preprocessing results of each modality through an attention layer, obtaining the key regions corresponding to each modality; through multiple parallel dynamic convolutional layers, extracting features of the key regions corresponding to each modality to obtain the extracted features corresponding to each modality; and respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features. The dual-channel decoupling module 804 is used to perform convolutional and pooling processing on the combined features to obtain a first processing result; and perform convolutional and mapping processing on the combined features to obtain a second processing result. The multi-modal fusion module 805 is used to perform multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features. The dynamic convolutional decoding module 806 is used to perform attention weighting on the multi-modal fusion features to obtain weighted features; and extract the receptive field through the weighted features to obtain the location information of the brain tumor.
[0042] In a possible implementation manner, the preprocessing module is specifically used to perform bias field correction on the brain magnetic resonance imaging of the multiple modalities; perform normalization on the corrected imaging; perform rigid registration on the normalized imaging; and identify the region of interest of the rigidly registered imaging to obtain the region of interest of each modality imaging.
[0043] In a possible implementation manner, the dynamic convolutional encoding module is used to randomly discard some neurons through a dropout layer, and extract features of the key regions corresponding to each modality through the remaining neurons to obtain the target extracted features corresponding to each modality; and the respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features includes: respectively performing feature fusion on the target extracted features corresponding to each modality to obtain the combined features.
[0044] In a possible implementation manner, the system further includes: The three-branch output module is used to judge the reliability of the location information of the brain tumor through voxel-level majority voting; when the reliability is greater than a preset threshold, determining the location information of the brain tumor as the calculation result and outputting it.
[0045] In a possible implementation manner, the system further includes: The loss function calculation module is used to calculate the current loss according to the output result and the preset validation set information. Adjust the dynamic convolutional layer according to the calculated current loss.
[0046] In a possible implementation, the multiple-modal brain magnetic resonance imaging includes: T1-weighted image, T1-enhanced image, T2-weighted image, fluid-attenuated inversion recovery image FLAIR.
[0047] It can be seen that through the method of the embodiments of the present application, multiple-modal brain magnetic resonance imaging can be obtained, then the obtained imaging is preprocessed, then the preprocessed imaging is subjected to dynamic convolution coding and feature extraction, and then the extracted features are subjected to multi-modal fusion, so as to perform attention weighting and receptive field extraction through the fused features to obtain the position information of the brain tumor. Therefore, through the parameter self-adaptation of dynamic convolution, the decoupled design of dual-channel attention, and the information complementarity of multi-modal fusion, the bottlenecks of traditional methods in terms of accuracy, efficiency, and robustness are systematically solved. And by combining the lightweight dynamic network with multi-modal information integration, while improving the performance of medical image segmentation, the computational efficiency and clinical practicability are taken into account, providing a new technical solution for the precise diagnosis and treatment of brain tumors.
[0048] The embodiments of the present application also provide an electronic device, as Figure 9 shown, including: A memory 901 for storing a computer program; A processor 902, when executing the program stored on the memory 901, implements the following steps: Obtain multiple-modal brain magnetic resonance imaging; Preprocess the multiple-modal brain magnetic resonance imaging to obtain the preprocessing result of each modality; Through an attention layer, respectively identify the key regions of the preprocessing results of each modality to obtain the key regions corresponding to each modality; through multiple parallel dynamic convolution layers, extract features from the key regions corresponding to each modality to obtain the extracted features corresponding to each modality; respectively perform feature fusion on the extracted features corresponding to each modality to obtain a combined feature; Perform convolution and pooling processing on the combined feature to obtain a first processing result; perform convolution and mapping processing on the combined feature to obtain a second processing result; Perform multi-modal fusion on the first processing result and the second processing result to obtain a multi-modal fusion feature; Perform attention weighting on the multi-modal fusion feature to obtain a weighted feature; through the weighted feature, perform receptive field extraction to obtain the position information of the brain tumor.
[0049] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0050] The communication interface is used for communication between the above electronic device and other devices.
[0051] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0052] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0053] In another embodiment provided by the present application, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above multi-modal segmentation methods for brain tumor imaging are implemented.
[0054] In another embodiment provided by the present application, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute any of the multi-modal segmentation methods for brain tumor imaging in the above embodiments.
[0055] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0056] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0057] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system, electronic device, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0058] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A multimodal segmentation method for brain tumor imaging, characterized in that: The method comprises: Acquire brain magnetic resonance imaging of multiple modalities; Preprocessing brain magnetic resonance imaging of multiple modalities to obtain preprocessing results of each modality; Through the attention layer, the key areas of the preprocessing results of each modality are identified respectively to obtain the key areas corresponding to each modality; through multiple parallel dynamic convolution layers, the key areas corresponding to each modality are extracted to obtain the extracted features corresponding to each modality; the extracted features corresponding to each modality are fused to obtain the combined features; Performing convolution and pooling processing on the combined features to obtain a first processing result; performing convolution and mapping processing on the combined features to obtain a second processing result; Performing multimodal fusion on the first processing result and the second processing result to obtain a multimodal fusion feature; Attention weighting is performed on the multimodal fusion features to obtain weighted features; receptive field extraction is performed through the weighted features to obtain location information of the brain tumor.
2. The method according to claim 1, characterized in that The preprocessing result includes the region of interest of each modality imaging, and the preprocessing of the brain magnetic resonance imaging of multiple modalities is performed to obtain the preprocessing result of each modality, including: Performing bias field correction on the brain magnetic resonance imaging of the multiple modalities; Normalize the corrected imaging; Perform rigid registration on the normalized images; The region of interest is identified for the rigidly registered imaging to obtain the region of interest for each imaging modality.
3. The method according to claim 1, characterized in that The method of extracting features of the key areas corresponding to each modality through multiple parallel dynamic convolutional layers to obtain extracted features corresponding to each modality includes: Randomly discard some neurons through the discard layer, and extract features of the key areas corresponding to each modality through the remaining neurons to obtain target extraction features corresponding to each modality; The extracting features corresponding to each mode are respectively fused to obtain combined features, including: Feature fusion is performed on the target extraction features corresponding to each modality to obtain the combined features.
4. The method according to claim 1, characterized in that: After extracting the receptive field through the weighted features to obtain the location information of the brain tumor, the method further includes: Determining the reliability of the location information of the brain tumor by voxel-level majority voting; When the reliability is greater than a preset threshold, the location information of the brain tumor is determined as a calculation result and output.
5. The method according to claim 4, characterized in that After determining the location information of the brain tumor as a calculation result and outputting it, the method further includes: Calculate the current loss based on the output result and the preset validation set information; The dynamic convolution layer is adjusted according to the calculated current loss.
6. The method according to claim 1, characterized in that The multiple modalities of brain magnetic resonance imaging include: T1 weighted image, T1 enhanced image, T2 weighted image, and fluid attenuated inversion recovery image FLAIR.
7. A multimodal segmentation system for brain tumor imaging, characterized in that: The system comprises: A multimodal input module for acquiring brain magnetic resonance imaging in multiple modalities; A preprocessing module is used to preprocess brain magnetic resonance imaging of multiple modalities to obtain preprocessing results of each modality; A dynamic convolutional coding module is used to identify key areas of the preprocessing results of each modality through an attention layer to obtain key areas corresponding to each modality; extract features of the key areas corresponding to each modality through multiple parallel dynamic convolutional layers to obtain extracted features corresponding to each modality; and fuse the extracted features corresponding to each modality to obtain combined features. A dual-channel decoupling module is used to perform convolution and pooling processing on the combined features to obtain a first processing result; and to perform convolution and mapping processing on the combined features to obtain a second processing result; A multimodal fusion module, used for performing multimodal fusion on the first processing result and the second processing result to obtain a multimodal fusion feature; The dynamic convolution decoding module is used to perform attention weighting on the multimodal fusion features to obtain weighted features; and extract the receptive field through the weighted features to obtain the location information of the brain tumor.
8. The system according to claim 7, characterized in that The system further comprises: The three-branch output module is used to determine the reliability of the location information of the brain tumor through voxel-level majority voting; when the reliability is greater than a preset threshold, the location information of the brain tumor is determined as a calculation result and output.
9. The system according to claim 8, characterized in that The system further comprises: A loss function calculation module, used to calculate the current loss based on the output result and preset verification set information; The dynamic convolution layer is adjusted according to the calculated current loss.
10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-6 when executing a program stored in a memory.
Citation Information
Patent Citations
MRI (Magnetic Resonance Imaging) brain tumor localization and intratumoral segmentation method based on deep cascaded convolution network
CN108492297A
Multi-modal brain tumor image segmentation method and device, electronic equipment and storage medium
CN117576387A
Multi-modal imaging tumor positioning method and system
CN119027504A
Multi-modal brain tumor image segmentation system based on sparse attention mechanism
CN119323583A
MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-modal feature fusion
CN119850961A
Cited By
Multi-mode cerebral arterial thrombosis medical image segmentation method, device and equipment
CN121353683A