A multi-modal segmentation method and system for brain tumor imaging
Through the multimodal segmentation method, brain magnetic resonance imaging of multiple modes is used for preprocessing and feature fusion, combined with dynamic convolution and attention weighting, automatic identification of brain tumor location is achieved, error problems caused by insufficient doctors' experience are solved, and recognition accuracy is improved.
Patent Information
- Application Number
- CN202510546969.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In the prior art, brain tumor diagnosis depends on the experience of doctors, resulting in identification errors and affecting the treatment effect.
The multimodal segmentation method is adopted to obtain brain magnetic resonance imaging of multiple modalities, preprocessing, key area recognition, feature extraction and fusion, combined with dynamic convolution and attention weighting, automatic identification of brain tumor location is achieved.
It improves the accuracy of brain tumor recognition, reduces the error of manual recognition, and is suitable for rapid clinical diagnosis.
Smart Images

Figure CN120070480B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical imaging technology, and in particular, to a multi-modal segmentation method and system for brain tumor imaging. Background Art
[0002] Currently, when identifying brain tumors, generally clinicians will combine various types of medical imaging images for diagnosis to further determine the possible location and signs of the brain tumor. For example, in the clinical process, magnetic resonance imaging (MRI) of the patient's brain is collected for the judgment and identification of brain tumors.
[0003] However, in the actual diagnosis process, since it is necessary to accurately identify the boundary between the lesion area and the rest of the image to identify the size and location of the brain tumor. When identifying manually, since the diagnosis result depends on the doctor's experience, when the doctor lacks experience, errors often occur, affecting subsequent treatment. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a multi-modal segmentation method and system for brain tumor imaging to solve the problem of errors in manual identification during the brain tumor diagnosis process. The specific technical solutions are as follows:
[0005] In the first aspect of the embodiments of the present application, first, a multi-modal segmentation method for brain tumor imaging is provided. The method includes:
[0006] Obtain magnetic resonance imaging of the brain in multiple modalities;
[0007] Preprocess the magnetic resonance imaging of the brain in multiple modalities to obtain the preprocessing result of each modality;
[0008] Through the attention layer, respectively identify the key regions of the preprocessing results of each modality to obtain the key regions corresponding to each modality; through multiple parallel dynamic convolution layers, extract features from the key regions corresponding to each modality to obtain the extracted features corresponding to each modality; respectively perform feature fusion on the extracted features corresponding to each modality to obtain combined features;
[0009] Perform convolution and pooling processing on the combined features to obtain a first processing result; perform convolution and mapping processing on the combined features to obtain a second processing result;
[0010] Perform multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features;
[0011] Perform attention weighting on the multi-modal fusion features to obtain weighted features; through the weighted features, perform receptive field extraction to obtain the location information of the brain tumor.
[0012] In a possible implementation manner, the preprocessing result includes the region of interest of each modality imaging. The preprocessing of the brain magnetic resonance imaging of multiple modalities to obtain the preprocessing result of each modality includes:
[0013] Perform bias field correction on the brain magnetic resonance imaging of the multiple modalities;
[0014] Perform normalization on the corrected imaging;
[0015] Perform rigid registration on the normalized imaging;
[0016] Perform identification of the region of interest on the rigidly registered imaging to obtain the region of interest of each modality imaging.
[0017] In a possible implementation manner, the feature extraction of the key area corresponding to each modality through multiple parallel dynamic convolution layers to obtain the extraction feature corresponding to each modality includes:
[0018] Randomly discard some neurons through a dropout layer, and perform feature extraction on the key area corresponding to each modality through the remaining neurons to obtain the target extraction feature corresponding to each modality;
[0019] The feature fusion of the extraction features corresponding to each modality respectively to obtain a combined feature includes:
[0020] Perform feature fusion on the target extraction features corresponding to each modality respectively to obtain the combined feature.
[0021] In a possible implementation manner, after the location information of the brain tumor is obtained by performing receptive field extraction through the weighted features, the method further includes:
[0022] Judge the reliability of the location information of the brain tumor through voxel-level majority voting;
[0023] When the reliability is greater than a preset threshold, determine the location information of the brain tumor as the calculation result and output it.
[0024] In a possible implementation manner, after the location information of the brain tumor is determined as the calculation result and output, the method further includes:
[0025] Calculate the current loss according to the output result and the preset validation set information;
[0026] Adjust the dynamic convolution layer according to the calculated current loss.
[0027] In a possible implementation manner, the multi-modal brain magnetic resonance imaging includes: T1-weighted image, T1-enhanced image, T2-weighted image, fluid-attenuated inversion recovery image FLAIR.
[0028] In a second aspect of the embodiments of the present application, a multi-modal segmentation system for brain tumor imaging is provided. The system includes:
[0029] A multi-modal input module for acquiring multi-modal brain magnetic resonance imaging;
[0030] A preprocessing module for preprocessing the multi-modal brain magnetic resonance imaging to obtain the preprocessing result of each modality;
[0031] A dynamic convolution encoding module for respectively identifying key regions of the preprocessing results of each modality through an attention layer to obtain key regions corresponding to each modality; extracting features of the key regions corresponding to each modality through a plurality of parallel dynamic convolution layers to obtain extraction features corresponding to each modality; respectively performing feature fusion on the extraction features corresponding to each modality to obtain combined features;
[0032] A dual-channel decoupling module for performing convolution and pooling processing on the combined features to obtain a first processing result; performing convolution and mapping processing on the combined features to obtain a second processing result;
[0033] A multi-modal fusion module for performing multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features;
[0034] A dynamic convolution decoding module for performing attention weighting on the multi-modal fusion features to obtain weighted features; extracting a receptive field through the weighted features to obtain the position information of the brain tumor.
[0035] In a possible implementation manner, the preprocessing module is specifically configured to perform bias field correction on the multi-modal brain magnetic resonance imaging; perform normalization on the corrected imaging; perform rigid registration on the normalized imaging; perform identification of regions of interest on the rigidly registered imaging to obtain regions of interest of each modality imaging.
[0036] In a possible implementation, the dynamic convolutional coding module is configured to randomly discard some neurons through a dropout layer, and extract features from the key regions corresponding to each modality through the remaining neurons to obtain the target extraction features corresponding to each modality; the separately performing feature fusion on the extraction features corresponding to each modality to obtain a combined feature includes: separately performing feature fusion on the target extraction features corresponding to each modality to obtain the combined feature.
[0037] In a possible implementation, the system further includes:
[0038] A three-branch output module, configured to judge the reliability of the position information of the brain tumor through voxel-level majority voting; when the reliability is greater than a preset threshold, determine the position information of the brain tumor as a calculation result and output it.
[0039] In a possible implementation, the system further includes:
[0040] A loss function calculation module, configured to calculate a current loss according to the output result and preset validation set information;
[0041] Adjust the dynamic convolutional layer according to the calculated current loss.
[0042] In a possible implementation, the multi-modal brain magnetic resonance imaging includes: T1-weighted image, T1-enhanced image, T2-weighted image, fluid-attenuated inversion recovery image FLAIR.
[0043] On the other hand, an embodiment of the present application further provides an electronic device, including:
[0044] A memory, configured to store a computer program;
[0045] A processor, configured to implement any of the above multi-modal segmentation methods for brain tumor imaging when executing the program stored on the memory.
[0046] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any of the above multi-modal segmentation methods for brain tumor imaging is implemented.
[0047] On the other hand, an embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the above multi-modal segmentation methods for brain tumor imaging.
[0048] Advantageous effects of the embodiments of the present application:
[0049] An embodiment of the present application provides a multi-modal segmentation method and system for brain tumor imaging. The method includes: obtaining magnetic resonance imaging of the brain in multiple modalities; preprocessing the magnetic resonance imaging of the brain in multiple modalities to obtain a preprocessing result for each modality; identifying key regions for each preprocessing result of each modality through an attention layer to obtain key regions corresponding to each modality; extracting features of the key regions corresponding to each modality through a plurality of parallel dynamic convolutional layers to obtain extracted features corresponding to each modality; respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features; performing convolution and pooling processing on the combined features to obtain a first processing result; performing convolution and mapping processing on the combined features to obtain a second processing result; performing multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features; performing attention weighting on the multi-modal fusion features to obtain weighted features; and extracting a receptive field through the weighted features to obtain the position information of the brain tumor. Through the method of the embodiment of the present application, automatic identification of the position of the brain tumor in brain imaging can be achieved, thereby solving the problem of errors caused by manual identification and improving the accuracy of brain tumor identification.
[0050] Of course, implementing any product or method of the present application does not necessarily require achieving all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0052] Figure 1a It is a schematic flowchart of a multi-modal segmentation method for brain tumor imaging provided by an embodiment of the present application;
[0053] Figure 1b It is a schematic flowchart of the dual-channel decoupling provided by an embodiment of the present application;
[0054] Figure 1c It is a schematic structural diagram of a multi-modal fusion module provided by an embodiment of the present application;
[0055] Figure 2 It is a schematic flowchart of a preprocessing provided by an embodiment of the present application;
[0056] Figure 3 It is a schematic flowchart of feature extraction for key regions provided by an embodiment of the present application;
[0057] Figure 4A schematic diagram of a process for voxel-level majority voting provided by an embodiment of the present application;
[0058] Figure 5 A schematic diagram of a process for dynamic convolution layer adjustment provided by an embodiment of the present application;
[0059] Figure 6 A schematic diagram of a process for generating a dynamic convolution kernel provided by an embodiment of the present application;
[0060] Figure 7 A schematic diagram of the system architecture corresponding to the multi-modal segmentation method for brain tumor imaging provided by an embodiment of the present application;
[0061] Figure 8 A schematic diagram of the structure of a multi-modal segmentation system for brain tumor imaging provided by an embodiment of the present application;
[0062] Figure 9 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the protection scope of the present application.
[0064] To solve the problem in the prior art that when identifying manually, since the diagnostic result depends on the doctor's experience, when the doctor has insufficient experience, errors often occur, affecting subsequent treatment
[0065] In the first aspect of the embodiments of the present application, first, a multi-modal segmentation method for brain tumor imaging is provided. Refer to Figure 1a , Figure 1a A schematic diagram of a process for the multi-modal segmentation method for brain tumor imaging provided by an embodiment of the present application. The method includes:
[0066] Step S11, obtaining magnetic resonance imaging of the brain in multiple modalities;
[0067] Step S12, preprocessing the magnetic resonance imaging of the brain in multiple modalities to obtain the preprocessing result of each modality;
[0068] Step S13: Through the attention layer, identify the key regions for the preprocessing results of each modality respectively to obtain the key regions corresponding to each modality; through multiple parallel dynamic convolution layers, extract features from the key regions corresponding to each modality to obtain the extracted features corresponding to each modality; perform feature fusion on the extracted features corresponding to each modality respectively to obtain combined features;
[0069] Step S14: Perform convolution and pooling processing on the combined features to obtain a first processing result; perform convolution and mapping processing on the combined features to obtain a second processing result;
[0070] Step S15: Perform multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features;
[0071] Step S16: Perform attention weighting on the multi-modal fusion features to obtain weighted features; through the weighted features, extract the receptive field to obtain the location information of the brain tumor.
[0072] Corresponding to the above step S11, when obtaining multi-modal brain magnetic resonance imaging, the brain MRI (Magnetic Resonance Imaging) of the patient can be obtained. Specifically, the brain magnetic resonance imaging in the embodiments of the present application is multi-modal imaging. In a possible implementation manner, the multi-modal brain magnetic resonance imaging includes: T1-weighted imaging, Contrast-enhanced T1-weighted imaging, T2-weighted imaging, and Fluid Attenuated Inversion Recovery (FLAIR). Among them, the T1-weighted image is based on the longitudinal relaxation time (T1 time) of the tissue and reflects the speed at which the tissue recovers longitudinal magnetization; the Contrast-enhanced T1-weighted image is T1-weighted imaging performed after intravenous injection of a gadolinium contrast agent, and the contrast agent enters the lesion area through the damaged blood-brain barrier; the T2-weighted image is based on the transverse relaxation time (T2 time) of the tissue and reflects the attenuation speed of transverse magnetization; FLAIR suppresses the signal of free water (such as cerebrospinal fluid) and highlights the signal of bound water (such as edema or lesion). Among them, the Contrast-enhanced T1-weighted image can highlight the tumor boundary, the T2-weighted image can show the edema area, and FLAIR can show the suppression of cerebrospinal fluid signal. In the embodiments of the present application, the recognition rate of low-contrast regions can be improved through multi-modal information complementarity. In the actual use process, the resolution of the obtained imaging can also be set for subsequent processing. In a possible implementation manner, the resolution of the imaging can be 0.5mm×0.5mm×0.5mm, 1mm×1mm×1mm, 0.5mm×0.5mm×1mm, etc.
[0073] Corresponding to the above step S12, preprocessing of multi-modal brain magnetic resonance imaging can be performed by various methods, such as registration, normalization, etc. Since in the actual use process, the brain magnetic resonance imaging obtained from different hospitals, devices, or patients may have deviations, in the embodiments of the present application, before identifying brain tumors, preprocessing the obtained multi-modal brain magnetic resonance imaging can standardize the obtained imaging, thereby facilitating subsequent processing and reducing errors in the recognition process. Specifically, when preprocessing multi-modal brain magnetic resonance imaging, operations such as resolution adjustment, image normalization, registration, and extraction of regions of interest can be performed to obtain the preprocessing results of each modality.
[0074] Corresponding to the above step S13, this step is mainly used for dynamic convolution feature extraction. Specifically, through the attention layer, that is, the Attention layer, the key regions of the preprocessing results of each modality are respectively identified to obtain the key regions corresponding to each modality. Among them, the key regions identified by the Attention layer can be regions of interest. When identifying the key regions, threshold segmentation can be used, such as single / multi-band thresholding method; as well as feature classification, histogram optimization and other methods for identification. Through this attention layer, limited attention can be concentrated on the key information, that is, the key regions corresponding to each modality, so as to save resources and quickly obtain the most effective information. Then, for the obtained key regions, through multiple parallel dynamic convolution layers, feature extraction is performed on the key regions corresponding to each modality. Then, through multiple parallel dynamic convolution layers, feature extraction is performed on the key regions corresponding to each modality to obtain the extracted features corresponding to each modality. It should be noted that when performing feature extraction on the key regions corresponding to each modality in this application, dynamic convolution layers are used. The weights of the convolution kernels in the dynamic convolution layers are not fixed, but are dynamically adjusted according to the input features of each time step, which means that the convolution kernels can change according to the input changes, so as to provide more flexible and adaptive feature extraction. Compared with the traditional fixed convolution kernel that cannot adapt to the dynamic changes of tumor morphology, the dynamic convolution kernel can adjust the receptive field according to local features. For example, the convolution kernel size is increased at the tumor boundary to capture richer context information. The fitting ability for complex morphologies is enhanced through dynamic parameter adaptability. Finally, by respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features, the fusion of features in multimodal imaging can be realized. Due to the different characteristics corresponding to each modality, through multimodal fusion, the recognition ability of tumor boundaries and heterogeneous features can be strengthened, and the accuracy of brain tumor recognition can be improved.
[0075] Corresponding to the above step S14, this step mainly extracts features through the COS-SOS (dual-channel decoupling) attention mechanism. Specifically, it can be processed in parallel by a channel attention branch and a spatial attention branch. Among them, the channel branch can perform convolution and pooling operations on the combined features to obtain a first processing result; the spatial attention branch can perform convolution and mapping operations on the combined features to obtain a second processing result. In one example, the channel branch uses 1×1×1 convolution and global pooling, and the spatial branch uses 1×1×1 convolution and the Sigmoid function (activation function) for mapping processing. In the embodiments of the present application, through the COS-SOS attention mechanism, the attention calculation in the channel and spatial dimensions can be separated, reducing redundancy and optimizing the feature weight allocation. The inventors have found that through dual-channel decoupling, the computational complexity is reduced from O (C²HWD) to O (CHWD), where O represents the algorithm complexity, C represents the number of channels, H represents the height, W represents the width, and D represents the depth; real-time segmentation can be achieved while maintaining the accuracy, which is suitable for clinical rapid diagnosis scenarios. See Figure 1b , Figure 1b is a schematic flowchart of the dual-channel decoupling provided by the embodiments of the present application. The input parameter X is processed in two paths. One path is subjected to convolution processing through a (1×1×1) convolutional layer, and after global pooling, matrix transformation, and normalization, it passes through calculation, that is, matrix dot product operation. Finally, after matrix transformation and activation function calculation, it passes through the ⊙ calculation, that is, the multiplication operator in the spatial dimension, to obtain the calculation result of the first path; the other path is subjected to convolution processing through a (1×1×1) convolutional layer, and after global pooling, matrix transformation, and normalization, it passes through calculation, that is, matrix dot product operation. Finally, after convolution, normalization, and activation function calculation, it passes through the ⊙ calculation, that is, the multiplication operator in the channel dimension, to obtain the calculation result of the second path; finally, through the ⊕ calculation, that is, calculating the direct sum, the calculation results of the two paths are fused.
[0076] Corresponding to the above step S15, this step can achieve multimodal fusion. Specifically, the first processing result and the second processing result can be subjected to multimodal fusion to obtain multimodal fusion features. Among them, when performing multimodal fusion, various methods can be used for fusion, such as feature splicing, superposition, weighted summation, etc. for feature fusion. Specifically, it can be set according to the actual situation. In the embodiments of the present application, it is preferably to use the method of feature splicing for fusion. Among them, in the embodiments of the present application, through multimodal fusion, the recognition rate of low-contrast regions can be improved through multimodal information complementarity. The inventor's research found that through the fusion of multiple modalities, it is convenient to correct errors in the subsequent voting process when a certain modality misjudges the necrotic region. Among them, in the cross-modal test, the fluctuation of DSC (Dice similarity coefficient) is <2%, which is 5.7% higher than that of the single-modal model; the segmentation sensitivity of low-contrast regions reaches 86.9%, effectively coping with the challenges of tumor heterogeneity. See Figure 1c , Figure 1c FIG. Figure 1c is a schematic structural diagram of the multimodal fusion module provided by the embodiment of the present application. When the brain magnetic resonance imaging of multiple modalities in the embodiment of the present application is four modalities, the four-modal input can be performed, and the images of the four modalities are input into a three-level fusion architecture. First, feature concatenation is performed to obtain a cascaded feature architecture of 4C×H×W×D, where C represents the number of channels, H represents the height, W represents the width, and D represents the depth; then attention weighting is performed to obtain weighted features; then dynamic convolution decoding is performed to obtain the segmentation maps of each modality, and then single-modal segmentation is performed according to the segmentation maps of various modalities to obtain a segmentation result structure of 4×H×W×D; finally, through decision fusion, after obtaining the fusion result, the final segmentation result is obtained.
[0077] Corresponding to the above step S16, attention weighting is performed on the multimodal fusion features. Through attention weighting, the relevance of each part of the input data to the current task can be calculated, such as vector dot product or similarity function, to assign weights to different positions, and a more focused context representation is generated after weighted summation. By performing receptive field extraction on the weighted features, the position information of the brain tumor can be obtained. Specifically, in the actual use process, the position information can be the point cloud information corresponding to the position of the brain tumor, and through this point cloud information, the position, size, and boundary of the brain tumor can be determined.
[0078] It can be seen that through the method of the embodiments of the present application, multi-modal brain magnetic resonance imaging can be obtained, and then the obtained imaging is preprocessed. Then, dynamic convolution coding and feature extraction are performed on the preprocessed imaging, and the extracted features are subjected to multi-modal fusion. Thus, attention weighting and receptive field extraction are performed through the fused features to obtain the location information of the brain tumor. Therefore, through the parameter self-adaptation of dynamic convolution, the decoupled design of dual-channel attention, and the information complementarity of multi-modal fusion, the bottlenecks of traditional methods in terms of accuracy, efficiency, and robustness are systematically solved. Moreover, by combining the lightweight dynamic network with the integration of multi-modal information, while improving the performance of medical image segmentation, computational efficiency and clinical practicability are taken into account, providing a new technical solution for the precise diagnosis and treatment of brain tumors.
[0079] In a possible implementation manner, referring to Figure 2 , the preprocessing result in step S12 includes the region of interest of each modal imaging. The preprocessing of the multi-modal brain magnetic resonance imaging to obtain the preprocessing result of each modality includes:
[0080] Step S121, perform bias field correction on the multi-modal brain magnetic resonance imaging;
[0081] Step S122, perform normalization on the corrected imaging;
[0082] Step S123, perform rigid registration on the normalized imaging;
[0083] Step S124, perform identification of the region of interest on the rigidly registered imaging to obtain the region of interest of each modal imaging.
[0084] Among them, for the multi-modal brain magnetic resonance imaging, bias field correction can be performed, and N4 bias field (N4BiasFieldCorrection) correction can be carried out to correct the signal intensity change caused by magnetic field inhomogeneity. Then, normalization can be performed on the corrected imaging, and Z-score normalization can be carried out. Through Z-score normalization, the corrected imaging can be unified to the same standard scale. Specifically, all pixels in the imaging can be unified to the interval [-1, 1], and the pixel values in the interval [-1, 1] corresponding to the corrected imaging can be obtained. Then, rigid registration can be performed on the normalized imaging. Specifically, rigid registration can be carried out through ANTS (Advanced Normalization Tools). During the registration process, the mutual information threshold can be set to >0.8. Finally, ROI (region of interest) identification can be performed on the rigidly registered imaging. Specifically, the Otsu threshold method (Otsu algorithm) can be used. This algorithm is an adaptive binary segmentation algorithm based on the gray distribution of the image. The optimal threshold can be determined by maximizing the between-class variance, and through three-dimensional connected component analysis, the region composed of pixels or voxels with similar properties in the image can be identified and segmented, so as to obtain the region of interest of each modal imaging. In a possible implementation manner, a fixed output size can be set to facilitate subsequent processing, such as 128×128×128 pixels, 224×224×128 pixels, etc.
[0085] It can be seen that through the method of the embodiment of the present application, before the identification of brain tumors, preprocessing can be performed on the obtained multi-modal brain magnetic resonance imaging, and standardization of the obtained imaging can be achieved, so as to facilitate subsequent processing and reduce errors in the identification process.
[0086] In a possible implementation manner, referring to Figure 3 , the feature extraction of the key area corresponding to each modality is performed through multiple parallel dynamic convolutional layers to obtain the extracted features corresponding to each modality, including:
[0087] Step S31, randomly discard some neurons through the dropout layer, and perform feature extraction on the key area corresponding to each modality through the remaining neurons to obtain the target extraction features corresponding to each modality;
[0088] The feature fusion of the extracted features corresponding to each modality is respectively performed to obtain the combined features, including:
[0089] Step S32, respectively perform feature fusion on the target extraction features corresponding to each modality to obtain the combined features.
[0090] In deep learning, when there are too many parameters and the training samples are relatively few, the model is prone to overfitting. Overfitting is a common problem in many deep learning and even machine learning algorithms, specifically manifested as high prediction accuracy on the training set and a significant drop in accuracy on the test set. Therefore, in the embodiments of the present application, by randomly discarding some neurons through a dropout layer, and then using the remaining neurons to extract features from the key regions corresponding to each modality, the target extraction features corresponding to each modality are obtained. Then, feature fusion is performed on the target extraction features corresponding to each modality respectively to obtain the combined features. During training, each neuron is retained with probability p, that is, it stops working with probability 1 - p. The neurons retained in each forward propagation are different, which can make the model less dependent on certain local features and have stronger generalization performance. In the actual use process, through the method of the embodiments of the present application, the convolutional weights are adjusted according to the feature weights obtained by the attention mechanism, and Dropout is added to improve the generalization ability and convergence speed of the model. On the BraTS2021 (Brain Tumor Segmentation 2021) dataset, the Dice (a set similarity metric function) coefficient of the tumor core (TC) reaches 91.2%, which is 4.1% higher than that of nnUNetv2 (the second-generation automated medical image segmentation neural network framework).
[0091] In a possible implementation manner, after obtaining the position information of the brain tumor by performing receptive field extraction through the weighted features, refer to Figure 4 , the method further includes:
[0092] Step S41, determining the reliability of the position information of the brain tumor through voxel-level majority voting;
[0093] Step S42, when the reliability is greater than a preset threshold, determining the position information of the brain tumor as the calculation result and outputting it.
[0094] In the embodiments of the present application, after obtaining the position information of the brain tumor by performing receptive field extraction through the weighted features, determining the reliability of the position information of the brain tumor through voxel-level majority voting can implement a three-level fusion architecture including feature concatenation, channel attention weighting (dynamically allocating modality weights), and voxel-level majority voting (three-dimensional spatial consistency constraint), integrating multi-modal MRI information. Among them, through voxel-level majority voting, the results of multiple predictions can be integrated to improve the segmentation accuracy and robustness. At each voxel (three-dimensional pixel) position, different prediction results are statistically counted, and the category with the highest frequency of occurrence is selected as the final output. Thus, combined with the multi-modal imaging in the present application, when a certain modality misjudges the necrotic area, the voting results of other modalities can correct the error and improve the recognition accuracy.
[0095] In a possible implementation, after determining the position information of the brain tumor as a calculation result and outputting it, refer to Figure 5 , the method further includes:
[0096] Step S51, calculating a current loss according to the output result and preset verification set information;
[0097] Step S52, adjusting the dynamic convolution layer according to the calculated current loss.
[0098] Refer to Figure 6 , the working principle of the dynamic convolution in the embodiment of the present application is that the input features complete basic feature extraction and normalization through a triple operation, and then enter the parallel convolution layer. By adaptively using different numbers of convolution kernels, differential feature maps are generated. In this structure, the input parameter X is weighted by attention through the attention layer, then convolution calculations are performed in parallel through multiple convolution layers, then convolution calculations are performed through the convolution layer, then partial neurons are discarded through the dropout layer, and then through normalization processing and activation function calculation, the output parameter Y is obtained. Refer to Figure 7 , after multi-modal input and input of multi-modal imaging; preprocessing the imaging through a preprocessing module; then performing convolution calculation on the preprocessed imaging through dynamic convolution encoding; then performing calculation extraction of multi-modal features through dual-channel decoupling; then performing feature fusion through multi-modal fusion; then determining the receptive field through dynamic convolution decoding; finally performing three-dimensional voxel-level majority voting and output through three-branch output; finally, after calculating through the loss function and determining the position information of the brain tumor as a calculation result and outputting it, the current loss can be calculated according to the output result and preset verification set information, and thus the dynamic convolution layer can be adjusted according to the calculated current loss. Specifically, the loss can be calculated through various loss functions, such as mean square error loss function, cross-entropy loss function, logarithmic loss function, etc. By adjusting the dynamic convolution layer according to the calculated current loss, the adjustment of the dynamic convolution layer can be realized, so that the receptive field can be adjusted according to local features. For example, the convolution kernel size is increased at the tumor boundary to capture richer context information, and the fitting ability for complex shapes is enhanced through dynamic parameter adaptation.
[0099] In the second aspect of the embodiment of the present application, a multi-modal segmentation system for brain tumor imaging is provided, refer to Figure 8 , the system includes:
[0100] A multi-modal input module 801, configured to obtain brain magnetic resonance imaging of multiple modalities;
[0101] The preprocessing module 802 is used to preprocess brain magnetic resonance imaging of multiple modalities to obtain the preprocessing results of each modality;
[0102] The dynamic convolution encoding module 803 is used to identify key regions of the preprocessing results of each modality through an attention layer to obtain the key regions corresponding to each modality; extract features of the key regions corresponding to each modality through multiple parallel dynamic convolution layers to obtain the extracted features corresponding to each modality; and perform feature fusion on the extracted features corresponding to each modality to obtain combined features;
[0103] The dual-channel decoupling module 804 is used to perform convolution and pooling processing on the combined features to obtain a first processing result; and perform convolution and mapping processing on the combined features to obtain a second processing result;
[0104] The multi-modal fusion module 805 is used to perform multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features;
[0105] The dynamic convolution decoding module 806 is used to perform attention weighting on the multi-modal fusion features to obtain weighted features; and extract the receptive field through the weighted features to obtain the location information of the brain tumor.
[0106] In a possible implementation manner, the preprocessing module is specifically used to perform bias field correction on the brain magnetic resonance imaging of multiple modalities; perform normalization on the corrected imaging; perform rigid registration on the normalized imaging; and identify regions of interest of the imaging after rigid registration to obtain the regions of interest of each modality imaging.
[0107] In a possible implementation manner, the dynamic convolution encoding module is used to randomly discard some neurons through a dropout layer, and extract features of the key regions corresponding to each modality through the remaining neurons to obtain the target extracted features corresponding to each modality; and the performing feature fusion on the extracted features corresponding to each modality to obtain combined features includes: performing feature fusion on the target extracted features corresponding to each modality to obtain the combined features.
[0108] In a possible implementation manner, the system further includes:
[0109] The three-branch output module is used to judge the reliability of the location information of the brain tumor through voxel-level majority voting; and when the reliability is greater than a preset threshold, determine the location information of the brain tumor as the calculation result and output it.
[0110] In a possible implementation manner, the system further includes:
[0111] A loss function calculation module, configured to calculate a current loss according to the output result and preset verification set information;
[0112] Adjust the dynamic convolution layer according to the calculated current loss.
[0113] In a possible implementation manner, the multiple-modal brain magnetic resonance imaging includes: T1-weighted image, T1-enhanced image, T2-weighted image, fluid-attenuated inversion recovery image FLAIR.
[0114] It can be seen that through the method of the embodiments of the present application, multiple-modal brain magnetic resonance imaging can be obtained, then the obtained imaging is preprocessed, then the preprocessed imaging is subjected to dynamic convolution encoding and feature extraction, and then the extracted features are subjected to multi-modal fusion, so as to perform attention weighting and receptive field extraction through the fused features to obtain the position information of the brain tumor. Thus, through the parameter self-adaptation of dynamic convolution, the decoupled design of dual-channel attention, and the information complementarity of multi-modal fusion, the bottlenecks of traditional methods in terms of accuracy, efficiency, and robustness are systematically solved. And by combining the lightweight dynamic network with multi-modal information integration, while improving the performance of medical image segmentation, the computational efficiency and clinical practicability are taken into account, providing a new technical solution for the precise diagnosis and treatment of brain tumors.
[0115] The embodiments of the present application also provide an electronic device, as Figure 9 shown, including:
[0116] A memory 901, configured to store a computer program;
[0117] A processor 902, configured to implement the following steps when executing the program stored on the memory 901:
[0118] Obtain multiple-modal brain magnetic resonance imaging;
[0119] Preprocess the multiple-modal brain magnetic resonance imaging to obtain a preprocessing result for each modality;
[0120] Through an attention layer, respectively identify key regions of the preprocessing result of each modality to obtain key regions corresponding to each modality; through a plurality of parallel dynamic convolution layers, extract features of the key regions corresponding to each modality to obtain extracted features corresponding to each modality; respectively perform feature fusion on the extracted features corresponding to each modality to obtain combined features;
[0121] Perform convolution and pooling processing on the combined features to obtain a first processing result; perform convolution and mapping processing on the combined features to obtain a second processing result;
[0122] Perform multimodal fusion on the first processing result and the second processing result to obtain multimodal fusion features;
[0123] Perform attention weighting on the multimodal fusion features to obtain weighted features; extract the receptive field through the weighted features to obtain the location information of the brain tumor.
[0124] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0125] The communication interface is used for communication between the above electronic device and other devices.
[0126] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0127] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0128] In another embodiment provided by the present application, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above multimodal segmentation methods for brain tumor imaging are implemented.
[0129] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute any of the multi-modal segmentation methods for brain tumor imaging in the above embodiments.
[0130] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0131] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0132] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system, electronic device, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0133] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A multi-modal segmentation method for brain tumor imaging, characterized in that, The method includes: Obtaining brain magnetic resonance imaging of multiple modalities; Preprocessing the brain magnetic resonance imaging of multiple modalities to obtain the preprocessing result of each modality; Through the attention layer, respectively identifying the key regions of the preprocessing results of each modality to obtain the key regions corresponding to each modality; through multiple parallel dynamic convolutional layers, extracting features from the key regions corresponding to each modality to obtain the extracted features corresponding to each modality; respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features; wherein, the key region is the region of interest; the step of extracting features from the key regions corresponding to each modality through multiple parallel dynamic convolutional layers to obtain the extracted features corresponding to each modality includes: randomly discarding some neurons through the dropout layer, and extracting features from the key regions corresponding to each modality through the remaining neurons to obtain the target extracted features corresponding to each modality; the step of respectively performing feature fusion on the extracted features corresponding to each modality to obtain combined features includes: respectively performing feature fusion on the target extracted features corresponding to each modality to obtain the combined features; Performing convolution and pooling processing on the combined features to obtain a first processing result; performing convolution and mapping processing on the combined features to obtain a second processing result; Performing multi-modal fusion on the first processing result and the second processing result to obtain multi-modal fusion features; Performing attention weighting on the multi-modal fusion features to obtain weighted features; extracting the receptive field through the weighted features to obtain the location information of the brain tumor.
2. The method according to claim 1, wherein The preprocessing result includes the region of interest of each modality imaging, and the step of preprocessing the brain magnetic resonance imaging of multiple modalities to obtain the preprocessing result of each modality includes: Performing bias field correction on the brain magnetic resonance imaging of multiple modalities; Normalizing the corrected imaging; Performing rigid registration on the normalized imaging; Identifying the region of interest of the imaging after rigid registration to obtain the region of interest of each modality imaging.
3. The method according to claim 1, wherein After the step of extracting the receptive field through the weighted features to obtain the location information of the brain tumor, the method further includes: Judging the reliability of the location information of the brain tumor through voxel-level majority voting; When the reliability is greater than a preset threshold, determining the location information of the brain tumor as the calculation result and outputting it.
4. The method according to claim 3, wherein After the step of determining the location information of the brain tumor as the calculation result and outputting it, the method further includes: Calculating the current loss according to the output result and the preset validation set information; Adjusting the dynamic convolutional layer according to the calculated current loss.
5. The method according to claim 1, wherein The brain magnetic resonance imaging of multiple modalities includes: T1-weighted image, T1-enhanced image, T2-weighted image, fluid-attenuated inversion recovery image FLAIR.
6. A multimodal segmentation system for brain tumor imaging, characterized in that, The system includes: A multi-modal input module for obtaining brain magnetic resonance imaging of multiple modalities; A preprocessing module for preprocessing the brain magnetic resonance imaging of multiple modalities to obtain the preprocessing result of each modality; A dynamic convolution encoding module is used to identify key regions of the preprocessing results of each modality through an attention layer, obtaining the key regions corresponding to each modality; through multiple parallel dynamic convolution layers, extract features from the key regions corresponding to each modality, obtaining the extracted features corresponding to each modality; and respectively perform feature fusion on the extracted features corresponding to each modality, obtaining combined features; wherein, the key regions are regions of interest. A two-channel decoupling module is used to perform convolution and pooling processing on the combined features, obtaining a first processing result; and perform convolution and mapping processing on the combined features, obtaining a second processing result. A multi-modal fusion module is used to perform multi-modal fusion on the first processing result and the second processing result, obtaining multi-modal fusion features. A dynamic convolution decoding module is used to perform attention weighting on the multi-modal fusion features, obtaining weighted features; and through the weighted features, extract the receptive field, obtaining the location information of the brain tumor. The dynamic convolution encoding module is used to randomly discard some neurons through a dropout layer, and extract features from the key regions corresponding to each modality through the remaining neurons, obtaining the target extracted features corresponding to each modality; and respectively perform feature fusion on the target extracted features corresponding to each modality, obtaining the combined features.
7. The system according to claim 6, wherein The system further includes: A three-branch output module is used to judge the reliability of the location information of the brain tumor through voxel-level majority voting; when the reliability is greater than a preset threshold, determine the location information of the brain tumor as the calculation result and output it.
8. The system according to claim 7, wherein The system further includes: A loss function calculation module is used to calculate the current loss according to the output result and the preset validation set information. Adjust the dynamic convolution layer according to the calculated current loss.
9. An electronic device, characterized in that, It includes: A memory for storing a computer program. A processor, when executing the program stored on the memory, implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Multi-modal imaging tumor positioning method and system
CN119027504A
MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-modal feature fusion
CN119850961A