High-grade glioma grading method and device based on multimodal completion and subregion segmentation
Full-modal MRI image data is generated through the pre-trained complementary subnet and multi-modal fusion is used to solve the problem of inaccurate glioma grading caused by MRI modal loss, and the accuracy of accurate grading and diagnosis of advanced gliomas is improved.
Patent Information
- Application Number
- CN202510142828.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In clinical diagnosis, abnormal MRI modal deletion leads to inaccurate glioma grading, especially when distinguishing between tertiary and quaternary gliomas, small differences in imaging characteristics or lack of modality may have a significant impact on diagnosis and treatment decisions.
The key information in the residual mode MRI image data set is captured through the pre-trained complementary subnet and the full mode MRI image data is generated; at the same time, the sub-regional segmentation subnet is used for multimodal fusion, and the weight is allocated to different modes based on the channel attention mechanism, and the glioma subregion characteristics are highlighted through the attention gating mechanism, so as to achieve fine segmentation of the glioma subregion and accurate grading of advanced gliomas.
The problem of MRI modal missing is solved, a complete imaging data foundation is provided, and the precise grading of advanced gliomas is achieved, which improves the accuracy of diagnosis and reliability of prognostic evaluation.
Smart Images

Figure CN119601180B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical image processing, and in particular to a method and device for grading high-grade gliomas based on multimodal completion and subregion segmentation. Background Art
[0002] Glioma is a type of primary brain tumor originating from glial cells, accounting for more than 80% of all malignant tumors. According to the classification of the World Health Organization (WHO), gliomas are divided into four grades, among which grade III and grade IV gliomas (high-grade gliomas) have higher invasiveness and poorer prognosis. Early and accurate diagnosis of the grade of glioma is crucial for formulating effective treatment plans and improving patient prognosis.
[0003] Magnetic resonance imaging (MRI) is the preferred imaging tool for diagnosing gliomas because it can provide high-resolution soft tissue contrast images. The advantages of MRI examinations include non-invasiveness, high resolution, and multi-plane imaging capabilities. However, MRI also has some limitations, such as allergic reactions to contrast agents, difficulty for patients to remain still during scanning, long scanning processes, and difficulty for patients to tolerate claustrophobic environments. These problems may lead to modality loss, which in turn affects the accuracy of diagnosis and prognostic assessment. In order to achieve an accurate diagnosis, it is usually necessary to combine the four MRI modalities of T1, T1ce, T2, and FLAIR. However, in clinical diagnosis, the problem of modality loss often occurs, resulting in the inability to accurately grade gliomas, especially when distinguishing between grade 3 and grade 4 gliomas. Slight differences in imaging features or the lack of modality may have a significant impact on diagnosis and treatment decisions.
[0004] Therefore, how to accurately distinguish between grade III and grade IV gliomas is of great significance for formulating individualized treatment plans and prognosis assessment. Summary of the invention
[0005] The embodiments of the present application provide a method and device for grading high-grade gliomas based on multimodal completion and subregion segmentation. The completion subnetwork is used to generate full-modality data based on the residual modality MRI image data set, providing a complete information basis for subsequent analysis. The subregion segmentation subnetwork is used to reasonably allocate weights to different modalities and highlight the characteristics of glioma subregions, thereby achieving fine segmentation of glioma subregions and accurately grading high-grade gliomas.
[0006] In a first aspect, an embodiment of the present application provides a method for grading high-grade gliomas based on multimodal completion and subregion segmentation, the method comprising:
[0007] Obtain a residual modality MRI image dataset of the same patient, and input the residual modality MRI image dataset into a pre-trained completion subnetwork to obtain a full-modality MRI image dataset, wherein the full-modality MRI image dataset includes MRI image data of four modalities: T1, T1ce, T2, and FLAIR, and the residual modality MRI image dataset is an MRI image dataset with missing modalities;
[0008] The full-modality MRI data set is input into a pre-trained sub-region segmentation sub-network to obtain a sub-region segmentation result. The sub-region segmentation sub-network includes a segmentation encoding unit and a segmentation decoding unit. The segmentation encoding unit performs multi-modal fusion on the full-modality MRI image data set to obtain a fusion result, and assigns different weights to different modalities based on a channel attention mechanism during the fusion process. The segmentation decoding unit increases the weight of the glioma sub-region feature in the fusion result based on an attention gating mechanism to obtain a sub-region segmentation result;
[0009] Glioma grading is performed based on the subregion segmentation results.
[0010] In a second aspect, an embodiment of the present application provides a high-grade glioma grading device based on multimodal completion and subregion segmentation, comprising:
[0011] A completion module is used to obtain a residual modality MRI image dataset of the same patient, and input the residual modality MRI image dataset into a pre-trained completion subnetwork to obtain a full-modality MRI image dataset, wherein the full-modality MRI image dataset includes MRI image data of four modalities: T1, T1ce, T2, and FLAIR, and the residual modality MRI image dataset is an MRI image dataset with missing modalities;
[0012] A segmentation module is used to input the full-modality MRI data set into a pre-trained sub-region segmentation sub-network to obtain a sub-region segmentation result. The sub-region segmentation sub-network includes a segmentation encoding unit and a segmentation decoding unit. The segmentation encoding unit performs multi-modal fusion on the full-modality MRI image data set to obtain a fusion result, and assigns different weights to different modalities based on a channel attention mechanism during the fusion process. The segmentation decoding unit increases the weight of the glioma sub-region feature in the fusion result based on an attention gating mechanism to obtain a sub-region segmentation result;
[0013] A grading module is used to grade glioma based on the subregion segmentation results.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute a method for grading high-grade gliomas based on multimodal completion and subregion segmentation.
[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process, and the process includes a method for grading high-grade gliomas based on multimodal completion and subregion segmentation.
[0016] The main contributions and innovations of the present invention are as follows:
[0017] The embodiment of the present application utilizes a pre-trained completion subnetwork to capture key information in the residual modality MRI image data set to perform modality completion to generate full-modality MRI image data, thereby solving the problem of missing MRI modalities in clinical practice; the segmentation subnetwork of the embodiment of the present application utilizes a channel attention mechanism through a segmentation encoding unit to assign reasonable weights to different modalities to achieve multimodal fusion; the segmentation decoding unit is based on an attention gating mechanism to highlight the characteristics of glioma subregions and achieve fine segmentation of glioma subregions, thereby helping doctors to more accurately grade high-grade gliomas.
[0018] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 is a flow chart of a method for grading high-grade gliomas based on multimodal completion and subregion segmentation according to an embodiment of the present application;
[0021] Figure 2 is a structural diagram of a completion sub-network according to an embodiment of the present application;
[0022] Figure 3 is a structural diagram of a feature compression module according to an embodiment of the present application;
[0023] Figure 4 is a structural diagram of a sub-region segmentation sub-network according to an embodiment of the present application;
[0024] Figure 5 It is a radar chart of a doctor's ability to identify internal features of a glioma in the absence of an MRI modality according to an embodiment of the present application;
[0025] Figure 6 is a statistical graph of brain glioma MRI segmentation performance evaluation according to an embodiment of the present application;
[0026] Figure 7is a schematic diagram of ROC performance comparison of tumor grading by doctors according to an embodiment of the present application;
[0027] Figure 8 is a structural block diagram of a high-grade glioma grading device based on multimodal completion and subregion segmentation according to an embodiment of the present application;
[0028] Fig. 9 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0030] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0031] In order to better understand this scheme, the four MRI modalities mentioned in this scheme are explained here:
[0032] T1 mode: T1 mode is T1-weighted imaging, which reflects the difference in T1 relaxation time (T1 value) between tissues. It has a high signal-to-noise ratio and good anatomical resolution. It can clearly display anatomical structures, such as the boundaries and morphology of brain gray matter, muscle, fat and other tissues, and helps to observe the morphological and structural changes of tissues.
[0033] T1ce modality: T1ce modality is based on T1-weighted imaging. Through intravenous injection of contrast agents (such as gadolinium), the contrast agents are taken up by vascular-rich tissues or lesions, thereby changing the relaxation time of the tissues and increasing the signal contrast between the lesions and the surrounding normal tissues. It can more clearly display the scope, morphology and internal structure of the lesions, especially for some small lesions or equal-signal lesions that are difficult to find on T1WI. The detection rate of lesions can be improved through enhanced scanning.
[0034] T2 modality: T2-weighted imaging reflects the difference in T2 relaxation time (T2 value) between tissues. It is very sensitive to fluid and edematous tissue and can clearly show the water content and distribution in the tissue. Therefore, it has advantages in showing the edema range, cystic lesions and inflammation of the lesions.
[0035] FLAIR mode: It is a special T2-weighted imaging sequence. Based on T2WI, it uses a long inversion time (TI) to suppress the signal of cerebrospinal fluid, so that the cerebrospinal fluid appears as a low signal on the image, while the diseased tissue (such as tumor, inflammation, infarction, etc.) still shows a high signal due to its long T2 value and is not suppressed, thereby improving the contrast between the lesion and the cerebrospinal fluid, and is more conducive to the discovery of lesions located around the ventricles, in the sulci and other parts adjacent to the cerebrospinal fluid.
[0036] Embodiment 1
[0037] The embodiment of the present application provides a high-grade glioma grading method based on multimodal completion and subregion segmentation, which generates full-modality data based on the residual modality MRI image data set through the completion subnetwork, providing a complete information basis for subsequent analysis, and reasonably assigning weights to different modalities through the subregion segmentation subnetwork and highlighting the characteristics of glioma subregions, thereby achieving fine segmentation of glioma subregions to accurately grade high-grade gliomas. Specifically, reference Figure 1 , the method comprising:
[0038] Obtain a residual modality MRI image dataset of the same patient, and input the residual modality MRI image dataset into a pre-trained completion subnetwork to obtain a full-modality MRI image dataset, wherein the full-modality MRI image dataset includes MRI image data of four modalities: T1, T1ce, T2, and FLAIR, and the residual modality MRI image dataset is an MRI image dataset with missing modalities;
[0039] The full-modality MRI data set is input into a pre-trained sub-region segmentation sub-network to obtain a sub-region segmentation result. The sub-region segmentation sub-network includes a segmentation encoding unit and a segmentation decoding unit. The segmentation encoding unit performs multi-modal fusion on the full-modality MRI image data set to obtain a fusion result, and assigns different weights to different modalities based on a channel attention mechanism during the fusion process. The segmentation decoding unit increases the weight of the glioma sub-region feature in the fusion result based on an attention gating mechanism to obtain a sub-region segmentation result;
[0040] Glioma grading is performed based on the subregion segmentation results.
[0041] In some specific embodiments, full-modality MRI image datasets of different patients are obtained from the public dataset BraTS2021 and the hospital clinical dataset, and these data are preprocessed to obtain a training dataset.
[0042] Furthermore, the diagnostic information of each MRI image data in the training data set is marked, and the four modal MRI image data of each patient in the training data set are paired in pairs to obtain multiple groups of paired data corresponding to each patient, and each group of paired data of each patient is continuously input into the completion subnetwork for iterative training, and an iteration stop condition is set. When the iteration stop condition is met, the training is completed to obtain a pre-trained completion subnetwork, wherein the completion subnetwork is a generation network, which is used to generate MRI images of the missing modality according to the input residual modality MRI image data set.
[0043] Specifically, since the completion subnetwork is a generative network, the result of each round of iterative training is discriminated by constructing a discriminator, and the iterative stopping condition is that the loss function is less than a set value or reaches a preset number of iterations.
[0044] Specifically, since the completion subnetwork in this solution is a generative model, it generates missing MRI image data of another modality based on MRI image data of any two or three modalities. Therefore, when training the completion subnetwork, by pairing the four modal MRI image data corresponding to the same patient, the completion subnetwork can learn the correlation features between any two modal MRI image data, thereby better generating MRI image data of other modalities.
[0045] Furthermore, this scheme pairs the MRI image data of the four modalities in pairs to obtain 12 sets of paired data corresponding to each patient. The 12 sets of paired data are T1ce-T1, T2-T1, FLAIR-T1, T1-T1ce, T2-T1ce, FLAIR-T1ce, T1-T2, T1ce-T2, FLAIR-T2, T1-FLAIR, T1ce-FLAIR, and T2-FLAIR.
[0046] In some specific embodiments, the pre-trained completion sub-network is tested by constructing a test data set to verify the training result of the completion sub-network. During the test, MRI image data of any three modalities are provided to the pre-trained completion sub-network, and the completion sub-network generates MRI image data of another modality. The generation effect of the completion sub-network is judged by comparing the MRI image data generated by the completion sub-network with the real MRI image data. In this solution, the generation effect of the full sub-network is judged based on the peak signal-to-noise ratio, the structural similarity index and the comprehensive index, as shown in Table 1:
[0047] Table 1 The generation effect of the completion sub-network
[0048]
[0049] Table 1 shows the mean and standard deviation of the peak signal-to-noise ratio, structural similarity index and comprehensive index. PSNR is the peak signal-to-noise ratio, SSIM is the structural similarity index, and CM is the comprehensive index. The comprehensive index is the mean value after the deviation of the peak signal-to-noise ratio and the structural similarity index is normalized. The peak signal-to-noise ratio, structural similarity index and comprehensive index can be used to comprehensively evaluate the performance of the completion sub-model in generating MRI images of different modalities.
[0050] It should be noted that 95% CI refers to the 95% confidence interval, which is a statistical concept that means that in the case of repeated sampling and calculation of confidence intervals, approximately 95% of the confidence intervals will contain the true population parameter value. Taking the 95% CI of PSNR as an example, assuming that the PSNR of the same set of images is measured multiple times or calculated by different methods to obtain a series of PSNR values, and then an interval range is calculated based on these values. This range has a 95% probability of containing the true PSNR value. Taking the 95% CI of PSNR as (39.818, 41.647) as an example, this means that there is a 95% confidence that the true PSNR value of the processed image and the original image is within the range of 39.818 to 41.647. In addition, the bold data in Table 1 are the best performing data in each group.
[0051] In some embodiments, the structure of the completion subnetwork is as follows: Figure 2 As shown, the completion subnetwork is constructed by a series of completion coding units, information bottleneck units and completion decoding units. The completion coding unit extracts the local hierarchical structure in the residual modality MRI image data set through multiple coding layers of different scales to obtain a completion coding result. The information bottleneck unit compresses and extracts the completion coding result to obtain a compressed result, and then uses the completion decoding unit to decode the compressed result to obtain a full modality MRI image data set.
[0052] Specifically, the complementary coding unit in this scheme is composed of four convolutional encoders with scales from large to small connected in series, so as to extract the local hierarchical structure in the residual modality MRI image dataset, wherein the residual modality MRI image dataset is input into the complementary coding unit in a multi-channel manner, and the input feature map is converted into the local hierarchical structure in the encoding process. Mapping to embedded latent feature map On, among them, is the number of channels, is the height, is the width of the feature map. During the training of the completion subnetwork, the convolutional encoder continuously learns the potential structure representation in the MRI image through the convolution operator.
[0053] Furthermore, the information bottleneck unit includes at least one feature compression module, and the structure of the feature compression module (ART) is as follows: Figure 3 As shown, the feature compression module includes a Transformer encoding layer, a channel compression layer and a residual output layer. The Transformer encoding layer processes the completion encoding result through a cascade of multi-head self-attention and a multi-layer perceptron to obtain a first feature map. The channel compression layer performs channel compression on the first feature map to obtain a second feature map. The residual output layer performs residual processing on the second feature map and outputs it to obtain a third feature map. The third feature map output by the last feature compression module is used as the compression result.
[0054] Specifically, before the complementary coding result is input into the Transformer coding layer, the complementary coding result is first downsampled, and then the downsampled result is divided into a plurality of non-overlapping blocks, and these non-overlapping blocks are flattened to obtain a flattened result, and the flattened result is projected to A plurality of feature embeddings and a position encoding of each feature embedding are obtained in the dimensional space, and the complementary encoding result is processed by a cascade of multi-head self-attention and multi-layer perceptron in the Transformer encoding layer to obtain a first feature map.
[0055] Exemplarily, the information bottleneck unit receives the j-th layer feature map As input, due to computational limitations, the desired feature map of the Transformer layer has a smaller resolution, so downsampling is performed through the downsampling block DS to reduce the resolution of the feature map, as shown in the following formula:
[0056]
[0057] Among them, DS is implemented as a stack of strided convolutional layers for downsampling. is the downsampled feature map, where , , , M is the downsampling factor, is the number of channels after downsampling, is the height after downsampling, is the width of the downsampled feature map.
[0058] Will Split into non-overlapping blocks of size (P, P), and then flatten these non-overlapping blocks into The dimensional vector is flattened and then projected to Multiple feature embeddings are obtained in the dimensional space, supplemented by learnable position encoding, and the formula is as follows:
[0059]
[0060] in, is feature embedding, is the pth feature embedding, For embedded projection, For learnable positional encodings.
[0061] After that, the feature embedding and position encoding are input to the Transformer encoding layer for layer normalization. The multi-head self-attention result is obtained from the multi-head self-attention mechanism of the layer, and the multi-head self-attention result is residually connected with the feature embedding and the position encoding to obtain a first residual result. The first residual result is normalized and output by a multi-layer perceptron. The output result of the multi-layer perceptron is residually connected with the first residual result to obtain a second residual result, which is the output of the Transformer encoding layer.
[0062] Specifically, the multi-head self-attention mechanism The output of the layer is:
[0063]
[0064]
[0065] Among them, LN represents normalization, MSA is a multi-head self-attention mechanism using s independent self-attention heads, and the formula is expressed as:
[0066]
[0067] in, is the sth attention head ( ), represents the learnable tensor projection attention head output, z is the input feature sequence, and the SA layer calculates the weighted combination of all elements of the input sequence z. ,in is the value, attention weight It is regarded as the pairwise similarity between query q and key k and is formulated as follows:
[0068]
[0069] Finally, the output sequence of the Transformer coding layer is inversely flattened and upsampled, and then residually connected with the complementary coding result received by the information bottleneck unit to obtain the first feature map. The formula for inversely flattening and upsampling the output sequence of the Transformer coding layer is as follows:
[0070]
[0071] in, is the result of inverse flattening, US means upsampling, is the result of upsampling.
[0072] Specifically, residual connection of the upsampling result with the feature map of the completion encoding result can better extract context information, thereby facilitating the completion of MRI image data of missing modalities.
[0073] In some embodiments, two parallel convolution branches with different convolution kernel sizes are used in the channel compression layer to process the first feature map, and the results of the two parallel convolution branches are residually connected to obtain the second feature map.
[0074] Specifically, the formula of the channel compression layer is as follows:
[0075]
[0076] in, is the second feature map, CC is the channel compression layer, Represents the first feature map.
[0077] Specifically, in the channel compression layer, one of the convolution branches is a separate 1*1 convolution layer, and the other convolution branch is two serially connected 3*3 convolution layers.
[0078] In some embodiments, the residual output layer is composed of two residual blocks connected in series. The two residual blocks have the same structure, which are connected in series by a 3*3 convolution layer, a batch normalization layer, and an activation function layer. The output result of the second residual block is residually connected with the second feature map and then output by the activation function to obtain a third feature map. The formula of the residual output layer is expressed as follows:
[0079]
[0080] in, represents the output of the feature compression module of the jth network layer, Represents the residual output layer of the CNN structure.
[0081] In some specific embodiments, the number of decoding layers in the complementary decoding unit is the same as the number of encoding layers in the complementary encoding unit, and the complementary decoding unit obtains multimodal images in different channels based on the compression results. That is to say, since the compression results include a large number of multimodal MRI image data-related features, the complementary decoding unit can restore the missing modality MRI image data based on these features to achieve completion.
[0082] In some specific embodiments, the structure of the sub-region segmentation sub-network is as follows: Figure 4 As shown, in the training stage, the full-modality MRI image dataset completed in the completion subnetwork or the full-modality MRI image dataset obtained from the public dataset is used as a training sample, the glioma subregions in the training samples are marked, and the subregion segmentation subnetwork is trained, and a segmentation loss function is constructed. When the segmentation loss function reaches the set target or converges, the training is completed to obtain a pre-trained subregion segmentation subnetwork.
[0083] Furthermore, glioma subregions include necrotic core region, edema region, and enhancing tumor region.
[0084] In some embodiments, the segmentation encoding unit is composed of a first encoder and multiple second encoders connected in series, the segmentation decoding unit is composed of multiple first decoders and a second decoder connected in series, and each second encoder is jump-connected to the first decoding unit of the corresponding scale, wherein the second encoder is composed of an encoding convolutional layer and a channel attention mechanism connected in series, and the first decoder is composed of an attention gating mechanism and a decoding convolutional layer connected in series.
[0085] Furthermore, the segmentation encoding unit is used to perform multimodal fusion on the full-modality MRI image data set. In the multimodal fusion process, different modalities will have different expressiveness. Only when different weights are assigned to different modal features can the network better focus on different sub-regions of the tumor. Therefore, this scheme uses the convolutional layer and channel attention in series to construct the second encoder, so as to learn the importance of each channel through global information. The formula is expressed as follows:
[0086]
[0087] in, It is a channel The feature map of is the ReLU activation function, is the Sigmoid activation function, is a learnable weight matrix, is the learned channel importance coefficient, To assign attention weights to each channel, that is, different modalities, so as to better identify the various sub-regions of the tumor.
[0088] Furthermore, the first decoder connected in series with the decoding convolutional layer by the attention gating mechanism can automatically focus on the region of interest in the fusion result, that is, the glioma subregion to complete the subregion segmentation. Specifically, the attention gating mechanism calculates the weight of each region in the fusion result and applies the weight to the fusion result so that the segmentation subnetwork can focus on the tumor subregion. The formula is as follows:
[0089]
[0090] in, is the weighted output feature map, is the input feature map, is the target feature from the decoder, is the Sigmoid activation function. Representation feature map And the target feature map The splicing, is the learned weight matrix. is an element-by-element multiplication operation, which represents the weighting of the feature map. From the formula, we can see that the feature map is controlled by the attention gating mechanism. Weighted, where the target feature map It provides key information about the tumor area, and finally the weighted feature map is sent to the subsequent layers for processing.
[0091] In some specific embodiments, the sub-region segmentation results can be used by clinicians or pre-trained classification models to grade gliomas.
[0092] To verify the feasibility of this scheme, 20 MRI images were randomly selected from the public dataset BraTS2021, divided into 4 groups, each containing 5 images, and each group was processed with single modality loss, that is, one of the modalities of T1, T1ce, T2, and FLAIR was missing. Subsequently, a radiologist with at least 5 years of experience in the field of glioma MRI was invited to use 3D Slicer software to finely outline the three main areas of glioma: necrotic tumor core (NCR), peritumoral edema (ED), and enhancing tumor (ET). The doctor first outlined the NCR, ED, and ET labels on the full-modality image (T1, T1ce, T2, and FLAIR are all complete) as the GROUND TRUTH; two weeks later, the doctor performed the same outline again on the image after single-modality loss processing. Based on GROUND TRUTH, the five indicators of DSC, HD95, Acc, Rec and Spec were calculated in the case of single modality loss, and these indicators were normalized by MIN-MAX to obtain the normalized Dice similarity coefficient (Norm_DSC), normalized HD95 (Norm_HD95), normalized accuracy (Norm_Acc), normalized recall (Norm_Rec) and normalized specificity (Norm_Spec). In addition, a radar chart of the doctor's ability to identify the internal features of gliomas in the absence of MRI modality was drawn, as shown in Figure 5 As shown, Figure 5 (a) shows the normalized performance index of the NCR region; Figure 5 (b) shows the normalized performance index of the ED region; Figure 5 (c) shows the normalized performance index of the ET region; Figure 5 (d) in the figure shows the overall normalized performance index after weighting the NCR, ED, and ET regions by the GROUD TRUTH area. These radar charts compare the differences in doctors' recognition ability in the absence of T1, T1ce, T2, and FLAIR modalities through five dimensions (Norm_DSC, Norm_HD95', Norm_Acc, Norm_Rec, Norm_Spec). Since the smaller the value of HD95, the better the segmentation effect, and the larger the value of other indicators (such as DSC, Acc, Rec, Spec), the better the performance, in order to unify the directionality in the radar chart (that is, the larger the indicators, the better), Norm_HD95 is converted from 1 to Norm_HD95 when drawing the radar chart to generate Norm_HD95'.
[0093] about Figure 5 In the NCR region ( Figure 5 In (a) of the figure, the absence of T1ce modality significantly reduced the doctors’ segmentation performance, especially in the dimensions of Norm_DSC (T1ce absence vs. T1 absence, P=0.010) and Norm_Rec (T1ce absence vs. T1 absence, P=0.031), showing that the doctors’ accuracy in outlining the NCR boundary dropped significantly. In contrast, the absence of T1 modality had less impact on NCR, and all indicators remained at a high level. This may be because the T1ce modality contains a lot of information related to the T1 modality, so when T1ce is present, the absence of T1 modality has limited impact on recognition. For the ED area ( Figure 5 In (b), the loss of T1ce and FLAIR modalities had the most significant impact on Norm_DSC (T1ce loss vs. T1 loss, P=0.010; FLAIR loss vs. T1 loss, P=0.018) and Norm_HD95 (T1ce loss vs. T1 loss, P=0.018; FLAIR loss vs. T1 loss, P=0.010), which was manifested by a significant decrease in segmentation accuracy and boundary consistency. ET region ( Figure 5 (c) is particularly affected by T1ce loss, with Norm_DSC (T1ce loss vs. T1 loss, P=0.010; T1ce loss vs. T2 loss, P=0.008) and Norm_Rec (T1ce loss vs. T1 loss, P=0.010; T1ce loss vs. T2 loss, P=0.008) both significantly decreased, indicating that T1ce modality plays a key role in identifying enhancement areas and core necrosis areas. Overall performance ( Figure 5 (d) in the figure combines the performance of the above regions and verifies the negative impact of FLAIR and T1ce modality loss on the overall recognition performance, especially in Norm_DSC (T1ce loss vs. T1 loss, P=0.010; FLAIR loss vs. T1 loss, P=0.010) and Norm_Rec (FLAIR loss vs. T1 loss, P=0.018; FLAIR loss vs. T1 loss, P=0.018). These radar maps clearly show that FLAIR and T1ce modalities are crucial for the accuracy and consistency of glioma recognition, especially in the ED and ET regions. Therefore, in clinical practice, maintaining the integrity of these modalities is crucial for high-quality tumor segmentation and diagnosis.
[0094] This solution converts Norm_HD95 from 1 to Norm_HD95 to generate Norm_HD95'. The calculation process of the Overall label is as follows:
[0095] First, calculate the GROUND TRUTH area of each region , and calculate the weight of each region based on its area ratio :
[0096]
[0097] Then, the DSC and HD95 indicators of each region are weighted averaged to obtain the Overall DSC and HD95 values. The calculation formula is:
[0098]
[0099]
[0100] Through these weighted indicators, the doctor's ability to recognize the internal features of glioma is comprehensively evaluated as shown in Table 2:
[0101] Table 2 Differential evaluation of doctors’ recognition performance of glioma lesion areas before and after single modality loss
[0102]
[0103] From the results in Table 2, it can be seen that the lack of T1ce modality has the most significant impact on doctors' identification of the internal features of gliomas, especially in the ED and ET areas. The HD95 value of Overall is as high as 12.502, indicating that the boundary consistency is significantly reduced and the recognition accuracy is greatly reduced. The lack of FLAIR modality also has a significant impact on the recognition of the ED area, with HD95 reaching 14.011, showing that in this case, there is a large difference between the glioma features outlined by the doctor and the GROUD TRUTH. In contrast, the lack of T1 modality has less impact on the NCR area, with Overall's DSC only slightly reduced and HD95 maintained at a low level (1.011), indicating that even in the absence of T1, doctors can still accurately identify the characteristics of this area. The main reason for this result is the presence of T1ce modality, which retains a lot of important information related to T1 modality.
[0104] Similarly, for the sub-region segmentation sub-network, the 1096 training sets containing NCR, ED, and ET labels were input into the multimodal segmentation network of this study in the form of multimodality (T1, T1ce, T2, and FLAIR) for training. After the training was completed, the 155 test sets were evaluated using the optimal parameters obtained. Finally, the performance of NCR, ED, ET, and Overall labels was measured by calculating the DSC and HD95 indicators, respectively, and Table 3 was obtained.
[0105] Table 3 Evaluation of multimodal sub-region segmentation performance of glioma MRI images
[0106]
[0107] Table 3 shows two indicators in the performance evaluation of multimodal sub-region segmentation of glioma MRI images: DSC (DiceSimilarity Coefficient) and HD95 (Hausdorff Distance 95%), which are used to evaluate the segmentation effect of different label areas (NCR, ED, ET and Overall). The table lists the DSC value and HD95 value corresponding to each label to quantify the performance of the segmentation algorithm in different areas. Among them, "Overall" represents a comprehensive evaluation of the overall segmentation performance of all label areas (NCR, ED, ET), reflecting the average performance of the segmentation algorithm on the entire image. The statistical graph of glioma MRI segmentation performance evaluation is shown in the figure. Figure 6 As shown in the figure, according to the evaluation results of the Dice similarity coefficient (DSC Mean), the segmentation accuracy of the enhanced tumor (ET) and peritumoral edema (ED) is relatively high, with DSC Mean of 0.877 (ET vs. NCR, P < 0.001) and 0.873 (ED vs. NCR, P < 0.001), respectively, showing that the segmentation results of these regions are highly consistent with the true labels. Combining the performance of all label regions, the overall DSC Mean is 0.878, which further verifies that the proposed algorithm can achieve relatively accurate segmentation in most cases. However, for the necrotic tumor core (NCR) region, the DSC Mean was only 0.750, indicating that the segmentation accuracy of this region was relatively insufficient, which may reflect the potential challenge of segmentation in this region. From the 95th percentile of the Hausdorff distance (HD95), the HD95 value of the ED region was significantly higher than that of other regions, reaching 25.319 (ED vs. NCR, P < 0.001; ED vs. ET, P < 0.001), which indicates that there is a large error in the segmentation boundary in the ED region, which may have an adverse effect on the overall accuracy of the segmentation. In contrast, the HD95 values of the ET and NCR regions were 12.15 and 13.37, respectively, indicating that the segmentation errors in these regions were relatively small and the boundary segmentation was more accurate. Although the overall HD95 value was 19.491, reflecting that there were still large segmentation errors in some cases, the segmentation performance of the ET and NCR regions was relatively ideal, showing that the proposed algorithm has high segmentation reliability and boundary accuracy in these key areas.
[0108] In addition, this scheme also evaluated the role of combining the completion subnetwork with the subregion segmentation subnetwork. 181 screened clinical brain glioma MRI image data were selected. The images of these cases generally had incomplete T1, T1ce, T2 and FLAIR modalities. First, the completion subnetwork was used to complete the modality, and then the subregion segmentation subnetwork was used to perform fine segmentation on the completed images to generate NCR, ED and ET labels. In order to systematically evaluate the role of modality completion and subregion segmentation in assisting doctors to identify glioma grades, the following three-step process was designed: Step 1. Under the original incomplete modality image data, ask the doctor to judge the patient's glioma grade (grade 3 or 4); Step 2. After the modality is completed, ask the doctor to make a grade judgment again; Step 3. When the modality is completed and has subregion segmentation labels, ask the doctor to make a third grade judgment. Each evaluation was conducted 15 days apart, and in each case the physician was asked to rate the confidence in grading using a Likert scale of 1 to 7 (1 for almost certain grade 3, 7 for almost certain grade 4) based on their overall impression and professional experience. Finally, we used the pathological results as the gold standard and compared the physicians’ diagnostic accuracy under these three conditions by drawing ROC curves, thus obtaining a schematic diagram of the ROC performance comparison of physicians’ tumor grading. Figure 7 As shown, Figure 7 (a) in the figure is the ROC curve of junior doctors' grading of high-grade gliomas under three conditions: modality missing (corresponding to modality 1), modality completion (corresponding to modality 2), modality completion + segmentation label (corresponding to modality 3)). Figure 7 (b) is a heat map of the AUC differences under three conditions for junior doctors. Figure 7 (c) is a heat map of the consistency of glioma grading scores of junior doctors under three conditions. Figure 7 (d) in the figure is the ROC curve of high-grade glioma grading by senior doctors under the same modality processing conditions (modality missing, modality completion, modality completion + segmentation label). Figure 7 (e) is a heat map of the AUC differences under three conditions for senior doctors. Figure 7 (f) is a heat map of the consistency of glioma grading scores of senior doctors under three conditions. Figure 7It can be seen that for junior doctors, the AUC was 0.667 (95% CI: 0.580-0.750) when the modality was missing, and the AUC after modality completion was 0.698 (95% CI: 0.619-0.779), and under the condition of modality completion combined with segmentation labels, the AUC increased to 0.722 (95% CI: 0.640-0.798). However, statistical analysis showed that these differences were not significant, but showed low consistency. For senior doctors, the AUCs were 0.718 (95% CI: 0.643-0.797), 0.842 (95% CI: 0.775-0.905), and 0.913 (95% CI: 0.858-0.956) under the conditions of modality missing, modality completion, and modality completion combined with segmentation labels, respectively. Among them, the AUC of modality completion was significantly improved compared with modality omission (P < 0.001), with lower consistency (kappa = 0.237), while the difference of modality completion + segmentation label was more significant compared with modality omission (P < 0.001), with higher consistency (kappa = 0.423). In addition, senior doctors performed better than junior doctors in grading under all modality processing conditions, especially under the condition of modality completion combined with segmentation label, showing higher grading accuracy (P < 0.001). These results show that modality completion and segmentation label have a significant effect on improving the ROC performance of doctors in the task of grading high-grade gliomas, especially for senior doctors, these processing methods can significantly enhance the accuracy of grading.
[0109] Embodiment 2
[0110] Based on the same idea, refer to Figure 8 , the present application also proposes a high-grade glioma grading device based on multimodal completion and subregion segmentation, comprising:
[0111] A completion module is used to obtain a residual modality MRI image dataset of the same patient, and input the residual modality MRI image dataset into a pre-trained completion subnetwork to obtain a full-modality MRI image dataset, wherein the full-modality MRI image dataset includes MRI image data of four modalities: T1, T1ce, T2, and FLAIR, and the residual modality MRI image dataset is an MRI image dataset with missing modalities;
[0112] A segmentation module is used to input the full-modality MRI data set into a pre-trained sub-region segmentation sub-network to obtain a sub-region segmentation result. The sub-region segmentation sub-network includes a segmentation encoding unit and a segmentation decoding unit. The segmentation encoding unit performs multi-modal fusion on the full-modality MRI image data set to obtain a fusion result, and assigns different weights to different modalities based on a channel attention mechanism during the fusion process. The segmentation decoding unit increases the weight of the glioma sub-region feature in the fusion result based on an attention gating mechanism to obtain a sub-region segmentation result;
[0113] A grading module is used to grade glioma based on the subregion segmentation results.
[0114] Embodiment 3
[0115] This embodiment also provides an electronic device, referring to Fig. 9 , comprises a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any of the above method embodiments.
[0116] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0117] Among them, the memory 404 may include a large capacity memory 404 for data or instructions. For example, but not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, the memory 404 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 404 may be inside or outside the data processing device. In a specific embodiment, the memory 404 is a non-volatile memory. In a specific embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM can be a fast page mode dynamic random access memory 404 (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0118] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .
[0119] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the high-grade glioma grading methods based on multimodal completion and subregion segmentation in the above embodiments.
[0120] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .
[0121] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.
[0122] The input and output device 408 is used to input or output information. In this embodiment, the input information may be a residual modality MRI image data set, etc., and the output information may be a sub-region segmentation result, a glioma grading result, etc.
[0123] Optionally, in this embodiment, the processor 402 may be configured to perform the following steps through a computer program:
[0124] Obtain a residual modality MRI image dataset of the same patient, and input the residual modality MRI image dataset into a pre-trained completion subnetwork to obtain a full-modality MRI image dataset, wherein the full-modality MRI image dataset includes MRI image data of four modalities: T1, T1ce, T2, and FLAIR, and the residual modality MRI image dataset is an MRI image dataset with missing modalities;
[0125] The full-modality MRI data set is input into a pre-trained sub-region segmentation sub-network to obtain a sub-region segmentation result. The sub-region segmentation sub-network includes a segmentation encoding unit and a segmentation decoding unit. The segmentation encoding unit performs multi-modal fusion on the full-modality MRI image data set to obtain a fusion result, and assigns different weights to different modalities based on a channel attention mechanism during the fusion process. The segmentation decoding unit increases the weight of the glioma sub-region feature in the fusion result based on an attention gating mechanism to obtain a sub-region segmentation result;
[0126] Glioma grading is performed based on the subregion segmentation results.
[0127] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0128] In general, various embodiments may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the boxes, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0129] Embodiments of the present invention may be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components configured to perform an embodiment when the program is run. One or more computer executable components may be at least one software code or a portion thereof. In addition, at this point, it should be noted that, for example, Fig. 9 Any block of the logic flow in the program may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. Physical media are non-transitory media.
[0130] Those skilled in the art should understand that the technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0131] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A high-grade glioma grading method based on multimodal completion and subregion segmentation, characterized in that: The following steps are involved: A residual modality MRI image dataset of the same patient is obtained, and the residual modality MRI image dataset is input into a pre-trained completion subnetwork to obtain a full-modality MRI image dataset, wherein the completion subnetwork is constructed by a series-connected completion coding unit, an information bottleneck unit, and a completion decoding unit, wherein the completion coding unit extracts the local hierarchical structure in the residual modality MRI image dataset through a plurality of coding layers of different scales to obtain a completion coding result, and the information bottleneck unit compresses and extracts the completion coding result to obtain a compression result, and then uses the completion decoding unit to decode the compression result to obtain a full-modality MRI image dataset, wherein the completion subnetwork is a generation network, which is used to generate MRI images of missing modalities according to the input residual modality MRI image dataset, and the full-modality MRI image dataset is MRI image data including four modalities of T1, T1ce, T2 and FLAIR, and the residual modality MRI image dataset is an MRI image dataset with modality missing; The full-modality MRI data set is input into a pre-trained sub-region segmentation sub-network to obtain a sub-region segmentation result. The sub-region segmentation sub-network includes a segmentation encoding unit and a segmentation decoding unit. The segmentation encoding unit performs multi-modal fusion on the full-modality MRI image data set to obtain a fusion result, and assigns different weights to different modalities based on a channel attention mechanism during the fusion process. The segmentation decoding unit increases the weight of the glioma sub-region feature in the fusion result based on an attention gating mechanism to obtain a sub-region segmentation result; Glioma grading is performed based on the subregion segmentation results.
2. The method for grading high-grade gliomas based on multimodal completion and subregion segmentation according to claim 1, characterized in that: The diagnostic information of each MRI image data in the training data set is marked, and the four modal MRI image data of each patient in the training data set are paired with each other to obtain multiple groups of paired data corresponding to each patient. Each group of paired data of each patient is continuously input into the completion subnetwork for iterative training, and the iteration stop condition is set. When the iteration stop condition is met, the training is completed to obtain the pre-trained completion subnetwork.
3. The method for grading high-grade gliomas based on multimodal completion and subregion segmentation according to claim 1, characterized in that: The information bottleneck unit includes at least one feature compression module, which includes a Transformer encoding layer, a channel compression layer and a residual output layer. The Transformer encoding layer processes the completion encoding result through a cascade of multi-head self-attention and a multi-layer perceptron to obtain a first feature map, the channel compression layer performs channel compression on the first feature map to obtain a second feature map, the residual output layer performs residual processing on the second feature map and outputs a third feature map, and the third feature map output by the last feature compression module is used as the compression result.
4. The method for grading high-grade gliomas based on multimodal completion and subregion segmentation according to claim 3, characterized in that: Before the complementary coding result is input into the Transformer coding layer, the complementary coding result is first downsampled, and then the downsampled result is divided into multiple non-overlapping blocks, and these non-overlapping blocks are flattened to obtain a flattened result, and the flattened result is projected to A plurality of feature embeddings and a position encoding of each feature embedding are obtained in the dimensional space, and the complementary encoding result is processed in the Transformer encoding layer through a cascade of multi-head self-attention and a multi-layer perceptron to obtain a first feature map.
5. The method for grading high-grade gliomas based on multimodal completion and subregion segmentation according to claim 3, characterized in that: In the channel compression layer, two parallel convolution branches with different convolution kernel sizes are used to process the first feature map, and the results of the two parallel convolution branches are residually connected to obtain the second feature map.
6. The method for grading high-grade gliomas based on multimodal completion and subregion segmentation according to claim 1, characterized in that: The segmentation encoding unit is composed of a first encoder and multiple second encoders connected in series, the segmentation decoding unit is composed of multiple first decoders and a second decoder connected in series, and each second encoder is jump-connected to the first decoding unit of the corresponding scale, wherein the second encoder is composed of an encoding convolutional layer and a channel attention mechanism connected in series, and the first decoder is composed of an attention gating mechanism and a decoding convolutional layer connected in series.
7. A high-grade glioma grading device based on multimodal completion and subregion segmentation, characterized in that: include: A completion module is used to obtain a residual modality MRI image dataset of the same patient, and input the residual modality MRI image dataset into a pre-trained completion subnetwork to obtain a full-modality MRI image dataset, wherein the completion subnetwork is constructed by a series of completion coding units, information bottleneck units, and completion decoding units, wherein the completion coding unit extracts the local hierarchical structure in the residual modality MRI image dataset through multiple coding layers of different scales to obtain a completion coding result, and the information bottleneck unit compresses and extracts the completion coding result to obtain a compression result, and then uses the completion decoding unit to decode the compression result to obtain a full-modality MRI image dataset, wherein the completion subnetwork is a generation network, and is used to generate MRI images of missing modalities according to the input residual modality MRI image dataset, and the full-modality MRI image dataset is MRI image data including four modalities of T1, T1ce, T2, and FLAIR, and the residual modality MRI image dataset is an MRI image dataset with modality missing; A segmentation module is used to input the full-modality MRI data set into a pre-trained sub-region segmentation sub-network to obtain a sub-region segmentation result. The sub-region segmentation sub-network includes a segmentation encoding unit and a segmentation decoding unit. The segmentation encoding unit performs multi-modal fusion on the full-modality MRI image data set to obtain a fusion result, and assigns different weights to different modalities based on a channel attention mechanism during the fusion process. The segmentation decoding unit increases the weight of the glioma sub-region feature in the fusion result based on an attention gating mechanism to obtain a sub-region segmentation result; A grading module is used to grade glioma based on the subregion segmentation results.
8. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute a high-grade glioma grading method based on multimodal completion and subregion segmentation as described in any one of claims 1-6.
9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes a high-grade glioma grading method based on multimodal completion and subregion segmentation according to any one of claims 1-6.
Citation Information
Patent Citations
Hippocampus subregion segmentation method and system based on ultrahigh field magnetic resonance image reconstruction
CN116071383A
MRI brain tumor segmentation method based on attention bottleneck fusion
CN118314350A