Multi-head semantic supervision brain tumor three-dimensional segmentation method and system

By introducing multi-head semantic supervision method and CLIP Prompt technology in three-dimensional medical image segmentation, combined with the prompt attention fusion module, the problems of channel information obfuscation and semantic information neglect in the existing technology are solved, and higher segmentation accuracy and efficiency are achieved.

CN120163983APending Publication Date: 2025-06-17CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325171.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing three-dimensional medical image segmentation method is prone to channel information confusion and prediction category errors when facing multiple segmentation areas, and ignores the role of semantic information, affecting the accuracy and efficiency of segmentation.

Method used

A three-dimensional segmentation method of brain tumors with multi-head semantic supervision is proposed, combining CLIP Prompt and the cues attention fusion module (PAF module), using multi-segment head network and three-dimensional visual language pre-training technology, a semantic supervision mechanism is introduced to improve segmentation accuracy.

Benefits of technology

It effectively avoids channel information confusion and incorrect prediction, improves the accuracy and efficiency of multi-objective segmentation, significantly improves the segmentation accuracy and efficiency, and achieves high matching of graphics and text features in three-dimensional space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163983A_ABST
    Figure CN120163983A_ABST
Patent Text Reader

Abstract

The invention provides a multi-head semantic supervision brain tumor three-dimensional segmentation method and system. Comprising the following steps: 1) a novel multi-head semantic supervision three-dimensional medical image segmentation method, wherein a network model of the method is composed of a three-dimensional body image codec, a multi-segmentation-head network and a prompt vector generation network (PVG module); and 2) in order to introduce a CLIP prompt driven segmentation method, designing a prompt vector generation network based on CLIP model text feature-image feature matching. And 3) in order to perform semantic supervision segmentation on the three-dimensional reconstruction feature image by using prompt information, fusing the prompt information of different tumor areas with the feature image, proposing a prompt attention fusion module (PAF module: Prompt attention fusion), and designing a multi-head semantic supervision driven multi-segmentation head network. And 4) a CLIP model is originally based on two-dimensional image training and is difficult to be directly applied to a three-dimensional image, the invention provides a three-dimensional visual language pre-training method, a two-dimensional CLIP text is embedded and mapped to a three-dimensional space, and the two-dimensional CLIP text and three-dimensional brain tumor features are subjected to comparative learning, so that the matching consistency of the image-text features in a combined space is achieved. And 5) in order to guarantee the training effect of the model, a stage-type training method is provided, three-dimensional visual language pre-training is carried out in the first stage of training, and model segmentation training is carried out in the second stage of training. And 6) in order to assist doctors to observe a plurality of target tumor parts more clearly during brain tumor disease diagnosis, the invention designs a brain tumor MRI three-dimensional segmentation prototype system which is simple to operate and complete in function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional medical image segmentation, and implements a multi-head semantically supervised three-dimensional brain tumor segmentation method and system in combination with CLIP Prompt. Background Art

[0002] Existing technology: As the incidence of brain tumors increases year by year, the clinical demand for accurate three-dimensional segmentation is becoming increasingly urgent. Traditional segmentation relies on radiologists to manually outline MRI images, which is time-consuming, highly subjective, and limited by differences in physician experience. It can easily lead to missed diagnosis of tiny tumors or misjudgment of boundaries, affecting surgical planning and efficacy evaluation. Deep learning methods have greatly promoted the development of automated segmentation. Three-dimensional volume data is directly input into the model, and three-dimensional segmentation predictions are output. Doctors can intuitively observe the target image area in each dimension, and then judge whether lesions or atrophy occur based on clinical experience.

[0003] At present, the most widely used architecture for 3D medical image segmentation is the 3D-UNet structure. Due to its flexible, optimized modular design and success in all medical image modes, it has attracted great attention from academic and industrial researchers. However, for some targets that occupy too small a volume in the entire medical image, have blurred edges and no fixed shape, the existing segmentation methods still have some problems: 1) The 3D-UNet structure uses a single segmentation head, which is prone to channel information confusion and prediction category errors when facing multiple segmentation regions. For example, the "enhanced tumor" channel mistakenly predicts the "tumor core" area, which needs to be avoided in medical diagnosis. 2) At present, most medical image segmentation research focuses on the model's ability to capture detailed features, while ignoring the role of semantic information. In medical multi-target segmentation, the anatomical relationship between targets can be used as a semantic cue information, such as the similarity and inclusion relationship between "whole tumor" and "tumor core".

[0004] Purpose: The development of visual language pre-training technology allows researchers to use the category information contained in the organ itself to obtain the corresponding high-quality visual features, that is, to visualize the anatomical relationship between multiple medical targets, so that specific text embeddings correspond to specific organ visual features. Therefore, this technology proposes a multi-head semantically supervised 3D brain tumor segmentation method, which combines multiple segmentation heads with multi-channel semantic cues to effectively avoid the confusion of channel information and incorrect prediction of a single segmentation head. At the same time, the introduction of the CLIP model can help supervise the segmentation task by matching text cues and image features, effectively improving the accuracy and efficiency of segmentation. Summary of the invention

[0005] Technical solution and invention key points: Based on the above problems, the present invention proposes a three-dimensional segmentation method and system for brain tumors with multi-head semantic supervision. In order to introduce the CLIP prompt-driven segmentation method, a prompt vector generation network based on the matching of text features and image features of the CLIP model is designed. In order to use the prompt information for semantic supervised segmentation of the three-dimensional reconstruction feature image, the prompt information of different tumor regions is fused with the feature image, and a prompt attention fusion module (PAF module: Prompt attention fusion) is proposed, and a multi-segmentation head network driven by multi-head semantic supervision is designed. The CLIP model was originally trained based on two-dimensional images and is difficult to be directly applied to three-dimensional images. The present invention proposes a three-dimensional vision-language pre-training method, which maps the two-dimensional CLIP text embedding to the three-dimensional space and performs contrastive learning with the three-dimensional brain tumor features to achieve the matching consistency of text and image features in the joint space. In order to ensure the training effect of the model, a staged training method is proposed. In the first stage of training, three-dimensional vision-language pre-training is carried out to obtain three-dimensional text embeddings. In the second stage, the segmentation model is trained.

[0006] In view of this, the specific technical solution adopted by the present invention is as follows:

[0007] A three-dimensional segmentation method for brain tumors with multi-head semantic supervision, comprising the following steps:

[0008] 1) First, preset the prompt text templates corresponding to different regions of the brain tumor, such as "tumor core tomography", "enhanced tumor tomography", etc.

[0009] 2) Adopt a two-stage training method. In the first stage, three-dimensional vision-language pre-training is carried out to obtain three-dimensional text embeddings. Refer to the contrastive learning method of the CLIP model to fine-tune the prompt text and the brain tumor segmentation labels, and then directly use this part of the network in the prompt vector generation network.

[0010] 3) In the second stage of training, that is, when training the segmentation model, first send the prompt text into the prompt vector generation network. In this network, the prompt text is converted into a two-dimensional text embedding by the CLIP text encoder, and then remapped to the three-dimensional space to adapt to the three-dimensional image data, and the obtained three-dimensional text embedding is reserved.

[0011] 4) Input the brain tumor MRI data to be segmented into the volume image encoder to obtain global image features, fuse them with the three-dimensional text embeddings, and obtain the prompt vector for semantic supervision (including the prompts of different tumor regions) through the perception layer. The present invention uses the prompt vector as a feature-guided prompt information to make the model pay more attention to the feature images of certain channels.

[0012] 5) Feed the global image features into the decoder for n - layer decoding to obtain a brain tumor feature map as rich as possible. To enhance the feature reconstruction effect, the multi - view cross - slice auxiliary module in the 3V3D model is introduced in the present invention, and finally a rich global feature map is obtained.

[0013] 6) Feed the feature guiding prompt information of different tumor regions and the global feature map into the PAF module in the multi - segmentation head network respectively. The prompt vector (feature guiding prompt information), as a supervision mechanism, can effectively prompt the model which features on which channels are more important for the corresponding organs, and supervise the segmentation of the corresponding tumor regions.

[0014] 7) Finally, after m - layer decoding in each segmentation head, the corresponding dense prediction of tumor segmentation is obtained.

[0015] 8) Calculate the loss function and update the parameters of the segmentation model through gradient propagation.

[0016] A three - dimensional brain tumor segmentation system, comprising:

[0017] 1) Model management module: This module integrates the segmentation method proposed in the present invention and the current mainstream segmentation methods, and is used for users to select a specific segmentation model and construct it.

[0018] 2) Upload image module: This module is used to upload local MRI images as the original images for segmentation.

[0019] 3) Execution segmentation module: This module is used for model inference and the process of performing segmentation.

[0020] 4) Result display module: This module is used for the display and saving of segmentation results.

[0021] In summary, the beneficial effects of the present invention are as follows:

[0022] 1) The three - dimensional brain tumor segmentation method with multi - head semantic supervision in the present invention effectively prevents situations such as channel information confusion and prediction category errors in conventional single - segmentation - head prediction, and at the same time considers the anatomical semantic relationship between multiple medical targets, improving the accuracy and efficiency of multi - target segmentation.

[0023] 2) CLIP Prompt is introduced. By combining different descriptions of tumor regions, the model can understand the different features of each region during the training process, which can significantly improve the segmentation accuracy and efficiency. The proposed PAF module effectively converts the semantic prompt information into a kind of channel - like attention mechanism information, prompts the beneficial channel features of the current segmentation head and increases their weights, and supervises the corresponding segmentation head to segment the corresponding tumor region.

[0024] 3) Extend the two-dimensional visual language pre-training method to the three-dimensional field, connect the text encoder of the CLIP model with a multi-layer perceptron, so that the two-dimensional text embedding is mapped to the three-dimensional space, and paired training can be carried out with the three-dimensional image features. Through contrastive learning training, the three-dimensional brain tumor text and image features have a high degree of matching in the joint space and can be applied to various downstream tasks such as segmentation and classification.

[0025] 4) Provide a three-dimensional brain tumor segmentation system, which integrates the multi-head semantic supervision segmentation method described in the present invention and the current mainstream three-dimensional segmentation methods, facilitating doctors to more clearly observe multiple target tumor parts (including the tumor core area, the overall tumor area, and the enhanced tumor area) during the diagnosis of brain tumor diseases and facilitating researchers to test the segmentation efficiency of the model. Description of the Drawings

[0026] Figure 1 It is the model framework diagram of the present invention;

[0027] Figure 2 It is the schematic diagram of the training method of the present invention;

[0028] Figure 3 It is the structure diagram of the multi-segmentation head network and the prompt vector generation network of the present invention;

[0029] Figure 4 It is the three-dimensional brain tumor segmentation system of the present invention. Detailed Embodiments

[0030] The present invention will be described in detail below in conjunction with embodiments. These embodiments are only illustrative and are not limited to the application scope of the present invention. The present invention is not limited to the following embodiments or implementation manners. Any modifications and deformations made without departing from the spirit of the present invention shall be included within the scope of the present invention.

[0031] The technical solution of the present invention will be described in detail below in conjunction with the drawings:

[0032] A three-dimensional brain tumor segmentation method with multi-head semantic supervision includes:

[0033] Step 1: First, preset the prompt text templates corresponding to different regions of the brain tumor. In the present invention, it is set in the mode of "organ name" + "tomographic scan", and the effects of different templates can be experimented in specific implementations.

[0034] Step 2: Adopt phased training. First is the first-stage training: three-dimensional visual language pre-training. The overall framework is as Figure 2As shown in the left half of the training method, it consists of a text encoder (encoder2) and a volumetric image encoder (encoder1). Through joint training with a large number of medical image-text pairs, the two features are brought closer in the joint space. At the same time, k layers of MLP are connected at the end of encoder2 (refer to Figure 3 as shown in the upper half) to introduce the two-dimensional text embedding into the three-dimensional space, facilitating paired training with the three-dimensional image features.

[0035] Step 3: In each training batch, the input sample pairs are {T1, Y1}...{T N , Y N}, where N represents the number of sample pairs in a batch, and T and Y represent the text and the segmentation label respectively. During the training process, the sample pairs pass through the two encoders to obtain N text embeddings and the segmentation label encodings Calculate the cosine similarity between each and , then a two-dimensional cosine similarity matrix is generated. The model is optimized by minimizing the similarity loss function on the positive sample pairs, and the objective function is as follows:

[0036]

[0037] where, represents the similarity between and , and τ is the score scale, which is a learnable parameter. Through contrastive learning training, the distance between the medical image-text features in the unified space is brought closer, and the model can better adapt to the downstream tasks of medical image processing.

[0038] Step 4: Then comes the second-stage training, that is, the training of the segmentation network. Refer to Figure 2 the right half. The weights of the text encoder pre-trained in the first stage are directly frozen and used as an important part of the prompt vector generation network. The generated prompt information will supervise the multi-segmentation head network for multi-object segmentation. First, the prompt text p t is converted into a two-dimensional text embedding by the CLIP text encoder. By remapping it to the three-dimensional space to adapt to the three-dimensional image data, the spare three-dimensional text embedding p e is obtained.

[0039] p e = MLP k (TextEncoder(p t )

[0040] Step 5: Input the brain tumor MRI data into the three-dimensional volumetric image encoder to obtain the global image feature x eHere, the encoder of 3DResUNet is adopted, which has a total of four downsampling modules and one dilated convolutional layer. The formula is as follows:

[0041]

[0042] Before x e is fed into the decoder, it is sent into the prompt vector generation network.

[0043] Step 6: The global image feature x e is sent into the prompt vector generation network and combined with the 3D text embedding p e to generate the prompt vector p v for different tumor regions, which is used to supervise the model for corresponding semantic region segmentation.

[0044]

[0045] Step 7: Return to the decoding process. First, x e is decoded for n layers, and the multi-view cross-slice auxiliary module is introduced. This module takes the feature changes between different slices as an attention to enhance the learning ability of the global edge features and calculates the cross-slice difference information as auxiliary information (for the specific method, refer to the 3V3D model). At the same time, skip connections are added to enhance the context connection to obtain the rich feature map f.

[0046] x = x + x skip

[0047] f = Interpolation n (x e )

[0048] Step 8: The feature map f is sent into the corresponding PAF module in the multi-segmentation head network and fused with the corresponding prompt vector p v The PAF module imitates the channel attention mechanism to transform p v into different channel weights w v representing the importance of the features of different channels of the feature map f for the corresponding tumor regions, to supervise the decoding of the corresponding segmentation head.

[0049]

[0050] Among them, i represents the i-th segmentation head route.

[0051] After that, the prompt attention weight is fused with the feature map. Specifically, let be the value at the k-th position of the feature supervision weight of the i-th segmentation target, and f i,k be the feature image on the k-th channel of the feature image of the i-th segmentation target. The fusion formula is as follows:

[0052]

[0053] Among them, N is the number of channels of f i and represents the feature image after class supervision, that is, the feature image of the i-th segmentation target. is constrained to the interval of 0-1 by the sigmoid function. Then when is closer to 1, it acts as an excitation for f i,k and when is closer to 0, it acts as an inhibition for f i,k

[0054] Step Nine: Finally, each segmentation head decodes m times using 3D ResUNet to obtain the segmentation regions of their respective tumor regions. The formula is as follows:

[0055]

[0056] Among them, Conv represents the 3D convolution operation, and S i is the segmentation prediction of the i-th tumor region.

[0057] Step Ten: Train the segmentation network using a composite loss function. The formula is as follows:

[0058]

[0059] Among them, c represents the number of segmentation heads (the number of tumor regions), and i represents the i-th tumor region. The diff loss is the multi-view cross-slice loss. Referring to the 3V3D model, the dice loss is used to measure the similarity metric of the pixel sets between the segmentation label and the segmentation prediction. The bce loss is the binary cross-entropy loss function, which is used to measure the accuracy of the segmented pixels. At the same time, a weight coefficient λ is added to balance the loss.

[0060] A three-dimensional brain tumor segmentation system, as Figure 4 shown, includes:

[0061] Model Management Module: This module integrates the segmentation method proposed in the present invention and the current mainstream segmentation methods, and is used for users to select a specific segmentation model and construct it. For example, when the method described in the present invention is selected, a three-dimensional brain tumor segmentation model with multi-head semantic supervision will be automatically constructed and the weights will be loaded, including a volume image codec, a prompt vector generation network, and a multi-segmentation head network.

[0062] Upload Image Module: This module is used to upload local MRI images as the original images for segmentation. As Figure 4 shown in the right half, when the user uploads an image, it can be displayed in a specific area.

[0063] ​Execution Segmentation Module: This module is used for model inference and the process of performing segmentation. The original image is input into the model, and the segmentation result is output end-to-end and displayed in the lower right area of the system interface, as Figure 4 shown.

[0064] Result Display Module: This module is used for the display and saving of the segmentation result, providing a three-dimensional display function of the segmentation result, which allows users to perform operations such as dragging, zooming, and rotating, and optionally save the local operations.

[0065] The conventional technologies and the solutions not described in detail in the above embodiments are well-known in the art, so they will not be elaborated here in detail. The above embodiments and / or experimental examples have described in detail the preferred embodiments of the present invention. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.

Claims

1. A multi-head semantically supervised 3D brain tumor segmentation method, characterized in that: The following steps are involved: 1) First, preset prompt text templates corresponding to different areas of brain tumors, such as "tumor core tomography", "enhanced tumor tomography", etc. 2) A two-stage training method is adopted. In the first stage, 3D visual language pre-training is performed to obtain 3D text embedding. The prompt text and brain tumor segmentation labels are fine-tuned by referring to the CLIP model for comparison of learning methods. This part of the network is directly used in the prompt vector generation network later. 3) During the second stage of training, i.e., the segmentation model training, the prompt text is first sent to the prompt vector generation network, in which the prompt text is converted into a two-dimensional text embedding by the CLIP text encoder, and then mapped to the three-dimensional space to adapt to the three-dimensional image data, and the obtained three-dimensional text embedding is used for backup. 4) Input the brain tumor MRI data to be segmented into the volume image encoder to obtain global image features, which are fused with the three-dimensional text embedding, and the prompt vector (including prompts of different tumor areas) for semantic supervision is obtained through the perception layer. The present invention uses the prompt vector as a feature-guided prompt information, so that the model pays more attention to the feature images of certain channels. 5) The global image features are sent to the decoder for n-layer decoding to obtain a brain tumor feature map that is as rich as possible. In order to enhance the feature reconstruction effect, the present invention introduces a multi-view cross-slice auxiliary module in the 3V3D model, and finally obtains a rich global feature map. 6) The feature-guided hint information of different tumor regions is sent to the PAF module in the multi-segmentation head network together with the global feature map. The hint vector, as a semantic supervision mechanism, can effectively prompt the model which channel features are more important for the corresponding organs and supervise the segmentation of the corresponding tumor region. 7) Finally, after m decodings at each segmentation head, the corresponding tumor segmentation dense prediction is obtained. 8) Calculate the loss function and propagate the gradient to update the parameters of the segmentation model.

2. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: Step 1) establishing corresponding prompt text templates for different areas of the brain tumor.

3. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: Step 2) 3D visual language pre-training. Pair the CLIP model text encoder with the 3DResUNet image encoder. Since 2D text embedding cannot be paired with 3D image features for learning, we add k layers of MLP and 3D interpolation layers after the text encoder to embed 2D text into 3D. Then, refer to the CLIP model to feed the preset prompt text and brain tumor segmentation labels for comparative learning training to improve the uniformity of brain tumor image and text features in the joint space.

4. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: In step 3), the prompt vector generation network mainly consists of a CLIP model text encoder, k+1 MLP layers, and a 3D interpolation layer. t Here it is converted into a three-dimensional text embedded in p e As a backup. p e =MLP k (TextEncoder(p t )。 5. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: In step 4), the volume image encoder uses 3DResUNet, and the global image feature x is obtained e The hint vector is fed into the network to generate the 3D text embedding p e Combined to generate the hint vector p for different tumor regions v , which is used to supervise the model to perform corresponding semantic region segmentation.

6. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: In step 5), the multi-view cross-slice auxiliary module of the 3V3D model is introduced in the first n layers of the decoding process. This module uses the feature changes between different slices as an attention enhancement model to learn the global edge features, hoping to obtain the richest feature map f possible.

7. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: Step 6) Send the feature map f to the PAF module of the corresponding segmentation head and the corresponding prompt vector p v The PAF module imitates the channel attention mechanism to v Transformed into different channel weights w v , represents the importance of the features of different channels of the feature map f for the corresponding tumor area, to supervise the decoding of the corresponding segmentation head. Specifically, is the value of the k-th position of the feature supervision weight of the i-th segmentation target, f i,k is the feature image on the kth channel of the feature image of the i-th segmentation target, and the fusion formula is as follows: Where N is f i The number of channels, Represents the feature image after category supervision, that is, the feature image of the i-th segmentation target. is constrained to the interval 0-1 by the sigmoid function, then when The closer it is to 1, the better the i,k The role of is to motivate, when The closer to 0, the better for f i,k The role is inhibitory.

8. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: Step 7) Finally, each segmentation head still uses 3DResUNet for m decoding to obtain the segmentation area of ​​each tumor area. The formula is as follows: Among them, 3DConv represents the 3D convolution operation, S i is the segmentation prediction of the i-th tumor region.

9. The multi-head semantically supervised brain tumor 3D segmentation method according to claim 1, characterized in that: Step 8) The composite loss function is as follows: Among them, diff loss is a multi-view cross-slice loss. Referring to the 3V3D model, dice loss is used to measure the pixel set similarity between the segmentation label and the segmentation prediction. bce loss is a binary cross entropy loss function used to measure the accuracy of the segmented pixels. At the same time, a weight coefficient λ is added to balance the loss. c represents the number of segmentation heads (the number of tumor regions), and i represents the i-th tumor region.

10. A three-dimensional brain tumor segmentation system, characterized in that: include: Model management module: This module integrates the segmentation method proposed in the present invention and the current mainstream segmentation method, and is used for users to select and build a specific segmentation model. Upload image module: This module is used to upload local MRI images as segmentation original images. Execute segmentation module: This module is used for model inference and performs segmentation process. Result display module: This module is used to display and save segmentation results.