Multi-modal Medical Image Fusion and Segmentation Method Guided by Textual Prompts

By introducing text prompts and CLIP models into the multimodal medical image fusion segmentation method, combined with the cross attention mechanism, the problem that existing methods fail to make full use of text information is solved, and a more accurate and flexible medical image segmentation effect is achieved.

CN118262115BActive Publication Date: 2025-05-27TIANJIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410439601.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-05-27
Estimated Expiration
2044-04-12

AI Technical Summary

Technical Problem

The existing multimodal medical image fusion segmentation methods mainly focus on image feature extraction, and fail to fully utilize text information to guide medical image segmentation, resulting in the improvement of segmentation accuracy and robustness in tasks such as brain tumor segmentation.

Method used

A multimodal medical image fusion segmentation method based on text prompt guidance was designed. By introducing a CLIP model and a cross-attention mechanism, text semantic information is fused with image features to guide the feature extraction and segmentation process.

Benefits of technology

It realizes more flexible and accurate segmentation results in medical image segmentation tasks, improves the performance of multimodal medical image fusion segmentation, especially in brain tumor segmentation tasks, and obtains more accurate and reliable segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118262115B_ABST
    Figure CN118262115B_ABST
Patent Text Reader

Abstract

The present invention proposes a multimodal medical image fusion segmentation method based on text prompt guidance, aiming to solve the challenges in medical image segmentation. This method makes full use of existing medical image datasets, combines key technologies such as image feature extraction, text prompt fusion and attention mechanism, and realizes accurate segmentation of multimodal medical images. By combining text prompts with multimodal image features, the accuracy and robustness of the segmentation results are improved, providing strong support for medical research and clinical practice. This method is not only applicable to medical image segmentation tasks such as brain tumors, but also can provide new ideas and possibilities for natural language processing and image processing in the field of medical images. The innovation of the present invention is to introduce text prompts into multimodal image segmentation, give full play to the guiding role of text information, and provide a new perspective and solution for medical image analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-modal medical image fusion and segmentation method guided by text prompts, belonging to the fields of computer vision and medical image processing. Background Art

[0002] Multi-modal fusion segmentation refers to a method of combining information from different data sources or different modalities to solve the problem of image segmentation. In the fields of computer vision and medical image processing, multi-modal fusion segmentation has been widely applied. Its main purpose is to utilize the complementary information between different modalities to improve the accuracy and robustness of image segmentation. Driven by deep learning technology, multi-modal fusion segmentation methods have made remarkable progress. Compared with traditional methods, multi-modal fusion segmentation methods based on deep learning can usually better capture the complex relationships between data and achieve end-to-end training and optimization. These methods usually consist of deep convolutional neural networks (CNNs) or other deep learning models and can process input data from multiple modalities simultaneously. Typical multi-modal fusion segmentation methods mainly include multi-scale feature fusion, multi-modal feature fusion, spatial feature fusion, and attention mechanism fusion.

[0003] At the current stage, multi-modal fusion networks mainly focus on extracting features of the images themselves and do not fully consider the idea of using text information to guide multi-modal image segmentation. CLIP (Contrastive Language-Image Pre-training) is a model based on contrastive learning that can understand images and texts simultaneously and map them into a common semantic space. However, although CLIP has achieved remarkable results in natural image processing tasks, current multi-modal fusion networks still mainly focus on extracting image features and do not fully utilize text information to guide medical image segmentation. First, CLIP is pre-trained on a large-scale dataset containing millions of image-text pairs. CLIP uses more than 4 million images and approximately 4 billion text segments for pre-training. This huge dataset enables CLIP to learn a wide range of visual and semantic concepts, thus demonstrating strong generalization ability in various tasks. In multi-modal image segmentation tasks, text information plays an important role. It can provide valuable context and semantic guidance for the model, help better understand the similarities and differences between modalities, and promote better fusion between modalities. Second, CLIP can describe task requirements more accurately. Compared with traditional deep learning tasks, CLIP can better understand the content of images from a semantic perspective and make more accurate segmentation judgments on repetitive prediction regions. This ability helps avoid the serious impact of small errors on the overall segmentation result.

[0004] For the brain tumor segmentation task, each patient's medical image consists of four modalities: T1, T1CE, T2, and FLAIR. These different modalities reflect different characteristics of the internal brain tissues. For example, T1CE better reveals the active area of ​​the tumor core, while FLAIR can more accurately show the edema area of ​​the brain tumor. Due to the limitations of a single modality, it is impossible to accurately segment multiple tissue types inside the tumor. By fusing multimodal information, we can enhance the overall segmentation of the tumor, thereby obtaining more accurate and reliable results. The current fusion network mainly focuses on the data feature extraction of the image itself, but due to the limitation of the size of the medical image dataset, this method may not be able to fully exert its segmentation performance. To overcome this limitation, we can use the powerful text capabilities of CLIP. In the modality feature extraction stage, text semantic information is introduced to guide the feature extraction process; at the same time, in the output stage, text prompts are used to describe the task requirements, making the segmentation results more flexible. Through the language description of each modality and category, we can understand the medical image more accurately and obtain more instructive segmentation results.

[0005] This method realizes the possibility of combining natural language with medical image segmentation in the field of medical images, and provides new ideas for medical image segmentation tasks such as brain tumor segmentation. Summary of the invention

[0006] The present invention aims to provide a multimodal medical image fusion segmentation method based on text prompt guidance, which segments various tumor categories by fusing multimodal medical images and text prompts.

[0007] In order to achieve the above object, the scheme of the present invention is:

[0008] Using the general image segmentation network as the basic framework, a multimodal medical image and text prompt fusion segmentation framework is designed, which can complete the text-guided multimodal feature fusion and specific segmentation requirements, and introduce the cross-attention mechanism for image-text fusion. The specific steps are as follows:

[0009] (1) Obtain the Brain Tumor Segmentation (BraTS) dataset, which is commonly used in multimodal medical images, with a total of 369 sets of 3D images, perform data preprocessing, and save them as numpy format images after normalization;

[0010] (2) Design an image-text fusion segmentation network. The image part includes multiple modality feature extraction encoders and a shared feature decoder, and the text part includes a pre-trained modality feature extractor and a category feature extractor.

[0011] (3) The modal feature extraction encoder extracts multi-level semantic information from the image by using the U-Net encoder;

[0012] (4) Use the CLIP model to extract the semantic information of the modal text, and linearly align the extracted semantic information of the modal text with the different-level image semantic information extracted by the U-Net encoder;

[0013] (5) Use the attention mechanism to perform cross-attention fusion on the image semantic information and text semantic information in each modality, and perform Concatenate data splicing on the fusion results of multiple modalities; the features after image-text fusion are obtained through a shared decoder to get a preliminary segmentation result;

[0014] (6) Use the CLIP model to extract the semantic information of the category text, align and fuse the deepest-level image semantic information extracted by the U-Net with the category text semantic information, transform it into model parameters, and guide the preliminary segmentation result to complete the specified segmentation requirements to obtain the final segmentation result.

[0015] (7) Conduct multiple experiments, and record the Dice coefficients and 95% Hausdorff distances of each category in the results of each experiment;

[0016] (8) Design ablation experiments to explore the contribution degree of each image module and text prompt module to the segmentation performance.

[0017] The beneficial effects of the present invention are as follows: The present invention proposes a multi-modal medical image fusion and segmentation method guided by text prompts. First, we obtain common multi-modal segmentation data sets and perform necessary data preprocessing on them. Then, combined with text prompts, we conduct segmentation guidance on multi-modal medical data, introduce the attention mechanism in Transformer, and design and complete a multi-modal fusion and segmentation network. The text prompts include modal prompts and category prompts, which are used to promote fusion and guide the segmentation process respectively. This design not only helps to extract deep features of each modality, but also can achieve more flexible and accurate segmentation results. Description of the Drawings

[0018] Figure 1 : Flow chart of text-prompted multi-modal fusion and segmentation.

[0019] Figure 2 : Flow chart of image-text fusion in each stage. Detailed Embodiments

[0020] Obtain the BraTS dataset in the.nii.gz format, with a total of 369 groups. Each group contains four modalities (T1, T1CE, T2, and FLAIR); it is divided into 3 categories, namely the edema region (ED), the enhancing tumor region (ET), and the necrotic region (NCR / NET). In terms of preprocessing, it is necessary to crop the original image data, remove the background area without brain tissue, and normalize the images. Save the processed data as numpy-format images for subsequent data reading and training. In terms of dataset division, divide the entire dataset into a training set (70%), a validation set (10%), and a test set (20%) to ensure the effectiveness and generalization ability of the trained model.

[0021] The text-prompt-guided multi-modal medical image fusion segmentation network first extracts features from multi-modal images. CLIP converts the text prompt into tensors, and then aligns these tensors semantically with the image features. In this process, key steps such as multi-modal fusion algorithms and loss functions are designed to ensure the effective combination of text prompts and image information and output accurate segmentation results. The model structure is as Figure 1 shown.

[0022] This segmentation network takes UNet as a reference and mainly consists of the following parts: image feature extraction, text feature extraction, modality fusion, and feature reconstruction. First, through multiple independent feature extraction networks consistent with the number of modalities, multi-scale image features of multiple modalities are extracted. Then, the modality text prompts corresponding to each modality are fused with the image features at different scales to achieve more comprehensive information acquisition and utilization. Next, after multi-scale feature reconstruction, the text and feature fusion generate weight parameters, and these weight parameters will be weighted into the reconstructed features to obtain the final output. Finally, calculate the Dice coefficient and 95% Hausdorff distance based on the classification results to evaluate the performance of the model.

[0023] Inside the segmentation network, use UNet encoder branches with the same number of modalities as the number of modalities to perform multi-level feature extraction on different modalities. In the modality text prompt stage, the image features extracted for each modality are fused with the text information corresponding to that modality through cross-attention. To better improve the fit between image features and text features, in addition to ordinary text prompts, we also introduce learnable modality text. The fusion is completed by leveraging the advantage of the long-range dependency modeling of Transformer, thereby obtaining a more accurate feature representation. The formula is as follows:

[0024]

[0025] Among them, I is the image feature, L is the text prompt feature, m is the specific modality, s is the different stage, and h and l respectively represent the invariant text prompt and the learnable text prompt. After the image-text feature fusion of each stage and each modality is completed, the features of multiple modalities are concatenated together using the Concatenate method. Then, the fusion is completed through a convolution operation to obtain a more comprehensive and rich feature representation. Except for the fusion feature of the deepest layer, the remaining fusion features will be used as skip connections and directly connected to the decoder part to utilize low-level and high-level information in the decoder to obtain a preliminary output result.

[0026] The category text prompt is used to guide the network to complete a specific category segmentation task. The deepest layer feature is aligned with the category text prompt feature through a max pooling operation and fused through a multi-layer perceptron method. The fused feature is transformed into weight parameters to perform weighted learning on the above preliminary result, thereby obtaining the final segmentation result. In the training process of this model, the Dice loss and cross-entropy loss are used as loss functions for the output result and the fusion features of each intermediate stage. The formulas are as follows:

[0027]

[0028]

[0029] Where N represents the number of voxels and C is the number of categories. For whether the i-th voxel point in the image belongs to the j-th category, it is represented by y ij If it belongs to the j-th category, then y ij = 1, otherwise it is y ij = 0, represents the probability that voxel i belongs to the j-th category.

[0030] The multi-modal medical image fusion segmentation method guided by text prompts can be applied to real medical research, realizing the connection between the natural language field and the medical image field, and more flexibly completing segmentation tasks of different categories according to task requirements, contributing a little to medical research.

[0031] It should be noted that the above description is only an embodiment of the present invention, which only explains the present invention and does not limit the scope of the patent of the present invention. Modifications that are merely obvious to those belonging to the technical concept of the present invention are also within the protection scope of the present invention.

Claims

1. A multimodal medical image fusion segmentation algorithm based on text prompt guidance, characterized in that: The steps include: (1) Obtain the Brain Tumor Segmentation (BraTS) dataset, which is commonly used in multimodal medical images, with a total of 369 sets of 3D images, perform data preprocessing, and save them as numpy format images after normalization; (2) Design an image-text fusion segmentation network. The image part includes multiple modality feature extraction encoders and a shared feature decoder, and the text part includes two pre-trained modality text feature extractors and a category text feature extractor. (3) The modal feature extraction encoder extracts multi-level semantic information from the image by using the U-Net encoder; (4) Use the CLIP model to extract modal text semantic information and linearly align the extracted modal text semantic information with the different levels of image semantic information extracted by the U-Net encoder; (5) The attention mechanism is used to cross-attend the image semantic information and text semantic information under each modality, and the fusion results of multiple modalities are concatenated; the features after image and text fusion are obtained through a shared decoder to obtain preliminary segmentation results; (6) Use the CLIP model to extract the semantic information of the category text, align and fuse the deepest semantic information of the image extracted by U-Net with the semantics of the category text, and convert it into model parameters to guide the preliminary segmentation results to complete the specified segmentation requirements and obtain the final segmentation results; (7) The performance of the model was evaluated by calculating the Dice coefficient and 95% Hausdorff distance of each category.

2. The multimodal medical image fusion segmentation algorithm based on text prompt guidance as claimed in claim 1, characterized in that: The image part multiple modality feature extraction encoder in step (2) extracts multi-scale image features of multiple modality images by formulating multiple independent feature extraction networks with the same number of modalities; each modality text feature extractor is pre-trained by CLIP and is used to extract semantic information of text sentences corresponding to different modalities; the category text feature extractor is pre-trained by CLIP and is used to extract semantic information in category text to guide the network to complete specific category segmentation tasks.

3. The multimodal medical image fusion segmentation algorithm based on text prompt guidance as claimed in claim 1, characterized in that: The two modal text feature extractors described in step (2) are used to extract semantic information and trainable text semantic encoding in the original modal text respectively. The original modal text guides the image feature extraction encoder to more accurately extract image semantic information, and the trainable text encoding is used to improve the fit between image and text features.

4. The multimodal medical image fusion segmentation algorithm based on text prompt guidance as claimed in claim 1, characterized in that: The multi-level image feature dimensions extracted by U-Net in step (4) are all different. Both prompts of the modal text need to be aligned with different modal feature dimensions through a linear layer for subsequent cross-attention fusion.

5. The multimodal medical image fusion segmentation algorithm based on text prompt guidance as claimed in claim 1, characterized in that: The cross-attention fusion described in step (5) corresponds to different stages of U-Net. In each stage, the image features of different modalities, the original modality text and the learned text encoding are set as a group, the image features in the family are set as query, the original modality text features are set as key, and the learnable text encoding is set as value to achieve cross-attention fusion.

6. The multimodal medical image fusion segmentation algorithm based on text prompt guidance as claimed in claim 1, characterized in that: In step (6), the deepest layer features are used to fuse category text hints and upsampled feature reconstruction, and the remaining feature layers are used to form skip connections. The preliminary results generated by the shared decoder are weighted by the image category text feature weights to obtain the final segmentation prediction.