Text prompt-based bidirectional progressive fusion medical image segmentation method and system
Patent Information
- Application Number
- CN202511043565.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-07-28
AI Technical Summary
[0004]本发明的目的在于解决现有基于文本提示的医学影像分割方法中存在的视觉编码器与文本编码器跨模态交互不足、解码器对精细结构识别能力有限等问题,并提供一种基于文本提示的双向渐进融合医学影像分割方法及系统
[0035] This invention introduces a bidirectional progressive fusion mechanism, constructing a hierarchical information interaction channel between the visual encoder and the text encoder. This enables bidirectional optimized transfer of anatomical semantic features and visual features, significantly improving cross-modal feature alignment accuracy and addressing the problem of insufficient information transfer in traditional unidirectional fusion architectures. It also enhances the model's ability to jointly understand medical images and text prompts. Furthermore, by incorporating a mask token-based dynamic convolutional kernel generation process and a semantic similarity-guided mechanism into the query decoder, it improves boundary and structure recognition capabilities, achieving more flexible region prediction and precise localization of visual regions by text prompts. This enhances the model's accuracy in segmenting complex anatomical structures. This invention improves the performance and generalization ability of text-prompt-based medical image segmentation models, providing more efficient and accurate automated segmentation solutions for clinical practice and promoting the practical application of medical image analysis technology in disease diagnosis and treatment.
Smart Images

Figure CN121169957B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing, and particularly relates to a bidirectional progressive fusion medical image segmentation method and system based on text prompts. Background Technology
[0002] In the field of medical information technology, medical image segmentation is a core technology for disease diagnosis, treatment planning, and disease progression monitoring. Its automation and accuracy directly impact the efficiency of clinical decision-making. Current medical image segmentation technology has evolved from manual slice-by-slice annotation to automated processing by deep learning models. However, significant technical bottlenecks remain: On the one hand, while traditional specialized segmentation models (such as nnU-Net and U-Mamba) demonstrate advantages in single anatomical structures or specific modalities, their "task-customized" design leads to a high degree of binding between their feature extraction networks and decision logic and specific datasets. This results in significant generalization limitations when facing cross-modal images (such as mixed CT / MRI scenes) and multi-anatomical region joint segmentation tasks. On the other hand, when the large-scale visual model SAM, pre-trained on natural images, is transferred to medical scenarios, it often requires points or boxes as input cues, making it an interactive segmentation model. In practical applications, it still requires extensive manual manipulation of annotation points or boxes to prompt the model, failing to meet the demands of high efficiency and accuracy in clinical practice. Furthermore, the domain differences between natural images and medical images in pixel statistical distribution and anatomical semantic systems lead to semantic understanding biases in the models.
[0003] In recent years, text-driven medical image segmentation technology has developed. However, existing models suffer from low fusion efficiency in the design of visual and text encoders, leading to insufficient information transmission and an inability to fully exploit the potential correlation between visual and textual information. In the query decoder, existing models underutilize multi-scale high-resolution features, limiting their boundary and structure recognition capabilities and making it difficult to accurately capture detailed information in complex scenes. Furthermore, their region prediction methods lack flexibility and cannot be dynamically adjusted according to different task requirements. In addition, text prompts lack an effective semantic similarity guidance mechanism when guiding visual region localization, making it difficult to accurately locate specific visual regions, thus affecting the overall performance of the model in segmentation tasks. Summary of the Invention
[0004] The purpose of this invention is to solve the problems of insufficient cross-modal interaction between the visual encoder and the text encoder and the limited ability of the decoder to recognize fine structures in existing text-based medical image segmentation methods, and to provide a bidirectional progressive fusion medical image segmentation method and system based on text prompts.
[0005] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:
[0006] In a first aspect, the present invention provides a bidirectional progressive fusion medical image segmentation method based on text prompts, which includes the following steps:
[0007] S1: Acquire medical images containing organs or lesions, and preprocess the medical images to obtain preprocessed medical images;
[0008] S2: The preprocessed medical images and medical text prompts are input into a pre-trained segmentation network based on the SAT model. The segmentation network uses 3D U-Net as the backbone network for visual feature extraction and a staged improved BERT model as the text encoder. Simultaneously, a query decoder is constructed within the segmentation network based on the Transformer decoder. First, the text encoder processes the medical text prompts, and then the visual encoder in 3D U-Net processes the preprocessed medical images. A bidirectional progressive fusion module aligns and fuses the text features output from each stage of the text encoder with the visual features output from each stage of the visual encoder, thereby achieving hierarchical interaction. The U-Net visual decoder progressively upsamples and fuses the multi-scale visual features output by the visual encoder to generate pixel-wise dense feature maps for segmentation mask prediction. The text features output by the final stage of the text encoder are used as text semantic embeddings, and the visual features output by the final stage of the visual encoder are used as visual semantic embeddings. The text semantic embeddings, visual semantic embeddings, and pixel-wise dense feature maps are input to the query decoder, which outputs multiple mask probability maps with the same number of target categories. An adaptive threshold segmentation algorithm is used to process the mask probability maps output by the query decoder, and finally, a 3D segmentation mask corresponding to the medical text prompt is generated.
[0009] Based on the above scheme, each step can be implemented in the following preferred manner.
[0010] As a preferred embodiment of the first aspect mentioned above, in step S2, the text encoder is based on the BERT model consisting of 12 identical Transformer layers, and the Transformer layers of the text encoder are divided according to the number of layers so that they correspond to the four stages of the visual encoder; wherein, the first stage includes Transformer layers 1-6, the second stage includes Transformer layers 7-8, the third stage includes Transformer layers 9-10, and the fourth stage includes Transformer layers 11-12.
[0011] As a preferred embodiment of the first aspect, in step S2, the visual encoder comprises four stages. The first three stages are each composed of two 3×3×3 convolutional layers with BatchNorm and ReLU activation functions, and a 2×2×2 max pooling layer cascaded in sequence. The last stage is composed of only two 3×3×3 convolutional layers with BatchNorm and ReLU activation functions cascaded in sequence.
[0012] As a preferred embodiment of the first aspect mentioned above, the specific process of realizing the interaction between the text encoder and the visual encoder through the bidirectional progressive fusion module in step S2 is as follows:
[0013] AS1: After the medical text prompts are input into the text encoder, they first pass through the Embedding module to generate the initial input embedding;
[0014] AS2: The initial input embedding first passes through the Transformer layer of the first stage of the text encoder, and the hidden state output by the first stage is used as the first text feature; the preprocessed medical image first passes through the first stage of the visual encoder path to obtain the first visual feature.
[0015] AS3: The first text feature and the first visual feature are processed through the first bidirectional progressive fusion module, and the updated first text feature and the updated first visual feature are output.
[0016] AS4: The updated first text feature passes through the Transformer layer of the second stage of the text encoder, and the hidden state output by the second stage is used as the second text feature; the updated first visual feature then passes through the second stage of the visual encoder path to obtain the second visual feature; the second text feature and the second visual feature are then passed through the second bidirectional progressive fusion module to output the updated second text feature and the updated second visual feature.
[0017] AS5: The updated second text feature passes through the Transformer layer of the third stage of the text encoder, and the hidden state output by the third stage is used as the third text feature; the updated second visual feature then passes through the third stage of the visual encoder path to obtain the third visual feature; the third text feature and the third visual feature are then passed through the third bidirectional progressive fusion module to output the updated third text feature and the updated third visual feature.
[0018] AS6: The updated third text feature passes through the Transformer layer of the fourth stage of the text encoder, and the hidden state output by the fourth stage is used as the fourth text feature; the updated third visual feature then passes through the fourth stage of the visual encoder path to obtain the fourth visual feature; the fourth text feature and the fourth visual feature are then passed through the fourth bidirectional progressive fusion module to output the updated fourth text feature and the updated fourth visual feature.
[0019] As a preferred embodiment of the first aspect mentioned above, the specific processing procedure in each bidirectional progressive fusion module in step S2 is as follows:
[0020] BS1: First, the visual features are dimensionally adjusted by flattening them into vector form, aligning them with the dimensions of the text features; the text features are then mapped to the same channel dimension as the visual features through a learnable linear layer.
[0021] BS2: The processed visual features and processed text features are each subjected to layer normalization and the contextual relationships within their respective modalities are captured through a self-attention mechanism to generate intermediate visual features and intermediate text features.
[0022] BS3: Subsequently, cross-attention fusion is performed. For text-guided visual fusion, the intermediate text features are used as queries, and the intermediate visual features are used as keys and values. Cross-attention is used to generate visual features that fuse text semantics. The visual features that fuse text semantics are then processed by a multilayer perceptron and residually connected with the intermediate text features to obtain the updated text features. For visual-guided text fusion, the intermediate visual features are used as queries, and the updated text features are used as keys and values. Cross-attention is used to generate text semantic representations that fuse visual features. The text semantic representations that fuse visual features are then processed by a multilayer perceptron and residually connected with the intermediate visual features to obtain the updated visual features.
[0023] As a preferred embodiment of the first aspect mentioned above, the specific processing flow of the query decoder in step S2 is as follows:
[0024] CS1: First, the text semantic embedding is concatenated with the learnable mask token to form a query vector sequence, which is then input into a series of 6 cascaded Transformer decoding layers. Through iterative updates of the 6 Transformer decoding layers, the text query is adapted to the visual context of the current image, generating an updated query vector.
[0025] CS2: Then, the dynamic convolution kernel generation stage begins, extracting kernels from the updated query vector. Each of the given mask tokens is mapped to a 1×1×1 dynamic convolutional kernel using a lightweight MLP, with the kernel parameters dynamically adjusted based on the input features. Indicates the number of target categories;
[0026] CS3: The text semantic embedding is mapped to the projected text semantic embedding through a linear projection layer. The cosine similarity between the projected text semantic embedding and each spatial location feature of the pixel-by-pixel dense feature map output by the visual decoder is calculated to form a text-visual semantic similarity map.
[0027] CS4: Dynamic convolutional kernels are applied to the pixel-wise dense feature map output by the visual decoder. Each dynamic convolutional kernel performs a spatially position-wise convolution operation on the pixel-wise dense feature map. The feature map obtained from the spatially position-wise convolution operation is activated by Sigmoid and then weighted and fused with the text-visual semantic similarity map. This is combined with semantic similarity to guide the generation of... A mask probability graph.
[0028] As a preferred embodiment of the first aspect mentioned above, in the query decoder, each Transformer decoding layer includes a multi-head self-attention mechanism and a cross-modal attention mechanism. The multi-head self-attention mechanism interacts with the query vector sequence; the cross-modal attention mechanism uses the query vector sequence as the query and the flattened visual semantic embedding as the key and value to perform attention calculation.
[0029] As a preferred embodiment of the first aspect mentioned above, the specific process for generating the 3D segmentation mask is as follows: The optimal segmentation threshold is estimated using the Otsu algorithm and used as the global threshold; a binary mask is then generated based on the global threshold through binarization; the value at each spatial location of the mask probability map is compared with the global threshold; if the value is greater than or equal to the global threshold, the binary mask at that spatial location is set to 1, otherwise it is set to 0; after obtaining the binary mask, it is post-processed; when the input medical text prompt is a single text prompt, a single-channel 3D segmentation mask is output; when the input medical text prompt contains… When there are multiple text prompts for multiple target categories, a multi-channel 3D segmentation mask is output.
[0030] As a preferred embodiment of the first aspect, during the training process, the segmentation network calculates the cross-entropy loss and dice loss between the predicted labels and the real labels for training data with real labels. Then, the cross-entropy loss and dice loss are added together to obtain the total loss of the segmentation network. With minimizing the total loss as the optimization objective, the segmentation network is continuously optimized using a stochastic gradient descent optimizer.
[0031] Secondly, the present invention provides a text-guided bidirectional progressive fusion medical image segmentation system, comprising:
[0032] The data acquisition module is used to acquire medical images containing organs or lesions, and to preprocess the medical images to obtain preprocessed medical images.
[0033] The results acquisition module inputs preprocessed medical images and medical text prompts into a pre-trained segmentation network based on the SAT model. The segmentation network uses 3D U-Net as the backbone for visual feature extraction and a staged improved BERT model as the text encoder. Simultaneously, a query decoder is constructed within the segmentation network based on the Transformer decoder. First, the text encoder processes the medical text prompts, and then the visual encoder in 3D U-Net processes the preprocessed medical images. A bidirectional progressive fusion module aligns and fuses the text features output from each stage of the text encoder with the visual features output from each stage of the visual encoder, thereby achieving hierarchical interaction. The U-Net visual decoder progressively upsamples and fuses the multi-scale visual features output by the visual encoder to generate pixel-wise dense feature maps for segmentation mask prediction. The text features output by the final stage of the text encoder are used as text semantic embeddings, and the visual features output by the final stage of the visual encoder are used as visual semantic embeddings. The text semantic embeddings, visual semantic embeddings, and pixel-wise dense feature maps are input to the query decoder, which outputs multiple mask probability maps with the same number of target categories. An adaptive threshold segmentation algorithm is used to process the mask probability maps output by the query decoder, and finally, a 3D segmentation mask corresponding to the medical text prompt is generated.
[0034] Compared with the prior art, the present invention has the following advantages:
[0035] This invention introduces a bidirectional progressive fusion mechanism, constructing a hierarchical information interaction channel between the visual encoder and the text encoder. This enables bidirectional optimized transfer of anatomical semantic features and visual features, significantly improving cross-modal feature alignment accuracy and addressing the problem of insufficient information transfer in traditional unidirectional fusion architectures. It also enhances the model's ability to jointly understand medical images and text prompts. Furthermore, by incorporating a mask token-based dynamic convolutional kernel generation process and a semantic similarity-guided mechanism into the query decoder, it improves boundary and structure recognition capabilities, achieving more flexible region prediction and precise localization of visual regions by text prompts. This enhances the model's accuracy in segmenting complex anatomical structures. This invention improves the performance and generalization ability of text-prompt-based medical image segmentation models, providing more efficient and accurate automated segmentation solutions for clinical practice and promoting the practical application of medical image analysis technology in disease diagnosis and treatment. Attached Figure Description
[0036] Figure 1 This is a flowchart of the steps of the method of the present invention;
[0037] Figure 2 This is a diagram illustrating the overall architecture of the method of the present invention;
[0038] Figure 3 This is a structural diagram of the bidirectional progressive fusion module of the method of the present invention;
[0039] Figure 4 This is a system block diagram of the present invention;
[0040] Figure 5 This is a schematic diagram of a computer electronic device according to the present invention. Detailed Implementation
[0041] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in various embodiments of the present invention can be combined accordingly without mutual conflict.
[0042] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include at least one of those features.
[0043] like Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned text-based bidirectional progressive fusion medical image segmentation method includes the following S1-S2 steps. The specific implementation process of each step will be described in detail below.
[0044] S1: Acquire medical images containing organs or lesions, and preprocess the medical images to obtain preprocessed medical images.
[0045] It should be noted that in this embodiment, 3D medical images (covering modalities such as CT, MRI, and PET) are standardized by resampling the original images to a uniform voxel spacing of 1×1×3mm. 3To eliminate spatial resolution differences between different devices, invalid backgrounds are removed by non-zero region cropping, retaining effective regions containing anatomical structures. For CT images, HU values are truncated to the range of [-500, 1000] to focus on key structures such as soft tissue and bone, followed by Z-score normalization to ensure pixel intensity follows a normal distribution with a mean of 0 and a standard deviation of 1. For MRI and PET images, 0.5% and 99.5% quantile cropping is used to filter outliers, followed by Z-score normalization to ensure consistent intensity distribution across different modalities. Furthermore, for different anatomical regions of CT images (such as soft tissue, lung, brain, and bone), corresponding window width and level settings are used (e.g., soft tissue window width 400, window level 40) to enhance the contrast of target structures, followed by uniform linear rescaling to the integer range of [0, 255] for easy model input.
[0046] S2: The preprocessed medical images and medical text prompts are input into a pre-trained segmentation network based on the SAT model. The segmentation network uses 3D U-Net as the backbone network for visual feature extraction and a staged improved BERT model as the text encoder. Simultaneously, a query decoder is constructed within the segmentation network based on the Transformer decoder. First, the text encoder processes the medical text prompts, and then the visual encoder in 3D U-Net processes the preprocessed medical images. A bidirectional progressive fusion module aligns and fuses the text features output from each stage of the text encoder with the visual features output from each stage of the visual encoder, thereby achieving hierarchical interaction. The U-Net visual decoder progressively upsamples and fuses the multi-scale visual features output by the visual encoder to generate pixel-wise dense feature maps for segmentation mask prediction. The text features output by the final stage of the text encoder are used as text semantic embeddings, and the visual features output by the final stage of the visual encoder are used as visual semantic embeddings. The text semantic embeddings, visual semantic embeddings, and pixel-wise dense feature maps are input to the query decoder, which outputs multiple mask probability maps with the same number of target categories. An adaptive threshold segmentation algorithm is used to process the mask probability maps output by the query decoder, and finally, a 3D segmentation mask corresponding to the medical text prompt is generated.
[0047] It should be noted that in this invention, the text encoder is based on the BERT model consisting of 12 identical Transformer layers, and the Transformer layers of the text encoder are divided according to the number of layers to correspond to the four stages of the visual encoder. Specifically, the first stage contains Transformer layers 1-6, the second stage contains Transformer layers 7-8, the third stage contains Transformer layers 9-10, and the fourth stage contains Transformer layers 11-12.
[0048] In this embodiment, the text encoder is used to generate textual semantic embeddings for medical terms. Its internal 12 Transformer layers employ the same structure, with each layer containing a multi-head self-attention module and a feedforward neural network. Specifically, the multi-head self-attention module divides the input into 12 attention heads, each independently calculating the query Q, key K, and value matrix V.
[0049]
[0050] In the formula: Attention represents attention calculation; softmax represents the softmax function; T represents transpose; d k As the key dimension, in this embodiment, d k =64.
[0051] Output after multi-head attention fusion The input is then normalized and connected to the residuals in a feedforward neural network. The feedforward neural network contains two linear transformation layers with GELU activation, maintaining an output dimension of 768. The final hidden state is output by the 12th Transformer layer. In the [CLS] position, the embedding As a semantic representation of the text as a whole.
[0052] It should be noted that in this invention, the visual encoder comprises four stages. The first three stages each consist of two 3×3×3 convolutional layers with BatchNorm and ReLU activation functions, followed by a 2×2×2 max-pooling layer cascaded sequentially. The last stage consists only of two 3×3×3 convolutional layers with BatchNorm and ReLU activation functions cascaded sequentially. The max-pooling layer here is used for downsampling, halving the spatial resolution and doubling the number of channels.
[0053] It should be noted that, in this invention, the specific process of realizing the interaction between the text encoder and the visual encoder through the bidirectional progressive fusion module is as follows:
[0054] AS1: After the medical text prompts are input into the text encoder, they first pass through the Embedding module to generate the initial input embedding.
[0055] In this embodiment, the Embedding module is an existing part of the BERT model and has not been modified. Specifically, the medical text prompt first uses the WordPiece word segmenter to decompose the text into sub-word units and adds special markers [CLS] and [SEP] to identify the beginning and end of sentences. Then, positional encoding is generated, and sequence order information is injected through sine and cosine functions, which are then superimposed with the word embeddings to form the initial input embedding. (L is the sequence length).
[0056] AS2: The initial input embedding first passes through the first stage Transformer layer of the text encoder, and the hidden state output from the first stage is used as the first text feature T1; the preprocessed medical image First, the visual encoder path passes through the first stage, resulting in the first visual feature V1. Here, H, W, and D represent the spatial dimensions of the medical image, respectively, and C represents the number of channels.
[0057] AS3: The first text feature and the first visual feature are processed through the first bidirectional progressive fusion module, and the updated first text feature is output. and updated first visual features
[0058] AS4: The updated first text feature passes through the Transformer layer of the second stage of the text encoder, and the hidden state output by the second stage is used as the second text feature T2; the updated first visual feature then passes through the second stage of the visual encoder path to obtain the second visual feature V2; the second text feature and the second visual feature are then passed through the second bidirectional progressive fusion module to output the updated second text feature. and updated second visual features
[0059] AS5: The updated second text feature passes through the Transformer layer of the third stage of the text encoder, and the hidden state output by the third stage is used as the third text feature T3; the updated second visual feature then passes through the third stage of the visual encoder path to obtain the third visual feature V3; the third text feature and the third visual feature are then passed through the third bidirectional progressive fusion module to output the updated third text feature. and updated third-vision features
[0060] AS6: The updated third text feature passes through the Transformer layer in the fourth stage of the text encoder, and the hidden state output from the fourth stage is used as the fourth text feature T4; the updated third visual feature then passes through the fourth stage of the visual encoder path to obtain the fourth visual feature V4; the fourth text feature and the fourth visual feature are then passed through the fourth bidirectional progressive fusion module to output the updated fourth text feature. and the updated fourth visual features
[0061] In this embodiment, without processing by the bidirectional progressive fusion module, the input image in the first stage of the visual encoder undergoes two 3×3×3 convolution processes, increasing the number of channels from C to 64, and outputting a feature map. The second stage involves increasing the number of channels to 128 using the same operation, and then outputting the feature map. The number of channels in the third stage is increased to 256, and the output feature map is improved. The fourth stage is the deepest layer of the visual encoder, with 512 channels, and the output... The multi-scale visual features {V1, V2, V3, V4} output from each of the above stages correspond to visual semantic representations at different resolutions. V1 retains more spatial details, while V4 extracts high-level semantic features for subsequent interaction with text features and feature recovery by the decoder. The first stage output of the text encoder represents shallow semantic features T1 (e.g., lexical structure), and the fourth stage outputs deep semantic features T4 (e.g., global concept representation). After processing by the bidirectional progressive fusion module, the updated visual features... and text features These are used as inputs to the visual encoder and text encoder in the next stage, respectively, forming a progressive transfer of cross-modal information. For example, the information after the first stage of fusion... and Each input is given to the encoder corresponding to the second stage, and so on, until the fourth stage outputs visual features containing complete cross-modal semantics. and text features
[0062] Through the above improvements, the output features of each stage of the text encoder are gradually integrated with the output features of each stage of the visual encoder. For example, the text features (such as word-level representations) of the first stage are integrated with shallow visual features (edges, textures) to enhance the description of anatomical structure shapes (such as "wedge"); the text features (global concepts) of the fourth stage are integrated with deep visual features (semantic regions) to enhance the understanding of anatomical relationships (such as "the adjacent relationship between the liver and gallbladder"). Finally, the text features output by the fourth stage of the text encoder... As a text semantic embedding, this embedding simultaneously contains text semantics, medical knowledge, and cross-modal visual spatial location information. For example, when the text prompt "myocardium" is entered, it not only encodes the semantics of "myocardium is the muscle tissue of the heart," but also, through fusion with visual features, implicitly contains its visual attributes such as position and shape in CT images, providing accurate cross-modal guidance for subsequent query decoding.
[0063] It should be noted that, in order to achieve deep interaction between visual and text features in this invention, a bidirectional progressive fusion mechanism of BiPVL-Seg is introduced in the visual encoder and text encoder stages. It realizes cross-modal information transmission through BiFusion blocks. As a result, four bidirectional progressive fusion modules are designed. Each bidirectional progressive fusion module is used to align and fuse the text features {T1,T2,T3,T4} of each stage of the text encoder with the visual features {V1,V2,V3,V4} of the same stage of the visual encoder stage by stage.
[0064] In this embodiment, as Figure 3 As shown, the specific processing steps in each bidirectional progressive fusion module are as follows:
[0065] BS1: First, visual feature V n Dimensional adjustments are made, and the vector form is transformed through a flattening operation, which is then compared with the text feature T. n Dimensional alignment; text features are mapped to the same channel dimension as visual features via a learnable linear layer, ensuring compatibility for cross-modal interactions.
[0066] In this embodiment, for visual features V of the visual encoder n (n = 1, 2, 3, 4) alignment requires text feature T n Convert to a representation compatible with visual features. For visual features First, flatten its spatial dimensions as (N n =H n ×W n ×D n This forms the processed visual features. Simultaneously, text features are also incorporated. (n = 1, 2, 3, 4) are sequentially processed through global pooling and linear projection layers to map to the processed text features, which have the same channel dimension as the processed visual features. Wherein, H... n W n D n C represents the spatial dimensions of visual features. n N represents the number of channels for visual features. n The length of the processed visual feature.
[0067] BS2: The processed visual features and processed text features are each subjected to layer normalization and self-attention mechanism to capture the contextual relationships within their respective modalities, generating intermediate visual features V′. n and intermediate text features T′ n .
[0068] In this embodiment, the self-attention operation is implemented through the calculation of the query, key, and value matrix, as shown in the formula:
[0069]
[0070] Where X represents the input feature, specifically the processed visual feature or the processed text feature.
[0071] BS3: Subsequently, cross-attention fusion is performed. For text-guided visual fusion, the intermediate text features are used as queries, and the intermediate visual features are used as keys and values. Cross-attention is used to generate visual features that fuse text semantics. These visual features are then processed by a multilayer perceptron (MLP) and residually connected with the intermediate text features to obtain the updated text features. For visual-guided text fusion, intermediate visual features are used as queries, and updated text features are used as keys and values. Cross-attention is used to generate a text semantic representation of the fused visual features. This text semantic representation of the fused visual features is then processed by a multilayer perceptron and combined with the intermediate visual features V′. n Perform residual connections to obtain updated visual features.
[0072]
[0073] Among them, MLP T and MLP V These represent the multilayer perceptron used in text-guided visual fusion and the multilayer perceptron used in visual-guided text fusion, respectively.
[0074] Therefore, the bidirectional progressive fusion module of this invention introduces cross-modal interaction at each stage of the visual encoder and text encoder, enabling visual features to gradually integrate with textual semantic information, while textual features simultaneously absorb visual spatial features. This enhances the model's understanding of the association between medical terms and their corresponding anatomical structures, particularly improving its ability to represent long-tailed categories and complex anatomical relationships. For example, when processing structures such as "myocardium," the semantic meaning of "located in the left ventricular wall" in the textual features can guide the visual features to focus on the heart region through the fusion mechanism, improving segmentation accuracy in cases of blurred boundaries.
[0075] It should be noted that the core structure of 3D U-Net consists of an encoder path (downsampling) and a decoder path (upsampling), achieving multi-scale feature fusion through skip connections. The visual decoder of this invention adopts the 3D U-Net decoder architecture to progressively upsample and fuse the multi-scale features output by the visual encoder, generating a pixel-by-pixel dense feature map for segmentation mask prediction. The input to the visual decoder is the four-level multi-scale visual features {V1, V2, V3, V4} output by the visual encoder. The deepest visual feature (corresponding to the fourth stage output) contains high-level semantic information; the rest are intermediate to shallow visual features, preserving more spatial details. The visual decoder uses the fourth visual feature V4 as the initial input, gradually restoring the spatial resolution through four decoding stages, while simultaneously fusing the skip connection features from the corresponding encoder stage. The specific process is as follows:
[0076] The first stage of the decoder uses fourth-vision features as input. It first upsamples these features using a 3D deconvolution layer with a kernel size of 2×2×2 and a stride of 2, doubling the spatial size. The number of channels is halved to 256, resulting in a feature map. Subsequently, feature map U4 is concatenated with the third visual feature output from the third stage of the visual encoder along the channel dimension to obtain the feature map. Next, the feature map C4 is processed by two 3×3×3 convolutional blocks. Each convolutional block contains batch normalization (BatchNorm3d), ReLU activation function, and convolutional layer to output the feature map. Complete the first stage of decoding.
[0077] The second stage of the decoder takes feature map D4 as input and upsamples it through 3D deconvolution (2×2×2, stride 2) to... The number of channels is halved to 128, resulting in a feature map. After concatenating with the second visual features output from the second stage of the visual encoder, a feature map is obtained. After processing with two 3×3×3 convolutional blocks (including BatchNorm3d and ReLU), the output feature map is obtained.
[0078] The third stage input feature map of the decoder Upsampled to 3D deconvolution (2×2×2, stride 2) The number of channels is halved to 64, resulting in a feature map. The feature map is obtained by concatenating the first visual feature output from the first stage of the visual encoder. After processing through two more 3×3×3 convolutional blocks, the feature map is output.
[0079] In the fourth stage of the decoder, the input feature map D2 is upsampled to the original input size H×W×D through 3D deconvolution (2×2×2, stride 2), and the number of channels is halved to 32, resulting in the feature map. Since the features from the zeroth stage of the visual encoder do not participate in skip connections, a 1×1×1 convolutional layer is used to directly reduce the channel dimensionality of U1, generating the final pixel-wise dense feature map. Where d′ = 64 is the feature dimension. This feature map contains a visual representation that integrates multi-scale semantics and spatial details, providing input for mask generation in the subsequent query decoder.
[0080] It should be noted that existing query decoders use textual semantic embeddings (such as the textual features of "myocardium") as queries, and the deep visual features output by the visual encoder are flattened and used as keys and values. These are then interacted through six Transformer decoding layers to generate a query vector. Simultaneously, the pixel-wise dense feature map output by the visual decoder is used to perform a pixel-level dot product operation with the query vector. Specifically, the query vector is mapped through a linear projection layer and then dot-producted with the feature vector at each spatial location of the pixel-wise dense feature map to form a pixel-level score map. This score map is then generated using a Sigmoid activation function to produce a mask probability map, which is then thresholded (e.g., 0.5) to obtain a binary segmentation mask. This process achieves semantic localization of visual regions using textual prompts, but the original architecture has three limitations: lack of multi-scale visual feature fusion, insufficient flexibility in mask generation, and limited text-visual alignment accuracy.
[0081] To address this, this invention redesigns the data processing flow in the query decoder based on the Transformer decoder, aiming to fuse textual semantic embeddings with visual features to generate a segmentation mask targeting the target structure. The input to the query decoder consists of three parts: first, the updated fourth textual feature (i.e., textual semantic embedding) output by the text encoder, which integrates visual semantics and contains semantic information about the dissected concepts and the spatial relationships of the corresponding visual features; second, the updated fourth visual feature (i.e., visual semantic embedding) output by the visual encoder; and third, the pixel-by-pixel dense feature map generated by the visual decoder.
[0082] In this invention, the improvements to the query decoder specifically include the following two aspects:
[0083] 1) To enhance the flexibility of region prediction, the mask generation module was restructured, and a dynamic convolution mechanism driven by masktoken and inspired by Text3DSAM was introduced. The specific steps are as follows:
[0084] After processing through all 6 Transformer decoding layers of the query decoder, the new... A learnable masktoken (in this embodiment, ), and initialized to The vector, concatenated with the text semantic embedding, serves as the query input to the query decoder. After processing through 6 Transformer decoding layers, each mask token is then mapped to 3D convolutional kernel parameters via a lightweight MLP.
[0085] For the pixel-wise dense feature map output by the visual decoder, each dynamic convolutional kernel performs a spatially position-wise convolution operation, and the result is generated after Sigmoid activation. An independent initial mask probability map Corresponding to The model can segment individual target structures (e.g., simultaneously segmenting the "liver" and "spleen"). This mechanism enables the model to segment instances, and the dynamic convolutional kernel can adapt to morphological changes in different anatomical structures (e.g., differences in kidney size).
[0086] 2) To enhance the accurate localization of visual regions by text prompts, a text-visual semantic similarity graph inspired by Text3DSAM is introduced, as implemented below:
[0087] Cross-modal similarity calculation is performed between text semantic embeddings and pixel-wise dense feature maps. First, the text semantic embeddings are mapped using linear projection. The cosine similarity between the mapped text semantic embeddings and each spatial location feature of the pixel-wise dense feature map is calculated to obtain a text-visual semantic similarity map. The text-visual semantic similarity map is then weighted and fused with an initial mask probability map generated by a dynamic convolutional kernel. The optimized mask probability map obtained after weighted fusion is used as the final generated mask probability map.
[0088] This operation forces the model to focus on regions that are highly semantically matched to the text, suppressing the activation of irrelevant regions (e.g., suppressing the response of the right kidney region when prompted with "left kidney").
[0089] After the above improvements, in this embodiment, the specific processing flow of the query decoder of the present invention is as follows:
[0090] CS1: First, text semantic embedding with a learnable mask token (initialized to...) (Vectors) are concatenated to form a query vector sequence The query is then fed into six cascaded Transformer decoding layers. Through iterative updates by these six layers, the text query adapts to the visual context of the current image, generating an updated query vector. This vector integrates textual semantics with the image-specific visual context. The target number of categories, such as when simultaneously segmenting "liver" and "spleen". ).
[0091] In the query decoder of this embodiment, each Transformer decoding layer includes a multi-head self-attention mechanism and a cross-modal attention mechanism. The multi-head self-attention mechanism interacts with the query vector sequence to capture the semantic association between different targets; the cross-modal attention mechanism uses the query vector sequence as the query and the flattened visual semantic embedding as the key and value to perform attention calculation.
[0092] CS2: Then, the dynamic convolution kernel generation stage begins, extracting kernels from the updated query vector. Each mask token is mapped to a 1×1×1 dynamic convolutional kernel by a lightweight MLP. Each dynamic convolutional kernel corresponds to a target category. The parameters of the dynamic convolutional kernel are dynamically adjusted according to the input features to adapt to the morphological differences of different anatomical structures.
[0093] The lightweight MLP in this embodiment is composed of two linear layers with GELU activation cascaded sequentially.
[0094] CS3: Embedding textual semantics through a linear projection layer Mapped to projected text semantic embedding The cosine similarity Sim(x,y,z) of each spatial location feature u(x,y,z) of the projected text semantic embedding and the pixel-by-pixel dense feature map output by the visual decoder is calculated using the following formula, thereby forming a text-visual semantic similarity map.
[0095]
[0096] Where · represents the dot product; ||·|| represents the norm; and (w,y,z) represents the spatial coordinates.
[0097] In this embodiment, the text-visual semantic similarity map is used to quantify the matching degree between text semantics and visual regions. The value at each spatial location in the similarity map represents the similarity between the projected text semantic embedding and the feature at that spatial location in the pixel-by-pixel dense feature map. The higher the value, the higher the matching degree between that location and the text semantics.
[0098] CS4: Dynamic convolutional kernels are applied to the pixel-wise dense feature map output by the visual decoder. Each dynamic convolutional kernel performs a spatially position-wise convolution operation on the pixel-wise dense feature map. The feature map obtained from the spatially position-wise convolution operation is activated by Sigmoid and then weighted and fused with the text-visual semantic similarity map. This is combined with semantic similarity to guide the generation of... A mask probability map
[0099] In this embodiment, the mask probability map of the k-th target category Its value at spatial location (x, y, z) is calculated as follows:
[0100]
[0101] Where σ is the Sigmoid activation function; This represents a dynamic convolutional kernel with target category k, and 64 represents the feature dimension of the visual decoder. λ represents the target category index; λ is the weight, which is 0.5 in this embodiment; * represents the 3D convolution operation.
[0102] In this embodiment, the mask probability map output by the query decoder corresponds to Given a medical text prompt (e.g., ["liver", "spleen", "kidney"]), the mask probability graph for the k-th channel. This represents the segmentation probability of the k-th class of targets.
[0103] like Figure 2 As shown, this process improves boundary accuracy through multi-scale feature fusion, enhances morphological adaptability through dynamic convolution kernels, and strengthens text-visual alignment through semantic similarity guidance. Compared with the original architecture, it significantly improves the segmentation of complex anatomical structures (such as lung segments) and small lesions (such as brain tumors).
[0104] It should be noted that, in this invention, the specific process for generating the 3D segmentation mask is as follows: The optimal segmentation threshold is estimated using the Otsu algorithm and used as the global threshold; based on the global threshold, a binary mask is generated by binarization; the value at each spatial location of the mask probability map is compared with the global threshold; if the value is greater than or equal to the global threshold, the binary mask at that spatial location is set to 1; otherwise, it is set to 0; after obtaining the binary mask, it is post-processed. When the input medical text prompt is a single text prompt, a single-channel 3D segmentation mask is output; when the input medical text prompt contains… When there are multiple text prompts for multiple target categories, a multi-channel 3D segmentation mask is output.
[0105] In this embodiment, based on the mask probability map output by the query decoder, an adaptive threshold segmentation algorithm integrating multi-objective parallel processing, threshold segmentation, and spatial alignment is used to generate a 3D segmentation mask corresponding to the text prompt. The specific process is as follows:
[0106] First, calculate the global threshold τ. k The optimal segmentation threshold is estimated using the Otsu algorithm, which is based on the gray-level histogram of the probability map and maximizes the inter-class variance between the foreground and background. The formula is as follows:
[0107]
[0108] Where ω1(τ) and ω2(τ) are the foreground and background pixel ratios, and μ1(τ) and μ2(τ) are the corresponding mean values; arg max represents the threshold that maximizes the inter-class variance.
[0109] Then, binarization is performed to generate a mask; the binary mask... Its value M at spatial location (x,y,z) k The calculation method for (x, y, z) is as follows:
[0110]
[0111] To improve the spatial continuity and anatomical rationality of the mask, the binary mask undergoes the following post-processing: A 3×3×3 cubic structuring element is used to remove small noise points and smooth boundaries; small internal holes are filled, and adjacent target regions are connected. Furthermore, all connected components are labeled, the volume of each connected component is calculated, connected components with a volume greater than a preset threshold are retained, and small-volume mis-segmented regions are removed.
[0112] When the input is a single text prompt, the output is a single-channel 3D segmentation mask. Strict spatial alignment with the original input medical image, with a one-to-one correspondence between pixel coordinates. For K targets, the output is a multi-channel 3D segmentation mask. Each channel corresponds to a target structure, and the channel order matches the order of the input text prompts. Finally, the obtained 3D segmentation mask is saved in NIfTI format (.nii.gz file), containing mask data and spatial reference information, suitable for various applications such as clinical diagnosis and surgical planning.
[0113] It should be noted that during the training process, the segmentation network calculates the cross-entropy loss between the predicted labels and the actual labels for training data with real labels. and dice loss The cross-entropy loss and the dice loss are then added together to obtain the total loss of the segmentation network. The stochastic gradient descent optimizer is used to continuously optimize the segmentation network with the goal of minimizing the total loss.
[0114] In this embodiment, the formula for calculating the binary cross-entropy loss (BCE Loss) is as follows:
[0115]
[0116] In this embodiment, the formula for calculating Dice Loss is as follows:
[0117]
[0118] in, y represents the number of target categories during training. iFor the real mask, y i (x,y,z) is the value of the actual mask at the spatial location (x,y,z); For the i-th type prediction mask, is the value of the i-th type prediction mask at spatial location (x,y,z); ∈ represents the hyperparameter for preventing division by zero errors. In this embodiment, ∈ = 1e-6.
[0119] The total loss is obtained by weighting and summing the cross-entropy loss and the dice loss, with the weighting coefficients empirically set to 1:1.
[0120]
[0121] By jointly optimizing the total loss, the model's segmentation accuracy for 3D medical images is improved.
[0122] The present invention will now demonstrate the application effect of the text-guided bidirectional progressive fusion medical image segmentation method described in S1 to S2 of the above embodiments on a specific dataset through a specific example, so as to facilitate understanding of the essence of the present invention.
[0123] Example
[0124] The specific implementation process of the text-based bidirectional progressive fusion medical image segmentation method used in this embodiment is as described above and will not be repeated here.
[0125] This embodiment conducted experiments on the publicly available AMOS22 CT dataset and MM WHS MRI dataset. To objectively evaluate the performance of the method in this embodiment, it was compared with current mainstream medical image segmentation methods, including traditional dedicated segmentation models, Transformer-based segmentation models, and general segmentation models. Performance metrics used in medical image segmentation included the Dice similarity coefficient (DSC) and normalized surface distance (NSD). DSC measures the degree of overlap between the segmented region and the ground truth annotation; a higher value indicates better overlap. NSD assesses the proximity of the segmentation boundary to the ground truth boundary; similarly, a higher value indicates better boundary consistency. The experimental results are shown in Table 1.
[0126] Table 1. Experimental results of the method of the present invention on two public datasets.
[0127]
[0128]
[0129] Therefore, the bidirectional progressive fusion medical image segmentation method based on text prompts proposed in this invention demonstrates its advantages in segmentation accuracy and boundary consistency through the improvement of the bidirectional progressive fusion mechanism and query decoder. Experimental results on public datasets verify that this method can effectively improve the segmentation performance of cross-modal feature alignment and complex anatomical structures, and has better practicality and generalization ability.
[0130] It should also be noted that the text-based bidirectional progressive fusion medical image segmentation method described in the above embodiments can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a text-based bidirectional progressive fusion medical image segmentation system corresponding to the text-based bidirectional progressive fusion medical image segmentation method provided in the above embodiments, such as... Figure 4 As shown, it includes:
[0131] The data acquisition module is used to acquire medical images containing organs or lesions, and to preprocess the medical images to obtain preprocessed medical images.
[0132] The results acquisition module inputs preprocessed medical images and medical text prompts into a pre-trained segmentation network based on the SAT model. The segmentation network uses 3D U-Net as the backbone for visual feature extraction and a staged improved BERT model as the text encoder. Simultaneously, a query decoder is constructed within the segmentation network based on the Transformer decoder. First, the text encoder processes the medical text prompts, and then the visual encoder in 3D U-Net processes the preprocessed medical images. A bidirectional progressive fusion module aligns and fuses the text features output from each stage of the text encoder with the visual features output from each stage of the visual encoder, thereby achieving hierarchical interaction. The U-Net visual decoder progressively upsamples and fuses the multi-scale visual features output by the visual encoder to generate pixel-wise dense feature maps for segmentation mask prediction. The text features output by the final stage of the text encoder are used as text semantic embeddings, and the visual features output by the final stage of the visual encoder are used as visual semantic embeddings. The text semantic embeddings, visual semantic embeddings, and pixel-wise dense feature maps are input to the query decoder, which outputs multiple mask probability maps with the same number of target categories. An adaptive threshold segmentation algorithm is used to process the mask probability maps output by the query decoder, and finally, a 3D segmentation mask corresponding to the medical text prompt is generated.
[0133] It is understood that the text-guided bidirectional progressive fusion medical image segmentation method described in S1-S2 above can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the text-guided bidirectional progressive fusion medical image segmentation method provided in the above embodiments. This product includes a computer program / instructions that, when executed by a processor, can implement the text-guided bidirectional progressive fusion medical image segmentation method as described in the above embodiments.
[0134] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the text-based bidirectional progressive fusion medical image segmentation method provided in the above embodiments, such as... Figure 5 As shown, it includes a memory and a processor;
[0135] The memory is used to store computer programs;
[0136] The processor is configured to implement the text-based bidirectional progressive fusion medical image segmentation method described in the above embodiments when executing the computer program.
[0137] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0138] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the text-based bidirectional progressive fusion medical image segmentation method provided in the above embodiments. The storage medium stores a computer program, which, when executed by a processor, can realize the text-based bidirectional progressive fusion medical image segmentation method in the above embodiments.
[0139] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by a processor and can perform the aforementioned steps S1 to S2.
[0140] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.
[0141] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0142] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.
[0143] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.
Claims
1. A text-based, bidirectional, progressive fusion medical image segmentation method, characterized in that, Includes the following steps: S1: Acquire medical images containing organs or lesions, and preprocess the medical images to obtain preprocessed medical images; S2: The preprocessed medical images and medical text prompts are input into a pre-trained segmentation network based on the SAT model. The segmentation network uses 3D U-Net as the backbone network for visual feature extraction and a staged improved BERT model as the text encoder. Simultaneously, a query decoder is constructed within the segmentation network based on the Transformer decoder. First, the text encoder processes the medical text prompts, and then the visual encoder in 3D U-Net processes the preprocessed medical images. A bidirectional progressive fusion module aligns and fuses the text features output from each stage of the text encoder with the visual features output from each stage of the visual encoder, thereby achieving hierarchical interaction. The U-Net visual decoder progressively upsamples and fuses the multi-scale visual features output by the visual encoder to generate pixel-wise dense feature maps for segmentation mask prediction. The text features output by the final stage of the text encoder are used as text semantic embeddings, and the visual features output by the final stage of the visual encoder are used as visual semantic embeddings. The text semantic embeddings, visual semantic embeddings, and pixel-wise dense feature maps are input into the query decoder, which outputs multiple mask probability maps with the same number of target categories. An adaptive threshold segmentation algorithm is used to process the mask probability maps output by the query decoder, and finally, a 3D segmentation mask corresponding to the medical text prompt is generated. In step S2, the specific process of realizing the interaction between the text encoder and the visual encoder through the bidirectional progressive fusion module is as follows: AS1: After the medical text prompts are input into the text encoder, they first pass through the Embedding module to generate the initial input embedding; AS2: The initial input embedding first passes through the Transformer layer of the first stage of the text encoder, and the hidden state output by the first stage is used as the first text feature; the preprocessed medical image first passes through the first stage of the visual encoder path to obtain the first visual feature. AS3: The first text feature and the first visual feature are processed through the first bidirectional progressive fusion module, and the updated first text feature and the updated first visual feature are output. AS4: The updated first text feature passes through the Transformer layer of the second stage of the text encoder, and the hidden state output by the second stage is used as the second text feature; the updated first visual feature then passes through the second stage of the visual encoder path to obtain the second visual feature. The second text feature and the second visual feature are then processed through the second bidirectional progressive fusion module to output the updated second text feature and the updated second visual feature. AS5: The updated second text feature passes through the Transformer layer of the third stage of the text encoder, and the hidden state output by the third stage is used as the third text feature; the updated second visual feature then passes through the third stage of the visual encoder path to obtain the third visual feature. The third text feature and the third visual feature are then processed through the third bidirectional progressive fusion module to output the updated third text feature and the updated third visual feature. AS6: The updated third text feature passes through the Transformer layer of the fourth stage of the text encoder, and the hidden state output by the fourth stage is used as the fourth text feature; the updated third visual feature then passes through the fourth stage of the visual encoder path to obtain the fourth visual feature. The fourth text feature and the fourth visual feature are then processed through the fourth bidirectional progressive fusion module to output the updated fourth text feature and the updated fourth visual feature. In step S2, the specific processing steps in each bidirectional progressive fusion module are as follows: BS1: First, the visual features are dimensionally adjusted by flattening them into vector form, aligning them with the dimensions of the text features; the text features are then mapped to the same channel dimension as the visual features through a learnable linear layer. BS2: The processed visual features and processed text features are each subjected to layer normalization and the contextual relationships within their respective modalities are captured through a self-attention mechanism to generate intermediate visual features and intermediate text features. BS3: Subsequently, cross-attention fusion is performed. For text-guided visual fusion, the intermediate text features are used as queries, and the intermediate visual features are used as keys and values. Cross-attention is used to generate visual features that fuse text semantics. The visual features that fuse text semantics are then processed by a multilayer perceptron and residually connected with the intermediate text features to obtain the updated text features. For visual-guided text fusion, the intermediate visual features are used as queries, and the updated text features are used as keys and values. Cross-attention is used to generate text semantic representations that fuse visual features. The text semantic representations that fuse visual features are then processed by a multilayer perceptron and residually connected with the intermediate visual features to obtain the updated visual features.
2. The bidirectional progressive fusion medical image segmentation method based on text prompts as described in claim 1, characterized in that, In step S2, the text encoder is based on the BERT model consisting of 12 identical Transformer layers, and the Transformer layers of the text encoder are divided according to the number of layers so that they correspond to the four stages of the visual encoder; wherein, the first stage contains Transformer layers 1-6, the second stage contains Transformer layers 7-8, the third stage contains Transformer layers 9-10, and the fourth stage contains Transformer layers 11-12.
3. The text-based bidirectional progressive fusion medical image segmentation method as described in claim 2, characterized in that, In step S2, the visual encoder consists of four stages. The first three stages are each composed of two 3×3×3 convolutional layers with BatchNorm and ReLU activation functions, and a 2×2×2 max pooling layer cascaded in sequence. The last stage is composed of only two 3×3×3 convolutional layers with BatchNorm and ReLU activation functions cascaded in sequence.
4. The bidirectional progressive fusion medical image segmentation method based on text prompts as described in claim 3, characterized in that, In step S2, the specific processing flow of the query decoder is as follows: CS1: First, the text semantic embedding is concatenated with the learnable mask token to form a query vector sequence, which is then input into a series of 6 cascaded Transformer decoding layers. Through iterative updates of the 6 Transformer decoding layers, the text query is adapted to the visual context of the current image, generating an updated query vector. CS2: Then, the dynamic convolution kernel generation stage begins, extracting kernels from the updated query vector. Each of the given mask tokens is mapped to a 1×1×1 dynamic convolutional kernel using a lightweight MLP, with the kernel parameters dynamically adjusted based on the input features. Indicates the number of target categories; CS3: The text semantic embedding is mapped to the projected text semantic embedding through a linear projection layer. The cosine similarity between the projected text semantic embedding and each spatial location feature of the pixel-by-pixel dense feature map output by the visual decoder is calculated to form a text-visual semantic similarity map. CS4: Dynamic convolutional kernels are applied to the pixel-wise dense feature map output by the visual decoder. Each dynamic convolutional kernel performs a spatially position-wise convolution operation on the pixel-wise dense feature map. The feature map obtained from the spatially position-wise convolution operation is activated by Sigmoid and then weighted and fused with the text-visual semantic similarity map. This is combined with semantic similarity to guide the generation of... A mask probability graph.
5. The bidirectional progressive fusion medical image segmentation method based on text prompts as described in claim 4, characterized in that, In the query decoder, each Transformer decoding layer contains a multi-head self-attention mechanism and a cross-modal attention mechanism. The multi-head self-attention mechanism interacts with the query vector sequence. The cross-modal attention mechanism uses the query vector sequence as the query and the flattened visual semantic embedding as the key and value to perform attention calculation.
6. The bidirectional progressive fusion medical image segmentation method based on text prompts as described in claim 4, characterized in that, The specific process for generating a 3D segmentation mask is as follows: The optimal segmentation threshold is estimated using the Otsu algorithm and used as the global threshold; based on the global threshold, a binary mask is generated through binarization. The value at each spatial location of the mask probability map is compared with the global threshold. If the value is greater than or equal to the global threshold, the binary mask value at that spatial location is 1; otherwise, it is 0. After obtaining the binary mask, it undergoes post-processing. When the input medical text prompt is a single text prompt, a single-channel 3D segmentation mask is output; when the input medical text prompt contains... When there are multiple text prompts for multiple target categories, a multi-channel 3D segmentation mask is output.
7. The bidirectional progressive fusion medical image segmentation method based on text prompts as described in claim 1, characterized in that, During the training process, the segmentation network calculates the cross-entropy loss and dice loss between the predicted labels and the real labels for training data with real labels. Then, the cross-entropy loss and dice loss are added together to obtain the total loss of the segmentation network. With minimizing the total loss as the optimization objective, the segmentation network is continuously optimized using a stochastic gradient descent optimizer.
8. A text-based, bidirectional progressive fusion medical image segmentation system, characterized in that, include: The data acquisition module is used to acquire medical images containing organs or lesions, and to preprocess the medical images to obtain preprocessed medical images. The results acquisition module inputs preprocessed medical images and medical text prompts into a pre-trained segmentation network based on the SAT model. The segmentation network uses 3D U-Net as the backbone for visual feature extraction and a staged improved BERT model as the text encoder. Simultaneously, a query decoder is constructed within the segmentation network based on the Transformer decoder. First, the text encoder processes the medical text prompts, and then the visual encoder in 3D U-Net processes the preprocessed medical images. A bidirectional progressive fusion module aligns and fuses the text features output from each stage of the text encoder with the visual features output from each stage of the visual encoder, thereby achieving hierarchical interaction. The U-Net visual decoder progressively upsamples and fuses the multi-scale visual features output by the visual encoder to generate pixel-wise dense feature maps for segmentation mask prediction. The text features output by the final stage of the text encoder are used as text semantic embeddings, and the visual features output by the final stage of the visual encoder are used as visual semantic embeddings. The text semantic embeddings, visual semantic embeddings, and pixel-wise dense feature maps are input into the query decoder, which outputs multiple mask probability maps with the same number of target categories. An adaptive threshold segmentation algorithm is used to process the mask probability maps output by the query decoder, and finally, a 3D segmentation mask corresponding to the medical text prompt is generated. The specific process of implementing the interaction between the text encoder and the visual encoder through the bidirectional progressive fusion module is as follows: AS1: After the medical text prompts are input into the text encoder, they first pass through the Embedding module to generate the initial input embedding; AS2: The initial input embedding first passes through the Transformer layer of the first stage of the text encoder, and the hidden state output by the first stage is used as the first text feature; the preprocessed medical image first passes through the first stage of the visual encoder path to obtain the first visual feature. AS3: The first text feature and the first visual feature are processed through the first bidirectional progressive fusion module, and the updated first text feature and the updated first visual feature are output. AS4: The updated first text feature passes through the Transformer layer of the second stage of the text encoder, and the hidden state output by the second stage is used as the second text feature; the updated first visual feature then passes through the second stage of the visual encoder path to obtain the second visual feature. The second text feature and the second visual feature are then processed through the second bidirectional progressive fusion module to output the updated second text feature and the updated second visual feature. AS5: The updated second text feature passes through the Transformer layer of the third stage of the text encoder, and the hidden state output by the third stage is used as the third text feature; the updated second visual feature then passes through the third stage of the visual encoder path to obtain the third visual feature. The third text feature and the third visual feature are then processed through the third bidirectional progressive fusion module to output the updated third text feature and the updated third visual feature. AS6: The updated third text feature passes through the Transformer layer of the fourth stage of the text encoder, and the hidden state output by the fourth stage is used as the fourth text feature; the updated third visual feature then passes through the fourth stage of the visual encoder path to obtain the fourth visual feature. The fourth text feature and the fourth visual feature are then processed through the fourth bidirectional progressive fusion module to output the updated fourth text feature and the updated fourth visual feature. The specific processing steps in each bidirectional progressive fusion module are as follows: BS1: First, the visual features are dimensionally adjusted by flattening them into vector form, aligning them with the dimensions of the text features; the text features are then mapped to the same channel dimension as the visual features through a learnable linear layer. BS2: The processed visual features and processed text features are each subjected to layer normalization and the contextual relationships within their respective modalities are captured through a self-attention mechanism to generate intermediate visual features and intermediate text features. BS3: Subsequently, cross-attention fusion is performed. For text-guided visual fusion, the intermediate text features are used as queries, and the intermediate visual features are used as keys and values. Cross-attention is used to generate visual features that fuse text semantics. The visual features that fuse text semantics are then processed by a multilayer perceptron and residually connected with the intermediate text features to obtain the updated text features. For visual-guided text fusion, the intermediate visual features are used as queries, and the updated text features are used as keys and values. Cross-attention is used to generate text semantic representations that fuse visual features. The text semantic representations that fuse visual features are then processed by a multilayer perceptron and residually connected with the intermediate visual features to obtain the updated visual features.
Citation Information
Patent Citations
Medical image segmentation method fusing multi-modal graphic and text information
CN119205800A