Medical image segmentation method based on text guidance and electronic equipment

By transforming the medical natural language processing data set and introducing a conjugated cross attention mechanism, the problem of poor modal alignment effect in medical image segmentation is solved, and a more efficient and accurate image segmentation effect is achieved.

CN120374970APending Publication Date: 2025-07-25TONGJI UNIV

Patent Information

Application Number
CN202510330601.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, transfer learning in medical image segmentation has the problem that the transfer learning effect is poor between different modes, transfer learning takes a long time, and the semantic alignment effect of the encoder decoder multimodal information in multimodal medical image segmentation is not ideal.

Method used

The medical natural language processing data set is transformed into a multi-classified data set, the text is encoded using a BERT encoder, and the modified medical image segmentation backbone network is input for pre-training, and the pre-training parameters are migrated to a multi-modal medical image segmentation network based on the conjugated cross attention mechanism, and modal fusion is performed through a multi-layer conjugated cross attention mechanism, and image processing and feature fusion are performed in combination with the Swin-Transformer block.

Benefits of technology

The pre-training time is shortened, the efficiency and accuracy of medical image segmentation is improved, the problem of poor modal alignment is solved, and more efficient and accurate image segmentation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374970A_ABST
    Figure CN120374970A_ABST
Patent Text Reader

Abstract

The invention relates to a medical image segmentation method based on text guidance and electronic equipment. The method comprises the following steps: transforming a medical natural language processing data set into a multi-classification data set; pre-training a medical image segmentation model by using the medical text: encoding the text in the transformed multi-classification data set by using a BERT encoder, and inputting the encoded text into a transformed medical image segmentation backbone network for pre-training, so that the model learns the internal relation in the semantics of the medical text; migrating the pre-training parameters to a multi-modal medical image segmentation network based on a conjugate cross attention mechanism for training; and inputting the medical image to be segmented and the medical text corresponding to the medical image to be segmented into the trained multi-modal medical image segmentation network, and outputting an image segmentation result. Compared with the prior art, the method has the advantages that the pre-training time is shortened, and the segmentation efficiency and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and in particular, to a method for medical image segmentation guided by text and an electronic device. Background Art

[0002] Medical images are the most important basis for doctors to diagnose patients' diseases and propose treatment plans. Medical image segmentation is the foundation of intelligent healthcare. Traditional transfer learning is applied to situations where there is less data for the target task but relatively rich data for similar pre-training tasks. For transfer learning in medical image segmentation, there are problems such as poor transfer learning effects between different modalities and long time consumption for transfer learning; for multi-modal medical image segmentation, there is a problem of poor semantic alignment of multi-modal information in the encoder-decoder.

[0003] There is a work BioBERT that uses a dataset of medical natural language tasks to transfer to BERT pre-trained in natural language to verify the effectiveness of BERT in medical natural language tasks, but it does not study transfer learning across image and text modalities in the medical field.

[0004] There is a work Swin-Unet, which is a deep learning model based on the Swin Transformer architecture and is specifically used for medical image segmentation tasks. It combines the advantages of Swin Transformer and U-Net, aiming to improve the performance of medical image segmentation. Swin-UNet adopts Swin Transformer and PatchMerging, Patch Expanding structures that take into account both global and local information. Swin-UNet adopts a hierarchical design and can adjust the depth and width of the model according to the requirements of the task to adapt to images with different resolutions and complexities. There is a work LViT that uses CNN and Transformer modules to fuse features in a double-U structure through a pixel-level attention module, retaining local image and text features, and this work improves the depth of feature fusion in text-guided medical image segmentation tasks. However, the first of the above works does not combine multi-modal information, and the second work still uses a convolutional structure to process image information without a unified structure for processing different modal data, resulting in complex design of the fusion module after feature extraction in the network and unsatisfactory fusion effects.

[0005] After retrieval, Chinese Patent Application Publication No. CN118262115A discloses a method for multi-modal medical image fusion segmentation guided by text prompts, including the following steps:

[0006] (1) Obtain the commonly used Brain Tumor Segmentation (BraTS) dataset in multimodal medical images, a total of 369 groups of 3D images, and perform data preprocessing. After normalization, save them as numpy-format images;

[0007] (2) Design an image-text fusion segmentation network. The image part includes multiple modal feature extraction encoders and a shared feature decoder, and the text part includes two pre-trained modal text feature extractors and a category text feature extractor;

[0008] (3) The modal feature extraction encoder extracts multi-level semantic information of the image by using a U-Net encoder;

[0009] (4) Use the CLIP model to extract modal text semantic information, and linearly align the extracted modal text semantic information with the different-level image semantic information extracted by the U-Net encoder;

[0010] (5) Use the attention mechanism to perform cross-attention fusion on the image semantic information and text semantic information in each modality, and perform Concatenate data splicing on the fusion results of multiple modalities; The features after image-text fusion are obtained through the shared decoder to get a preliminary segmentation result;

[0011] (6) Use the CLIP model to extract category text semantic information, align and fuse the deepest-level image semantic information extracted by the U-Net with the category text semantic information, convert it into model parameters, and guide the preliminary segmentation result to complete the specified segmentation requirements to obtain the final segmentation result.

[0012] (7) Evaluate the performance of the model by calculating the Dice coefficient and 95% Hausdorff distance for each category. There are problems in the existing patent application that the dataset is not transformed and classified, and the segmentation network is not pre-trained, resulting in poor performance in terms of segmentation result accuracy and speed.

[0013] How to achieve the effectiveness and efficiency of transfer learning between different modalities, thereby improving the accuracy of medical image segmentation, has become a technical problem to be solved. Summary of the Invention

[0014] The purpose of the present invention is to provide a text-guided medical image segmentation method and an electronic device to overcome the defects existing in the above-mentioned prior art.

[0015] The purpose of the present invention can be achieved by the following technical solutions:

[0016] According to one aspect of the present invention, a text-guided medical image segmentation method is provided, and the method includes the following:

[0017] Step 1: Transform the medical natural language processing dataset into a multi-classification dataset;

[0018] Step 2, perform pre-training on the medical image segmentation model using the obtained medical text: Encode the text in the transformed multi-classification dataset using the BERT encoder and input it into the transformed medical image segmentation backbone network for pre-training, enabling the model to learn the internal connections in the medical text semantics;

[0019] Step 3, transfer the pre-training parameters in Step 2 to the multi-modal medical image segmentation network based on the conjugate cross-attention mechanism for training;

[0020] Step 4, input the medical image to be segmented and its corresponding medical text into the trained multi-modal medical image segmentation network, and output the image segmentation result.

[0021] Preferably, in Step 1, the transformation process of the medical natural language processing dataset includes:

[0022] For the discriminative task of named entity recognition, convert the category labels of the words in the dataset sentences into classification categories;

[0023] For the discriminative task of relation extraction, concatenate the words with relations in the dataset at the end of the sentence, change the label to the number corresponding to the category, and the number of relations in the original dataset is the number of numbers;

[0024] For the generative task of text question answering, for paragraphs with question answers, change the answer to a yes / no question and concatenate it at the end of the paragraph, and mark it as 1; for paragraphs without answers, change the question to a yes / no question and concatenate it at the end of the paragraph, and change the label to 0.

[0025] Preferably, the medical image segmentation backbone network is a network that removes the image embedding layer in the medical image segmentation model and adds a tokenizer, word vector embedding, and MLP layer.

[0026] Preferably, combine multiple tasks in the multi-classification dataset for the pre-training.

[0027] Preferably, the multi-modal medical image segmentation network is a network that removes the image embedding layer, text feature extraction layer, and multi-modal fusion layer in the medical image segmentation model and adds a word vector embedding and MLP layer.

[0028] Preferably, the multi-modal medical image segmentation network uses the text features of the frozen BERT encoder and the image features encoded by the medical image segmentation baseline model Swin-Unet, and performs modal fusion through multi-layer conjugate cross-attention;

[0029] The multi-modal medical image segmentation network includes a text processing module, an image processing module, and a feature fusion module;

[0030] The text processing module includes a BERT encoder and a text end-to-end feature extractor based on multi-layer conjugate cross-attention, which is used to process the input text features;

[0031] The image processing module includes multiple consecutive Swin-Transformer blocks, which are used to process the input image features;

[0032] The feature fusion module is used to perform multi-modal fusion on the processing results of text features and image features.

[0033] More preferably, the Swin-Transformer block includes a Swin-Transformer encoder and a Swin-Transformer decoder, where the Swin-Transformer encoder extracts image semantics, and the Swin-Transformer decoder block uses the image semantics and image features for image segmentation.

[0034] More preferably, the multi-layer conjugate cross-attention includes consecutive cross-attention layers, convolutional layers, SoftMax layers, and Concate layers, and an skip connection structure is applied to transfer the multi-modal fusion information from the encoder to the decoder;

[0035] The multi-layer conjugate cross-attention includes three parts. The first part uses the medical diagnosis features extracted by the frozen text encoder as Q, and the medical image features output by the image encoder with the multi-layer Swin-Transformer as the structure as K and V for cross-attention calculation to obtain the image features T1 of interest in medical text; the second part uses the image feature T1 as the Q of the attention mechanism, and the image feature as K and V for cross-attention calculation to obtain the fusion feature T2; the third part splices the multi-modal fusion features obtained in the previous two parts with the original image features to obtain the multi-modal information T3 transmitted by the final skip connection structure.

[0036] Preferably, the multi-modal medical image segmentation network is based on the SoftDice function loss SoftDice and the binary cross-entropy function loss CE for training, specifically:

[0037]

[0038] where N is the number of image pixels, y true is the label, and y pred is the probability predicted by the model as the label.

[0039] According to a second aspect of the present invention, there is provided an electronic device including a memory and a processor, where a computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1) The present invention transforms the task dataset into a multi-classification dataset, pre-trains a medical image segmentation model using medical texts in a medical context to obtain pre-training parameters that are better than the state of randomly initialized parameters, and uses the text encoding in the multi-classification dataset to pre-train the transformed medical image segmentation backbone network. Since the magnitude of text data is much smaller than that of image data, the time and video memory overheads in the pre-training stage are much smaller than those in other image tasks, and similar effects can be achieved, greatly shortening the pre-training time; the pre-training parameters are migrated to a multi-modal medical image segmentation network based on the conjugate cross-attention mechanism for training, enabling the model to deeply understand the text-image pair information, thus taking into account both segmentation efficiency and accuracy.

[0042] 2) Based on the transformed multi-classification dataset, the present invention transforms the medical image segmentation backbone network, deletes the image embedding layer, adds a tokenizer, word vector embedding, and MLP layer to better adapt to the multi-classification dataset, encodes the text therein using an encoder and inputs it into the transformed medical image segmentation backbone network for pre-training, enabling the model to learn the internal connections in the medical text semantics, thereby achieving more efficient and accurate image segmentation.

[0043] 3) The present invention introduces a conjugate cross-attention mechanism on the basis of a medical image segmentation model for deep multi-modal information fusion, and applies a skip connection structure to transfer the multi-modal fusion information from the encoder to the decoder, enabling the text features and image features to be aligned layer by layer, solving the problem of poor alignment effect in the prior art, and supplementing the part of information lost in the encoder stage through multi-modal feature fusion, thereby improving the segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flow chart of the method for training text-guided medical image segmentation in the present invention;

[0045] Figure 2 It is a schematic diagram of the preprocessing process of the medical natural language processing dataset in the present invention;

[0046] Figure 3 It is a schematic diagram of the structure of the medical image segmentation network based on conjugate cross-attention in the present invention;

[0047] Figure 4 It is a schematic diagram of the structure of the conjugate cross-attention mechanism in the present invention. Detailed implementation manners

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] The present invention application provides a text-guided medical image segmentation method for the problems existing in the foregoing prior art. The method pre-trains a modified medical image segmentation backbone network using a modified medical natural language processing task, and proposes a conjugate cross-attention mechanism in the migration stage for deep fusion of multi-modal features.

[0050] Embodiment 1

[0051] This embodiment relates to a text-guided medical image segmentation method. The method pre-trains a medical image segmentation backbone network using a modified medical natural language task dataset (i.e., processing task), and loads the parameters in the pre-training stage in the migration stage to fine-tune the parameters of a complete single-modal or multi-modal medical image segmentation network, ultimately improving the accuracy of the model for medical image segmentation.

[0052] As Figure 1 , the method includes the following steps:

[0053] S1. The first stage: Preprocess the medical natural language processing dataset and transform it into a multi-classification dataset;

[0054] S2. The second stage: Obtain word vectors from the transformed multi-classification dataset through a tokenizer and a BERT encoder, and complete the word vectors according to the input format of the image task;

[0055] S3. The second stage: Remove the image embedding layer in the segmentation model, add a tokenizer, a word vector embedding, and an MLP layer to adapt to the multi-classification task, and obtain a medical image segmentation backbone network. Use the dataset obtained in S2 for pre-training, and various tasks can be combined; classify the output results to enable the model to learn the internal connections in the text semantics;

[0056] S4. The third stage: Obtain the medical image to be segmented and its corresponding text description, and preprocess the original data;

[0057] S5. The third stage: Transfer the pre-trained parameters to a medical image segmentation network based on the conjugate cross-attention mechanism (such as the CCMA UNet model) for training. Using data transfer learning with different modalities that have similarities can improve the speed of pre-training and increase the performance of model parameters on downstream tasks and the training convergence speed. Among them, the medical image segmentation network can be: Delete the word vector embedding layer and the MLP layer from the pre-trained medical image segmentation backbone network obtained in S3, and add an image embedding layer, a text feature extraction layer, and a feature fusion layer to obtain a multi-modal medical image segmentation model for training medical image segmentation tasks. The medical image segmentation network can also be: Delete the word vector embedding layer and the MLP layer from the pre-trained medical image segmentation backbone network obtained in S3, and add an image embedding layer to obtain a single-modal medical image segmentation model.

[0058] In the first stage, the process of transforming the medical natural language task dataset is as Figure 2 , including: For the discriminative task of named entity recognition (NER), convert the category labels of the words in the dataset sentences into classification categories; for the discriminative task of relation (RE) extraction, concatenate the words with existing relations in the dataset at the end of the sentence, and change the label to the corresponding category number. The number of relations in the original dataset is the number of numbers. For the generative task of text question answering (QA), for those with question answers in the paragraph, change the answer to a general question and concatenate it at the end of the paragraph, and mark the label as 1. For those without answers in the paragraph, change the question to a general question and concatenate it at the end of the paragraph, and change the label to 0.

[0059] Specifically, for the medical named entity recognition (NER) task dataset, each sentence in each abstract is taken as a group. After word embedding by inputting into the BERT encoder, padding with 0 is performed, and the labels of the words in the sentence are converted into categories;

[0060] For the medical text question answering (QA) task dataset, after changing the inferable answer in the abstract to a general question, set the label as 1; if the question is not mentioned in the abstract, change the question to a general question and set the label as 0;

[0061] For the medical relation (RE) extraction task dataset, concatenate the two words in the relation at the end of the sentence. If the relation is promotion, then the label is 1. If the relation is neutral, then the label is 2. If the relation is inhibition, then the label is 0.

[0062] In the dataset transformation stage, for discriminative medical natural language processing tasks, use the original labels to construct categories; for generative task datasets, use the data structure of the original dataset to construct categories.

[0063] The medical image segmentation training method based on text guidance is as Figure 1As shown, an example of dataset transformation in the pre-training stage is as follows Figure 2 As shown. The multi-modal medical image segmentation network in the transfer learning stage is as follows Figure 3 As shown, the model mainly includes three modules: text processing, image processing, and feature fusion. The structure is as follows:

[0064] The text processing module contains a BERT encoder, an end-to-end text feature processing structure formed by connecting the first part of each layer of conjugate cross-attention. It can be expressed as:

[0065]

[0066] Where the subscript i represents the level of the conjugate cross-attention structure, and the superscript data represents the part of the conjugate cross-attention structure. T i 1 represents the text feature in the first part of the conjugate cross-attention at the i-th level, and V i 1 represents the image feature in the first part of the conjugate cross-attention at the i-th level. CrossAttention() represents the conjugate cross-attention operation.

[0067] The image processing module consists of Swin-Transformer blocks, including sliding windows, window multi-head self-attention (Multi-Head Self-Attention), feed-forward neural networks (Feed-Forward Neural Network), etc. Taking the image features as the input of the Swin-Transformer blocks, layer normalization is used at the input end of each layer of the Swin-Transformer to stabilize the training process.

[0068] The multi-modal medical image segmentation network includes a U-shaped encoder-decoder structure, continuous Swin-Transformer blocks, a skip connection structure based on conjugate cross-attention, a BERT encoder, and an end-to-end text feature extractor based on conjugate cross-attention.

[0069] Using the text features of the frozen BERT encoder and the picture features encoded by the medical image segmentation baseline model Swin-Unet for modality fusion through multi-layer conjugate cross-attention, including:

[0070] The image feature encoder and decoder based on Swin-Transformer blocks, including extracting image features using a convolutional neural network-based image feature embedding layer, extracting image semantics using continuous Swin-Transformer encoder blocks, and using image semantics and image features for image segmentation tasks by continuous Swin-Transformer decoder blocks.

[0071] The multi - layer conjugate cross - attention mechanism includes consecutive cross - attention layers, convolutional layers, SoftMax layers, and Concate layers. The multi - layer conjugate cross - attention mechanism is applied to the skip - connection structure to transfer multi - modal fusion information from the encoder to the decoder.

[0072] Based on the encoder - decoder structure, a skip - connection mechanism for multi - modal fusion is added to transfer the multi - modal information after deep fusion. The feature extractor of the text modality is frozen to reduce the number of parameters.

[0073] The skip - connection structure of conjugate cross - attention includes a cross - attention mechanism with text features fused with image information as Q, image features as K and V; a cross - attention mechanism with image features guided by text regions of interest as Q, image features as K and V; a 1x1 channel attention mechanism, and a layer for splicing multi - modal fusion features and image features.

[0074] The conjugate cross - attention structure is as Figure 4 shown. In the first part, the text - guided visual features are obtained through the cross - attention mechanism in this paper N T is the number of image features, C i is the dimension of image features, as shown in the formula:

[0075]

[0076] Q is the image feature, K and V are the text features, and dim is the dimension of the image feature.

[0077] In the second part, Q represents the text - guided visual features, and K and V represent the image features:

[0078]

[0079] where represents the text feature in the conjugate cross - attention of the second part at the i - th level, represents the text feature in the conjugate cross - attention of the first part at the (i + 1) - th level.

[0080] In the second part, the image features representing the same semantics in the text are extracted through dot - product (Dropout). Generally speaking, the visual features of interest in the text features are extracted. The output of the first part serves as the input to the next - layer conjugate cross - attention mechanism and the text - information prompt input of this layer.

[0081] In the second part, the roles of visual information and visual information enriched with text semantics are exchanged in the cross - attention mechanism, with visual information as Q, and visual information enriched with text semantics as K and V, as shown in the formula.

[0082]

[0083] The second part not only integrates visual information and text-guided visual information, but also maintains the same dimension as the input visual features in terms of structure. The above two parts constitute a conjugate attention mechanism for visual features.

[0084] In the third part, the fused features are used as the input to pass through a 1x1 channel attention convolutional layer and a batch normalization layer, and the original features in the encoder are concatenated with the fused feature V i 3 to transmit semantic information to the decoder of the corresponding layer. Through these three-part operations, overall, it can be regarded as a visual feature passing through a multi-head attention mechanism guided by text features. This structure is used to supplement the part of information lost in the encoder stage in the U-shaped segmentation network, as shown in the formula.

[0085] V i 4 = concatenate(conv(V i 3 ), V i 1 )

[0086] For the multi-layer conjugate cross-attention mechanism, the first part includes using the medical diagnosis features extracted by the frozen text encoder as Q, the medical image features encoded by the image encoder with a multi-layer Swin-Transformer structure as K, and V to perform cross-attention calculation to obtain the image features T1 of interest in the diagnostic text; the second part uses the image features T1 as Q of the attention mechanism, the image features as K, and V to perform cross-attention calculation to obtain the fused feature T2; the third part concatenates the multi-modal fused features obtained from the above two parts with the original image features to obtain the multi-modal information T3 transmitted by the final skip connection structure.

[0087] Multiple cross-attention modules or conjugate cross-attention layers can be stacked according to medical image segmentation tasks of different orders of magnitude. Select different numbers of conjugate multi-layer cross-attention basic modules according to the model depth.

[0088] Embodiment 2

[0089] This embodiment also relates to a text-guided medical image segmentation method, which will be described below with an actual application as an example, including the following steps:

[0090] Step1, preprocess the medical image natural language processing dataset;

[0091] Among them, the NCBI disease dataset is used for the named entity recognition task, the BioASQ dataset is used for the medical text question answering task, and the EU-ADR dataset is used for the medical relation extraction task.

[0092] Step2, transform the backbone network of medical image segmentation, delete the image embedding layer and the upsampling layer, add a tokenizer, word vector embedding, and MLP layer to adapt to the multi-classification task, and pre-train the transformed backbone network of medical image segmentation using the preprocessed data. For the single-modal medical image segmentation network, remove the image embedding module and add the word vector embedding and MLP module; for the multi-modal medical image segmentation network, remove the image embedding module, text feature extraction module, and multi-modal fusion module, and add the word vector embedding and MLP module.

[0093] Step3, load the parameters of the pre-trained medical image backbone network into the multi-modal fusion network based on the conjugate cross-attention mechanism; the pre-trained network in the transfer stage has an encoder for processing the text modality and a multi-modal fusion structure compared to the backbone network of medical image segmentation.

[0094] Step4, divide the image and text datasets according to a preset ratio, which are randomly divided into 7:1:2 for training, validation, and testing respectively;

[0095] Step5, input the training set of the dataset into the medical image segmentation model to train the medical image segmentation model;

[0096] The parameters of the model are updated using the Adam optimization strategy, and an early stopping strategy is used, with the threshold set at 50 epochs. For the batch size of one training, due to the limitation of video memory, it is set to 40.

[0097] Step6, input the test set of the dataset into the trained medical image segmentation model to obtain the segmentation results.

[0098] Among them, the segmentation results are evaluated by the Dice similarity coefficient (DICE) and the mean intersection over union (IOU).

[0099] After multiple rounds of experiments, the results of different pre-training tasks are shown in Table 1. For the results on different pre-training tasks, then we load the pre-trained model parameters into the multi-modal image segmentation network and conduct experiments on the QaTa-Covid19+ dataset and the SIIMPneumothorax dataset. After training for 50 epochs and 120 epochs respectively, the results of different pre-training tasks in the transfer stage are shown in Table 2.

[0100] Table 1

[0101] Pre-training task Round Time overhead IOU(%) DICE(%) NER 40 3.04h 83.96 90.35 QA 64 1.06h 31.40 61.34 RE 40 0.266h 93.37 96.16 ACDC - 10+h - 79.13

[0102] Among them, ACDC is a picture dataset, representing pre-training with pictures.

[0103] Table 2

[0104]

[0105] Meanwhile, we will directly train the conjugate cross-attention network proposed in stage three, compare it with a variety of single-modal and multi-modal medical image segmentation networks, prove the effectiveness of the model of this application for multi-modal feature fusion processing, and obtain the results in Table 3.

[0106]

[0107]

[0108] The pre-training stage includes:

[0109] Use the encoder to encode the transformed text, and after inputting it into the pre-training model, output the classification result, that is, use the result output after the text vector passes through the transformed medical image segmentation backbone network for classification, so that the model learns the internal connection in the text semantics, and its parameters are helpful for the training of the medical image segmentation network.

[0110] The transformed medical image segmentation backbone network adopts a network that removes the image embedding module and the upsampling module and adds a tokenizer, word vector embedding, and MLP module.

[0111] Compared with pre-training using medical images, the pre-training time is shortened by an order of magnitude.

[0112] Among them, the transfer learning stage includes: loading the parameters of the medical image segmentation backbone network obtained in the pre-training stage into a multi-modal medical image segmentation network based on the cross-attention mechanism.

[0113] Train the medical image segmentation model based on the loss function to obtain a trained medical image segmentation model, where the loss function includes the SoftDice function loss SoftDice and the binary cross-entropy function loss CE .

[0114]

[0115] Among them, N is the number of image pixels, y true is the label, and y pred is the probability that the model predicts as the label.

[0116] Example 3

[0117] The electronic device of the present invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0118] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a magnetic disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0119] The processing unit executes the various methods and processes described above. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute the method by any other suitable means (e.g., by means of firmware).

[0120] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0121] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0122] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] As described above, the foregoing are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A text-guided medical image segmentation method, characterized in that The method includes the following: Step 1: Transform the medical natural language processing dataset into a multi-classification dataset; Step 2, use the obtained medical text for pre-training of the medical image segmentation model: Encode the text in the transformed multi-classification dataset using the BERT encoder and input it into the transformed medical image segmentation backbone network for pre-training, so that the model learns the internal connections in the medical text semantics; Step 3, transfer the pre-training parameters in Step 2 to the multi-modal medical image segmentation network based on the conjugate cross-attention mechanism for training; Step 4, input the medical image to be segmented and its corresponding medical text into the trained multi-modal medical image segmentation network, and output the image segmentation result.

2. The method for text-guided medical image segmentation according to claim 1, wherein In the said Step 1, the transformation process of the medical natural language processing dataset includes: For the discriminative task of named entity recognition, convert the category labels of the words in the dataset sentences into classification categories; For the discriminative task of relation extraction, splice the words with relations in the dataset at the end of the sentence, change the label to the number corresponding to the category, and the number of relations in the original dataset is the number of the numbers; For the generative task of text question answering, for those with question answers in the paragraph, change the answer to a general question and splice it at the end of the paragraph, and mark it as 1; for those without answers in the paragraph, change the question to a general question and splice it at the end of the paragraph, and change the label to 0.

3. A method for text-guided medical image segmentation according to claim 1, wherein The said medical image segmentation backbone network is a network that removes the image embedding layer in the medical image segmentation model and adds a tokenizer, word vector embedding, and MLP layer.

4. A text-guided medical image segmentation method according to claim 1, characterized in that Combine multiple tasks in the multi-classification dataset for the said pre-training.

5. A method for text-guided medical image segmentation according to claim 1, characterized in that, The said multi-modal medical image segmentation network is a network that removes the word vector embedding and MLP layer in the medical image segmentation model and adds an image embedding layer, a text feature extraction layer, and a feature fusion layer.

6. The method for text-guided medical image segmentation according to claim 1, wherein The said multi-modal medical image segmentation network uses the text features of the frozen BERT encoder and the picture features encoded by the medical image segmentation baseline model Swin-Unet, and performs modal fusion through multi-layer conjugate cross-attention; The said multi-modal medical image segmentation network includes a text processing module, an image processing module, and a feature fusion module; The said text processing module contains a BERT encoder and a text end-to-end feature extractor based on multi-layer conjugate cross-attention, and is used to process the input text features; The said image processing module includes multiple consecutive Swin-Transformer blocks and is used to process the input image features; The said feature fusion module is used to perform multi-modal fusion on the processing results of the text features and the image features.

7. A text-guided medical image segmentation method according to claim 6, characterized in that, The said Swin-Transformer block includes a Swin-Transformer encoder and a Swin-Transformer decoder, where the Swin-Transformer encoder extracts the image semantics, and the Swin-Transformer decoder block uses the image semantics and the image features for image segmentation.

8. A method for text-guided medical image segmentation according to claim 6, characterized in that, The multi-layer conjugate cross-attention includes consecutive cross-attention layers, convolutional layers, SoftMax layers, and Concate layers, and applies a skip connection structure to transmit multi-modal fusion information from the encoder to the decoder; The multi-layer conjugate cross-attention includes three parts. The first part uses the medical diagnosis features extracted by the frozen text encoder as Q, and the medical image features output by the image encoder with the structure of the multi-layer Swin-Transformer as K and V for cross-attention calculation to obtain the image features T1 of interest in medical texts. The second part uses the image feature T1 as the Q of the attention mechanism, and the image features as K and V for cross-attention calculation to obtain the fusion feature T2. The third part splices the multi-modal fusion features obtained in the first two parts with the original image features to obtain the multi-modal information T3 transmitted by the final skip connection structure.

9. A method for text-guided medical image segmentation according to claim 1, wherein The described multi-modal medical image segmentation network is based on the loss of the SoftDice function SoftDice and the loss of the binary cross-entropy function CE for training, specifically as follows: where N is the number of image pixels, y true is the label, and y pred is the probability that the model predicts as the label.

10. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-modal medical image fusion segmentation method based on text prompt guidance

    CN118262115A

Cited By

  • Medical multi-modal model training method and device, electronic equipment and storage medium

    CN120853191A