Image segmentation method and device
Through the visual language model, the image and text are encoded into the same feature space, combined with the image decoder, the problem of insufficient segmentation accuracy in medical image segmentation is solved, and high-precision segmentation and automated labeling of unprocessed image modes are realized.
Patent Information
- Application Number
- CN202410511601.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-04-26
AI Technical Summary
The existing medical image segmentation methods have the problem of insufficient segmentation accuracy in small sample learning and general segmentation methods, especially in unprocessed medical image modalities or segmentation tasks, which are difficult to obtain high-precision segmentation results.
The visual language model is adopted to encode the training sample images, image segmentation masks and text segmentation instructions into the same feature space. Through the combination of visual language model training and image decoder, multimodal feature vector decoding of images and text is realized, and segmentation accuracy is improved.
It improves the accuracy of medical image segmentation, can perform high-precision segmentation of general medical images and unprocessed medical image modalities, simplifies the training needs of markers, and realizes automated marking.
Smart Images

Figure CN118334660B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of computer vision, natural language processing, and deep learning technology. Background Art
[0002] Medical image segmentation is a very important medical image processing technology. Currently, methods based on deep neural networks are usually adopted to establish a segmentation model. By learning the manual segmentation annotations of medical images, high-precision segmentation results can be output.
[0003] The commonly used medical image segmentation methods mainly include the following two: few-shot learning methods and general segmentation methods. Few-shot learning methods are a class of methods designed for task scenarios with a small amount of available training samples. If there is only one sample for each category, it is called single-shot learning. These methods can achieve more ideal fitting performance than the naive neural network training method under the condition of the same amount of available data. The vision segmentation base model is a representative of general segmentation methods. Through approximately 11 million natural images and 1.1 billion mask data of different targets, the segmentation of the target of interest in almost all natural images is realized. It allows users to interactively obtain the segmentation results of the region of interest by clicking or drawing a box. Summary of the Invention
[0004] The embodiments of the present disclosure propose an image segmentation method, apparatus, device, storage medium, and program product.
[0005] In a first aspect, the embodiments of the present disclosure propose a method for training a vision language model, including: encoding a training sample image, a training sample image segmentation mask, and a training sample text segmentation instruction into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group, and a training sample text feature vector group; inputting the training sample image feature vector group and the training sample text feature vector group into an initialized vision language model to obtain a training sample multimodal feature vector group; calculating a first loss based on the training sample multimodal feature vector group, the training sample image segmentation mask feature vector group, and the training sample text feature vector group; and adjusting the parameters of the initialized vision language model based on the first loss to obtain a trained vision language model.
[0006] In a second aspect, embodiments of the present disclosure propose an image segmentation method, including: encoding an image and a text segmentation instruction into the same feature space to obtain an image feature vector group and a text feature vector group; inputting the image feature vector group and the text feature vector group into a vision-language model to obtain a multi-modal feature vector group, where the vision-language model is trained with a training sample image feature vector group and a training sample text feature vector group as inputs and a training sample image segmentation mask feature vector group and a training sample text feature vector group as supervision; decoding the multi-modal feature vector group to obtain a text segmentation description and an image segmentation mask.
[0007] In a third aspect, embodiments of the present disclosure propose a vision-language model training device, including: an encoding module configured to encode a training sample image, a training sample image segmentation mask, and a training sample text segmentation instruction into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group, and a training sample text feature vector group; a first input module configured to input the training sample image feature vector group and the training sample text feature vector group into an initialized vision-language model to obtain a training sample multi-modal feature vector group; a first calculation module configured to calculate a first loss based on the training sample multi-modal feature vector group, the training sample image segmentation mask feature vector group, and the training sample text feature vector group; and a first adjustment module configured to adjust parameters of the initialized vision-language model based on the first loss to obtain a trained vision-language model.
[0008] In a fourth aspect, embodiments of the present disclosure propose an image segmentation device, including: a first encoding module configured to encode an image and a text segmentation instruction into the same feature space to obtain an image feature vector group and a text feature vector group; an input module configured to input the image feature vector group and the text feature vector group into a vision-language model to obtain a multi-modal feature vector group, where the vision-language model is trained with a training sample image feature vector group and a training sample text feature vector group as inputs and a training sample image segmentation mask feature vector group and a training sample text feature vector group as supervision; and a decoding module configured to decode the multi-modal feature vector group to obtain a text segmentation description and an image segmentation mask.
[0009] In a fifth aspect, embodiments of the present disclosure propose an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect or the second aspect.
[0010] Sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the method described in the first aspect or the second aspect.
[0011] Seventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method described in the first aspect or the second aspect when executed by a processor.
[0012] An embodiment of the present disclosure provides an image segmentation method, which uses a vision-language model for image segmentation, improving the accuracy of image segmentation. When applied to the field of medical image segmentation, it can not only segment general medical images, but also segment medical image modalities or segmentation tasks that have not been processed by the vision-language model, thereby obtaining more accurate segmentation results. It can be applied to an automated annotation platform, enabling annotators to output corresponding segmentation results only through text guidance without strict training, nor the need to read the image to be annotated and find the target area of interest.
[0013] The key or important features of the embodiments of the present disclosure do not limit the scope of the present disclosure either. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more obvious. The drawings are used to better understand the solution and do not limit the present disclosure. Among them:
[0015] Figure 1 is a flowchart of an embodiment of the vision-language model training method according to the present disclosure;
[0016] Figure 2 is a flowchart of another embodiment of the vision-language model training method according to the present disclosure;
[0017] Figure 3 is a flowchart of an embodiment of the image segmentation method according to the present disclosure;
[0018] Figure 4 is a flowchart of another embodiment of the image segmentation method according to the present disclosure;
[0019] Figure 5 is a flowchart of another embodiment of the image segmentation method according to the present disclosure;
[0020] Figure 6 is a scene diagram that can implement the image segmentation method of the embodiment of the present disclosure;
[0021] Figure 7It is a schematic structural diagram of an embodiment of a visual language model training device according to the present disclosure;
[0022] Figure 8 It is a schematic structural diagram of an embodiment of an image segmentation device according to the present disclosure;
[0023] Figure 9 It is a block diagram of an electronic device for implementing the image segmentation method of the embodiments of the present disclosure. Detailed implementation manners
[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0026] Figure 1 Flow 100 of an embodiment of a visual language model training method according to the present disclosure is shown. The visual language model training method includes the following steps:
[0027] Step 101, encoding a training sample image, a training sample image segmentation mask, and a training sample text segmentation instruction into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group, and a training sample text feature vector group.
[0028] In this embodiment, the execution subject of the visual language model training method can obtain training samples. Encoding the training sample image, the training sample image segmentation mask, and the training sample text segmentation instruction in the training samples into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group, and a training sample text feature vector group.
[0029] Among them, the execution subject of the visual language model training method is usually a server. The server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services) or as a single software or software module. No specific limitation is made here.
[0030] Among them, the training samples can include training sample images, training sample image segmentation masks, and training sample text segmentation instructions. The training sample images can be images containing the target. The training sample image segmentation masks can be used to extract the target in the training sample images. Multiplying the training sample image segmentation mask by the training sample image can obtain the target image. Among them, the pixel values of the target in the target image remain unchanged, and the pixel values outside the target are all 0. The training sample text segmentation instructions can be used to indicate the segmentation of the training sample images.
[0031] Generally, encoding the training sample images and the training sample image segmentation masks from the RGB (Red Green Blue) pixel value space to the feature semantic space can obtain the training sample image feature vector group and the training sample image segmentation mask feature vector group. Among them, the training sample image feature vector group can include the features of the training sample images, and the training sample image segmentation mask feature vector group can include the features of the training sample image segmentation masks, both of which are vector groups of T v ,q] in the form of T v with a length of q. Among them, the encoding formula can be as follows:
[0032] F img = Enc vision (X img );
[0033] F mask = Enc vision (X mask ).
[0034] Among them, F img is the training sample image feature vector group. X img is the training sample image. F mask is the training sample image segmentation mask feature vector group. X mask is the training sample image segmentation mask. Enc vision () is the image encoding operation. T v and q are positive integers.
[0035] Generally, encoding the training sample text segmentation instructions into the feature semantic space to establish the mapping relationship from the word list to the features. Each common Chinese character and word corresponds to a feature vector, so that the training sample text feature vector group can be obtained. Among them, the training sample text feature vector group can include the features of the training sample text segmentation instructions, which is a vector group of T t ,p] in the form of T t with a length of p. Among them, the encoding formula can be as follows:
[0036] F text = Enc text (Xtext )。
[0037] Among them, F text is the training sample text feature vector group. X text is the training sample text segmentation instruction. Enc text () is the text encoding operation. T t and p are positive integers.
[0038] It should be noted that the feature semantic space of the text in this embodiment is aligned with the feature semantic space of the image to establish the correlation between the image and the text in the feature semantic space.
[0039] Step 102: Input the training sample image feature vector group and the training sample text feature vector group into the initialized vision-language model to obtain the training sample multi-modal feature vector group.
[0040] In this embodiment, the above execution subject can input the training sample image feature vector group and the training sample text feature vector group into the initialized vision-language model to obtain the training sample multi-modal feature vector group.
[0041] Among them, the vision-language model (Vision-Language Models, VLM) can be a large-scale neural network model established based on the Transformer structure, including but not limited to LLaMA, Vicuna, etc. The initialized vision-language model can be the model after initializing the parameters of the vision-language model. The training sample multi-modal feature vector group can include both the features of the training sample image and the features of the training sample text segmentation instruction. That is, it contains the features of both the image and the text modalities at the same time.
[0042] Step 103: Calculate the first loss based on the training sample multi-modal feature vector group, the training sample image segmentation mask feature vector group, and the training sample text feature vector group.
[0043] In this embodiment, the above execution subject can calculate the first loss based on the training sample multi-modal feature vector group, the training sample image segmentation mask feature vector group, and the training sample text feature vector group. Among them, the first loss can be used to characterize the difference between the training sample image segmentation mask feature vector group and the training sample text feature vector group, and the training sample multi-modal feature vector group.
[0044] In some embodiments, the mean squared error is used to calculate the loss function between the multi-modal feature vector group of the training samples and the image segmentation mask feature vector group of the training samples, and the first mean squared error loss can be obtained. The cross-entropy is used to calculate the loss function between the multi-modal feature vector group of the training samples and the text feature vector group of the training samples, and the cross-entropy loss can be obtained. The first mean squared error loss and the cross-entropy loss are weighted and summed to obtain the first loss. The calculation formula can be as follows:
[0045] L img =L MSE (F O [i:i+T v ,F mask );
[0046] L text =CrossEntropy(F O [j],F text );
[0047] L = mL img +nL text .
[0048] Among them, L img is the first mean squared error loss. F O is the multi-modal feature vector group of the training samples, including T o feature vectors. i is the image start flag in the multi-modal feature vector group of the training samples. F O [i:i+T v includes T consecutive groups of multi-modal feature vectors of the training samples after the image start flag i. F v is the image segmentation mask feature vector group of the training samples. L mask () is the mean squared error function. L MSE is the cross-entropy loss. j ≤ i or j ≥ i + T text , F v [j] includes the multi-modal feature vectors of the training samples before the image start flag i and the multi-modal feature vectors of the training samples after i + T O . F v is the text feature vector group of the training samples. CrossEntropy() is the cross-entropy loss function. L is the first loss, m is the weight of the first mean squared error loss L img , and n is the weight of the cross-entropy loss L text . dec
[0049] Step 104: Based on the first loss, adjust the parameters of the initialized vision-language model to obtain the trained vision-language model.
[0050] In this embodiment, the above-mentioned execution entity may adjust the parameters of the vision-language model based on the first loss until the first loss is relatively small and the model converges, obtaining a trained vision-language model.
[0051] In some embodiments, in the case where it is necessary to use the vision-language model to segment medical images for a target segmentation task, the training sample images may be medical images, such as magnetic resonance images, computed tomography images, fundus color photographs, etc. The target segmentation task may be a task of segmenting a target part. The target may be, for example, the eyeball, hand, foot, etc. In this way, the trained vision-language model can be applied to the field of medical image segmentation, not only having the general medical image segmentation ability, but also being able to segment medical image modalities or segmentation tasks that the model has not processed, thereby obtaining more accurate segmentation results.
[0052] The embodiments of the present disclosure provide a vision-language model training method. The trained vision-language model has the ability of image segmentation, improving the accuracy of image segmentation. When applied to the field of medical image segmentation, it not only has the general medical image segmentation ability, but also can segment medical image modalities or segmentation tasks that the model has not processed, thereby obtaining more accurate segmentation results.
[0053] Continue to refer to Figure 2 Flow 200 of another embodiment of the vision-language model training method according to the present disclosure is shown. The vision-language model training method includes the following steps:
[0054] Step 201, encoding the training sample images, the training sample image segmentation masks, and the training sample text segmentation instructions into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group, and a training sample text feature vector group.
[0055] Step 202, inputting the training sample image feature vector group and the training sample text feature vector group into the initialized vision-language model to obtain a training sample multi-modal feature vector group.
[0056] Step 203, calculating a first loss based on the training sample multi-modal feature vector group, the training sample image segmentation mask feature vector group, and the training sample text feature vector group.
[0057] Step 204, adjusting the parameters of the initialized vision-language model based on the first loss to obtain a trained vision-language model.
[0058] In this embodiment, the specific operations of steps 201-204 have been described in detail in steps 101-104 of the embodiment shown in Figure 1 and will not be elaborated here.
[0059] Step 205: Input the training sample image segmentation mask feature vector group into the initialized image decoder to obtain a reconstructed image segmentation mask.
[0060] In this embodiment, the execution subject of the visual language model training method may input the training sample image segmentation mask feature vector group into the initialized image decoder to obtain a reconstructed image segmentation mask.
[0061] Among them, the image decoder can decode the image feature vectors output by the visual language model into images or segmentation masks that can be read by humans. The image decoder can be a diffusion generation model such as the Stable Diffusion model or the Control Net model, or a non-random reconstruction model such as VQVAE, as long as it can achieve the decoding ability from features to images. The initialized image decoder can be a model with initialized parameters of the image decoder.
[0062] Step 206: Calculate a second loss based on the reconstructed image segmentation mask and the training sample image segmentation mask.
[0063] In this embodiment, the above execution subject may calculate a second loss based on the reconstructed image segmentation mask and the training sample image segmentation mask. Among them, the second loss can be used to characterize the difference between the reconstructed image segmentation mask and the training sample image segmentation mask.
[0064] In some embodiments, the mean square error is used to calculate the loss function between the reconstructed image segmentation mask and the training sample image segmentation mask to obtain the second loss. The calculation formula can be as follows:
[0065] L dec =L MSE (Dec img (F mask ),X mask )。
[0066] Among them, L dec is the second loss. F mask is the training sample image segmentation mask feature vector group. X mask is the training sample image segmentation mask. Dec img () is the decoding operation. L MSE () is the mean square error function.
[0067] Step 207: Adjust the parameters of the initialized image decoder based on the second loss to obtain a trained image decoder.
[0068] In this embodiment, the above execution subject may adjust the parameters of the initialized image decoder based on the second loss until the second loss is relatively small and the model converges, to obtain a trained image decoder.
[0069] In some embodiments, the models involved in the image segmentation method may include an image encoder, a text encoder, a vision-language model, an image decoder, and a text decoder. Among them, the text decoder does not require additional training and can perform argmax calculations and look up tables based on the encoding vocabulary. Argmax (arguments of the maxima) is used to calculate the subscript index of the maximum value element for a one-dimensional array (vector). For example, for the array [1, 5, 4, 3, 2], the maximum element value is 5, and the corresponding subscript (starting from 0) is 1, so argmax() is 1. The image encoder and the text encoder can be CLIP (Contrastive Language-Image Pre-training) models trained using medical images or publicly available natural images. The CLIP model is a multi-modal model based on contrastive learning. The training data of the CLIP model is text-image pairs: an image and its corresponding text description. Through contrastive learning, the CLIP model can learn the matching relationship between text-image pairs. The image decoder can use the reconstruction task from the training sample image segmentation mask feature vector group output by the image encoder to the training sample image segmentation mask as the training target. That is, the trained image decoder can map the vectors in the feature space back to the image space. When training the vision-language model, there is no need to load the image and text decoders in the memory to save computing resources. For text output, the cross-entropy is used to calculate the loss function between the output vector and the text label. For image output, the mean squared error is used to calculate the loss function between the output vector and the encoded features.
[0070] The embodiments of the present disclosure provide a vision-language model training method. The trained vision-language model has the ability of image segmentation, improving the accuracy of image segmentation. The trained image decoder can map the vectors in the feature space back to the image space.
[0071] Further referring to Figure 3 , which shows a flow 300 of an embodiment of the image segmentation method according to the present disclosure. The image segmentation method includes the following steps:
[0072] Step 301, encoding an image and a text segmentation instruction into the same feature space to obtain an image feature vector group and a text feature vector group.
[0073] In this embodiment, the execution subject of the image segmentation method can encode an image and a text segmentation instruction into the same feature space to obtain an image feature vector group and a text feature vector group.
[0074] Among them, the execution entity of the image segmentation method is usually a server. The server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (such as those used to provide distributed services) or as a single software or software module. No specific limitation is made here.
[0075] It should be noted that the execution entity of the image segmentation method and the execution entity of the visual language model training method can be the same server or different servers. In the case of the same server, after the server trains the visual language model, it can directly use the visual language model for image segmentation. In the case of different servers, after the first server trains the visual language model, it can deploy the visual language model to the second server. Use the visual language model deployed on the second server for image segmentation.
[0076] Among them, the image can be an image containing a target. The text segmentation instruction can be used to indicate the segmentation of the image.
[0077] Generally, by encoding the image from the RGB pixel value space to the feature semantic space, an image feature vector group can be obtained. Among them, the image feature vector group can include the features of the image and is a vector group of T v , q] with a length of q for T v vectors. Among them, the encoding formula can be as follows:
[0078] F img = Enc vision (X img ).
[0079] Among them, F img is the image feature vector group. X img is the image. Enc vision () is the image encoding operation. T v and q are positive integers.
[0080] Generally, by encoding the text segmentation instruction into the feature semantic space, a mapping relationship from the vocabulary to the feature is established. Each common word corresponds to a feature vector, and thus a text feature vector group can be obtained. Among them, the text feature vector group can include the features of the text segmentation instruction and is a vector group of T t , p] with a length of p for T t vectors. Among them, the encoding formula can be as follows:
[0081] F text = Enc text (X text ).
[0082] Among them, F text is a text feature vector group. X text is a text segmentation instruction. Enc text () is a text encoding operation. T t and p are positive integers.
[0083] It should be noted that the feature semantic space of the text in this embodiment is aligned with the feature semantic space of the image to establish the correlation between the image and the text in the feature semantic space.
[0084] Step 302: Input the image feature vector group and the text feature vector group into the vision-language model to obtain a multi-modal feature vector group.
[0085] In this embodiment, the above-mentioned execution subject can input the image feature vector group and the text feature vector group into the vision-language model to obtain a multi-modal feature vector group.
[0086] Among them, the vision-language model can be a large-scale neural network model established based on the transformer structure, including but not limited to LLaMA, Vicuna, etc. The vision-language model is trained with the training sample image feature vector group and the training sample text feature vector group as the input, and the training sample image segmentation mask feature vector group and the training sample text feature vector group as the supervision. The training process can refer to Figure 1 the embodiments shown, which will not be elaborated here.
[0087] It should be noted that in the case of unlabeled samples, the vision-language model can rely on its own knowledge for segmentation and text output. At this time, the multi-modal feature vector group F o = Model(F text , F img ). Among them, among them, F text is the text feature vector group. F img is the image feature vector group.
[0088] Step 303: Decode the multi-modal feature vector group to obtain a text segmentation description and an image segmentation mask.
[0089] In this embodiment, the above-mentioned execution subject can decode the multi-modal feature vector group to obtain a text segmentation description and an image segmentation mask.
[0090] Generally, the multi-modal feature vector group is decoded into text and segmentation masks that can be read by humans. Among them, the decoding formula can be as follows:
[0091] i = argmin(DeC text (F o ) = <imstart>);
[0092] P mask = Dec img (F o [i:i + T v );
[0093]
[0094] where i is the start flag of the image in the multi-modal feature vector group (such as <imstart>)。F O [i] is the i-th multi-modal feature vector. Dec text () is the text decoding operation. P mask is the image segmentation mask. F o [i:i+T v includes the consecutive T v groups of multi-modal feature vectors after the image start flag i. Dec img () is the image decoding operation. j ≤ i or j ≥ i+T v , F O [j] includes the multi-modal feature vectors before the image start flag i and the multi-modal feature vectors after i+T v .
[0095] In some embodiments, by binarizing the image segmentation mask based on a preset threshold, a binarized image segmentation mask can be obtained, thereby simplifying the segmentation mask.
[0096] The embodiments of the present disclosure provide an image segmentation method, which uses a vision-language model for image segmentation, improving the accuracy of image segmentation. When applied to the field of medical image segmentation, it can not only segment general medical images, but also segment medical image modalities or segmentation tasks that have not been processed by the vision-language model, thereby obtaining more accurate segmentation results. It can be applied to an automated annotation platform, enabling annotators to output corresponding segmentation results only through text guidance without strict training, nor reading the image to be annotated and finding the target area of interest.
[0097] Further referring to Figure 4 , which shows the flow 400 of another embodiment of the image segmentation method according to the present disclosure. The image segmentation method includes the following steps:
[0098] Step 401, input an image into an image encoder with a transformation layer added to obtain a group of image feature vectors.
[0099] In this embodiment, the execution subject of the image segmentation method can input an image into an image encoder with a transformation layer added to obtain a group of image feature vectors.
[0100] Generally, the image encoder can encode the image from the RGB pixel value space to the feature semantic space, and a group of image feature vectors can be obtained. Among them, the group of image feature vectors can include the features of the image, which is a group of T v ,q vectors of length q. Among them, the encoding formula can be as follows: v
[0101] F img = EnC vision (X img ).
[0102] Among them, F img is a set of image feature vectors. X img is an image. Enc vision () is an image encoding operation. T v and q are positive integers.
[0103] It should be noted that the feature semantic space of the text encoder in this embodiment is aligned with the feature semantic space of the image encoder to establish the correlation between images and texts in the feature semantic space. For example, adding a transformation layer to the output of the image encoder can make the outputs of the two encoders in the same space.
[0104] Step 402: Input the text segmentation instruction into the text encoder to obtain a set of text feature vectors.
[0105] In this embodiment, the above-mentioned execution entity can input the text segmentation instruction into the text encoder to obtain a set of text feature vectors.
[0106] Generally, the text encoder can encode the text segmentation instruction into the feature semantic space to establish the mapping relationship from the vocabulary to the features. Each common word corresponds to a feature vector, so that a set of text feature vectors can be obtained. Among them, the set of text feature vectors can include the features of the text segmentation instruction, and is a set of T t , p vectors of length p in the form of [T t . Among them, the encoding formula can be as follows:
[0107] F text = Enc text (X text ).
[0108] Among them, F text is a set of text feature vectors. X text is the text segmentation instruction. Enc text () is the text encoding operation. T t and p are positive integers.
[0109] Step 403: Input the set of image feature vectors and the set of text feature vectors into the vision-language model to obtain a set of multi-modal feature vectors.
[0110] In this embodiment, the above-mentioned execution entity can input the set of image feature vectors and the set of text feature vectors into the vision-language model to obtain a set of multi-modal feature vectors.
[0111] Among them, the vision-language model can be a large-scale neural network model based on the Transformer architecture, including but not limited to LLaMA, Vicuna, etc. The vision-language model is trained with the training sample image feature vector group and the training sample text feature vector group as the input, and the training sample image segmentation mask feature vector group and the training sample text feature vector group as the supervision. The training process can refer to Figure 1 the embodiments shown, which will not be elaborated here.
[0112] It should be noted that in the case of unlabeled samples, the vision-language model can rely on its own knowledge for segmentation and text output. At this time, the multi-modal feature vector group F o = Model(F text , F img ). Among them, F text is the text feature vector group. F img is the image feature vector group.
[0113] Step 404: Input the multi-modal feature vector group into the text decoder and the image decoder to obtain the text segmentation description and the image segmentation mask.
[0114] In this embodiment, the above execution subject can input the multi-modal feature vector group into the text decoder and the image decoder to obtain the text segmentation description and the image segmentation mask.
[0115] Generally, the image decoder can decode the image feature vector output by the vision-language model into an image or a segmentation mask that can be read by humans. The image decoder can be a diffusion generation model such as the Stable Diffusion model or the Control Net model, or a non-random reconstruction model such as VQVAE, as long as it can achieve the decoding ability from features to images. The text decoder can decode the text feature vector output by the vision-language model into text that can be read by humans. The text decoder can perform argmax calculation and look up the table based on the encoding vocabulary. Argmax calculates the subscript index of the maximum value element for a one-dimensional array (vector). For example, for the array [1, 5, 4, 3, 2], the maximum element value is 5, and the corresponding subscript (starting from 0) is 1, so argmax() is 1.
[0116] In some embodiments, the decoding steps of the text decoder and the image decoder are as follows:
[0117] First, use the text decoder to decode the multi-modal feature vectors in the multi-modal feature vector group in sequence until the image start flag in the multi-modal feature vector group is decoded.
[0118] Then, an image decoder decodes consecutive preset numbers of groups of multimodal feature vectors after the image start flag to obtain an image segmentation mask.
[0119] Among them, the image decoder is trained with the reconstruction task from the training sample image segmentation mask feature vector group to the training sample image segmentation mask as the training objective. The training process can refer to Figure 2 the embodiments shown, which will not be elaborated here.
[0120] Finally, a text decoder continues to decode the remaining multimodal feature vectors to obtain a text segmentation description.
[0121] Among them, the decoding formula can be as follows:
[0122] =argmin(Dec text (F o [i])= <imstart>);
[0123] P mask = Dec img (F o [i:i + T v );
[0124]
[0125] where i is the start flag of the image in the multi-modal feature vector group (such as <imstart>)。F O [i] is the i-th multi-modal feature vector. Dec text () is the text decoding operation. P mask is the image segmentation mask. F o [i:i+T v includes consecutive T v groups of multi-modal feature vectors after the image start flag i. Dec img () is the image decoding operation. j ≤ i or j ≥ i+T v , F O [j] includes the multi-modal feature vectors before the image start flag i and the multi-modal feature vectors after i+T v .
[0126] Embodiments of the present disclosure provide an image segmentation method that uses a vision-language model for image segmentation, improving the accuracy of image segmentation. When applied to the field of medical image segmentation, it can not only segment general medical images, but also segment medical image modalities or segmentation tasks that have not been processed by the vision-language model, thereby obtaining more accurate segmentation results. It can be applied to an automated annotation platform, enabling annotators to output corresponding segmentation results only through text guidance without strict training or the need to read the image to be annotated and find the target area of interest.
[0127] Further referring to Figure 5 , which shows a flow 500 of another embodiment of the image segmentation method according to the present disclosure. The image segmentation method includes the following steps:
[0128] Step 501, encoding an image, a text segmentation instruction, a labeled sample image, a labeled sample image segmentation mask, and a labeled sample text segmentation instruction into the same feature space to obtain an image feature vector group, a text feature vector group, a labeled sample image feature vector group, a labeled sample image segmentation mask feature vector group, and a labeled sample text feature vector group.
[0129] In this embodiment, the execution subject of the image segmentation method can encode an image, a text segmentation instruction, a labeled sample image, a labeled sample image segmentation mask, and a labeled sample text segmentation instruction into the same feature space to obtain an image feature vector group, a text feature vector group, a labeled sample image feature vector group, a labeled sample image segmentation mask feature vector group, and a labeled sample text feature vector group.
[0130] Among them, the image can be an image containing the target. The text segmentation instruction can be used to indicate the segmentation of the image. The annotation sample can include an annotation sample image, an annotation sample image segmentation mask, and an annotation sample text segmentation instruction. The annotation sample image can be an image containing the target. The annotation sample image segmentation mask can be used to extract the target in the annotation sample image. Multiplying the annotation sample image segmentation mask by the annotation sample image can obtain the target image. Among them, the pixel values of the target in the target image remain unchanged, and the pixel values outside the target are all 0. The annotation text segmentation instruction can be used to indicate the segmentation of the annotation sample image.
[0131] Generally, encoding the image, the annotation sample image, and the annotation sample image segmentation mask from the RGB pixel value space to the feature semantic space can obtain an image feature vector group, an annotation sample image feature vector group, and an annotation sample image segmentation mask feature vector group. Among them, the image feature vector group can include the features of the image, the annotation sample image feature vector group can include the features of the annotation sample image, and the annotation sample image segmentation mask feature vector group can include the features of the annotation sample image segmentation mask, all of which are in the form of [T v ,q] of T v vector groups of length q. Among them, the encoding formula can be as follows:
[0132] F img = Enc vision (X img );
[0133] F mask = Enc vision (X mask ).
[0134] Among them, F img is the image feature vector group or the annotation sample image feature vector group. X img is the image or the annotation sample image. F mask is the annotation sample image segmentation mask feature vector group. X mask is the annotation sample image segmentation mask. Enc vision () is the image encoding operation. T v and q are positive integers.
[0135] Generally, encoding the text segmentation instruction and the annotation sample text segmentation instruction into the feature semantic space to establish a mapping relationship from the word list to the feature. Each common word corresponds to a feature vector, so that a text feature vector group and an annotation sample text feature vector group can be obtained. Among them, the text feature vector group can be the features including the text segmentation instruction, and the annotation sample text feature vector group can be the features including the annotation sample text segmentation instruction, all of which are in the form of [T t ,p] of T t A vector group of length p. Among them, the encoding formula can be as follows:
[0136] F text = Enc text (X text ).
[0137] Among them, F text is the text feature vector group or the labeled text feature vector group. X text is the text segmentation instruction or the labeled sample text segmentation instruction. Enc text () is the text encoding operation. T t and p are positive integers.
[0138] In some embodiments, the image encoder can encode the image from the RGB pixel value space to the feature semantic space, and an image feature vector group can be obtained. The text encoder can encode the text segmentation instruction into the feature semantic space and establish a mapping relationship from the vocabulary to the feature. For example, the labeled sample image and the labeled sample image segmentation mask are respectively input into the image encoder with a transformation layer added, and a labeled sample image feature vector group and a labeled sample image segmentation mask feature vector group are obtained; the labeled sample text segmentation instruction is input into the text encoder to obtain a labeled sample text feature vector group.
[0139] It should be noted that the feature semantic space of the text encoder in this embodiment is aligned with the feature semantic space of the image encoder to establish the correlation between the image and the text in the feature semantic space. For example, adding a transformation layer to the output of the image encoder can make the outputs of the two encoders in the same space.
[0140] Step 502, input the image feature vector group, the text feature vector group, the labeled sample image feature vector group, the labeled sample image segmentation mask feature vector group, and the labeled sample text feature vector group into the vision-language model for context learning to obtain a multi-modal feature vector group.
[0141] In this embodiment, the above-mentioned execution subject can input the image feature vector group, the text feature vector group, the labeled sample image feature vector group, the labeled sample image segmentation mask feature vector group, and the labeled sample text feature vector group into the vision-language model for context learning to obtain a multi-modal feature vector group.
[0142] Among them, the vision-language model can be a large-scale neural network model established based on the Transformer structure, including but not limited to LLaMA, Vicuna, etc. The vision-language model is trained with the training sample image feature vector group and the training sample text feature vector group as the input and the training sample image segmentation mask feature vector group and the training sample text feature vector group as the supervision. The training process can refer to Figure 1 The embodiments shown are not described in detail here.
[0143] It should be noted that in the case of labeled samples, the labeled sample image feature vector group, the labeled sample image segmentation mask feature vector group, and the labeled sample text feature vector group can be input into the vision-language model as prompts for context learning, thereby enhancing the model's segmentation performance. Learning of labeled samples can be achieved without adjusting the model parameters, so it can be quickly applied to new segmentation tasks or unseen imaging modalities. At this time, the multi-modal feature vector group output by the vision-language model where F text includes the text feature vector group and all labeled sample text feature vector groups. is the first labeled sample image feature vector group. is the first labeled sample image segmentation mask feature vector group. is the second labeled sample image feature vector group. is the second labeled sample image segmentation mask feature vector group. is the image feature vector group.
[0144] Step 503: Decode the multi-modal feature vector group to obtain a text segmentation description and an image segmentation mask.
[0145] In this embodiment, the above execution entity can decode the multi-modal feature vector group to obtain a text segmentation description and an image segmentation mask.
[0146] Generally, the multi-modal feature vector group is decoded into text and a segmentation mask that can be read by humans. Among them, the decoding formula can be as follows:
[0147] i = argmin(Dec text (F o [i]) = 〈imstart〉);
[0148] P mask = Dec img (F o [i:i + T v );
[0149]
[0150] where i is the image start flag in the multi-modal feature vector group (such as <imstart>)。F O [i] is the i-th multimodal feature vector. Dec text () is the text decoding operation. P mask is the image segmentation mask. F o [i:i+T v includes consecutive T v groups of multimodal feature vectors after the image start flag i. Dec img () is the image decoding operation. j ≤ i or j ≥ i + T v , F O [j] includes the multimodal feature vectors before the image start flag i and the multimodal feature vectors after i + T v .
[0151] In some embodiments, the image and the annotated sample image are medical images for a target segmentation task, such as nuclear magnetic resonance images, computed tomography images, fundus color photographs, etc. The target segmentation task may be a task of segmenting a target part. The target may be, for example, an eyeball, a hand, a foot, etc. In this way, the trained vision-language model can be applied to the field of medical image segmentation, not only having the general medical image segmentation ability, but also being able to segment medical image modalities or segmentation tasks that the model has not processed before, thereby obtaining more accurate segmentation results.
[0152] The embodiments of the present disclosure provide an image segmentation method based on context learning. Even for medical image modalities or segmentation tasks that the model has not learned, relatively ideal segmentation performance can be achieved through very few annotated samples. Based on the context learning method, the model can learn the segmentation tasks to be processed through a small number of annotated samples without adjusting parameters. Therefore, after the model is trained, for medical image modalities or segmentation tasks that have not been learned, relatively ideal segmentation performance can also be achieved through very few annotated samples.
[0153] The embodiments of the present disclosure provide an image segmentation method based on context learning, which mainly includes the following two advantages:
[0154] First, solve the segmentation tasks outside the training data distribution in a way that does not require training. As long as the model for out-of-distribution task segmentation based on context information. When the model is trained, when facing a new task, only a small amount of example data needs to be given, and the model can reason about the new task.
[0155] Second, text-driven segmentation of the target region without prior knowledge. Guide the segmentation of the target region through text. This way of segmentation solves the problem of insufficient prior knowledge and can segment the target region without giving any hint of geometric information.
[0156] For ease of understanding, Figure 6 A scenario diagram is shown that can implement the image segmentation method of the embodiments of the present disclosure.
[0157] As Figure 6 shown, the model involved in the image segmentation method may include an image encoder 601, a text encoder 602, a vision-language model 603, an image decoder 604, and a text decoder 605.
[0158] After the model training is completed, the user inputs a nuclear magnetic resonance image and a text segmentation instruction "This is a modal imaging. Please segment the instances in the image". At this time, there are already a small number of labeled samples. The labeled samples include fundus color photos, fundus color photo segmentation masks, and a text segmentation instruction "Please segment the instances according to the context prompt". The labeled samples can be used as input to the model as a prompt for context learning to enhance the model's segmentation performance. Specifically, the nuclear magnetic resonance image, the fundus color photo, and the fundus color photo segmentation mask are input into the image encoder 601 for encoding. The text segmentation instruction "This is a modal imaging. Please segment the instances in the image" and the text segmentation instruction "Please segment the instances according to the context prompt" are input into the text encoder 602 for encoding. The encoding results of the two encoders are input into the vision-language model 603. The output result of the vision-language model 603 is input into the image decoder 604 and the text decoder 605 for decoding to obtain the text segmentation description "This is the instance segmentation result" and the nuclear magnetic resonance image segmentation mask.
[0159] The image segmentation method based on context learning provided by the embodiments of the present disclosure can be applied to the field of medical image segmentation. It not only has the ability to segment general medical images, but also for medical image modalities or segmentation tasks that the model has not seen before. Through a small number of data pairs composed of labeled image-masks, the model is further prompted with segmentation requirements, and then more accurate segmentation results can be obtained.
[0160] For example, only nuclear magnetic resonance images and computed tomography images are used during model training, and fundus color photos are not used. However, during inference, it is required to segment the optic disc structure of the fundus color photo. When using the image segmentation method based on context learning provided by the embodiments of the present disclosure for segmentation, as the number of samples added for context learning increases, the segmentation performance and accuracy of the model are also gradually improved.
[0161] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, an embodiment of a vision-language model training device is provided by the present disclosure. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0162] As Figure 7 As shown in the figure, the visual language training device 700 of this embodiment may include: an encoding module 701, a first input module 702, a first calculation module 703, and a first adjustment module 704. Among them, the encoding module 701 is configured to encode the training sample image, the training sample image segmentation mask, and the training sample text segmentation instruction into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group, and a training sample text feature vector group; the first input module 702 is configured to input the training sample image feature vector group and the training sample text feature vector group into the initialized visual language model to obtain a training sample multi-modal feature vector group; the first calculation module 703 is configured to calculate a first loss based on the training sample multi-modal feature vector group, the training sample image segmentation mask feature vector group, and the training sample text feature vector group; the first adjustment module 704 is configured to adjust the parameters of the initialized visual language model based on the first loss to obtain a trained visual language model.
[0163] In this embodiment, in the visual language model training device 700: the specific processing of the encoding module 701, the first input module 702, the first calculation module 703, and the first adjustment module 704 and the technical effects brought by them can be respectively referred to Figure 1 the relevant descriptions of steps 101-104 in the corresponding embodiment, which will not be elaborated here.
[0164] In some optional implementation manners of this embodiment, the first calculation module 703 is further configured to: use the mean square error to calculate the loss function between the training sample multi-modal feature vector group and the training sample image segmentation mask feature vector group to obtain a first mean square error loss; use the cross entropy to calculate the loss function between the training sample multi-modal feature vector group and the training sample text feature vector group to obtain a cross entropy loss; perform a weighted sum of the first mean square error loss and the cross entropy loss to obtain a first loss.
[0165] In some optional implementation manners of this embodiment, the visual language model training device 700 further includes: a second input module, configured to input the training sample image segmentation mask feature vector group into the initialized image decoder to obtain a reconstructed image segmentation mask; a second calculation module, configured to calculate a second loss based on the reconstructed image segmentation mask and the training sample image segmentation mask; a second adjustment module, configured to adjust the parameters of the initialized image decoder based on the second loss to obtain a trained image decoder.
[0166] In some optional implementation manners of this embodiment, the second calculation module is further configured to: use the mean square error to calculate the loss function between the reconstructed image segmentation mask and the training sample image segmentation mask to obtain a second loss.
[0167] In some alternative implementation manners of this embodiment, the training sample image is a medical image.
[0168] With further reference to Figure 8 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image segmentation device. This device embodiment corresponds to Figure 3 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0169] As shown in Figure 8 , the image segmentation device 800 in this embodiment may include: a first encoding module 801, an input module 802, and a decoding module 803. Among them, the first encoding module 801 is configured to encode an image and a text segmentation instruction into the same feature space to obtain an image feature vector group and a text feature vector group; the input module 802 is configured to input the image feature vector group and the text feature vector group into a vision-language model to obtain a multi-modal feature vector group, where the vision-language model is trained with a training sample image feature vector group and a training sample text feature vector group as inputs and a training sample image segmentation mask feature vector group and a training sample text feature vector group as supervision; the decoding module 803 is configured to decode the multi-modal feature vector group to obtain a text segmentation description and an image segmentation mask.
[0170] In this embodiment, in the image segmentation device 800: the specific processing of the first encoding module 801, the input module 802, and the decoding module 803 and the technical effects brought thereby can be respectively referred to Figure 3 the relevant descriptions of steps 301-303 in the corresponding embodiment, which will not be elaborated here.
[0171] In some alternative implementation manners of this embodiment, the first encoding module 801 is further configured to: input the image into an image encoder with a transformation layer added to obtain an image feature vector group; input the text segmentation instruction into a text encoder to obtain a text feature vector group.
[0172] In some alternative implementation manners of this embodiment, the decoding module 803 is further configured to: input the multi-modal feature vector group into a text decoder and an image decoder to obtain a text segmentation description and an image segmentation mask.
[0173] In some alternative implementation manners of this embodiment, the decoding module 803 is further configured to: sequentially decode the multimodal feature vectors in the multimodal feature vector group by using a text decoder until an image start flag in the multimodal feature vector group is decoded; decode a continuous preset number of groups of multimodal feature vectors after the image start flag by using an image decoder to obtain an image segmentation mask, where the image decoder is trained with the reconstruction task from a training sample image segmentation mask feature vector group to a training sample image segmentation mask; and continue to decode the remaining multimodal feature vectors by using the text decoder to obtain a text segmentation description.
[0174] In some alternative implementation manners of this embodiment, the image segmentation device 800 further includes: a processing module configured to perform binarization processing on the image segmentation mask based on a preset threshold to obtain a binarized image segmentation mask.
[0175] In some alternative implementation manners of this embodiment, the image segmentation device 800 further includes: a second encoding module configured to encode an annotated sample image, an annotated sample image segmentation mask, and an annotated sample text segmentation instruction into the same feature space to obtain an annotated sample image feature vector group, an annotated sample image segmentation mask feature vector group, and an annotated sample text feature vector group; and the input module 802 is further configured to: input the image feature vector group, the text feature vector group, the annotated sample image feature vector group, the annotated sample image segmentation mask feature vector group, and the annotated sample text feature vector group into a vision-language model for context learning to obtain a multimodal feature vector group.
[0176] In some alternative implementation manners of this embodiment, the image and the annotated sample image are medical images of a target segmentation task.
[0177] In some alternative implementation manners of this embodiment, the second encoding module is further configured to: respectively input the annotated sample image and the annotated sample image segmentation mask into an image encoder with a transformation layer added to obtain an annotated sample image feature vector group and an annotated sample image segmentation mask feature vector group; and input the annotated sample text segmentation instruction into a text encoder to obtain an annotated sample text feature vector group.
[0178] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0179] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0180] Figure 9 FIG. 0 shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0181] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0182] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0183] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the image segmentation method. For example, in some embodiments, the image segmentation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image segmentation method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the image segmentation method by any other suitable means (e.g., by means of firmware).
[0184] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0185] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0186] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0187] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0188] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0189] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0190] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution provided by this disclosure can be achieved, and no limitation is imposed herein.
[0191] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.< / imstart> < / imstart> < / imstart> < / imstart> < / imstart>
Claims
1. A visual language model training method, comprising: Encoding the training sample image, the training sample image segmentation mask and the training sample text segmentation instruction into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group and a training sample text feature vector group; Inputting the training sample image feature vector group and the training sample text feature vector group into an initialized visual language model to obtain a training sample multimodal feature vector group; Calculating a first loss based on the training sample multimodal feature vector group, the training sample image segmentation mask feature vector group, and the training sample text feature vector group; The parameters of the initialized visual language model are adjusted based on the first loss to obtain a trained visual language model.
2. The method according to claim 1, wherein: The calculating a first loss based on the training sample multimodal feature vector group, the training sample image segmentation mask feature vector group and the training sample text feature vector group comprises: Calculating the loss function between the training sample multimodal feature vector group and the training sample image segmentation mask feature vector group using a mean square error to obtain a first mean square error loss; Calculating the loss function between the training sample multimodal feature vector group and the training sample text feature vector group using cross entropy to obtain a cross entropy loss; The first loss is obtained by weighted summing the first mean square error loss and the cross entropy loss.
3. The method according to claim 1, wherein: The method further comprises: Inputting the training sample image segmentation mask feature vector group into the initialized image decoder to obtain a reconstructed image segmentation mask; Calculating a second loss based on the reconstructed image segmentation mask and the training sample image segmentation mask; The parameters of the initialized image decoder are adjusted based on the second loss to obtain a trained image decoder.
4. The method according to claim 3, wherein: The calculating the second loss based on the reconstructed image segmentation mask and the training sample image segmentation mask comprises: The loss function between the reconstructed image segmentation mask and the training sample image segmentation mask is calculated using a mean square error to obtain the second loss.
5. The method according to any one of claims 1 to 4, wherein: The training sample images are medical images.
6. An image segmentation method, comprising: Encode the image and text segmentation instructions into the same feature space to obtain an image feature vector group and a text feature vector group; Inputting the image feature vector group and the text feature vector group into a visual language model to obtain a multimodal feature vector group, wherein the visual language model is trained with the training sample image feature vector group and the training sample text feature vector group as input and the training sample image segmentation mask feature vector group and the training sample text feature vector group as supervision; The multimodal feature vector group is decoded to obtain a text segmentation description and an image segmentation mask.
7. The method according to claim 6, wherein: The step of encoding the image and text segmentation instructions into the same feature space to obtain an image feature vector group and a text feature vector group includes: Inputting the image into an image encoder with a transform layer added thereto to obtain the image feature vector group; The text segmentation instruction is input into a text encoder to obtain the text feature vector group.
8. The method according to claim 6, wherein: Decoding the multimodal feature vector group to obtain a text segmentation description and an image segmentation mask includes: The multimodal feature vector group is input into a text decoder and an image decoder to obtain the text segmentation description and the image segmentation mask.
9. The method according to claim 8, wherein: The step of inputting the multimodal feature vector group into a text decoder and an image decoder to obtain the text segmentation description and the image segmentation mask comprises: Decoding the multimodal feature vectors in the multimodal feature vector group in sequence by using the text decoder until an image start mark in the multimodal feature vector group is decoded; Using the image decoder to decode a preset number of groups of multimodal feature vectors after the start mark of the image to obtain the image segmentation mask, wherein the image decoder is trained with the reconstruction task of training sample image segmentation mask feature vector groups to training sample image segmentation masks as the training target; The text decoder is used to continue decoding the remaining multimodal feature vectors to obtain the text segmentation description.
10. The method according to claim 9, wherein: The method further comprises: The image segmentation mask is binarized based on a preset threshold to obtain a binary image segmentation mask.
11. The method according to claim 7, wherein: The method further comprises: Encoding the labeled sample image, the labeled sample image segmentation mask, and the labeled sample text segmentation instruction into the same feature space to obtain a labeled sample image feature vector group, a labeled sample image segmentation mask feature vector group, and a labeled sample text feature vector group; and The step of inputting the image feature vector group and the text feature vector group into a visual language model to obtain a multimodal feature vector group includes: The image feature vector group, the text feature vector group, the annotated sample image feature vector group, the annotated sample image segmentation mask feature vector group and the annotated sample text feature vector group are input into the visual language model for context learning to obtain the multimodal feature vector group.
12. The method according to claim 11, wherein: The image and the labeled sample image are medical images for a target segmentation task.
13. The method according to claim 11, wherein: The method of encoding the labeled sample image, the labeled sample image segmentation mask and the labeled sample text segmentation instruction into the same feature space to obtain a labeled sample image feature vector group, a labeled sample image segmentation mask feature vector group and a labeled sample text feature vector group includes: Inputting the labeled sample image and the labeled sample image segmentation mask into the image encoder with the added transformation layer respectively, to obtain the labeled sample image feature vector group and the labeled sample image segmentation mask feature vector group; The labeled sample text segmentation instruction is input into the text encoder to obtain the labeled sample text feature vector group.
14. A visual language model training device, comprising: An encoding module is configured to encode the training sample image, the training sample image segmentation mask and the training sample text segmentation instruction into the same feature space to obtain a training sample image feature vector group, a training sample image segmentation mask feature vector group and a training sample text feature vector group; A first input module is configured to input the training sample image feature vector group and the training sample text feature vector group into an initialized visual language model to obtain a training sample multimodal feature vector group; A first calculation module is configured to calculate a first loss based on the training sample multimodal feature vector group, the training sample image segmentation mask feature vector group and the training sample text feature vector group; The first adjustment module is configured to adjust the parameters of the initialized visual language model based on the first loss to obtain a trained visual language model.
15. The device according to claim 14, wherein: The first computing module is further configured to: Calculating the loss function between the training sample multimodal feature vector group and the training sample image segmentation mask feature vector group using a mean square error to obtain a first mean square error loss; Calculating the loss function between the training sample multimodal feature vector group and the training sample text feature vector group using cross entropy to obtain a cross entropy loss; The first loss is obtained by weighted summing the first mean square error loss and the cross entropy loss.
16. The device according to claim 14, wherein: The device also includes: A second input module is configured to input the training sample image segmentation mask feature vector group into an initialized image decoder to obtain a reconstructed image segmentation mask; A second calculation module is configured to calculate a second loss based on the reconstructed image segmentation mask and the training sample image segmentation mask; The second adjustment module is configured to adjust the parameters of the initialized image decoder based on the second loss to obtain a trained image decoder.
17. The device according to claim 16, wherein: The second computing module is further configured to: The loss function between the reconstructed image segmentation mask and the training sample image segmentation mask is calculated using a mean square error to obtain the second loss.
18. The device according to any one of claims 14 to 17, wherein: The training sample images are medical images.
19. An image segmentation device, comprising: A first encoding module is configured to encode the image and text segmentation instructions into the same feature space to obtain an image feature vector group and a text feature vector group; An input module is configured to input the image feature vector group and the text feature vector group into a visual language model to obtain a multimodal feature vector group, wherein the visual language model is trained with the training sample image feature vector group and the training sample text feature vector group as input and the training sample image segmentation mask feature vector group and the training sample text feature vector group as supervision; The decoding module is configured to decode the multimodal feature vector group to obtain a text segmentation description and an image segmentation mask.
20. The device according to claim 19, wherein The first encoding module is further configured to: Inputting the image into an image encoder with a transform layer added thereto to obtain the image feature vector group; The text segmentation instruction is input into a text encoder to obtain the text feature vector group.
21. The device according to claim 19, wherein The decoding module is further configured to: The multimodal feature vector group is input into a text decoder and an image decoder to obtain the text segmentation description and the image segmentation mask.
22. The device according to claim 21, wherein The decoding module is further configured to: Decoding the multimodal feature vectors in the multimodal feature vector group in sequence by using the text decoder until an image start mark in the multimodal feature vector group is decoded; Using the image decoder to decode a preset number of groups of multimodal feature vectors after the start mark of the image to obtain the image segmentation mask, wherein the image decoder is trained with the reconstruction task of training sample image segmentation mask feature vector groups to training sample image segmentation masks as the training target; The text decoder is used to continue decoding the remaining multimodal feature vectors to obtain the text segmentation description.
23. The device according to claim 22, wherein: The device also includes: The processing module is configured to perform binarization processing on the image segmentation mask based on a preset threshold value to obtain a binary image segmentation mask.
24. The device according to claim 20, wherein: The device also includes: A second encoding module is configured to encode the labeled sample image, the labeled sample image segmentation mask and the labeled sample text segmentation instruction into the same feature space to obtain a labeled sample image feature vector group, a labeled sample image segmentation mask feature vector group and a labeled sample text feature vector group; and The input module is further configured to: The image feature vector group, the text feature vector group, the annotated sample image feature vector group, the annotated sample image segmentation mask feature vector group and the annotated sample text feature vector group are input into the visual language model for context learning to obtain the multimodal feature vector group.
25. The device according to claim 24, wherein: The image and the labeled sample image are medical images for a target segmentation task.
26. The device according to claim 24, wherein: The second encoding module is further configured to: Inputting the labeled sample image and the labeled sample image segmentation mask into the image encoder with the added transformation layer respectively, to obtain the labeled sample image feature vector group and the labeled sample image segmentation mask feature vector group; The labeled sample text segmentation instruction is input into the text encoder to obtain the labeled sample text feature vector group.
27. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5 or 6-13.
28. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method of any one of claims 1-5 or 6-13.
29. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1-5 or 6-13.
Citation Information
Patent Citations
Semantic segmentation model training method and semantic segmentation method and device
CN114648638A
Text visual question and answer method and device, computer equipment and storage medium
CN117033609A