Image segmentation method, model, model training method and image segmentation system

By combining text encoder and image encoder in the image segmentation model, and using attention module and splicing module to perform multiple feature fusion, the problem of low target and unseen category recognition accuracy in complex scenarios is solved, and efficient open vocabulary semantic segmentation is achieved.

CN119992550AActive Publication Date: 2025-05-13SHENZHEN YISHIHUOLALA TECH CO LTD

Patent Information

Application Number
CN202510465181.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art has low recognition accuracy of targets and unseen categories in complex scenarios, making it difficult to fully utilize the visual-language understanding ability of the diffusion model, affecting the segmentation ability in zero-sample environment and in complex scenarios.

Method used

The text embedding features and image potential features are obtained through the text encoder and image encoder of the image segmentation model, and the attention module and the splicing module are used for multiple fusions, visual features are extracted and text embedding features are aligned and fused, and the noise-free mask latent features are generated, and the predicted target mask is generated through the latent space decoder.

Benefits of technology

It improves the zero-sample segmentation ability, enhances the ability to capture target boundaries and semantic information, and improves the segmentation accuracy and generalization ability of unknown categories in complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992550A_ABST
    Figure CN119992550A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an image segmentation method, a model, a model training method and an image segmentation system. The method comprises the following steps: acquiring a to-be-segmented image containing a target object and text information corresponding to the to-be-segmented image; encoding the text information through a text encoder of the image segmentation model to generate text embedded features, and encoding the image through an image encoder of the image segmentation model to generate image potential features; and processing the image potential features through an attention module and a splicing module of the image segmentation model based on the text embedding features according to a preset fusion frequency until a noisy mask potential feature generated during the last fusion is obtained, and decoding through a submerged space decoder of the image segmentation model after denoising to obtain the final fusion. Generating a prediction target mask used for segmenting the target object; and segmenting the target object in the image based on the predicted target mask. According to the invention, the accuracy of image segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image segmentation method, a model, a model training method and an image segmentation system. Background Art

[0002] Open-Vocabulary Semantic Segmentation (OVSS) is an important computer vision task that aims to semantically annotate each pixel in an image based on a text description. Unlike traditional semantic segmentation methods, OVSS needs to have the ability to recognize unseen categories to adapt to diverse targets in open environments. Although multimodal models based on text-image matching perform well in single-target classification, they are difficult to accurately understand the relationship between objects in an image in complex scenes, which limits their application in pixel-level segmentation tasks. In addition, although text-to-image diffusion models have shown strong capabilities in generating high-quality images in recent years, most OVSS studies only use these diffusion models as fixed feature extractors and fail to fully utilize the model's rich visual-language understanding capabilities, resulting in limited performance of diffusion models in pixel-level semantic alignment and open category recognition, affecting their segmentation capabilities in zero-shot environments and complex scenes. Therefore, it is necessary to provide an image segmentation method, model, model training method, and image segmentation system. Summary of the invention

[0003] In view of the shortcomings of the prior art described above, an object of the present invention is to provide an image segmentation method, a model, a model training method and an image segmentation system, which improve the problem of low recognition accuracy of targets and unseen categories in complex scenes when performing open vocabulary semantic segmentation tasks in the prior art.

[0004] To achieve the above-mentioned purpose and other related purposes, the present invention provides an image segmentation method, comprising: obtaining an image to be segmented containing a target object and text information corresponding thereto; wherein the text information is used to characterize the category of the target object in the image; encoding the text information through a text encoder of an image segmentation model to generate text embedding features, and encoding the image through an image encoder of the image segmentation model to generate image latent features; according to a preset number of fusions, based on the text embedding features, processing the image latent features through an attention module and a splicing module of the image segmentation model within a preset number of fusions, until a noisy mask generated at the last fusion is obtained. The method comprises the following steps: first, the image latent features are fused with the noisy mask latent features through the splicing module of the image segmentation model, and the visual features are extracted from the fused features obtained after fusion through the attention module, and the visual features are aligned and fused with the text embedding features to update the noisy mask latent features at the next fusion; second, the noisy mask latent features at the first fusion are random noises generated by the noise generator of the image segmentation model; and the target object in the image is segmented based on the predicted target mask.

[0005] In one embodiment of the present invention, there are K predicted target masks generated by the latent space decoder, where K is a positive integer greater than 1, and the processed mask latent features are decoded to generate a predicted target mask for segmenting the target object, including: decoding the processed mask latent features to obtain K initial predicted target masks; calculating the mean of the K initial predicted target masks as the final predicted target mask.

[0006] In one embodiment of the present invention, the image latent features and the noisy mask latent features are fused through the splicing module of the image segmentation model, and the visual features are extracted from the fused features obtained after the fusion through the attention module, and the fused features are aligned and fused with the text embedding features to update the noisy mask latent features at the next fusion, including: fusing the image latent features and the noisy mask latent features through the splicing module of the image segmentation model to obtain fused features; extracting visual features of corresponding scales from the fused features based on downsampling networks of different scales through the diffusion coding unit of the attention module; wherein the encoder module includes a cascade of multiple downsampling networks with decreasing scales in sequence; and using the attention unit of the attention module to convert the minimum scale into a content-based attention mechanism. The visual features of the attention module are cross-modally fused and aligned with the text embedding features to obtain fused aligned features; the intermediate features and the visual features of the corresponding scale are fused in each upsampling unit through the diffusion decoding unit of the attention module to obtain the intermediate features input to the next upsampling unit, and the prediction noise is generated based on the intermediate features obtained from the last upsampling unit; wherein the initial intermediate features are the fused aligned features, and the decoder module includes a plurality of cascaded upsampling units with increasing scales, and the scale of the upsampling unit corresponds to the scale of the downsampling network; the noisy masked latent features at the time of current fusion are denoised based on the predicted noise through the denoising unit of the attention module to obtain the denoised masked latent features, and the denoised masked latent features are used as the noisy masked latent features for the next fusion.

[0007] In one embodiment of the present invention, the fusing of the image latent features and the noisy mask latent features includes: superimposing the image latent features and the noisy mask latent features to generate a joint feature; and splicing the joint feature and the noisy mask latent features to generate a fused feature.

[0008] In one embodiment of the present invention, the fusing of the image latent features and the noisy mask latent features includes: performing splicing processing on the image latent features and the noisy mask latent features to generate a fused feature.

[0009] In an embodiment of the present invention, the random noise is Gaussian noise.

[0010] In one embodiment of the present invention, the attention unit of the attention module performs cross-modal feature fusion and alignment on the minimum-scale visual features and the text embedding features based on the content attention mechanism to obtain fused and aligned features, including: generating a query matrix, a key matrix and a first value matrix based on the minimum-scale visual features, and correspondingly generating an intermediate matrix and a second value matrix based on the text embedding features; calculating the self-attention weights of the visual features based on the query matrix and the key matrix according to the scaled dot product attention mechanism; calculating the interactive attention weights of the text features to the visual features based on the key matrix and the intermediate matrix according to the scaled dot product attention mechanism; and performing weighted summation on the first value matrix and the second value matrix according to the self-attention weight and the interactive attention weight to obtain fused and aligned features.

[0011] In one embodiment of the present invention, an image segmentation model is also provided, which includes: an image encoder, a noise generator, a text encoder, a splicing module, an attention module and a latent space decoder, wherein: the output end of the image encoder is connected to the first input end of the splicing module; the output end of the noise generator is connected to the second input end of the splicing module; the output end of the text encoder is connected to the first input end of the attention module; the output end of the splicing module is connected to the second input end of the attention module; wherein the attention module is used to fuse the image potential features obtained after the image encoder encodes the image and the random noise generated by the noise generator to extract visual features, and align and fuse the visual features and the text embedding features obtained by the text encoder encoding the text information, and obtain denoising mask potential features after denoising; the output end of the attention module is connected to the input end of the latent space decoder; wherein the latent space decoder is used to decode the denoising mask potential features to generate a predicted target mask for segmenting the target object in the image.

[0012] In one embodiment of the present invention, a training method for an image segmentation model is also provided, the training method comprising: obtaining an image containing a target object and text information corresponding thereto, and a target mask of the target object in the image; wherein the text information is used to characterize the category of the target object in the image; inputting the text information into a text encoder of a semantic segmentation model for encoding to generate text embedding features, inputting the image into an image encoder of the semantic segmentation model for encoding to generate image latent features, and inputting the target mask into a mask encoder of the semantic segmentation model for encoding to generate mask latent features; according to a preset number of fusions, The attention module of the semantic segmentation model is fine-tuned and trained at each fusion to obtain the trained attention module of the semantic segmentation model; wherein, at each fusion: the masked latent features are noised and fused with the image latent features through the noise addition module and the splicing module of the semantic segmentation model, visual features are extracted from the fused features obtained after fusion through the attention module of the semantic segmentation model, the visual features and the text embedding features are aligned and fused based on the attention mechanism to generate noisy masked latent features, and the attention module of the semantic segmentation model is trained through the noisy masked latent features.

[0013] In one embodiment of the present invention, an image segmentation system is further provided, the image segmentation system comprising: an image encoder, a mask encoder, a noise generator, a text encoder, a noise summation module, a splicing module and an attention module, wherein: the output end of the mask encoder is connected to the first input end of the noise summation module; wherein the mask encoder is used to encode the target mask of the target object in the image into a mask potential feature; the output end of the noise generator is connected to the second input end of the noise summation module; wherein the noise generator is used to generate random noise; the output end of the image encoder is connected to the first input end of the splicing module; Wherein, the image encoder is used to encode the image to generate image latent features; the output end of the noise addition module is connected to the second input end of the splicing module; the output end of the text encoder is connected to the first input end of the attention module; wherein, the text encoder is used to encode text information to generate text embedding features; the output end of the splicing module is connected to the second input end of the attention module; wherein, the attention module is used to extract visual features from the fused features generated by the splicing module, and align and fuse the visual features and the text embedding features through the attention mechanism to generate noisy masked latent features.

[0014] As described above, an image segmentation method, model, model training method and image segmentation system of the present invention have the following beneficial effects: the text information of the target object is input into the text encoder, the text information is encoded to generate text embedding features, so that the model can understand the image content and adapt to unseen categories to improve the zero-sample segmentation ability. In addition, the image is encoded into image potential features by the image encoder. After encoding, according to the preset number of fusions, the image potential features and the noisy mask potential features are fused, and the visual information is extracted, and it is cross-modally aligned with the text embedding features, and the mask features are continuously optimized to enhance the model's ability to capture target boundaries and semantic information. At the time of initial fusion, the noisy mask potential features are initialized by random noise, and after a multi-step iterative diffusion denoising process, clear target mask features are gradually extracted, and the latent space decoder is decoded to generate a high-quality predicted target mask, and the original image is accurately segmented accordingly. Compared with traditional segmentation methods, the present invention can effectively remove irrelevant interference in a complex background environment, and at the same time, combined with cross-modal information fusion, improve the generalization ability and segmentation accuracy of unseen categories, and achieve more efficient open vocabulary semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A schematic diagram of a flow chart of an image segmentation method provided by an embodiment of the present invention; Figure 2 A schematic diagram of the overall architecture of an image segmentation model provided by an embodiment of the present invention; Figure 3 A schematic flow chart of a method for training an image segmentation model provided by an embodiment of the present invention; Figure 4 Shown is a schematic diagram of the overall architecture of an image segmentation system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0017] It should be noted that the illustrations provided in the following embodiments are only used to illustrate the basic concept of the present invention in a schematic manner, and thus the illustrations only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0018] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0019] Open vocabulary semantic segmentation (OVSS) is a task that aims to semantically annotate each pixel in an image based on natural language descriptions. Its goal is to break through the reliance of traditional semantic segmentation methods on a fixed set of categories and achieve effective recognition of unseen categories. Currently, methods based on large-scale pre-trained image and text models such as CLIP show good generalization capabilities in open category recognition, but there are still problems of misjudgment and insufficient understanding when dealing with spatial relationships between objects in complex scenes. In addition, the text-to-image diffusion model has gradually become the focus of multimodal research due to its good image generation quality, strong semantic guidance ability, and good cross-modal alignment characteristics. However, existing methods usually use the diffusion model as a static feature extractor, and do not effectively tune its internal visual language modeling module (such as U-Net), which cannot fully release the potential of the diffusion model. Therefore, the current OVSS field still faces the problem of insufficient generalization ability in zero-sample scenarios.

[0020] In view of the above situation, the present invention provides an image segmentation method, inputs the text information of the target object into a text encoder, encodes the text information to generate text embedding features, so that the model can understand the image content and adapt to unseen categories to improve the zero-sample segmentation ability. In addition, the image is encoded into image potential features by the image encoder. After encoding, according to the preset number of fusions, the image potential features and the noisy mask potential features are fused, and the visual information is extracted, and it is cross-modally aligned with the text embedding features, and the mask features are continuously optimized to enhance the model's ability to capture target boundaries and semantic information. At the time of initial fusion, the noisy mask potential features are initialized by random noise, and after a multi-step iterative diffusion denoising process, clear target mask features are gradually extracted, and finally decoded by the latent space decoder to generate a high-quality predicted target mask, and the original image is accurately segmented accordingly. Compared with traditional segmentation methods, the present invention can effectively remove irrelevant interference in a complex background environment, and at the same time, combined with cross-modal information fusion, improve the generalization ability and segmentation accuracy of unseen categories, and achieve more efficient open vocabulary semantic segmentation.

[0021] like Figure 1 As shown, the image segmentation method provided by the present invention comprises the following steps: S11, obtaining an image to be segmented containing a target object and corresponding text information; wherein the text information is used to characterize the category of the target object in the image.

[0022] The text information is used to characterize the category of the target object in the image, thereby providing support for the subsequent alignment of visual and semantic information. For example, the text information may be "a photo of {*}, a photo of more {*}, ...," where "{*}" is a rough category descriptor (e.g., "sky," "trees," "buildings," etc.). In order to make the acquired image adapt to the input requirements of the subsequent image segmentation model, the image may be preprocessed, wherein the preprocessing method includes but is not limited to normalizing the image, adjusting the image resolution, etc.

[0023] S12, encoding the text information through the text encoder 1 of the image segmentation model to generate text embedding features, and encoding the image through the image encoder 2 of the image segmentation model to generate image potential features.

[0024] like Figure 2 As shown, for RGB images, given an image x and its text information , The text information is input into the text encoder 1 for encoding to convert the text information into a high-dimensional embedding feature to obtain a text embedding feature, which is used as a guiding factor to denoise the generated noisy latent mask feature to update the noisy latent mask feature for subsequent image segmentation. Among them, the text encoder 1 can be any Transformer-based encoder, including but not limited to CLIP, BERT, etc. Preferably, in order to better capture cross-modal semantic relationships, the text encoder 1 is CLIP. The image is input into the image encoder 2, the image features are extracted and mapped to a low-dimensional latent space to reduce the computational complexity and generate image latent features. The image encoder 2 includes but is not limited to a variational autoencoder (VAE), a ResNet or a convolutional neural network, etc., as long as the image can be mapped to a low-dimensional latent space. Preferably, in order to effectively learn the latent representation of the image and be applicable to generative tasks, the image encoder 2 is an encoder of a variational autoencoder.

[0025] S13, according to the preset number of fusions, based on the text embedding features, the image potential features are processed by the attention module 5 and the splicing module 4 of the image segmentation model within the preset number of fusions, until the noisy mask potential features generated at the last fusion are obtained, and after denoising, they are decoded by the latent space decoder 6 of the image segmentation model to generate a predicted target mask for segmenting the target object; wherein, at each fusion: The image latent features and the noisy mask latent features are fused through the stitching module 4 of the image segmentation model, and the visual features are extracted from the fused features obtained after the fusion through the attention module 5, and the visual features are aligned and fused with the text embedding features to update the noisy mask latent features during the next fusion; wherein the noisy mask latent features during the first fusion are the random noise generated by the noise generator 3 of the image segmentation model.

[0026] Continue as Figure 2 As shown, in each fusion process, the splicing module 4 of the image segmentation model converts the image potential features Noisy masked latent features corresponding to the current fusion number t The fused visual features are extracted through the attention module 5 and aligned with the text embedding features to update the noisy masked latent features required for the next fusion time t-1 When fused for the first time, the noisy mask latent feature is the random noise generated by noise generator 3 In order to achieve more efficient and accurate image segmentation, the random noise is Gaussian noise. The above fusion process continues until the preset fusion times are reached, and the noisy mask potential features after the last fusion are obtained. . The mask latent feature After denoising, it is input to the latent space decoder 6 of the image segmentation model for decoding to generate a predicted target mask , to achieve accurate segmentation of the target object. It should be noted that in order to speed up the reasoning process, the present application performs non-Markov sampling by recalibrating the number of steps to obtain random noise.

[0027] Continue as Figure 2 As shown, in one embodiment of the present invention, S13 includes S131 to S135 (not shown in the figure): S131. The image latent features and the noisy mask latent features are fused through the splicing module 4 of the image segmentation model to obtain fused features.

[0028] At each fusion, the image latent features and the noisy mask latent features are fused through the splicing module 4 of the image segmentation model to obtain the fused features. At each fusion (i.e. at the current time step t): Figure 2 As shown, the image potential features and the noisy masked latent features at the current time step t Through the splicing module 4, the splicing fusion is performed to obtain the fusion features of the current time step The fusion features obtained after splicing It contains both the visual information of the original image and the noise information of the mask feature during the current fusion.

[0029] Optionally, the present invention can obtain fusion features by the following two different methods, which are described below by taking the t-th fusion (i.e., the t-th time step) as an example: Method 1: converting the image potential features and noisy masked latent features Perform superposition processing to generate joint features; combine the joint features and the noisy masked latent features Perform splicing to generate fusion features ,Right now =cat( , ). Specifically, the image latent features and the noisy masked latent feature at the current time step t The joint feature is generated by adding them element by element. Splice along the channel dimension to obtain the final fusion feature This fusion method not only retains the original mask information, but also enhances its interaction with the image features, so that the text information can better guide the subsequent denoising process, thereby providing a richer feature expression for the subsequent attention mechanism to improve the accuracy of target object segmentation. Method 2: Concatenate the image latent features and the noisy mask latent features to generate fused features. Specifically, the image latent features and the noisy masked latent feature at the current time step t Splice along the channel dimension to generate fusion features ,Right now =cat( , ). This adjustment helps to speed up the convergence of the model and significantly accelerate the training process without prediction errors. The fused features obtained in this way not only contain the local and global visual features of the image, but also retain the noise characteristics of the mask, so that the subsequent attention mechanism can optimize the mask more accurately. After obtaining the fused features in the above way, the number of input channels is doubled to fit the expanded input .

[0030] S132. Through the diffusion coding unit of the attention module 5, based on downsampling networks of different scales, visual features of corresponding scales are extracted from the fused features; wherein the diffusion coding unit includes a cascade of multiple downsampling networks with decreasing scales in sequence.

[0031] After the fused features are input into the diffusion coding unit, the features are propagated downward layer by layer through multiple cascaded downsampling networks. The feature map output by each downsampling network is used as the input of the next downsampling network. Through this layer-by-layer propagation method, the size of the feature map can be gradually reduced to extract higher-level semantic information. Specifically, for the initial downsampling network, basic information such as edges and textures of the feature map can be extracted. As the feature map propagates downward layer by layer, the intermediate-scale downsampling network pays more attention to the overall shape of the target object and the semantic information of the local area. When the feature enters the deeper downsampling network, the scale gradually decreases. The model focuses on extracting high-level semantic information, strengthening the ability to distinguish the target category, and focusing on key areas through the attention mechanism. These visual features of different scales provide richer feature representations for the subsequent semantic alignment and decoding processes to improve the segmentation accuracy and generalization ability of the model.

[0032] S133. Through the attention unit of the attention module 5, the minimum-scale visual features and text embedding features are cross-modally fused and aligned based on the content attention mechanism to obtain fused and aligned features.

[0033] Through the attention unit, the smallest-scale visual features are cross-modally aligned with the text embedding features. Since the smallest-scale visual features provide global semantic information, and the text embedding features give a semantic description of the target object, the text information and semantic information are aligned through the attention mechanism, so that the model can more accurately understand the correspondence between text information in the visual space.

[0034] In an optional embodiment, S133 includes: generating a query matrix, a key matrix and a first value matrix based on the minimum scale visual features, and generating an intermediate matrix and a second value matrix based on the text embedding features; based on the query matrix and the key matrix, according to the scaled dot product attention mechanism, calculating the self-attention weights of the visual features; based on the key matrix and the intermediate matrix, according to the scaled dot product attention mechanism, calculating the interactive attention weights of the text features to the visual features; based on the self-attention weights and the interactive attention weights, performing weighted summation on the first value matrix and the second value matrix to obtain a fused alignment feature.

[0035] Specifically, the query matrix Q, key matrix K and first value matrix are generated based on the minimum scale visual features. , where the query matrix represents the most representative part of the visual features of the minimum scale, the key matrix is ​​used to provide a reference for the query matrix to calculate the attention distribution, and the first value matrix is ​​used for weighted aggregation during attention calculation. The text embedding features are generated by linear projection to generate the intermediate matrix I and the second value matrix , where the middle matrix is ​​used to match the visual features so that the text and the visual features are mapped and associated, and the second value matrix can store the text embedding features to provide basic information for subsequent fusion. Using the scaled dot product attention mechanism, as shown in formula (1), the self-attention weight of the visual feature is calculated : (1) Among them, d is the feature dimension, Q is the query matrix, and K is the key matrix. The self-attention weight can help the model determine which areas in the image are more important, so as to better understand the internal structure and content of the image and realize subsequent image segmentation. Similarly, according to the key matrix and the intermediate matrix, the interactive attention weight is calculated using formula (2): : (2) Where I is the intermediate matrix. Combined with the calculated attention weights, the visual features and text features are weighted as shown in formula (3) to obtain the fused alignment feature : (3) in, and They are the first value matrix and the second value matrix respectively. Through this query-key-value mechanism, the visual features can be better aligned with the text semantics, thereby improving the accuracy of subsequent semantic segmentation. Through the above process, the gap between text embedding and visual features can be narrowed, and the feature advantages of both text and image modalities can be fully utilized to improve the accuracy of image segmentation.

[0036] S134. Through the diffusion decoding unit of the attention module 5, the intermediate features and the visual features of the corresponding scale are fused in each upsampling unit to obtain the intermediate features input to the next upsampling unit, and the prediction noise is generated based on the intermediate features obtained by the last upsampling unit; wherein the initial intermediate features are fused alignment features, and the decoder module includes a plurality of cascaded upsampling units with increasing scales in sequence, and the scale of the upsampling unit corresponds to the scale of the downsampling network.

[0037] In each upsampling unit, the intermediate features are upsampled to gradually increase their resolution. The upsampled features are fused with the visual features of the encoder at the corresponding scale to retain the key information of the target area and enhance the model's ability to express local structures. The above process is carried out in multiple upsampling units in sequence to restore image features layer by layer. The last upsampling unit generates the final intermediate features and calculates the predicted noise through the diffusion process for further denoising or generating target masks. Through this method of gradually increasing resolution and decoding layer by layer, the image segmentation model can gradually restore the mask features from the overall contour to the edge details, so that the semantic information remains consistent during the restoration process, so that the image segmentation model can more accurately restore the shape and boundaries of the target object in high-resolution images.

[0038] S135. De-noising the noisy masked latent features at the time of current fusion based on the predicted noise is performed through the denoising unit of the attention module 5 to obtain the denoised masked latent features, and use them as the noisy masked latent features at the time of next fusion.

[0039] like Figure 2 As shown, the denoising unit of the attention module 5 uses the predicted noise of the current time step to denoise the noisy mask potential features to remove unnecessary noise interference and extract clearer target mask information. The above process follows the reverse diffusion process in the diffusion model: the denoised mask potential features generated at each time step are used as the noisy mask input of the next time step, iterated in sequence, and gradually approached the final true target mask to improve the accuracy of image segmentation.

[0040] like Figure 2 As shown, in one embodiment of the present invention, there are K predicted target masks generated by the latent space decoder 6, where K is a positive integer greater than 1, and the processed mask latent features are decoded to generate a predicted target mask for segmenting the target object, including: First, the processed masked latent features are decoded to obtain K initial predicted target masks.

[0041] The masked latent features are gradually upsampled and reconstructed through the latent space decoder 6 to restore them to the resolution of the original input image, and K initial predicted target masks are obtained.

[0042] Then the mean of the K initial predicted target masks is calculated as the final predicted target mask.

[0043] Since the original image is copied into K channels during encoding, after decoding, the K initial prediction target masks need to be averaged to obtain the final prediction target mask. For example, K is 3, that is, an RGB image can be simulated, the original image is copied into three channels, and after decoding, the three initial prediction target masks are averaged to obtain the final prediction target mask.

[0044] S14. Segment the target object in the image based on the predicted target mask.

[0045] The predicted target mask provides the probability distribution of the category or target area to which each pixel belongs. The model classifies the image at the pixel level based on the mask. The target mask can accurately determine the boundary area of ​​the target object, thereby completing the segmentation of the target object in the image. It is understandable that since the category of the target object in the image may not be unique, the predicted target masks of target objects of different categories can be mapped to different depth values ​​according to the preset mapping rules, and these depth values ​​are combined to form a single depth map for segmenting the target object.

[0046] like Figure 2 As shown, the present invention also provides an image segmentation model, including: an image encoder 2, a noise generator 3, a text encoder 1, a splicing module 4, an attention module 5 and a latent space decoder 6, wherein: The output end of the image encoder 2 is connected to the first input end of the splicing module 4; The output end of the noise generator 3 is connected to the second input end of the splicing module 4; The output terminal of the text encoder 1 is connected to the first input terminal of the attention module 5; The output end of the splicing module 4 is connected to the second input end of the attention module 5; wherein the attention module 5 is used to extract visual features from the fused features generated by the splicing module 4, and align and fuse the visual features with the text embedding features obtained by encoding the text information by the text encoder 1, and obtain the denoised mask latent features after denoising; The output end of the attention module 5 is connected to the input end of the latent space decoder 6; wherein the latent space decoder 6 is used to decode the denoising mask latent features to generate a predicted target mask for segmenting the target object in the image.

[0047] like Figure 2 As shown, the output end of the image encoder 2 is connected to the second input end of the splicing module 4, which is used to convert the image potential features The output end of the noise generator 3 is connected to the second input end of the splicing module 4 to generate the noise. The two input ends of the attention module 5 are connected to the splicing module 4 and the text encoder 1 respectively, and are used to receive the text embedding features and the fusion features after splicing. , and after content attention processing, the noisy masked latent features are obtained , after denoising, we get the noisy masked potential features of the next fusion times t-1 , repeat the above process until the last fusion number of noisy masked latent features are obtained, and after denoising, the predicted target mask is obtained through the latent space decoder 6 .

[0048] like Figure 3 and Figure 4 As shown, the present invention also provides a training method for an image segmentation model, the training method comprising: S31, acquiring an image containing a target object and corresponding text information, as well as a target mask of the target object in the image; wherein the text information is used to characterize the category of the target object in the image.

[0049] The text information is used to characterize the target object category in the image and provide guidance for subsequent semantic alignment. The target mask is the pixel-level annotation of the target object in the image, which is used to supervise the image segmentation model to learn accurate segmentation boundaries.

[0050] S32, input the text information into the text encoder 1 of the semantic segmentation model for encoding to generate text embedding features, input the image into the image encoder 2 of the semantic segmentation model for encoding to generate image latent features, and input the target mask into the mask encoder 7 of the semantic segmentation model for encoding to generate mask latent features.

[0051] The text encoder 1 of the semantic segmentation model is used to encode text information (such as a picture of the sky), and the text embedding features of the text information are obtained through word segmentation, embedding mapping, attention calculation and other steps. The image encoder 2 is used to extract features from the image x and generate low-dimensional image potential features. Similarly, the target mask y is encoded by the mask encoder 7 to generate a low-dimensional masked latent feature . Among them, the image encoder 2 and the mask encoder 7 can be encoders of variational autoencoders.

[0052] S33, according to the preset number of fusions, fine-tune the attention module 5 of the semantic segmentation model at each fusion to obtain a trained attention module 5 of the semantic segmentation model; wherein, at each fusion: The masked latent features are subjected to noise processing by the noise generator 3 of the semantic segmentation model, and the noisy masked latent features and the image latent features are fused by the noise addition module 8 and the splicing module 4 of the semantic segmentation model. The visual features are extracted from the fused features obtained after fusion by the attention module 5 of the semantic segmentation model, and the visual features and the text embedding features are aligned and fused based on the attention mechanism to generate noisy masked latent features, and the attention module 5 of the semantic segmentation model is trained by the noisy masked latent features.

[0053] In each fusion, the mask latent features are input into the noise generator 3 of the semantic segmentation model for noise addition to simulate the noisy mask features in the diffusion process. The noise addition module 8 and the splicing module 4 are used to fuse the noisy mask latent features with the image latent features to obtain richer visual information. The key visual features are extracted from the image and fused with the text embedding features through the attention mechanism to learn a more accurate representation of the target area. Based on this process, the noisy masked latent features are generated. That is, an image containing prediction noise, and using the attention module 5 of the image segmentation model, continuously optimize its performance at different time steps, and finally train an efficient attention module 5 to improve the accuracy and generalization ability of the segmentation task. It can be understood that the attention module 5 of the present invention is constructed on the basis of a pre-trained text-to-image LDM to effectively utilize the image prior knowledge in the pre-trained dataset. By making minimal modifications to the model components, it is converted into an open vocabulary semantic segmentation model by adding the splicing module 4 and changing the attention mechanism. It should be further explained that the text encoder 1, image encoder 2, mask encoder 7 and latent space decoder 6 of the present invention do not change parameters during the above training process. In addition, in order to simulate RGB images, the target mask is copied to three channels. This step is crucial because the encoder is initially configured to process 3-channel (RGB) inputs, while the target mask has only a single channel. The present invention can reconstruct the target mask from the latent code without modifying the latent space or VAE. During the inference process, the values ​​on the three channels are averaged after decoding the latent code of the mask at the end of the inverse process to obtain the predicted target mask. In addition, in the training stage of the present invention, the random noise is multi-resolution noise.

[0054] The present invention also provides an image segmentation system, comprising: an image encoder 2, a mask encoder, a noise generator 3, a text encoder 1, a noise addition module, a splicing module 4 and an attention module 5, wherein: the output end of the mask encoder is connected to the first input end of the noise addition module; wherein the mask encoder is used to encode the target mask of the target object in the image into a mask potential feature; the output end of the noise generator 3 is connected to the second input end of the noise addition module; wherein the noise generator 3 is used to generate random noise; the output end of the image encoder 2 is connected to the first input end of the splicing module 4; wherein the image The encoder 2 is used to encode the image to generate image latent features; the output end of the noise addition module is connected to the second input end of the splicing module 4; the output end of the text encoder 1 is connected to the first input end of the attention module 5; wherein, the text encoder 1 is used to encode text information to generate text embedding features; the output end of the splicing module 4 is connected to the second input end of the attention module 5; wherein, the attention module 5 is used to extract visual features from the fused features generated by the splicing module 4, and align and fuse the visual features and the text embedding features through the attention mechanism to generate noisy masked latent features.

[0055] like Figure 4 As shown, the output end of the mask encoder and the output end of the noise generator 3 are connected to the first input end and the second input end of the noise summing module respectively, and the mask encoder encodes the real target mask into the mask potential feature , the image encoder 2 converts the image x Encoded as image latent features , the noise summation module 8 will mask the potential features and random noise At the tth fusion (i.e., the tth time step), the splicing module 4 combines the fused features and image latent features Splicing. The spliced ​​features And the text embedding feature is input into the attention module 5 to obtain the noisy masked latent feature . In summary, the present invention discloses an image segmentation method, a model, a model training method and an image segmentation system. The text information of the target object is input into a text encoder, and the text information is encoded to generate text embedding features, so that the model can understand the image content and adapt to unseen categories to improve the zero-sample segmentation ability. In addition, the image is encoded into image latent features by the image encoder. After encoding, the image latent features and the noisy mask latent features are fused according to the preset number of fusions, and the visual information is extracted, which is cross-modally aligned with the text embedding features, and the mask features are continuously optimized to enhance the model's ability to capture target boundaries and semantic information. At the time of initial fusion, the noisy mask latent features are initialized by random noise, and after a multi-step iterative diffusion denoising process, clear target mask features are gradually extracted, and finally decoded by the latent space decoder to generate a high-quality predicted target mask, and the original image is accurately segmented based on this.

[0056] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. An image segmentation method, characterized in that: The segmentation method comprises: Acquire an image to be segmented containing a target object and text information corresponding thereto; wherein the text information is used to characterize the category of the target object in the image; Encoding the text information through a text encoder of the image segmentation model to generate text embedding features, and encoding the image through an image encoder of the image segmentation model to generate image potential features; According to a preset number of fusions, based on the text embedding features, the image potential features are processed by the attention module and the splicing module of the image segmentation model within a preset number of fusions until the noisy mask potential features generated at the last fusion are obtained, and after denoising, they are decoded by the latent space decoder of the image segmentation model to generate a predicted target mask for segmenting the target object; wherein, at each fusion: The image latent features and the noisy mask latent features are fused through the splicing module of the image segmentation model, and visual features are extracted from the fused features obtained after fusion through the attention module, and the visual features are aligned and fused with the text embedding features to update the noisy mask latent features at the next fusion; wherein the noisy mask latent features at the first fusion are random noises generated by the noise generator of the image segmentation model; The target object in the image is segmented based on the predicted target mask.

2. The image segmentation method according to claim 1, characterized in that: The latent space decoder generates K predicted target masks, where K is a positive integer greater than 1, and decodes the processed mask latent features to generate a predicted target mask for segmenting the target object, including: Decode the processed masked latent features to obtain K initial predicted target masks; Calculate the mean of the K initial predicted target masks as the final predicted target mask.

3. The image segmentation method according to claim 1, characterized in that: The image potential features and the noisy mask potential features are fused through the splicing module of the image segmentation model, and visual features are extracted from the fused features obtained after fusion through the attention module, and the visual features are aligned and fused with the text embedding features to update the noisy mask potential features at the next fusion, including: The image potential feature and the noisy mask potential feature are fused through a splicing module of an image segmentation model to obtain a fused feature; Extracting visual features of corresponding scales from the fused features through the diffusion coding unit of the attention module based on downsampling networks of different scales; wherein the encoder module includes a cascade of multiple downsampling networks of decreasing scales in sequence; By using the attention unit of the attention module, the minimum-scale visual feature and the text embedding feature are cross-modal feature fused and aligned based on the content attention mechanism to obtain a fused alignment feature; Through the diffusion decoding unit of the attention module, the intermediate features and the visual features of the corresponding scale are fused in each upsampling unit to obtain the intermediate features input to the next upsampling unit, and the prediction noise is generated based on the intermediate features obtained by the last upsampling unit; wherein the initial intermediate features are the fused alignment features, and the decoder module includes a plurality of cascaded upsampling units with increasing scales, and the scale of the upsampling unit corresponds to the scale of the downsampling network; The denoising unit of the attention module performs denoising on the noisy masked latent features at the time of current fusion based on the predicted noise to obtain denoised masked latent features, which are used as the noisy masked latent features at the time of next fusion.

4. The image segmentation method according to claim 1, characterized in that: The fusing of the image potential features and the noisy mask potential features comprises: Superimposing the image potential feature and the noisy mask potential feature to generate a joint feature; The joint feature and the noisy masked latent feature are concatenated to generate a fused feature.

5. The image segmentation method according to claim 1, characterized in that: The fusing of the image potential features and the noisy mask potential features includes: splicing the image potential features and the noisy mask potential features to generate a fused feature.

6. The image segmentation method according to claim 1, characterized in that: The random noise is Gaussian noise.

7. The image segmentation method according to claim 1, characterized in that: The attention unit of the attention module performs cross-modal feature fusion and alignment on the minimum-scale visual feature and the text embedding feature based on the content attention mechanism to obtain a fused alignment feature, including: Generate a query matrix, a key matrix and a first value matrix based on the visual features of the minimum scale, and generate an intermediate matrix and a second value matrix based on the text embedding features; Based on the query matrix and the key matrix, calculating the self-attention weight of the visual feature according to a scaled dot product attention mechanism; Based on the key matrix and the intermediate matrix, calculating the interactive attention weight of the text feature to the visual feature according to the scaled dot product attention mechanism; According to the self-attention weight and the interactive attention weight, the first value matrix and the second value matrix are weighted summed to obtain a fused alignment feature.

8. An image segmentation model, characterized in that: The image segmentation model includes: an image encoder, a noise generator, a text encoder, a splicing module, an attention module and a latent space decoder, wherein: The output end of the image encoder is connected to the first input end of the splicing module; The output end of the noise generator is connected to the second input end of the splicing module; The output of the text encoder is connected to the first input of the attention module; The output end of the splicing module is connected to the second input end of the attention module; wherein the attention module is used to extract visual features from the fused features generated by the splicing module, align and fuse the visual features with the text embedding features obtained by encoding the text information by the text encoder, and obtain the denoised masked latent features after denoising; The output of the attention module is connected to the input of the latent space decoder; wherein the latent space decoder is used to decode the denoising mask latent features to generate a predicted target mask for segmenting the target object in the image.

9. A method for training an image segmentation model, characterized in that: The training method comprises: Acquire an image containing a target object and text information corresponding thereto, as well as a target mask of the target object in the image; wherein the text information is used to characterize the category of the target object in the image; Input the text information into the text encoder of the semantic segmentation model for encoding to generate text embedding features, input the image into the image encoder of the semantic segmentation model for encoding to generate image latent features, and input the target mask into the mask encoder of the semantic segmentation model for encoding to generate mask latent features; According to the preset number of fusions, the attention module of the semantic segmentation model is fine-tuned and trained at each fusion to obtain a trained attention module of the semantic segmentation model; wherein, at each fusion: The masked latent features are noised and fused with the image latent features through the noise addition module and the splicing module of the semantic segmentation model, visual features are extracted from the fused features obtained after fusion through the attention module of the semantic segmentation model, the visual features and the text embedding features are aligned and fused based on the attention mechanism to generate noisy masked latent features, and the attention module of the semantic segmentation model is trained through the noisy masked latent features.

10. An image segmentation system, characterized in that: The image segmentation system comprises: an image encoder, a mask encoder, a noise generator, a text encoder, a noise addition module, a splicing module and an attention module, wherein: The output end of the mask encoder is connected to the first input end of the noise summing module; wherein the mask encoder is used to encode the target mask of the target object in the image into a mask potential feature; The output end of the noise generator is connected to the second input end of the noise adding module; wherein the noise generator is used to generate random noise; The output end of the image encoder is connected to the first input end of the splicing module; wherein the image encoder is used to encode the image to generate image potential features; The output end of the noise adding module is connected to the second input end of the splicing module; The output end of the text encoder is connected to the first input end of the attention module; wherein the text encoder is used to encode text information to generate text embedding features; The output end of the splicing module is connected to the second input end of the attention module; wherein the attention module is used to extract visual features from the fused features generated by the splicing module, and align and fuse the visual features and the text embedding features through the attention mechanism to generate noisy masked latent features.

Citation Information

Patent Citations

  • Image processing method and device, computer, storage medium and program product

    CN117252947A

  • Zero sample image segmentation model training method and device based on multiple modes

    CN117788981A

  • Image generation model processing method and device, equipment, storage medium and product

    CN118115622A

  • Speech synthesis method and device based on artificial intelligence, terminal equipment and medium

    CN119068861A

  • Infrared small target detection method based on diffusion model

    CN119313913A

Cited By

  • Multi-target countermeasure attack method and device based on text description

    CN120510487A

  • Open vocabulary semantic segmentation method based on diffusion model

    CN120953614A

  • Liver minimally invasive surgery video segmentation method and system, computer equipment and storage medium

    CN121010935A