Semantic segmentation method, model training method, model and system
By designing a semantic segmentation model including image encoder, noise generator, text encoder, splicing module, attention module and latent space decoder, the problem of low target recognition accuracy in open vocabulary semantic segmentation tasks in complex scenarios is solved, and accurate and stable image segmentation in open categories and zero-sample scenarios is achieved.
Patent Information
- Application Number
- CN202510465185.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In the open vocabulary semantic segmentation task in complex scenarios, the target recognition accuracy is low and the visual-language understanding ability of the diffusion model is not fully utilized.
A semantic segmentation model including an image encoder, a noise generator, a text encoder, a splicing module, an attention module and a latent space decoder were designed. Through the stitching module, the latent image features are fused with the latent image with noise mask latent features, the attention module realizes cross-modal semantic alignment of the image and text, and finally generates the predicted target mask through the latent space decoder.
It improves the accuracy and stability of the semantic segmentation model in open category and zero-sample scenarios, improves segmentation accuracy, and can achieve accurate image segmentation in complex scenarios.
Smart Images

Figure CN119992551A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a semantic segmentation method, a model training method, a model and a system. Background Art
[0002] Open-Vocabulary Semantic Segmentation (OVSS) is an important computer vision task that aims to semantically annotate each pixel in an image based on text descriptions. Unlike traditional semantic segmentation methods, OVSS needs to have the ability to recognize unseen categories to adapt to diverse targets in open environments. Although multimodal models based on text-image matching perform well in single-target classification, it is difficult to accurately understand the relationship between objects in the image in complex scenes, which limits its application in pixel-level segmentation tasks. In addition, although text-to-image diffusion models have shown strong capabilities in generating high-quality images in recent years, most OVSS studies only use these diffusion models as fixed feature extractors and lack deep fusion designs between modules, such as the lack of dynamic information interaction between image encoders, mask modeling modules and attention mechanisms. Therefore, the rich visual-language understanding capabilities of the model are not fully utilized, resulting in performance bottlenecks in pixel-level semantic alignment and unseen category recognition. Therefore, it is necessary to provide a semantic segmentation method, a model training method, a model and a system. Summary of the invention
[0003] In view of the shortcomings of the prior art described above, the purpose of the present invention is to provide a semantic segmentation method, a model training method, a model and a system, which improve the problem of low target recognition accuracy in complex scenarios when the prior art performs open vocabulary semantic segmentation tasks.
[0004] To achieve the above-mentioned purpose and other related purposes, the present invention provides a semantic segmentation model, which includes: an image encoder, a noise generator, a text encoder, a splicing module, an attention module and a latent space decoder, wherein: the output end of the image encoder is connected to the first input end of the splicing module; the output end of the noise generator is connected to the second input end of the splicing module; the output end of the text encoder is connected to the first input end of the attention module; the output end of the splicing module is connected to the second input end of the attention module; wherein the attention module is used to extract visual features from the fusion features generated by the splicing module, and align and fuse the visual features and the text embedding features obtained by encoding the text information by the text encoder, and obtain denoising mask latent features after denoising; the output end of the attention module is connected to the input end of the latent space decoder; wherein the latent space decoder is used to decode the denoising mask latent features to generate a predicted target mask for segmenting the target object in the image.
[0005] In one embodiment of the present invention, the splicing module includes: a superposition unit, a first input end of which is connected to the output end of the image encoder, and a second input end of which is connected to the output end of the noise generator; wherein the superposition unit is used to weightedly fuse the image potential features and the noisy mask potential features generated by the noise generator to generate a joint feature; a splicing unit, a first input end of which is connected to the output end of the superposition unit, and a second input end of which is connected to the output end of the noise generator; wherein the splicing unit is used to splice the joint feature and the noisy mask potential features to generate a fused feature.
[0006] In one embodiment of the present invention, the splicing module includes: a splicing unit, a first input end of which is connected to the output end of the image encoder, and a second input end of which is connected to the output end of the noise generator; wherein the splicing unit is used to splice the image latent features and the noisy masked latent features generated by the noise generator to generate fused features.
[0007] In one embodiment of the present invention, the attention module includes: a diffusion coding unit, whose input end is connected to the output end of the splicing module, and is used to extract visual features of corresponding scales from the fused features generated by the splicing module based on downsampling networks of different scales; wherein the diffusion coding unit includes a cascade of multiple downsampling networks with decreasing scales in sequence; a content attention unit, whose first input end is connected to the output end of the text encoder, and whose second input end is connected to the output end of the diffusion coding unit, and is used to perform cross-modal feature fusion and alignment of the minimum scale visual features and the text embedding features based on the content attention mechanism , obtaining fused alignment features; a diffusion decoding unit, a first input end of which is connected to the output end of the content attention unit, and a second input end of which is connected to the output end of the diffusion coding unit, for fusing the intermediate features with the visual features of the corresponding scale in each upsampling network to obtain the intermediate features input to the next upsampling network, and generating an image containing prediction noise based on the intermediate features obtained by the last upsampling network; wherein the initial intermediate features are the fused alignment features, and the diffusion decoding unit includes a plurality of cascaded upsampling networks with increasing scales in sequence, and the scale of the upsampling network corresponds to the scale of the downsampling network.
[0008] In one embodiment of the present invention, the content attention unit includes: a matrix generation subunit, a first input end of which is connected to the output end of the text encoder, and a second input end of which is connected to the output end of the diffusion coding unit, for generating a query matrix, a key matrix and a first value matrix based on the minimum scale visual feature, and correspondingly generating an intermediate matrix and a second value matrix based on the text embedding feature; a self-attention weight subunit, an input end of which is connected to the first output end of the matrix generation subunit, for calculating the self-attention weight of the visual feature based on the query matrix and the key matrix according to the scaled dot product attention mechanism; an interaction attention weight subunit , whose input end is connected to the second output end of the matrix generation subunit, and is used to calculate the interactive attention weight of the text feature to the visual feature based on the key matrix and the intermediate matrix according to the scaled dot product attention mechanism; a feature fusion subunit, whose first input end is connected to the output end of the self-attention weight subunit, whose second input end is connected to the output end of the interactive attention weight subunit, and whose third input end is connected to the third output end of the matrix generation subunit, and is used to perform weighted summation of the first value matrix and the second value matrix according to the self-attention weight and the interactive attention weight to obtain a fused alignment feature.
[0009] In one embodiment of the present invention, the attention module also includes: a denoising unit, whose input end is connected to the output end of the diffusion decoding unit, for denoising the image containing predicted noise generated each time within a preset denoising number of times, and using the denoised image as the noisy masked latent feature generated by the noise generator, so as to extract visual features again.
[0010] In one embodiment of the present invention, the image is a K-channel image, wherein K is a positive integer greater than 1, and the latent space decoder includes: a decoding unit, whose input end is connected to the output end of the attention module, and is used to obtain K initial predicted target masks based on the image containing prediction noise; a mask generation unit, whose input end is connected to the output end of the decoding unit, and is used to perform weighted averaging on the K initial predicted target masks, and use the weighted average as the final predicted target mask.
[0011] In one embodiment of the present invention, a semantic segmentation method is further provided, the semantic segmentation method comprising: obtaining an image to be segmented containing a target object and text information corresponding thereto; wherein the text information is used to characterize the category of the target object in the image; The text information is encoded by the text encoder of the image segmentation model to generate text embedding features, and the image is encoded by the image encoder of the image segmentation model to generate image latent features; according to a preset diffusion time step, based on the text embedding features, the image latent features are processed by the attention module and the splicing module of the image segmentation model at the preset diffusion time step until the noisy mask latent features generated at the last fusion are obtained, and after denoising, they are decoded by the latent space decoder of the image segmentation model to generate a predicted target mask for segmenting the target object; wherein, at each fusion: the image latent features and the noisy mask latent features are fused by the splicing module of the image segmentation model, and the visual features are extracted from the fused features obtained after fusion by the attention module, and the visual features are aligned and fused with the text embedding features to update the noisy mask latent features at the next fusion; wherein the noisy mask latent features at the first fusion are random noise generated by the noise generator of the image segmentation model; the target object in the image is segmented based on the predicted target mask.
[0012] In one embodiment of the present invention, a semantic segmentation system is further provided, the semantic segmentation system comprising: an image encoder, a mask encoder, a noise generator, a text encoder, a noise summation module, a splicing module and an attention module, wherein: the output end of the mask encoder is connected to the first input end of the noise summation module; wherein the mask encoder is used to encode the target mask of the target object in the image into a mask potential feature; the output end of the noise generator is connected to the second input end of the noise summation module; wherein the noise generator is used to generate random noise; the output end of the image encoder is connected to the first input end of the splicing module; Wherein, the image encoder is used to encode the image to generate image latent features; the output end of the noise addition module is connected to the second input end of the splicing module; the output end of the text encoder is connected to the first input end of the attention module; wherein, the text encoder is used to encode text information to generate text embedding features; the output end of the splicing module is connected to the second input end of the attention module; wherein, the attention module is used to extract visual features from the fused features generated by the splicing module, and align and fuse the visual features and the text embedding features through the attention mechanism to generate noisy masked latent features.
[0013] In one embodiment of the present invention, a training method for a semantic segmentation model is also provided, the training method comprising: obtaining an image containing a target object and text information corresponding thereto, and a target mask of the target object in the image; wherein the text information is used to characterize the category of the target object in the image; inputting the text information into a text encoder of a semantic segmentation model for encoding to generate text embedding features, inputting the image into an image encoder of the semantic segmentation model for encoding to generate image latent features, and inputting the target mask into a mask encoder of the semantic segmentation model for encoding to generate mask latent features; and performing the diffusion process according to a preset diffusion time step. , at each fusion, the attention module of the semantic segmentation model is fine-tuned and trained to obtain the attention module of the trained semantic segmentation model; wherein, at each fusion: the mask latent feature is noised and fused with the image latent feature through the noise addition module and the splicing module of the semantic segmentation model, the visual feature is extracted from the fused feature obtained after the fusion through the attention module of the semantic segmentation model, the visual feature and the text embedding feature are aligned and fused based on the attention mechanism to generate a noisy mask latent feature, and the attention module of the semantic segmentation model is trained through the noisy mask latent feature.
[0014] As described above, a semantic segmentation method, a model training method, a model and a system of the present invention have the following beneficial effects: the image latent features generated by the image encoder and the noisy mask latent features generated by the noise generator are fused through the splicing module to generate fused features. The attention module takes the fused features as input, and combines them with the text embedding features generated by the text encoder to achieve semantic alignment of cross-modal information between the image and the text through the content attention mechanism, thereby helping to improve the detection and recognition capabilities of the semantic segmentation model for the target semantic position. The denoised mask latent features finally obtained are decoded by the latent space decoder, so that the predicted target mask consistent with the structure of the target object can be restored to improve the segmentation accuracy. The semantic segmentation model can achieve accurate and stable image segmentation in open categories and zero-sample scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic diagram of the overall structure of a semantic segmentation model provided by an embodiment of the present invention; Figure 2 A schematic diagram of the overall structure of a splicing module provided by an embodiment of the present invention; Figure 3 A schematic diagram of the overall structure of a splicing module provided by another embodiment of the present invention; Figure 4 A schematic diagram of the overall structure of a content attention module provided by an embodiment of the present invention; Figure 5A schematic diagram of the overall structure of a content attention unit provided by an embodiment of the present invention; Figure 6 A schematic diagram of the overall structure of a latent space decoder provided by an embodiment of the present invention; Figure 7 A schematic diagram of the overall architecture of a semantic segmentation model provided by an embodiment of the present invention; Figure 8 A flowchart of a semantic segmentation method provided by an embodiment of the present invention; Fig. 9 A schematic diagram of the overall architecture of an image segmentation system provided by an embodiment of the present invention; Fig.10 A flowchart of a method for training a semantic segmentation model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0017] It should be noted that the illustrations provided in the following embodiments are only used to schematically illustrate the basic concept of the present invention, and thus the illustrations only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0018] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0019] Open vocabulary semantic segmentation (OVSS) is a challenging computer vision task, where the goal is to annotate each pixel in an image based on a textual description. Traditional semantic segmentation methods rely on predefined categories, which makes it difficult to recognize unseen categories during the inference phase, thus limiting their applicability in open settings. In contrast, humans have strong visual generalization capabilities and can recognize and understand a large number of unseen categories. To simulate this ability, in recent years, many studies have been devoted to exploring OVSS to achieve pixel-level segmentation of a wide range of semantic categories, including those that did not appear in the training data.
[0020] Many open vocabulary semantic segmentation approaches leverage the strong generalization capabilities of text-image discriminative models trained on large Internet datasets. While these pre-trained models perform well at classifying individual objects or pixels, they may not be fully suited for scene-level analysis. Recent work has shown that models like CLIP often misjudge spatial relationships between objects. Therefore, the lack of spatial and relational understanding of text-image discriminative models has become a barrier to open vocabulary semantic segmentation. In recent years, diffusion models trained on Internet-scale datasets for text-to-image generation have brought revolutionary changes in the field of image synthesis. These models offer superior image quality, adaptability, composability, and semantic guidance through text input. Notably, diffusion models facilitate the image generation process by creating cross-attention between text embeddings and their internal visual representations. This suggests that diffusion model representations are potentially unique and connected to higher-level semantic concepts expressible in language. Although there have been some attempts to use diffusion models for open vocabulary semantic segmentation, these works usually freeze stable diffusion models and use them alone as a means of feature extraction. They do not fine-tune and train the U-Net network in their model, which results in failure to fully exploit the rich visual-language capabilities of stable diffusion.
[0021] In addition, the current progress in the field of OVSS mainly relies on the improvement of model capacity, but existing methods still have limitations when dealing with unfamiliar images or new categories, and their visual language understanding is usually limited by training data. This has prompted researchers to explore more powerful prior knowledge to improve the zero-sample generalization ability of the model. In recent years, text-to-image diffusion models have shown strong capabilities in generating images with diverse open vocabulary descriptions. Therefore, a question worthy of attention is whether these diffusion models can provide more comprehensive visual-linguistic prior knowledge for OVSS tasks, thereby enhancing the generalization ability of the model on unseen categories. However, there is still a lack of a complete framework that can simultaneously parse object instances and scene semantics and make full use of the capabilities of diffusion models to improve the performance of OVSS in open environments.
[0022] In view of the above situation, the present invention provides a semantic segmentation model, which fuses the image latent features generated by the image encoder and the noisy mask latent features generated by the noise generator through a splicing module to generate fused features. The attention module takes the fused features as input, and combines them with the text embedding features generated by the text encoder to achieve semantic alignment of cross-modal information between image and text through the content attention mechanism, thereby helping to improve the detection and recognition ability of the semantic segmentation model for the target semantic position. The denoised mask latent features finally obtained are decoded by the latent space decoder, so that the predicted target mask consistent with the structure of the target object can be restored to improve the segmentation accuracy. The semantic segmentation model can achieve accurate and stable image segmentation in open categories and zero-sample scenarios.
[0023] The semantic segmentation model is used in the inference stage of the model. The semantic segmentation model takes the diffusion model as the core and gradually optimizes the mask potential features through diffusion modeling to achieve semantic-guided image segmentation. The semantic segmentation model is used to input an image to be segmented into the semantic segmentation model in the inference stage, segment the image, and combine the text information of the image to semantically understand and process the image, so as to accurately segment the target object from the image.
[0024] like Figure 1 As shown in the figure, the semantic segmentation model includes: image encoder, noise generator, text encoder, splicing module, attention module and latent space decoder. The modules are connected in order according to their functions to form an image segmentation framework based on diffusion mechanism and guided by text semantics. The following is a detailed introduction of each module: Specifically, the text information is input into the text encoder, and the text embedding features generated after encoding are input into the attention module through its output end and the first input end of the attention module. The text information is used to characterize the category of the target object in the image, thereby providing support for the subsequent alignment of visual and semantic information. Exemplarily, the text information may be "a photo of {*}, a photo of more {*}, ...," where "{*}" is a rough category descriptor (e.g., "sky", "trees", "buildings", etc.).
[0025] The image containing the target object is input into the image encoder, the image is encoded, the image features are extracted and mapped to a low-dimensional latent space, thereby generating image latent features with rich semantic information to reduce the complexity of subsequent calculations. The output end of the image encoder is connected to the first input end of the splicing module to send the image latent features to the splicing module so that the image latent features can be fused with the mask-related features in the subsequent diffusion modeling process.
[0026] The output end of the noise generator is connected to the second input end of the splicing module to provide mask-related features at each diffusion time step as input to the attention module. Specifically, in the first diffusion time step, the noise generator generates and outputs random noise as the initial noisy masked latent features; and in the subsequent diffusion time step, the noisy masked latent features after denoising by the attention module in the previous diffusion time step are received, and the noisy masked latent features of the current time step are obtained according to the preset diffusion scheduling rules as input to the splicing module. Optionally, the random noise is multi-resolution noise to adapt to the multi-scale structure of the image potential features and improve the image segmentation model's ability to process different semantic information.
[0027] The splicing module fuses the image latent features with the corresponding mask-related features according to the different diffusion time steps. Specifically, at the first diffusion time step, the splicing module fuses the image latent features generated by the image encoder with the random noise output by the noise generator to obtain the fused features of the first diffusion time step; at the non-first diffusion time step, the splicing module fuses the image latent features with the noisy mask latent features of the current diffusion time step to generate the fused features of the corresponding diffusion time step. After obtaining the fused features, the splicing module sends the fused features of the current diffusion time step to the attention module through the second input terminal of the attention module, so that the attention module further extracts semantic features to achieve cross-modal alignment with text features.
[0028] like Figure 2 As shown, in one embodiment of the present invention, the splicing module includes: a superposition unit and a splicing unit, which work together to achieve feature fusion processing between image potential features and noisy mask potential features. The superposition unit has two input ends and one output end, and the first input end of the superposition unit is connected to the output end of the image encoder to receive the image potential features obtained after the image encoder encodes the input image. The second input end of the superposition unit is connected to the output end of the noise generator, and is used to receive the random noise generated by the noise generator in the first diffusion time step, and receive the noisy mask potential features generated according to the denoising result of the previous time step and the diffusion scheduling rule in the subsequent diffusion time step. The superposition unit performs weighted fusion of the image potential features with the random noise or noisy mask potential features to generate a joint feature, which contains both the contextual features of the image and the gradual change features of the mask from noise to the target shape during the diffusion process. The weighted fusion method includes but is not limited to element-by-element weighted fusion or linear superposition processing.
[0029] Continue as Figure 2As shown, the splicing module also includes a splicing unit, a first input end of the splicing unit is connected to the output end of the superposition unit, for receiving the above-mentioned joint feature, and a second input end is connected to the output end of the noise generator, for synchronously receiving the noisy masked latent feature generated by the current diffusion time step. The splicing unit splices the joint feature with the mask feature of the current diffusion time step by splicing the channel dimension to generate a fusion feature rich in semantic information.
[0030] Furthermore, if Figure 3 As shown, the splicing module may also only include a splicing unit, which is used to realize the fusion of the image potential features and the noisy mask potential features. The splicing unit has two input ends, the first input end is connected to the output end of the image encoder, and is used to receive the image potential features obtained after the image encoder encodes the original image. The second input end is connected to the output end of the noise generator, and is used to receive the random noise generated by the noise generator in the first diffusion time step or the noisy mask potential features iteratively obtained during the diffusion process. Specifically, the splicing unit splices the image potential features and the noisy mask potential features according to the channel dimension to generate a fusion feature with richer semantic information. The fusion feature not only contains the visual semantic information of the image, but also integrates the noise representation of the mask, which helps the subsequent attention module to perform more effective semantic alignment and feature extraction, thereby improving the accuracy and semantic consistency of mask reconstruction.
[0031] Through the above process, the fused features output by the splicing module are input into the attention module to participate in the cross-modal alignment and feature optimization process, thereby ensuring that the model can gradually improve the accuracy and boundary integrity of the target mask in each diffusion time step, thereby improving the mask reconstruction performance of the diffusion model under semantic guidance.
[0032] As a major component of the diffusion model, the attention module is used to extract features and semantically guide the input fusion features in each diffusion time step. Specifically, based on the attention mechanism, the attention module extracts multi-scale visual features from the fusion features generated by the splicing module, and aggregates the visual features of each scale to generate the fusion alignment features of the current diffusion time step. The obtained fusion alignment features contain both the spatial semantic information of the image and the category information of the target object provided by the text embedding features to achieve the combination of image and text information. In the aggregation process, the text embedding features are used as semantic guidance conditions to guide the attention module to focus on the areas highly related to the text description in the visual features to achieve the alignment and fusion of cross-modal features such as text-image. After the aggregation is completed, the attention module generates the prediction noise of the diffusion time step based on the fusion alignment features, and performs reverse denoising on the prediction noise based on the diffusion mechanism to obtain the denoised mask latent features of the current diffusion time step. The denoised mask latent features are passed to the splicing module as input so that they can be fused with the image latent features in the next diffusion time step to complete the next round of diffusion denoising process. The above iterative denoising process is repeated until all diffusion time steps are completed, and the denoising mask potential features obtained in the last diffusion time step are input into the latent space decoder to generate the predicted mask of the target object to achieve the segmentation of the target object in the image.
[0033] like Figure 4As shown, in one embodiment of the present invention, the attention module includes: a diffusion coding unit, a content attention unit and a diffusion decoding unit, and each module works together to achieve text-guided image mask prediction and target object segmentation. Among them, the input end of the diffusion coding unit is connected to the output end of the splicing module, and is used to receive the fusion feature formed by the fusion of the image potential feature and the noisy mask potential feature. The diffusion coding unit includes a cascade of multiple downsampling networks with decreasing scales in sequence. By downsampling layer by layer, the multi-scale semantic information of the image is extracted from the fusion feature, thereby obtaining a multi-level visual feature representation from low-level local details to high-level global semantics, so as to provide semantic support for the subsequent attention mechanism and decoding process. The content attention unit is used to achieve cross-modal semantic alignment, and its first input end is connected to the output end of the text encoder to receive the text embedding feature, and the second input end is connected to the output end of the diffusion coding unit to receive the multi-scale visual feature generated after downsampling the fusion feature. The content attention unit uses the content attention mechanism to guide the text embedding features by weighting through the attention weight according to the correspondence between the visual features of the minimum scale and the text embedding features, thereby generating a fused alignment feature that integrates the text semantic information and the visual information, thereby improving the recognition ability of the semantic segmentation model for the target object in the image. The diffusion decoding unit has two input terminals, the first input terminal of which is connected to the output terminal of the content attention unit, and is used to use the fused alignment feature output by the content attention unit as the initial intermediate feature. The second input terminal is connected to the output terminal of the diffusion coding unit, and is used to synchronously obtain the visual features of the corresponding scale when the diffusion coding unit downsamples. During the decoding process, the diffusion coding unit gradually restores the size of the feature map to the size before downsampling through multiple upsampling networks with increasing scales. In each upsampling process, the intermediate features corresponding to the upsampling network are fused with the visual features of the corresponding scale, and the fused features are used as the intermediate features corresponding to the next upsampling network. Repeat the upsampling until the intermediate features output by the last upsampling network are obtained, and the intermediate features are used as an image containing prediction noise, which contains the structural information of the target mask, so that it can be restored to the final segmentation mask after subsequent decoding. It should be noted that the scale of the upsampling network in the content attention module is consistent with the corresponding downsampling network, so that the features can be effectively aligned between encoding and decoding to ensure the coherence of semantic information transmission and improve the accuracy and robustness of subsequent target mask generation.
[0034] Furthermore, the content attention unit is used to achieve deep alignment between image visual features and text semantic features, such as Figure 5As shown, the content attention unit includes a matrix generation subunit, a self-attention weight subunit, an interactive attention weight subunit, and a feature fusion subunit. The subunits work together to complete the cross-modal feature fusion process to improve the accuracy of subsequent target mask generation. Specifically, the matrix generation subunit has two input terminals, a first input terminal connected to the output terminal of the text encoder to receive the text embedding features generated by the text encoder, and a second input terminal connected to the output terminal of the diffusion coding unit to receive multi-scale visual features. The matrix generation subunit generates a query matrix Q, a key matrix K, and a first value matrix based on the minimum scale visual features. , where the query matrix Q represents the most representative part of the visual features of the minimum scale, the key matrix K is used to provide a reference for calculating the attention distribution of the query matrix Q, and the first value matrix Used for weighted aggregation during attention calculation. The text embedding features are generated through linear projection to generate the intermediate matrix I and the second value matrix , where the intermediate matrix I is used to match visual features to associate text with visual features, and the second value matrix Text embedding features can be stored to provide text information for subsequent fusion processing.
[0035] The self-attention weight subunit is used to calculate the attention relationship within the visual feature, and its input is connected to the first output of the matrix generation subunit. Thus, according to the scaled dot product attention mechanism, the self-attention weight of the visual feature is calculated using the generated query matrix Q and key matrix K. , where d is the feature dimension. The self-attention weight can help the model determine which areas in the image are more important, so as to better understand the internal structure and content of the image and realize subsequent image segmentation.
[0036] The input end of the interaction attention weight subunit is connected to the second output end of the matrix generation subunit, which calculates the interaction attention weight using the key matrix K and the intermediate matrix I according to the scaled dot product attention mechanism. .
[0037] The feature fusion subunit is connected to the third output terminal of the self-attention weight subunit, the interactive attention weight subunit and the matrix generation subunit respectively to receive , , and The first value matrix of the feature fusion subunit and the second value matrix Perform double attention weighted fusion to obtain fused alignment features Through this query-key-value mechanism, visual features can be better aligned with text semantics, thereby improving the accuracy of subsequent semantic segmentation. Through the above process, the gap between text embedding and visual features can be narrowed, the advantages of text and image modalities can be fully utilized, and the accuracy of image segmentation can be improved.
[0038] In addition, the attention module also includes a denoising unit (not shown in the figure), which is used to implement iterative optimization of the mask potential features within multiple diffusion time steps. The input end of the denoising unit is connected to the output end of the diffusion decoding unit to receive the image containing prediction noise generated by the diffusion decoding unit at each diffusion time step. The denoising unit performs reverse denoising on the image containing prediction noise obtained each time within a preset number of denoising times (such as T steps). Specifically, the denoising unit uses the inverse process of the diffusion model to gradually restore the image containing prediction noise to the denoised mask potential features that are closer to the actual mask structure. The denoised data is used as a new image containing prediction noise for subsequent decoding processing.
[0039] Furthermore, the image is a K-channel image, where K is a positive integer greater than 1. In order to achieve accurate semantic segmentation of the multi-channel image, Figure 6 As shown, the latent space decoder in the present application includes a decoding unit and a mask generation unit, which work together to decode the denoised image containing prediction noise into a prediction target mask. Among them, the input end of the decoding unit is connected to the output end of the attention module, which is used to receive the image that is denoised by integrating visual and textual semantics (that is, the image containing prediction noise denoised by the denoising unit). The decoding unit restores the input image to the resolution of the original input image through upsampling and feature reconstruction, and obtains K initial prediction target masks.
[0040] The input end of the mask generation unit is connected to the output end of the decoding unit for receiving the K initial prediction masks. Since the original image is copied into K channels during encoding, after decoding, the K initial prediction target masks obtained need to be averaged to obtain the final prediction target mask. Exemplarily, K is 3, that is, an RGB image can be simulated, the original image is copied into three channels, and after decoding, the 3 initial prediction target masks are averaged to obtain the final prediction target mask.
[0041] like Figure 7 As shown, it shows the overall processing process of the semantic segmentation model of this application. The output end of the image encoder is connected to the second input end of the splicing module to convert the image potential features The output end of the noise generator is connected to the second input end of the splicing module to generate the noise. The two input ends of the attention module are connected to the splicing module and the text encoder respectively to receive the text embedding features and the fusion features after splicing. , and after content attention processing, the noisy masked latent features are obtained , after denoising, we get the noisy masked potential feature of the next diffusion time step t-1 , repeat the above process until the noisy masked latent features of the last diffusion time step are obtained, and after denoising, the predicted target mask is obtained through the latent space decoder .
[0042] like Figure 8 As shown, the present invention also provides a semantic segmentation method, the method comprising: S81. Acquire an image to be segmented containing a target object and corresponding text information; wherein the text information is used to characterize the category of the target object in the image.
[0043] S82. Encode the text information through a text encoder of the image segmentation model to generate text embedding features, and encode the image through an image encoder of the image segmentation model to generate image latent features.
[0044] S83. According to a preset diffusion time step, based on the text embedding feature, the image latent feature is processed at a preset diffusion time step through the attention module and the stitching module of the image segmentation model until the noisy masked latent feature generated at the last fusion is obtained, and after denoising, it is decoded through the latent space decoder of the image segmentation model to generate a predicted target mask for segmenting the target object; wherein, at each fusion: the image latent feature is fused with the noisy masked latent feature through the stitching module of the image segmentation model, and the visual feature is extracted from the fused feature obtained after the fusion through the attention module, and it is aligned and fused with the text embedding feature to update the noisy masked latent feature at the next fusion; wherein the noisy masked latent feature at the first fusion is the random noise generated by the noise generator of the image segmentation model.
[0045] S84: Segment the target object in the image based on the predicted target mask.
[0046] The text information of the target object is input into the text encoder, and the text information is encoded to generate text embedding features, so that the model can understand the image content and adapt to unseen categories to improve the zero-shot segmentation ability. In addition, the image is encoded into image latent features through the image encoder. After encoding, the image latent features and the noisy mask latent features are fused according to the preset diffusion time step, and the visual information is extracted. It is cross-modally aligned with the text embedding features, and the mask features are continuously optimized to enhance the model's ability to capture the target boundary and semantic information. At the initial fusion, the noisy mask latent features are initialized by the random noise generated by the noise generator. After a multi-step iterative diffusion denoising process, the target mask features with clear and accurate boundaries are gradually extracted, and the predicted target mask is generated by the latent space decoder for decoding, and the target object in the original image is accurately segmented based on this.
[0047] The present application also provides a semantic segmentation system, which includes: an image encoder, a mask encoder, a noise generator, a text encoder, a noise addition module, a splicing module and an attention module, wherein: the output end of the mask encoder is connected to the first input end of the noise addition module; wherein the compression encoder is used to encode the target mask of the target object in the image into a mask potential feature; the output end of the noise generator is connected to the second input end of the noise addition module; wherein the noise generator is used to generate random noise; the output end of the image encoder is connected to the first input end of the splicing module; wherein the image encoder is used to encode the image to generate image potential features; the output end of the noise addition module is connected to the second input end of the splicing module; the output end of the text encoder is connected to the first input end of the attention module; wherein the text encoder is used to encode text information to generate text embedding features; the output end of the splicing module is connected to the second input end of the attention module; wherein the attention module is used to extract visual features from the fused features generated by the splicing module, and align and fuse the visual features and the text embedding features through the attention mechanism to generate noisy mask potential features.
[0048] like Fig. 9 As shown, the output end of the mask encoder and the output end of the noise generator are connected to the first input end and the second input end of the noise summation module respectively, and the mask encoder encodes the real target mask into the mask potential feature , the image encoder encodes the image x into image latent features , the noise summation module will mask the latent features and random noise In the tth fusion (i.e., the tth diffusion time step), the splicing module combines the fused features and image latent features Splicing. The spliced features And the text embedding features are input into the attention module to obtain the noisy masked latent features .
[0049] like Fig.10 As shown, the present application also provides a method for training a semantic segmentation model, the training method comprising: S101, acquiring an image containing a target object and text information corresponding thereto, as well as a target mask of the target object in the image; wherein the text information is used to characterize the category of the target object in the image; S102, inputting the text information into a text encoder of a semantic segmentation model for encoding to generate text embedding features, inputting the image into an image encoder of the semantic segmentation model for encoding to generate image latent features, and inputting the target mask into a mask encoder of the semantic segmentation model for encoding to generate mask latent features; S103. According to the preset diffusion time step, the attention module of the semantic segmentation model is fine-tuned and trained at each fusion to obtain the trained attention module of the semantic segmentation model; wherein, at each fusion: the mask latent feature is noised through the noise addition module and the splicing module of the semantic segmentation model, and then fused with the image latent feature; the visual feature is extracted from the fused feature obtained after the fusion through the attention module of the semantic segmentation model; the visual feature and the text embedding feature are aligned and fused based on the attention mechanism to generate a noisy mask latent feature; and the attention module of the semantic segmentation model is trained through the noisy mask latent feature.
[0050] like Fig. 9 and Fig.10 As shown, in the training phase, an image containing a target object is obtained, and the text information and target mask corresponding to the image are obtained. Among them, the text information is used to describe the category of the target object in the image, which is used to provide a basis for subsequent semantic guidance. The target mask is used to mark the actual position of the target object in the image as a label for model training. The text information is input into the text encoder of the semantic segmentation model to extract its semantic features and obtain text embedding features; the image x is input into the image encoder to extract image features and map the image features to the latent space to generate image latent features Similarly, the target mask is fed into the mask encoder to generate the masked latent features The semantic segmentation model is iteratively trained according to the preset diffusion time step T. In each diffusion time step, the noise summation module is used to mask the latent features. Add noise to get the feature corresponding to the diffusion time step t . With the image latent features Input them into the splicing module for fusion processing to obtain fusion features The fusion features Input to the attention module, extract fusion features through the attention unit The multi-scale visual features are then cross-modally aligned with the text embedding features based on the content attention mechanism to generate intermediate features with semantic guidance capabilities. The intermediate features are used as the noisy masked latent features at the current time step. , which is used to supervise the training of the parameters of the attention module under the diffusion time step. Among them, the image encoder and the mask encoder can be variational autoencoders. It should be noted that the text encoder, image encoder, mask encoder and latent space decoder of the present invention do not change parameters during the above training process. In addition, in order to simulate RGB images, the target mask is copied to three channels. This step is crucial because the encoder is originally configured to process 3-channel (RGB) inputs, while the target mask has only a single channel. The present invention can reconstruct the target mask from the latent code without modifying the latent space or VAE. During the inference process, the predicted target mask is obtained by averaging the values on the three channels after decoding the latent code of the mask at the end of the inverse process. In addition, in the training stage of the present invention, the random noise generated by the noise generator at the first diffusion time step is multi-resolution noise.
[0051] For the specific limitations of the training system for the semantic segmentation model, please refer to the limitations of the training method for the semantic segmentation model above, which will not be repeated here. Each module in the above-mentioned training system for the semantic segmentation model can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware format, or can be stored in the memory in the computer device in software format, so that the processor can call the operations corresponding to the above modules.
[0052] It should be noted that, in order to highlight the innovative part of the present invention, the present embodiment does not introduce modules that are not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other modules in the present embodiment.
[0053] In summary, the present invention discloses a semantic segmentation method, a model training method, a model and a system, which fuse the image potential features generated by the image encoder and the noisy mask potential features generated by the noise generator through a splicing module to generate fused features. The attention module takes the fused features as input, and combines the text embedding features generated by the text encoder to achieve semantic alignment of cross-modal information between the image and the text through the content attention mechanism, thereby helping to improve the detection and recognition ability of the semantic segmentation model for the target semantic position. The denoised mask potential features finally obtained are decoded by the latent space decoder, so that the predicted target mask consistent with the structure of the target object can be restored to improve the segmentation accuracy. The semantic segmentation model can achieve accurate and stable image segmentation in open categories and zero-sample scenarios. Therefore, the present invention effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.
[0054] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A semantic segmentation model, characterized in that: The semantic segmentation model includes: an image encoder, a noise generator, a text encoder, a splicing module, an attention module and a latent space decoder, wherein: The output end of the image encoder is connected to the first input end of the splicing module; The output end of the noise generator is connected to the second input end of the splicing module; The output of the text encoder is connected to the first input of the attention module; The output end of the splicing module is connected to the second input end of the attention module; wherein the attention module is used to extract visual features from the fused features generated by the splicing module, align and fuse the visual features with the text embedding features obtained by encoding the text information by the text encoder, and obtain the denoised masked latent features after denoising; The output of the attention module is connected to the input of the latent space decoder; wherein the latent space decoder is used to decode the denoising mask latent features to generate a predicted target mask for segmenting the target object in the image.
2. The semantic segmentation model according to claim 1, characterized in that The splicing module comprises: A superposition unit, wherein a first input end thereof is connected to an output end of the image encoder, and a second input end thereof is connected to an output end of the noise generator; wherein the superposition unit is used to perform weighted fusion of the image latent feature and the noisy mask latent feature generated by the noise generator to generate a joint feature; A splicing unit, whose first input end is connected to the output end of the superposition unit, and whose second input end is connected to the output end of the noise generator; wherein the splicing unit is used to splice the joint feature and the noisy masked latent feature to generate a fused feature.
3. The semantic segmentation model according to claim 1, characterized in that The splicing module comprises: A splicing unit, wherein a first input end thereof is connected to an output end of the image encoder, and a second input end thereof is connected to an output end of the noise generator; wherein the splicing unit is used to splice the image latent features and the noisy masked latent features generated by the noise generator to generate fused features.
4. The semantic segmentation model according to claim 1, characterized in that The attention module comprises: A diffusion coding unit, whose input end is connected to the output end of the splicing module, is used to extract visual features of corresponding scales from the fusion features generated by the splicing module based on downsampling networks of different scales; wherein the diffusion coding unit includes a cascade of multiple downsampling networks with decreasing scales in sequence; A content attention unit, a first input end of which is connected to the output end of the text encoder, and a second input end of which is connected to the output end of the diffusion coding unit, for performing cross-modal feature fusion and alignment of the minimum-scale visual features and the text embedding features based on a content attention mechanism to obtain a fused and aligned feature; A diffusion decoding unit, a first input end of which is connected to the output end of the content attention unit, and a second input end of which is connected to the output end of the diffusion coding unit, is used to fuse the intermediate features with the visual features of the corresponding scale in each upsampling network to obtain the intermediate features input to the next upsampling network, and generate an image containing prediction noise based on the intermediate features obtained by the last upsampling network; wherein the initial intermediate features are the fused alignment features, and the diffusion decoding unit includes a plurality of cascaded upsampling networks with increasing scales in sequence, and the scale of the upsampling network corresponds to the scale of the downsampling network.
5. The semantic segmentation model according to claim 4, characterized in that The content attention unit comprises: A matrix generation subunit, a first input end of which is connected to the output end of the text encoder, a second input end of which is connected to the output end of the diffusion coding unit, and is used to generate a query matrix, a key matrix and a first value matrix based on the minimum scale visual feature, and to generate an intermediate matrix and a second value matrix based on the text embedding feature; a self-attention weight subunit, whose input end is connected to the first output end of the matrix generation subunit, and is used to calculate the self-attention weight of the visual feature based on the query matrix and the key matrix according to the scaled dot product attention mechanism; an interactive attention weight subunit, whose input end is connected to the second output end of the matrix generation subunit, and is used to calculate the interactive attention weight of the text feature to the visual feature based on the key matrix and the intermediate matrix according to the scaled dot product attention mechanism; A feature fusion subunit, whose first input end is connected to the output end of the self-attention weight subunit, whose second input end is connected to the output end of the interaction attention weight subunit, and whose third input end is connected to the third output end of the matrix generation subunit, is used to perform weighted summation of the first value matrix and the second value matrix according to the self-attention weight and the interaction attention weight to obtain a fused alignment feature.
6. The semantic segmentation model according to claim 4, characterized in that The attention module also includes: A denoising unit, whose input end is connected to the output end of the diffusion decoding unit, is used to denoise the image containing predicted noise generated each time within a preset denoising number of times, and use the denoised image as the noisy masked latent feature generated by the noise generator to extract visual features again.
7. The semantic segmentation model according to claim 1, characterized in that The image is a K-channel image, where K is a positive integer greater than 1, and the latent space decoder includes: A decoding unit, whose input end is connected to the output end of the attention module, and is used to obtain K initial predicted target masks according to the image containing prediction noise; The mask generation unit has an input end connected to the output end of the decoding unit and is used to perform weighted averaging on the K initial prediction target masks and use the weighted average as the final prediction target mask.
8. A semantic segmentation method, characterized in that: The semantic segmentation method comprises: Acquire an image to be segmented containing a target object and text information corresponding thereto; wherein the text information is used to characterize the category of the target object in the image; Encoding the text information through a text encoder of the image segmentation model to generate text embedding features, and encoding the image through an image encoder of the image segmentation model to generate image potential features; According to a preset diffusion time step, based on the text embedding feature, the image potential features are processed by the attention module and the splicing module of the image segmentation model at the preset diffusion time step until the noisy mask potential features generated at the last fusion are obtained, and after denoising, they are decoded by the latent space decoder of the image segmentation model to generate a predicted target mask for segmenting the target object; wherein, at each fusion: The image latent features and the noisy mask latent features are fused through the splicing module of the image segmentation model, and visual features are extracted from the fused features obtained after fusion through the attention module, and the visual features are aligned and fused with the text embedding features to update the noisy mask latent features at the next fusion; wherein the noisy mask latent features at the first fusion are random noises generated by the noise generator of the image segmentation model; The target object in the image is segmented based on the predicted target mask.
9. A semantic segmentation system, characterized in that: The semantic segmentation system comprises: an image encoder, a mask encoder, a noise generator, a text encoder, a noise summation module, a splicing module and an attention module, wherein: The output end of the mask encoder is connected to the first input end of the noise summing module; wherein the mask encoder is used to encode the target mask of the target object in the image into a mask potential feature; The output end of the noise generator is connected to the second input end of the noise adding module; wherein the noise generator is used to generate random noise; The output end of the image encoder is connected to the first input end of the splicing module; wherein the image encoder is used to encode the image to generate image potential features; The output end of the noise adding module is connected to the second input end of the splicing module; The output end of the text encoder is connected to the first input end of the attention module; wherein the text encoder is used to encode text information to generate text embedding features; The output end of the splicing module is connected to the second input end of the attention module; wherein the attention module is used to extract visual features from the fused features generated by the splicing module, and align and fuse the visual features and the text embedding features through the attention mechanism to generate noisy masked latent features.
10. A method for training a semantic segmentation model, characterized in that: The training method comprises: Acquire an image containing a target object and text information corresponding thereto, as well as a target mask of the target object in the image; wherein the text information is used to characterize the category of the target object in the image; Input the text information into the text encoder of the semantic segmentation model for encoding to generate text embedding features, input the image into the image encoder of the semantic segmentation model for encoding to generate image latent features, and input the target mask into the mask encoder of the semantic segmentation model for encoding to generate mask latent features; According to the preset diffusion time step, the attention module of the semantic segmentation model is fine-tuned and trained at each fusion to obtain the trained attention module of the semantic segmentation model; wherein, at each fusion: The masked latent features are noised and fused with the image latent features through the noise addition module and the splicing module of the semantic segmentation model, visual features are extracted from the fused features obtained after fusion through the attention module of the semantic segmentation model, the visual features and the text embedding features are aligned and fused based on the attention mechanism to generate noisy masked latent features, and the attention module of the semantic segmentation model is trained through the noisy masked latent features.
Citation Information
Patent Citations
Image processing method and device, computer, storage medium and program product
CN117252947A
Image generation model processing method and device, equipment, storage medium and product
CN118115622A
Traffic degraded image restoration method and device, terminal equipment and medium
CN119048401A
Speech synthesis method and device based on artificial intelligence, terminal equipment and medium
CN119068861A
Remote sensing image anaphora segmentation method and system
CN119380033A
Cited By
Medical image segmentation method and device, equipment and medium
CN121564003A
A medical image segmentation method, apparatus, device and medium
CN121564003B
Self-adaptive condition enhancement method and device for face generation, equipment and storage medium
CN121661435A