Target segmentation method, electronic equipment, readable medium and program product
By performing feature extraction and noise addition processing on the query image, combined with cascade model and category proxy optimization, the segmentation accuracy problem caused by category differences in small sample segmentation is solved, and higher segmentation accuracy and lower computing cost are achieved.
Patent Information
- Application Number
- CN202510231587.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
The existing small sample segmentation method is difficult to achieve accurate segmentation of target categories in the image to be segmented when the sample data is not high or there is a deviation, especially affected by intra-class variation and inter-class similarity problems.
By extracting the query image features, generating the first region semantics, and performing noise addition and denoising based on the semantics, using a cascade model to perform step by step denoising, combining a category token and a category agent for training loss optimization, improving segmentation accuracy.
It effectively avoids the reduction in segmentation accuracy caused by category differences, improves the segmentation accuracy of target category objects, and reduces the calculation cost.
Smart Images

Figure CN120163835A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and particularly to an object segmentation method, an electronic device, a readable medium and a program product. Background Art
[0002] Few-shot segmentation (FSS) is a pixel-level segmentation technology that uses a limited number of sample data to segment object category objects in the image to be segmented. Currently, some few-shot segmentation methods are usually implemented based on an image-mask decoding framework. This framework extracts the features of the image to be segmented, matches them with the features of a limited number of sample data, and decodes the matched features in the image to be segmented through a decoder to generate a mask for characterizing the segmentation result, thereby realizing the segmentation of object category objects.
[0003] However, the above solution highly depends on the quality of the sample data (such as whether the sample data is representative of the object category, etc.). When the quality of the sample data is not high or there are biases, the segmentation accuracy will be greatly reduced. For example, due to the within-class variation problem, the features of the object category objects in the image to be segmented may not be correctly matched with the features of the object category objects in the sample image. Another example is that due to the between-class similarity problem, other category objects in the image to be segmented may be wrongly segmented as object category objects. Thus, in the scenario of few-shot segmentation, it is difficult to accurately segment the object category objects in the image to be segmented. Summary of the Invention
[0004] In order to solve the problem that it is difficult to accurately segment the object category objects in the image to be segmented in the scenario of few-shot segmentation, embodiments of the present application provide an object segmentation method, an electronic device, a readable medium and a program product.
[0005] In a first aspect, an embodiment of the present application provides a target segmentation method, including: obtaining a query image, where the query image includes a first type of object; obtaining at least one sample data, where the number of at least one sample data meets a quantity condition, and each sample data includes a sample image and a segmentation result of the sample image, the sample image includes a first type of object, and the segmentation result corresponds to the first type of object; performing feature extraction on the query image to obtain a first image feature of the query image; based on the sample image and the segmentation result of the sample image, obtaining a second image feature, a foreground feature, and a background feature of each sample image; based on a first similarity information between the first image feature and the foreground feature and a second similarity information between the first image feature and the background feature, determining a first regional semantics of the query image, where the first regional semantics is used to represent whether each pixel in the query image belongs to the first type of object or the probability of belonging to the first type of object; performing noise addition processing on the query image according to the first regional semantics to obtain a noisy query image, where the noise level of noise addition to the foreground region in the noisy query image is less than the noise level of noise addition to the background region in the query image, the foreground region belongs to the first type of object, and the background region does not belong to the first type of object; performing feature extraction on the noisy query image to obtain a third image feature of the noisy query image; performing denoising processing on the third image feature through a denoising model to obtain a denoising feature of the query image; and performing segmentation decoding on the denoising feature to predict a segmentation result of the query image corresponding to the first type of object.
[0006] In this way, through the sample data corresponding to the image to be segmented, the first regional semantics of the image to be segmented is generated, and noise addition processing and denoising processing are performed on the image to be segmented according to the first regional semantics, so as to obtain the segmentation result of the target category object in the image to be segmented. The problem of the decrease in segmentation accuracy caused by category differences (such as intra-class variation and inter-class similarity) is avoided.
[0007] In a possible implementation of the above first aspect, the segmentation result of the sample image is a mask. Based on the sample image and the segmentation result of the sample image, obtaining a second image feature, a foreground feature, and a background feature of each sample image includes: performing feature extraction on the sample image to obtain a second image feature; based on the mask, performing pooling processing through a pooling model to obtain a foreground feature; and performing a clustering operation based on the complement code of the mask and the second image feature to obtain a background feature.
[0008] In a possible implementation of the above first aspect, performing noise addition processing on the query image according to the first regional semantics to obtain a noisy query image includes: obtaining an original noise, where the original noise corresponds to a pixel matrix of the same size as the first regional semantics; determining a second regional semantics based on the first regional semantics, where the second regional semantics is opposite to the first regional semantics; obtaining an adjusted noise by multiplying the original noise by the second regional semantics; and performing noise addition processing on the query image based on the adjusted noise to obtain a noisy query image.
[0009] In this way, the query image can be denoised by the generated first-region semantics, which can be used to characterize the possibility of the object belonging to the target category in the query image, improving the accuracy of segmenting the object of the target category.
[0010] In a possible implementation of the first aspect above, the denoising model includes n cascaded models, and through the denoising model, the third image feature is denoised to obtain the denoised feature of the query image, including: through n cascaded models, the third image feature is denoised based on the second image feature to obtain the denoised feature of the query image, where the second image feature is the input enhancement feature of the first cascaded model, the third image feature is the input denoised feature of the first cascaded model, the output denoised feature of the k-th cascaded model is the input denoised feature of the (k + 1)-th cascaded model, and the output enhancement feature of the k-th cascaded model is the input enhancement feature of the (k + 1)-th cascaded model; the output denoised feature of the n-th cascaded model is the denoised feature of the query image.
[0011] In this way, by gradually denoising the third image feature of the noisy query image, the accuracy of denoising is improved, and thus the accuracy of image segmentation is improved.
[0012] In a possible implementation of the first aspect above, each cascaded model includes a self-attention model, a cross-attention model, a feed-forward neural network model, and a noise prediction model. Denoising the third image feature through n cascaded models includes: obtaining the input denoised feature and input enhancement feature of the k-th cascaded model; enhancing the input enhancement feature through the self-attention model of the k-th cascaded model to obtain the output enhancement feature of the k-th cascaded model; denoising the input denoised feature through the self-attention model of the k-th cascaded model to obtain the first intermediate feature of the k-th cascaded model; denoising the first intermediate feature of the k-th cascaded model based on the output enhancement feature of the k-th cascaded model through the cross-attention model of the k-th cascaded model to obtain the second intermediate feature of the k-th cascaded model; performing non-linear processing on the second intermediate feature of the k-th cascaded model through the feed-forward neural network model of the k-th cascaded model to obtain the third intermediate feature of the k-th cascaded model; predicting the background noise of the third intermediate feature of the k-th cascaded model through the noise prediction model of the k-th cascaded model to obtain the predicted noise of the k-th cascaded model, where the background noise does not correspond to the first-class object; denoising the third intermediate feature of the k-th cascaded model based on the predicted noise of the k-th cascaded model to obtain the output denoised feature of the k-th cascaded model.
[0013] In this way, by means of the self-attention model, cross-attention model, feed-forward neural network model, and noise prediction model to perform step-by-step denoising processing on the third image feature of the noisy query image, the accuracy of denoising is improved, and thus the accuracy of image segmentation is improved.
[0014] In a possible implementation of the above first aspect, the denoising model includes a class token and a class proxy, where the class token is used to learn the features of the first type of object, and the class proxy includes a first class proxy corresponding to the first type of object. By using the denoising model to perform denoising processing on the third image feature, the denoised feature of the query image is obtained, including: by using the denoising model, performing denoising processing on the third image feature based on the class proxy, and learning the features of the first type of object based on the class token to obtain the denoised feature of the query image and the updated class token; determining the training loss based on the training parameters, where the training parameters include the updated class token, class proxy, and segmentation result; and updating the denoising model based on the training loss.
[0015] In a possible implementation of the above first aspect, the training loss includes a discriminative attribute learning loss, a random cross-proxy regularization loss, and a segmentation loss. The class proxy further includes a second class proxy corresponding to the second type of object, and the second type of object is different from the first type of object. Determining the training loss based on the training parameters includes: projecting the updated class token into the proxy space to obtain a class token with the same dimension as the first class proxy; obtaining the discriminative attribute learning loss based on the class token with the same dimension as the first class proxy and the first class proxy; pairing the first class proxy with the second class proxy to obtain the random cross-proxy regularization loss; obtaining the segmentation loss by comparing the segmentation result with the reference result; and obtaining the final training loss according to the discriminative attribute learning loss, random cross-proxy regularization loss, and segmentation loss.
[0016] In this way, updating the parameters of the denoising model according to the final training loss further improves the accuracy of image segmentation.
[0017] In a second aspect, an embodiment of the present application provides an electronic device, including one or more processors; one or more memories; and one or more programs are stored in one or more memories. When the one or more programs are executed by the one or more processors, the device executes the target segmentation method according to any one of the first aspect.
[0018] In a third aspect, an embodiment of the present application provides a computer-readable medium, on which instructions are stored. When the instructions are executed on an electronic device, the electronic device executes the target segmentation method according to any one of the first aspect.
[0019] Fourthly, an embodiment of the present application provides a computer program product, which includes computer instructions. When executed by an electronic device, the electronic device executes the target segmentation method according to any one of the first aspect. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following provides an illustration of the drawings.
[0021] Figure 1 According to some embodiments of the present application, a schematic diagram of few-shot segmentation based on an image-mask decoding framework is shown;
[0022] Figure 2 According to some embodiments of the present application, a schematic diagram of few-shot segmentation based on a noise-mask denoising framework is shown;
[0023] Figure 3 According to some embodiments of the present application, a schematic flowchart of a target segmentation method is shown;
[0024] Figure 4 According to some embodiments of the present application, a schematic structural diagram of a denoising model including three cascaded models is shown;
[0025] Figure 5 According to some embodiments of the present application, a schematic diagram of the influence of two parameters of a category object on the segmentation performance is shown;
[0026] Figure 6 According to some embodiments of the present application, a schematic diagram of the average interaction ratio comparison between the proxy attribute learning method and other methods is shown;
[0027] Figure 7 According to some embodiments of the present application, a schematic structural diagram of an electronic device applicable to a target segmentation method is shown. Detailed Embodiments
[0028] To facilitate those skilled in the art to understand the solutions in the embodiments of the present application, the following first explains some concepts and terms related to the embodiments of the present application.
[0029] 1. Few-shot Segmentation
[0030] Few-shot segmentation refers to the technology of segmenting target category objects from an image to be segmented through a deep learning model by using a small amount of sample data (for example, 1 to 5 pieces of sample data for each category). Taking the image-mask decoding framework model as an example, the image-mask decoding framework extracts features from the image to be segmented, matches them with the features in a small amount of sample data, segments the target category objects in the image to be segmented, and predicts and generates a corresponding query mask (or called a prediction mask). This query mask characterizes the pixel-level distribution area of the target category objects in the image to be segmented, thus completing the segmentation of the target category objects from the image to be segmented. However, due to the small total number of sample data, usually less than 5 pieces of sample data for each category, in the few-shot segmentation task, the small number of sample data has a great impact on the accuracy of the segmentation task.
[0031] In the few-shot segmentation scenario, the sample data set can correspond to the support set, which includes at least one category of sample data, and the number of sample data meets the quantity condition (usually 1 to 5 pieces of sample data for each category). The image of the sample data can be the support image in the support set. The support set includes the support image and the support mask. Among them, the support image is used for training or assisting the segmentation task. For example, it can be an RGB image. The support mask can be used as the annotation information of the support image to represent the segmentation result of the target category objects in the support image. For example, the support mask can be a black-and-white binary image, where the white area represents the area occupied by the target category objects in the support image, and the black area represents the background area in the support image (that is, the area other than the area occupied by the target category objects).
[0032] The image to be segmented can correspond to the query image in the query set. The query set includes the query image and the query mask. Among them, the query mask can be used as the annotation information of the query image to represent the segmentation result of the target category objects in the query image. For example, the support mask can be a black-and-white binary image, where the white area represents the area occupied by the target category objects in the query image, and the black area represents the background area in the query image (that is, the area other than the area occupied by the target category objects).
[0033] In some embodiments, the support set can be abstractly defined as S = (I s , M s ). Where S represents the support set, I s represents the support image, and M s represents the support mask. The query set is abstractly defined as Q = (I q , M q ). Where Q represents the query set, I q represents the query image, and M q represents the query mask. And the features extracted from the support image (abbreviated as support features) can be represented as F s, the features extracted from the query image (hereinafter referred to as query features) can be represented as F q . Taking the image-mask decoding framework model as an example, by extracting the feature F in the query image q , and matching it with the segmentation result in the support image, the target category object in the query image I q is segmented, and the query mask is predicted for the query image I q
[0034] It can be understood that in few-shot learning, the support set S can also be extended to K support pairs, that is, the above support set contains K pairs of support images and corresponding support masks. For example, in 5-shot learning, the support set S can be extended to 5 support pairs, that is, the number of sample data can be 5.
[0035] 2. Intra-class variation
[0036] Intra-class variation refers to the visual feature differences of objects of the same category in different images or different scenarios, which makes it difficult for the few-shot segmentation model to accurately match the features of the target category, resulting in the situation where the target category object is not correctly segmented.
[0037] For example, cats of the same category have differences in features such as color, texture, and pose among different samples. The few-shot segmentation model is difficult to learn the sample features corresponding to the category cat with a small number of samples, so it cannot correctly match the features of the cat in the image to be segmented with the features corresponding to the category cat in the sample image. Eventually, the segmentation result is inaccurate and the area occupied by the cat in the image to be segmented cannot be correctly identified. For example, in the few-shot segmentation model, the sample is a Garfield cat and the image to be segmented is a Ragdoll cat. Since the Garfield cat and the Ragdoll cat have differences in color, texture, and pose, it may not be possible to segment the area occupied by the Ragdoll cat in the image to be segmented based on the Garfield cat in the sample.
[0038] 3. Inter-class similarity
[0039] Inter-class similarity means that objects of different categories may have highly similar features visually, such as similar shapes, colors, textures, etc., which makes it difficult for the few-shot segmentation model to distinguish, resulting in incorrect segmentation.
[0040] For example, Ragdoll cats and dogs have similarities in features such as color, texture, and pose. When the target category object in the image to be segmented is a cat segmentation task and the image to be segmented contains a dog at the same time, the few-shot segmentation model may misclassify the part of the area occupied by the dog in the image to be segmented into the category cat, resulting in an incorrect segmentation of the target category object.
[0041] 4. Diffusion models
[0042] The diffusion model is a generative model, consisting of two stages: the forward process and the reverse process. Among them, the forward process can also be called the noise-adding process. In this stage, Gaussian noise is gradually added to the original sample data until the distribution of the original sample data approaches the standard Gaussian distribution (i.e., the original sample data becomes pure Gaussian noise data). Among them, the forward process can be abstractly defined as:
[0043]
[0044] where \(x\) t represents the noisy data at time step \(t\), \(x_0\) represents the original sample data, \(\epsilon\) t represents the noise that follows a Gaussian distribution (i.e., the above-mentioned Gaussian noise), represents the cumulative coefficient of the noise attenuation factor, and can be abstractly defined as:
[0045]
[0046] where \(\alpha\) i represents the noise attenuation factor, \(\beta\) i represents the coefficient used to control the noise schedule in the forward process, and \(\beta\) i \(\in(0,1)\).
[0047] The reverse process can be called the denoising process. In this stage, the original sample data is gradually restored or data similar to the distribution of the original sample data is generated by gradually removing the noise. Among them, the reverse process can be abstractly defined as:
[0048]
[0049] where \(x\) t-1 is the reconstructed noisy data at time step \(t - 1\), \(x\) t represents the noisy data at time step \(t\), represents the cumulative coefficient of the noise attenuation factor, \(\epsilon\) θ \((x\) t , t)\) represents the denoising network for denoising processing, \(z\) represents the random noise that follows a Gaussian distribution, represents the coefficient used to reconstruct the sample data after adjustment, and can be abstractly defined as:
[0050]
[0051] where \(\beta\) t represents the coefficient used to reconstruct the sample data in the reverse process.
[0052] Based on this, from a noisy data \(x\) tStarting from the sequence x t →x t-1 →…→x0, it can be gradually iteratively reconstructed, and finally the original sample data x0 can be obtained or data close to the original sample data x0 can be generated. It can be understood that the noisy data x t can be pure Gaussian noise data, which is not restricted here.
[0053] The following two small-sample segmentation schemes will be explained.
[0054] Figure 1 According to some embodiments of the present application, a schematic diagram of small-sample segmentation based on an image-mask decoding framework is shown. As Figure 1 shown, when completing the segmentation task under the image-mask decoding framework, first extract the image features of the input image I to be segmented, and then match the extracted image features with the sample features of the preset sample data to generate a matched result. The decoder decodes the matched result to generate a mask M (mask) for characterizing the segmentation result, so as to realize the segmentation of the target category object from the image I to be segmented.
[0055] Among them, matching and decoding the extracted image features with the sample features of the preset sample data can be completed by various learning methods. For example, it can be through a pixel-level feature matching method to match the features of each pixel in the image to be segmented with the corresponding pixel features in the sample data one by one to determine the category to which each pixel in the image belongs, and then decode the mask. It can also match and decode the features extracted from the image to be segmented with the preset sample features through other learning methods, which is not specifically limited here.
[0056] However, in the image-mask decoding framework, since the image features extracted from the image to be segmented are directly matched with the sample features of the preset sample data, it highly depends on the quality of the sample data. In the small-sample segmentation task, when the number of sample data is limited and there are easily category difference problems (such as intra-class variation problems, inter-class similarity problems), the segmentation accuracy will be greatly reduced.
[0057] Figure 2 According to some embodiments of the present application, a schematic diagram of small-sample segmentation based on a noise-mask denoising framework is shown. As Figure 2 shown, the noise-mask denoising framework can use the generation mechanism of the diffusion model for the image, inject the image I to be segmented into the denoising process in the form of a condition, guide the denoising direction of the pure Gaussian noise data ∈, generate a mask M for characterizing the segmentation result, and then realize the segmentation of the target category object.
[0058] However, the above noise-masking denoising framework requires the diffusion model to perform complex image feature fusion design. At the same time, denoising on the basis of pure noise data to generate a segmentation result representing the target category object will result in a large computational power consumption of the noise-masking denoising framework model, thus leading to a high computational cost for segmenting the target category object.
[0059] To solve the above problems, the present application provides a target segmentation method. Specifically, the method can extract the first image feature of the query image by performing feature extraction on the query image; obtain the second image feature, foreground feature, and background feature of each sample image based on the sample image and the segmentation result of the sample image; determine the first regional semantics of the query image based on the first similarity information between the first image feature and the foreground feature and the second similarity information between the first image feature and the background feature. The first regional semantics represent whether each pixel in the query image belongs to the target category object or the probability of belonging to the target category object. For example, the greater the probability, the greater the possibility that the query image belongs to the target category object. And perform noise addition processing (such as adding Gaussian noise) on the query image based on the first regional semantics to obtain a noisy image highlighting the target category object; extract the noisy image feature through a pre-trained backbone network model to obtain the third image feature of the noisy query image, and perform denoising processing on the third image feature through a denoising model to obtain the denoised feature of the query image. Finally, decode the denoised feature of the query image through a decoder to obtain the segmentation result of the target category object corresponding to the query image.
[0060] In this way, the first regional semantics of the image to be segmented are generated through the sample data corresponding to the image to be segmented, and noise addition processing and denoising processing are performed on the image to be segmented according to the first regional semantics to obtain the segmentation result of the target category object of the image to be segmented. The problem of decreased segmentation accuracy caused by category differences (such as intra-class variation and inter-class similarity) is avoided. At the same time, denoising is performed on the noisy image rather than pure noise data, reducing the computational cost of segmenting the target category object.
[0061] The target segmentation method mentioned in the present application will be described below.
[0062] In the embodiments of the present application, the image to be segmented corresponds to the query image, and the sample image corresponds to the support image. Figure 3 According to some embodiments of the present application, a flowchart of a target segmentation method is shown. It can be understood that Figure 3 The execution subject of the illustrated target segmentation method can be an electronic device.
[0063] S101: Obtain a query image and at least one sample data.
[0064] In some embodiments, the query images in the query set may correspond to the images to be segmented, and include a first type of object, where the first type of object is an object of the target category. For example, in the query image, the first type of object is the category of cat, and the second type of object may be the category of dog, where the category of cat is the object of the target category.
[0065] In some embodiments, the sample data set includes at least one type of sample data, where the quantity of at least one type of sample data meets the quantity condition. Each sample data includes a sample image and the segmentation result of the sample image. The sample image includes a first type of object, and the segmentation result corresponds to the first type of object.
[0066] In some embodiments, the query image may include a first type of object and a segmentation result, where the segmentation result of the query image corresponds to the first type of object. For example, the first type of object is the category of cat, and the segmentation result of the category of cat may be a black-and-white binary image, where the white area represents the area occupied by the category of cat in the query image, and the black area represents the background area in the query image (i.e., the area other than the area occupied by the category of cat).
[0067] It can be understood that the first type of object in the query image and the first type of object in the sample data both belong to the objects of the target category. For example, the first type of object in the query image may be a Garfield cat, and the first type of object in the sample data may be a Ragdoll cat. The first type of object in the query image and the first type of object in the sample data both belong to the category of cat.
[0068] S102: Extract features from the query image to obtain the first image feature of the query image.
[0069] In some embodiments, for the query image I q perform feature extraction to obtain the query feature F q (corresponding to the first image feature of the query image).
[0070] For example, the query image is of the category of cat and the category of dog. Based on the features of the category of cat and the category of dog in terms of color, texture, and pose, the features of the category of cat and the category of dog in the query image are extracted.
[0071] S103: Based on the sample image and the segmentation result of the sample image, obtain the second image feature, foreground feature, and background feature of each sample image.
[0072] It can be understood that the foreground feature and background feature mentioned in the embodiments of the present application may correspond to the first type of object.
[0073] In some embodiments, the foreground can be the presence of a first type of object in the image or the probability of the presence of a first type of object. The greater the probability, the greater the likelihood of the presence of a first type of object. The background can be the absence of a first type of object in the image or the probability of the absence of a first type of object. The greater the probability, the greater the likelihood of the absence of a first type of object. For example, in an image with a cat and a dog, if the first type of object is the cat, then the cat is the foreground, and the dog and the solid-color background (such as the sky, road, trees, etc.) are the background.
[0074] It can be understood that the second image feature of the sample image mentioned in the embodiments of the present application can be the support feature F of the sample image s . For example, if the sample images are images of the category cat and the category dog, based on the characteristics of the category cat and the category dog in terms of color, texture, and pose, the characteristics belonging to the category cat and the category dog in the sample images are extracted. Among them, the characteristics of the category cat and the category dog are the second image features.
[0075] Based on the segmentation result (corresponding to the mask) of the first type of object in the sample image, the foreground feature p of the sample image is obtained through pooling models FG .
[0076] The other regions outside the region occupied by the segmentation result of the first type of object in the sample image (corresponding to the complement of the mask) are clustered into multiple superpixels, and the background feature p of the sample image is obtained through clustering operations using the second image feature BG . The process of obtaining the background feature p of the sample image BG can be abstractly expressed as the following expression:
[0077] {p BG} = Cluster(F s ⊙ (1 - M s )) Formula (5)
[0078] where {p BG} represents the background feature of the sample image, Cluster represents the clustering operation, F s represents the second image feature of the sample image, ⊙ represents multiplication, M s represents the mask (corresponding to the segmentation result of the first type of object in the sample image), and (1 - M s ) is the complement of the mask.
[0079] S104: Determine the first regional semantics of the query image based on the first similarity information between the first image feature and the foreground feature and the second similarity information between the first image feature and the background feature.
[0080] In some embodiments, the first region semantics is used to represent whether each pixel in the query image belongs to the first type of object or the probability of belonging to the first type of object.
[0081] Based on the foreground feature p of the sample image FG and the query feature F q (corresponding to the first image feature of the query image), an initial foreground mask is obtained Obtaining the initial foreground mask The process can be abstractly represented by the following expression:
[0082]
[0083] where represents the initial foreground mask, cos represents the cosine similarity, and F q represents the query feature (corresponding to the first image feature of the query image), and p FG represents the foreground feature of the sample image.
[0084] The foreground mask M FG is used to estimate the probability that each pixel in the query image belongs to the foreground region (corresponding to the first similarity information). Specifically, the foreground mask can be obtained through an intra-query propagation (IQP) operation. Among them, the intra-query propagation is achieved by multiplying the query feature (corresponding to the first image feature of the query image) F q by itself, and the process of obtaining the foreground mask M FG can be abstractly represented by the following expression:
[0085]
[0086] where M FG represents the foreground mask, Normalize means scaling this term to [0,1], the self-multiplication of F q represents the intra-query propagation, represents the initial foreground mask.
[0087] The background mask M BG is used to estimate the probability that each pixel in the query image belongs to the background region (corresponding to the second similarity information). Specifically, it can be obtained by calculating the similarity between the query feature (corresponding to the first image feature of the query image) F q and the background features of each sample image and selecting the background feature with the highest similarity for matching to obtain the background mask M BG . The process of obtaining the background mask M BG can be abstractly represented by the following expression:
[0088]
[0089] Among them, M BG represents the background mask in the query image, Normalize represents scaling this item to [0, 1], max represents finding the maximum value, cos represents finding the cosine similarity, and F q represents the query feature, represents the background features of each sample image.
[0090] In some embodiments, the above foreground mask M FG and background mask M BG can be represented by a multi-valued image (such as a heat map) of the probability that each pixel in the sample image belongs to the first type of object, which will not be elaborated here.
[0091] Finally, by performing a difference truncation operation on the foreground mask M FG and background mask M BG the first region semantics of the query image is determined. It can be understood that the first region semantics calculated through the difference truncation operation can characterize the probability that each corresponding pixel in the query image belongs to the first type of object. For example, pixels with a larger difference have a greater probability of belonging to the first type of object, and pixels with a smaller difference have a smaller probability of belonging to the first type of object. The process of obtaining the first region semantics can be abstracted into the following expression:
[0092] M prior = max(0, M FG - M BG ) Formula (9)
[0093] Among them, M prior represents the first region semantics, max represents finding the maximum value, M FG represents the foreground mask, and M BG represents the background mask.
[0094] In this way, the generated first region semantics can be used to characterize the possibility of the object belonging to the target category in the query image, improving the accuracy of segmenting the object of the target category.
[0095] S105: Perform noise addition processing on the query image according to the first region semantics to obtain a noisy query image.
[0096] After generating the above first region semantics, noise addition processing can be performed on the query image based on the first region semantics. Specifically, the process of performing noise addition processing on the query image is as follows:
[0097] Establish a raw Gaussian noise ∈ ∈ R that is consistent with the size or pixel matrix dimension of the query image H×W(i.e., the original noise), sample the semantics of the first region. For example, an interpolation algorithm can be used to increase the number of pixels in the query image so that the dimension of the pixel matrix of the semantics of the first region is consistent with the original Gaussian noise. H×W is the dimension of the pixel matrix that makes up the query image, i.e., the size of the query image, where H represents the height of the query image and W represents the width of the query image.
[0098] Adjust the noise of the query image by multiplying the original Gaussian noise by the semantics of the second region. Among them, the semantics of the second region is opposite to that of the first region and can represent the possibility that each corresponding pixel in the query image belongs to the background region. By multiplying the original Gaussian noise by the semantics of the second region, the adjusted noise is obtained. The adjusted noise can be adjusted accordingly according to the possibility of each pixel belonging to the background region. For example, the noise level of adding noise to the foreground region in the noisy query image is less than the noise level of adding noise to the background region in the query image. Among them, the foreground region belongs to the first type of object, and the background region does not belong to the first type of object. The adjusted noise obtained can be abstractly expressed as the following expression:
[0099]
[0100] Among them, represents the adjusted noise, ∈ represents the original Gaussian noise, ⊙ represents multiplication, Upsample represents upsampling, M prior represents the semantics of the first region, (1 - Upsample(M prior )) represents the upsampled semantics of the second region.
[0101] Furthermore, perform image perturbation processing (i.e., noise addition processing) on the query image based on the adjusted noise to obtain a noisy query image Obtain a noisy query image It can be abstractly expressed as the following expression:
[0102]
[0103] Among them, represents the noisy query image, I q represents the query image, represents the noise attenuation accumulation coefficient, τ represents the perturbation timestamp coefficient, which is used to control the intensity of the image perturbation processing, represents the adjusted noise.
[0104] In this way, by adding noise to the query image, the high dependence on sample data during image segmentation is reduced, thereby reducing the class difference problems (such as within-class variation problems and between-class similarity problems) that occur during image segmentation, and further improving the accuracy of image segmentation.
[0105] S106: Extract features from the noisy query image to obtain the third image feature of the noisy query image.
[0106] In some embodiments, the noisy query image can be feature-extracted. Specifically, a pre-trained backbone network can be used to extract features from the noisy query image to obtain the third image feature of the noisy query image as the input for the subsequent denoising process. For example, the third image feature can be denoted as x τ .
[0107] S107: Denoise the third image feature through a denoising model to obtain the denoised feature of the query image.
[0108] In some embodiments, the denoising model can include n cascaded models. For example, n can be 3.
[0109] In some embodiments, the cascaded model includes a self-attention model, a cross-attention model, a feed-forward neural network model, and a noise prediction model.
[0110] In the embodiments of the present application, the denoising model is described by taking the denoising model including 3 cascaded models as an example.
[0111] Figure 4 According to some embodiments of the present application, a schematic structural diagram of the denoising model including 3 cascaded models is shown. As Figure 4 shown, the second image feature is the input enhancement feature of the first cascaded model, and the third image feature is the input denoising feature of the first cascaded model; the output denoising feature of the first cascaded model is the input denoising feature of the second cascaded model, and the output enhancement feature of the first cascaded model is the input enhancement feature of the second cascaded model; the output denoising feature of the second cascaded model is the input denoising feature of the third cascaded model, and the output enhancement feature of the second cascaded model is the input enhancement feature of the third cascaded model; the output denoising feature of the third cascaded model is the denoised feature of the query image.
[0112] The following describes how to obtain the enhancement feature of the cascaded model.
[0113] The second image feature of the sample image is enhanced through a self-attention model to obtain the enhancement feature of the first cascaded model. Among them, the self-attention model is used to identify the first type of object in the sample image and enhance the first type of object.
[0114] It can be understood that in some embodiments, during the denoising process, obtaining enhanced features can be achieved through a vision transformer (ViT) encoder. For example, in three embodiments where n = 3, the second image feature of the sample image passes through the self-attention model (SA) in the first cascaded model to obtain the enhanced feature of the first cascaded model; the enhanced feature of the first cascaded model passes through the self-attention model in the second cascaded model to obtain the enhanced feature of the second cascaded model; the enhanced feature of the second cascaded model passes through the self-attention model in the third cascaded model to obtain the enhanced feature of the third cascaded model. It can be represented by the following formula:
[0115]
[0116] Where l represents the number of layers of the cascaded model. represents the enhanced feature of the l-th cascaded model, represents the enhanced feature of the (l - 1)-th cascaded model, and SelfAttn represents the self-attention model processing.
[0117] Next, taking l = 1, which represents the output enhanced feature after self-attention processing of the first-level cascaded model as an example, the process of obtaining the denoised feature output by the first cascaded model will be described.
[0118] The third image feature of the noisy query image is denoised through the self-attention model to obtain the feature after self-attention model processing (corresponding to the first intermediate feature), which can be represented by the following formula:
[0119] x′ l = SelfAttn(x l ) + x l Formula (13)
[0120] Where, x l represents the third image feature of the noisy query image, x′ l represents the feature after self-attention model processing (corresponding to the first intermediate feature), and SelfAttn represents the self-attention model processing.
[0121] Based on the output enhanced feature F l s after self-attention processing of the first-level cascaded model, the feature x′ l (corresponding to the first intermediate feature) after self-attention model processing is denoised through the cross-attention model to obtain the feature x″ l (corresponding to the second intermediate feature), which can be represented by the following formula:
[0122]
[0123] Among them, x″ l represents the feature after cross-attention processing (corresponding to the second intermediate feature), represents the output enhanced feature after self-attention model processing in the first-level cascade model, and CrossAttn represents cross-attention model processing.
[0124] Through the feed-forward neural network model, the feature x″ after cross-attention processing l (corresponding to the second intermediate feature) is subjected to non-linear denoising processing to obtain the feature x″′ processed by the feed-forward neural network model l (corresponding to the third intermediate feature), which can be expressed by the following formula:
[0125] x″′ l = FFN(x″ l ) + x″ l Formula (15)
[0126] Among them, x″′ l represents the feature processed by the feed-forward neural network model (corresponding to the third intermediate feature), and x″ l represents the feature after cross-attention processing (corresponding to the second intermediate feature), and FFN represents feed-forward neural network model processing.
[0127] The background noise of the feature processed by the feed-forward neural network model (corresponding to the third intermediate feature) is predicted by the noise prediction model. Among them, the noise prediction model can be a convolutional block model, which is used to predict the background noise in the feature processed by the feed-forward neural network model (corresponding to the third intermediate feature) that does not correspond to the first type of object, and the predicted noise can be expressed by the following formula:
[0128] ∈ = NoisePred(x″′ l ) = ReLU(Conv(x″′ l )) Formula (16)
[0129] Among them, ∈ represents the predicted noise, x″′ l represents the feature processed by the feed-forward neural network model (corresponding to the third intermediate feature), NoisePred represents noise prediction processing, Conv represents convolutional operation, and ReLU represents activation operation.
[0130] Based on the predicted noise of the noise prediction model, the feature processed by the feed-forward neural network model (corresponding to the third intermediate feature) is denoised to obtain the output denoised feature of the first-level cascade model, which can be expressed by the following formula:
[0131] x l+1 = x″′ l - ∈ Equation (17)
[0132] where ∈ represents the prediction noise, and x l+1 represents the denoised feature output by the first cascaded model, and x″′ l represents the feature processed by the feedforward neural network model (corresponding to the third intermediate feature).
[0133] It can be understood that in the second cascaded model, the denoised feature output by the first cascaded model is the input denoised feature of the second cascaded model, and the enhanced feature output by the first cascaded model is the input enhanced feature of the second cascaded model.
[0134] In the third cascaded model, the denoised feature output by the second cascaded model is the input denoised feature of the third cascaded model, and the enhanced feature output by the second cascaded model is the input enhanced feature of the third cascaded model, obtaining the denoised feature output by the third cascaded model. Among them, the denoised feature output by the third cascaded model is the denoised feature of the query image.
[0135] It can be understood that the above three cascaded models are only examples, and other numbers of cascaded models can also be used to denoise the third image feature of the noisy query image to obtain the denoised feature of the query image. For example, it can be four cascaded models, five cascaded models, and no specific limitation is made here.
[0136] In this way, by denoising the third image feature of the noisy query image layer by layer through the denoising model, the denoising accuracy is improved, and thus the image segmentation accuracy is improved.
[0137] S108: Perform segmentation decoding on the denoised feature to predict the segmentation result of the query image corresponding to the first type of object.
[0138] Perform segmentation decoding on the denoised feature through the decoder to predict the segmentation result of the query image corresponding to the first type of object.
[0139] In this way, the segmentation of the target category object in the query image is completed, avoiding the problem of reduced segmentation accuracy caused by category differences (such as intra-class variation and inter-class similarity). At the same time, denoising is performed on the basis of the noisy query image rather than pure noise data, reducing the computational cost of segmenting the target category object.
[0140] In some embodiments, to improve the accuracy of the segmentation result, the above denoising model can be trained to obtain an updated denoising model. The updated denoising model denoises the query image, which can further improve the classification accuracy of the denoising model, and thus obtain a more accurate image segmentation result. For example, the backpropagation process is performed according to the training loss to update the parameters of the denoising model, thereby obtaining an updated denoising model.
[0141] In some embodiments, the training loss includes discriminative attribute learning loss, stochastic cross-agent regularization loss, and segmentation loss. The training loss is determined by training parameters, where the training parameters include the updated class token, class agent, and segmentation result.
[0142] First, the process of calculating the discriminative attribute learning loss is introduced below.
[0143] Before the denoising model performs denoising processing, the original class token z can be inserted. And the foreground feature p of the sample image can be added to the original class token z FG , as a class prior, to guide the learning of the original class token z. It can be expressed by the following formula:
[0144] z′ = z + p FG Formula (18)
[0145] where z′ represents the class token added with the foreground feature p FG of, z represents the original class token, and p FG represents the foreground feature of the sample image.
[0146] Subsequently, the class token z′ and the embedding token enter the denoising model together. Among them, the embedding token is the third image feature of the noisy query image, which can be represented as x for example τ .
[0147] As mentioned above, the denoising model can denoise the third image feature of the noisy query image based on the second image feature of the sample image to obtain the denoised feature of the query image. In the embodiment where the class token z′ and the embedding token enter the denoising model together, during the denoising process of the denoising model, through the conditional interaction between the class token z′ and the embedding token, the class token z′ can gradually learn the feature information of the foreground object (i.e., the first type of object) to obtain an updated class token The above process can be expressed by the following formula:
[0148]
[0149] where, is the updated class token, z′ is the class token, and x τ is the embedding token, and F sis the second image feature of the sample image, is the denoising feature of the query image, and Denoise represents the denoising process.
[0150] The updated class token output is projected into the proxy space through a three-layer multi-layer perceptron (MLP), so that the updated class token has the same spatial dimension as the first-class proxy s c where the first-class proxy s c is a type of class proxy, representing the overall features of the first-type objects in the sample image. By calculating the updated class token projected into the three-layer perceptron and the Euclidean distance of the first-class proxy s c the discriminative attribute learning loss is obtained used to represent the difference between the features of the first-type objects in the query image and the first-class proxy s c and can be expressed by the following formula:
[0151]
[0152] where, is the discriminative attribute learning loss, is the updated class token projected into the proxy space through the perceptron, and s c is the first-class proxy.
[0153] The process of calculating the random cross-proxy regularization loss is introduced below.
[0154] To avoid the problem of inter-class similarity caused by the proximity or overlap between different classes in the sample image, the embodiment of the present application also proposes to introduce a cross-proxy regularization loss to represent the distance loss of proximity or overlap between classes. Among them, σ is the distance boundary value for maintaining the degree of freedom between classes.
[0155] In some embodiments, the parameters of the denoising model include class proxies. The class proxies include class proxies corresponding to multiple class objects. For example, it includes the first-class proxy s c corresponding to the first-type objects, and may also include the second-class proxy corresponding to the second-type objects. To calculate the distance loss of proximity or overlap between classes, it is necessary to traverse all combinations of classes and calculate the distance loss between classes. For example, a combination may include the class proxy corresponding to the object class c1 and the class proxy corresponding to the object class c2. Among them, the object class c1 and the object class c2 are any two classes among the multiple classes. The formula for calculating the distance loss between classes can be expressed as:
[0156]
[0157] Among them, represents the distance loss between categories. C is the total number of categories in the category proxies for multiple categories. and represent the category proxies of object category c1 and object category c2 respectively. σ represents the distance boundary value between object category c1 and object category c2.
[0158] Based on the above formula (21), the time complexity for calculating the distance loss between all categories can be determined to be O(C 2 d 2 ).
[0159] It can be understood that the above calculation of the distance loss between categories is an optional scheme. In some other embodiments, the final training loss may not include the distance loss between categories.
[0160] In some embodiments, to reduce the time complexity, a random cross-proxy regularization loss is introduced (that is, randomly select a category proxy from non-target category proxies to pair with the first category proxy s c as the distance loss between categories. Among them, the category object corresponding to the non-target category proxy and the first category object corresponding to the first category proxy s c are different).
[0161] It can be understood that the first category proxy s c is the target category proxy, which is a parameter for identifying the target category object in the denoising model. The formula for calculating the random cross-proxy regularization loss is as follows:
[0162]
[0163] Among them, is the random cross-proxy regularization loss, s c is the first category proxy, s z is the non-target category proxy, and σ represents the distance boundary value between the target category proxy s c and the non-target category proxy s z .
[0164] Randomly sampling a category from the non-target category proxy s z can be done by random sampling, and the sampling probability of obtaining the non-target category proxy can be:
[0165]
[0166] Among them, sc is the first - category proxy, s i is a randomly selected non - target - category proxy.
[0167] Based on the above formulas (22) and (23), it can be determined that the time complexity of calculating the random cross - proxy regularization loss is O(d 2 ).
[0168] In some embodiments, through the foregoing S101 - S108, the segmentation result of the predicted query image corresponding to the first - type object can be obtained, and according to the difference between the segmentation result of the predicted query image corresponding to the first - type object and the reference result, the segmentation loss is obtained
[0169] The process of calculating the final training loss is introduced below.
[0170] Training loss is composed of the sum of the segmentation loss the discriminative attribute learning loss and the random cross - proxy regularization loss and can be expressed by the following formula:
[0171]
[0172] where, represents the training loss, represents the segmentation loss, represents the discriminative attribute learning loss, represents the random cross - proxy regularization loss.
[0173] It can be understood that after obtaining the training loss , the back - propagation process can be performed based on the training loss to update the parameters of the denoising model, and the updated denoising model is obtained.
[0174] Taking the parameters in the denoising model as the first - category proxy as an example, a specific embodiment of updating the first - category proxy is introduced below.
[0175] In the embodiment of training the trained denoising model according to the initial denoising model, the target - category proxy in the initial denoising model can be randomly initialized parameters and is continuously updated with each round of training.
[0176] During the training of the denoising model, the smoothing factor α can be determined based on the training loss, and then the first - category proxy is updated based on the smoothing factor α. The process of updating the first - category proxy can be expressed by the following formula:
[0177]
[0178] where, s cRepresents the first category agent, Represents the updated category token projected into the agent space through the perceptron, used to make the category token and the first category agent s c be in the same dimension, and α represents adjusting the first category agent s according to the final training loss c 's smoothing factor, used to control the update speed of the first category agent s c
[0179] In some embodiments, the final training loss can be compared with a threshold, and the smoothing factor α is determined according to the comparison result. For example, if the final training loss is less than a certain threshold, it is predicted that the first category agent s c is relatively accurate, and at this time the α value is close to 1; if the final training loss is greater than a certain threshold, it is predicted that this first category agent is not accurate enough, and at this time the α value is close to 0.
[0180] It can be understood that in the next denoising process, the updated denoising model can be used to denoise the third image feature. For example, in the process of using a feedforward neural network model for denoising, the updated feedforward neural network model can be used for non-linear processing to obtain the features processed by the feedforward neural network model.
[0181] In this way, according to the segmentation loss the discriminative attribute learning loss and the random cross-agent regularization loss the final training loss is obtained, and the parameters of the denoising model are updated according to the final training loss, further realizing accurate target segmentation.
[0182] Next, in combination with Table 1, the target segmentation method (such as the DeFSS method) in the embodiments of the present application is compared with other methods (such as HSNet
[18] , DCAMA
[17] , HDMNet
[12] , SCCAN
[25] , MSI[9], RiFeNet[1], ABCB
[30] , etc.) in PASCAL-5 i The segmentation effects of the dataset are described contrastively. As shown in Table 1, under the conditions of using ResNet50 and ResNet101 as the backbone networks respectively, the segmentation effect performances of different target segmentation methods in single-sample learning (that is, each category only uses 1 sample image for learning) and 5-sample learning (each category uses 5 sample images for learning) are evaluated.
[0183] Among them, 0-fold, 1-fold, 2-fold, and 3-fold respectively represent in PASCAL-5 i Experiments were conducted on four different subsets (folds) of the dataset. Each time, three subsets were used for training, and the remaining one subset was used for testing. The mean intersection over union (MIoU) is an indicator to measure the accuracy of target segmentation. The higher the value, the better the segmentation effect. Among them, mIoU is the average of the results of subsets 0, 1, 2, and 3.
[0184] As shown in Table 1, when ResNet50 is used as the backbone network, the method of DeFSS achieves an average intersection over union of 71.0 in the single-sample task, which is higher than other methods (such as 70.6 for ABCB and 69.4 for HDMNet); in the five-sample task, the average intersection over union of the DeFSS method is 73.8, slightly higher than that of ABCB (73.6); when ResNet101 is used as the backbone network, DeFSS reaches 72.4 in the single-sample task and 74.8 in the five-sample task, also outperforming most target segmentation methods.
[0185] It can be seen from Table 1 that five-sample training provides more reference information than single-sample training, and the segmentation performance of all methods in the five-sample setting is higher than that in the single-sample setting. The single-sample performance of the target segmentation method (such as the DeFSS method) in the embodiments of this application has a greater improvement compared to the five-sample, indicating that it performs better in the case of scarce sample data.
[0186] Table 1
[0187]
[0188] Table 2 is a schematic table of the segmentation performance of the target segmentation method (such as the method of DeFSS) proposed in the embodiments of this application and other methods on the COCO-20 i dataset. Among them, the COCO-20 i dataset contains more categories and more complex image backgrounds. As shown in Table 2, the comparison is carried out under the condition of using ResNet50 and ResNet101 as the backbone network, and the performance of different target segmentation methods in single-sample learning (that is, only 1 sample image per category is used for learning) and 5-sample learning (5 sample images per category are used for learning) is evaluated respectively.
[0189] As shown in Table 2, when ResNet50 is used as the backbone network, the mean intersection over union (mIoU) of the DeFSS method in the one-shot task reaches 50.8, which is higher than other methods (such as 50.0 for ABCB and 50.0 for HDMNet); in the five-shot task, the mIoU of the DeFSS method is 56.7, higher than other methods (such as 55.1 for ABCB and 56.1 for HDMNet). When ResNet101 is used as the backbone network, the DeFSS method reaches 52.0 in the one-shot task and 59.0 in the five-shot task, also outperforming most object segmentation methods.
[0190] Table 2
[0191]
[0192]
[0193] As can be seen from Table 2, the object segmentation method (such as the DeFSS method) in the embodiments of the present application has higher processing performance for data segmentation than most object segmentation methods.
[0194] From the comparison between Table 1 and Table 2, it can be seen that since the COCO-20^i dataset contains more complex image backgrounds, the object segmentation method (such as the DeFSS method) proposed in the present application has a more significant improvement on the COCO-20^i dataset, indicating that it can better adapt to complex scenarios and has stronger generalization ability. At the same time, the one-shot performance improvement of the object segmentation method (such as the DeFSS method) proposed in the present application on the COCO-20^i dataset is more obvious, indicating that the denoising performance has stronger adaptability in small-sample tasks.
[0195] The following describes the comparison of segmentation performance for different noise addition methods and methods for generating the semantics of the first region.
[0196] Table 3 shows the comparison of the effects of different perturbation methods and prior generators on image segmentation performance. Among them, the perturbation method corresponds to the method of noise addition processing. The method of the first region semantics corresponds to different prior generators. As shown in Table 3, when the image perturbation method is removed, the mean intersection over union (mIoU) is 68.8 at this time; noise is added by adding a linear noise-image combination (corresponding to a linear combination), and the mIoU is 69.8 at this time; PFENet is used as the method for generating the semantics of the first region and combined with adaptive perturbation, and the mIoU is 70.1 at this time; SCCAN is used as the method for generating the semantics of the first region and combined with adaptive perturbation, and the mIoU is 70.4 at this time. Using the object segmentation method (such as DeFSS) proposed in the present application as the method for generating the semantics of the first region and combined with adaptive noise addition, the mIoU reaches 71.0 at this time, and the segmentation performance is the best among all methods.
[0197] Table 3
[0198] Perturbation method Prior generator mIoU None / 68.8 Linear combination / 69.8 Adaptive PFENet 70.1 Adaptive SCCAN 70.4 Adaptive DeFSS 71.0
[0199] As can be seen from Table 3, the adaptive perturbation method performs better in all tests compared to the linear combination and non-perturbation methods. The adaptive perturbation method can better segment foreground objects in complex scenarios. At the same time, the method of generating the first region semantics has a significant impact on the performance of the small-sample model. In particular, the DeFSS method, as a method of generating the first region semantics from foreground / background features, can significantly improve the performance of image segmentation.
[0200] Table 4 is a schematic table showing the influence of the average intersection over union at different perturbation scales. It can be seen that when τ = 400, the average intersection over union is the largest, that is, the performance of image segmentation is the best. It can be concluded that a moderate noise perturbation achieves a balance: on the one hand, it introduces an appropriate amount of randomness to help find the global optimal segmentation performance during the noise addition process; on the other hand, it better preserves the information of the image to be segmented to support the association and matching of the image to be segmented.
[0201] Table 4
[0202] Perturbation timestamp τ mIoU 0 68.8 200 70.5 400 71.0 600 69.3 800 67.7
[0203] Table 5 shows the results of the component ablation experiment of the denoising model. The denoising model includes a denoising network and discriminative attribute learning (corresponding to class tokens, class proxies), and the influence of different components on the average intersection over union performance is evaluated.
[0204] As shown in Table 5 below, when the conditional ViT encoder and the noise prediction model (noisepred) are not introduced, and discriminative attribute learning is not added, the average intersection over union is 68.1;
[0205] When the conditional ViT encoder is introduced to represent enhanced features, the average intersection over union is 69.6;
[0206] When the conditional ViT encoder is introduced to represent enhanced features and the noise prediction model is added for double denoising, the average intersection over union is 70.0;
[0207] When the conditional ViT encoder and the noise prediction model are introduced for double denoising and discriminative attribute learning is added to optimize the denoising model through attribute learning, the average intersection over union is 70.6;
[0208] When the conditional ViT encoder and the noise prediction model are introduced for double denoising and one-hot code labels are used for classification learning, the average intersection over union is 65.7;
[0209] While introducing the conditional ViT encoder and the noise prediction model for dual denoising, discriminative attribute learning and stochastic cross-agent regularization are added to optimize the denoising model, with the mean intersection over union being 71.0.
[0210] Table 5
[0211]
[0212] As can be seen from Table 5, when introducing the conditional ViT encoder and the noise prediction model for dual denoising, adding discriminative attribute learning and stochastic cross-agent regularization to optimize the denoising model, the mean intersection over union reaches the best effect, and the segmentation effect is the best.
[0213] Figure 5 Schematic diagram of the influence of two parameters (class agent and regularization distance boundary σ) of the class object on the segmentation performance. Figure 5 The left subfigure shows the influence of the class agent (corresponding to the complexity of the class) on the mean intersection over union performance. As Figure 5 shown in the left subfigure, as the class agent increases, the mean intersection over union shows a trend of increasing first and then decreasing. When using ResNet50 as the backbone network, the optimal dimension of the class agent is 1024, and the segmentation performance of the model is the best at this time. For ResNet101 as the backbone network, the optimal dimension of the class agent is 2048, and the segmentation performance of the model reaches the best at this time.
[0214] From Figure 5 the left subfigure, it can be seen that different backbone networks have differences in segmentation ability. ResNet101 is a larger network with stronger semantic representation ability, so a more complex class agent (dimension of 2048) can obtain better segmentation performance; while ResNet50 is smaller, and choosing a lower class agent (dimension of 1024) can instead obtain the best performance.
[0215] Figure 5 The right subfigure shows the influence of the regularization distance boundary σ on the mean intersection over union performance. As Figure 5 shown in the right subfigure, the mean intersection over union shows a trend of increasing first and then decreasing with the change of σ, but its change range is small, indicating that the regularization distance boundary σ has a small influence on the segmentation performance. Figure 5 As can be seen from the right subfigure, a smaller σ is more beneficial to ResNet101. A smaller distance boundary can avoid excessive contraction of the distance between classes, so that the model can maintain sufficient differences between classes and adapt to complex tasks. It provides a looser constraint, enabling a stronger backbone network to effectively distinguish the differences between different classes.
[0216] Figure 6Shows a schematic diagram of the average interaction ratio comparison of the method for learning proxy attributes (corresponding to attribute learning) and other methods under different sample conditions. By comparing three methods: category-based attribute learning (the target segmentation method exemplified in this application), one-hot code-based classification learning, and label text-based conditional learning. From Figure 6 It can be seen that the test is carried out by reducing the number of base classes. As the number of base classes decreases, the segmentation performance of the small-sample segmentation model also decreases. When the number of base classes decreases, the segmentation performance of all three methods decreases. However, the attribute learning method shows the smallest decrease, indicating its best generalization ability.
[0217] Exemplarily, Figure 7 According to some embodiments of the present application, a schematic diagram of the structure of an electronic device 100 is shown.
[0218] As Figure 7 shown, the electronic device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output (I / O) device 105, and a system control logic 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output device 105. Among them:
[0219] The processor 101 can be used to control the electronic device 100 to execute the target segmentation method of the present application. Among them, the processor 101 can include one or more processing units. For example, it can include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, or a processing module or processing circuit of a field programmable gate array (FPGA). The processor 101 can include one or more single-core or multi-core processors.
[0220] The processor 101 can be used to control the electronic device 100 to execute the target segmentation method of this application. Among them, the processor 101 can include one or more processing units. For example, it can include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, or a processing module or processing circuit of a field programmable gate array (FPGA). The processor 101 can include one or more single-core or multi-core processors. In some embodiments, the processor 101 can be used to execute the target segmentation method in the embodiments of the present invention.
[0221] The system memory 102 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory 102 is used to temporarily store data and / or instructions.
[0222] The non-volatile memory 103 can include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 can include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 can also be a removable storage medium, such as a secure digital (SD) memory card, etc.
[0223] In particular, the system memory 102 and the non-volatile memory 103 can respectively include: a temporary copy and a permanent copy of the instruction 107. The instruction 107 can include: when executed by the processor 101, enabling the electronic device 100 to implement the target segmentation method provided in the embodiments of this application.
[0224] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the electronic device 100, and thus communicating with any other suitable device via one or more networks. In some embodiments, the communication interface 104 may be integrated with other components of the electronic device 100. For example, the communication interface 104 may be integrated in the processor 101. In some embodiments, the electronic device 100 may communicate with other devices via the communication interface 104.
[0225] The input / output device 105 may include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. The user may interact with the electronic device 100 via the input / output device 105.
[0226] The system control logic 106 may include any suitable interface controller to provide any suitable interface to other modules of the electronic device 100. For example, in some embodiments, the system control logic 106 may include one or more memory controllers to provide an interface connected to the system memory 102 and the non-volatile memory 103.
[0227] In some embodiments, at least one of the processors 101 may be packaged with the logic of one or more controllers for the system control logic 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may also be integrated with the logic of one or more controllers for the system control logic 106 on the same chip to form a SoC.
[0228] It can be understood that Figure 7 The structure of the illustrated electronic device 100 is only an example. In other embodiments, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0229] In some embodiments, the embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores at least one computer program instruction, at least one segment of program, a code set, or an instruction set. The at least one computer program instruction, at least one segment of program, a code set, or an instruction set is loaded and executed by the electronic device to implement the target segmentation method provided by each of the above method embodiments.
[0230] In some embodiments, the embodiments of the present application also provide a computer program product. The program product includes computer instructions. When executed by the electronic device, the electronic device executes the target segmentation method provided by each of the above method embodiments.
[0231] Embodiments of the mechanisms disclosed in this application may be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application may be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0232] The program code may be applied to the input instructions to perform the various functions described in this application and generate output information. The output information may be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0233] The program code may be implemented in a high-level procedural language or an object-oriented programming language in order to communicate with the processing system. When necessary, the program code may also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any particular programming language. In either case, the language may be a compiled language or an interpreted language.
[0234] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more transitory or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or via other computer-readable media. Thus, machine-readable media may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, CD-ROMs, magneto-optical disks, read only memory (ROM), random access memory (RAM), erasable programmable read only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms via the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0235] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0236] It should be noted that each unit / module mentioned in the device embodiments of this application is a logical unit / module. Physically, a logical unit / module may be a physical unit / module, a part of a physical unit / module, or may be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / module themselves is not the most important. The combination of the functions implemented by these logical units / module is the key to solving the technical problems proposed in this application. In addition, to highlight the innovative part of this application, the above device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that there are no other units / modules in the above device embodiments.
[0237] It should be noted that in the examples and the description of this patent, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0238] Although this application has been illustrated and described by reference to certain embodiments thereof, those of ordinary skill in the art should understand that various changes may be made therein in form and detail without departing from the spirit and scope of this application.
Claims
1. A target segmentation method, characterized in that: The method comprises: Acquire a query image, where the query image includes a first category of objects; Acquire at least one sample data, wherein the number of the at least one sample data satisfies a quantity condition, each of the sample data comprises a sample image and a segmentation result of the sample image, the sample image comprises a first category of objects, and the segmentation result corresponds to the first category of objects; Performing feature extraction on the query image to obtain a first image feature of the query image; Based on the sample images and the segmentation results of the sample images, obtaining a second image feature, a foreground feature, and a background feature of each of the sample images; Determining, based on first similarity information between the first image feature and the foreground feature and second similarity information between the first image feature and the background feature, a first region semantics of the query image, wherein the first region semantics is used to characterize whether each pixel in the query image belongs to the first category of objects or a probability of belonging to the first category of objects; Noising the query image according to the first region semantics to obtain a noisy query image, wherein a noise level of a foreground region in the noisy query image is less than a noise level of a background region in the query image, the foreground region belongs to the first category of objects, and the background region does not belong to the first category of objects; Performing feature extraction on the noisy query image to obtain a third image feature of the noisy query image; Performing denoising processing on the third image feature by using a denoising model to obtain a denoising feature of the query image; The denoising features are segmented and decoded to predict a segmentation result of the query image corresponding to the first category of objects.
2. The method according to claim 1, characterized in that The segmentation result of the sample image is a mask, The obtaining, based on the sample image and the segmentation result of the sample image, the second image feature, the foreground feature and the background feature of each sample image comprises: Extracting features from the sample image to obtain the second image features; Based on the mask, performing pooling processing through a pooling model to obtain the foreground feature; A clustering operation is performed based on the complement of the mask and the second image feature to obtain the background feature.
3. The method according to claim 1, characterized in that The step of performing noise processing on the query image according to the first region semantics to obtain a noisy query image includes: Acquire original noise, wherein the original noise and the first region semantically correspond to a pixel matrix of the same size; Determine a second region semantics based on the first region semantics, wherein the second region semantics is opposite to the first region semantics; Obtaining adjusted noise by multiplying the original noise with the second region semantics; The query image is subjected to the noise adding process based on the adjusted noise to obtain the noisy query image.
4. The method according to claim 3, characterized in that The denoising model includes n cascade models, and the denoising model is used to perform denoising processing on the third image feature to obtain the denoising feature of the query image, including: By using the n cascade models, the third image feature is denoised based on the second image feature to obtain a denoised feature of the query image, wherein the second image feature is an input enhancement feature of the first cascade model, and the third image feature is an input denoised feature of the first cascade model; The output denoising features of the k-th cascade model are the input denoising features of the k+1-th cascade model, and the output enhancement features of the k-th cascade model are the input enhancement features of the k+1-th cascade model; The output denoising features of the nth cascade model are the denoising features of the query image.
5. The method according to claim 4, characterized in that Each cascade model includes a self-attention model, a cross-attention model, a feedforward neural network model, and a noise prediction model. The denoising process of the third image feature by using the n cascade models includes: Obtain the input denoising features and input enhancement features of the kth cascade model; The input enhancement feature is enhanced by the self-attention model of the k-th cascade model to obtain the output enhancement feature of the k-th cascade model; Performing denoising processing on the input denoising feature through the self-attention model of the k-th cascade model to obtain a first intermediate feature of the k-th cascade model; By using the cross attention model of the kth cascade model, based on the output enhancement feature of the kth cascade model, the first intermediate feature of the kth cascade model is denoised to obtain the second intermediate feature of the kth cascade model; Performing nonlinear processing on the second intermediate feature of the kth cascade model through a feedforward neural network model of the kth cascade model to obtain a third intermediate feature of the kth cascade model; Predicting the background noise of the third intermediate feature of the kth cascade model by using the noise prediction model of the kth cascade model to obtain the predicted noise of the kth cascade model, wherein the background noise does not correspond to the first type of object; Based on the prediction noise of the k-th cascade model, the third intermediate feature of the k-th cascade model is denoised to obtain the output denoised feature of the k-th cascade model.
6. The method according to claim 5, characterized in that The denoising model includes a category token and a category agent, wherein the category token is used to learn the characteristics of a first category of objects, and the category agent includes a first category agent corresponding to the first category of objects. The step of performing denoising processing on the third image feature by using a denoising model to obtain the denoised feature of the query image includes: Performing denoising processing on the third image features based on the category agent through the denoising model, and learning the features of the first category of objects based on the category tokens, to obtain denoised features of the query image and updated category tokens; determining a training loss based on training parameters, wherein the training parameters include the updated category token, the category proxy, and the segmentation result; The denoising model is updated based on the training loss.
7. The method according to claim 6, characterized in that The training loss includes a discriminative attribute learning loss, a random cross-agent regularization loss, and a segmentation loss, the category agent also includes a second category agent corresponding to a second category of objects, and the second category of objects is different from the first category of objects, The determining of the training loss based on the training parameters includes: Projecting the updated category tokens into the proxy space to obtain category tokens with the same dimension as the first category proxy; Obtaining a discriminative attribute learning loss based on the category token having the same dimension as the first category agent and the first category agent; Pairing the first category of agents with the second category of agents to obtain a random cross-agent regularization loss; By comparing the segmentation result with a reference result, a segmentation loss is obtained; A final training loss is obtained according to the discriminative attribute learning loss, the random cross-agent regularization loss, and the segmentation loss.
8. An electronic device, characterized in that: The device comprises one or more processors; one or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the device executes the target segmentation method described in any one of claims 1 to 7.
9. A computer-readable medium, characterized in that The readable medium stores instructions, which, when executed on an electronic device, enable the electronic device to execute the object segmentation method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The program product includes computer instructions, and when executed by an electronic device, the electronic device performs the object segmentation method according to any one of claims 1 to 7.