Image segmentation mask refinement using diffusion model
By inputting the initial image segmentation mask to the trained diffusion model, and using the iterative process of the diffusion model to generate a refined mask, the problem of insufficient accuracy and details of the image segmentation mask in the prior art is solved, and higher image segmentation accuracy and applicability are achieved.
Patent Information
- Application Number
- CN202411816611.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-12-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to generate accurate and detailed segmentation masks during image segmentation, especially when processing high-resolution images and correcting errors in rough masks, the computational complexity and memory usage are high, and the methods are usually specific to a particular image segmentation algorithm or model.
By inputting the initial image segmentation mask into the trained diffusion model, the pixel values of the mask pixels are changed over a series of iteration cycles to generate a refined image segmentation mask. The diffusion model is trained through the forward and reverse diffusion stages to achieve refinement and error correction of the mask.
The accuracy of image segmentation is improved, especially when describing complex boundaries, and this technology is independent of the model and is suitable for various segmentation models and algorithms, improving the overall quality of segmentation.
Smart Images

Figure CN120219409A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more particularly, to image segmentation mask refinement using diffusion models. Background Art
[0002] Image segmentation is a computer vision task that aims to divide a digital image into multiple segments based on the content of the image. This technology has a wide range of applications, including image processing, medical imaging, autonomous vehicle navigation, etc. In some cases, the image segmentation process includes generating an image segmentation mask to clearly label different parts of the image - for example, to distinguish objects from the background or identify specific features within the image. Summary of the Invention
[0003] According to one aspect of the present disclosure, there is provided a computing system. The computing system includes a processor and a storage device that stores instructions that can be executed by the processor to receive an initial image segmentation mask for an image. The initial image segmentation mask is input into a diffusion model that is trained to change the pixel values of a plurality of mask pixels of the image segmentation mask, thereby generating a refined image segmentation mask for the image. The refined image segmentation mask is output.
[0004] The summary of the invention is provided to introduce some concepts in a simplified form, which will be further described in the following detailed description. The summary of the invention is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. In addition, the claimed subject matter is not limited to embodiments that solve any or all of the disadvantages noted in any part of the present disclosure. Brief Description of the Drawings
[0005] Figure 1 A schematic diagram showing an example computing system implementing a diffusion model for image segmentation mask refinement.
[0006] Figure 2 Schematically illustrates the forward diffusion stage and the reverse diffusion stage of the diffusion model.
[0007] Figure 3 Schematically illustrates the iterative generation of a refined image segmentation mask from an initial image segmentation mask.
[0008] Figure 4A and Figure 4B Illustrates an example algorithm for training and inference using a diffusion model for image segmentation mask refinement.
[0009] Figure 5 Illustrates an example method for image segmentation mask refinement.
[0010] Figure 6 An example computing system is shown schematically. Detailed Description
[0011] Some image segmentation methods include generating an image segmentation mask. For the purposes of the present disclosure, an image segmentation mask refers to a digital data structure of pixel values including a plurality of mask pixels, each mask pixel corresponding to an image pixel of the segmented image. These masks are typically binary, such that mask pixels having one value (e.g., 1) represent an object or region of interest, while mask pixels having another value (e.g., 0) represent the background. In general, however, an image segmentation mask can be used to distinguish any suitable number of different regions, objects, and / or other segments within an image, and any suitable pixel values can be used to represent these different segments.
[0012] Image segmentation masks are generated in a variety of different ways. As an example, an image segmentation mask can be generated by thresholding (e.g., based on pixel color values), edge detection, prediction by a suitable machine learning (ML) and / or artificial intelligence (AI) model, and / or other suitable means. However, generating an accurate and detailed segmentation mask (e.g., a mask that accurately represents the edges between different objects or regions in an image, even if those edges are blurred or contain fine details) can be challenging, time-consuming, and computationally expensive. This challenge becomes more severe as image resolution increases, potentially requiring significant computational complexity and memory usage to achieve high accuracy. Thus, existing segmentation algorithms typically generate masks at a lower resolution, which results in lower accuracy.
[0013] Due to the challenges of directly predicting an accurate and detailed mask, some methods focus on refining a "coarse" mask. A coarse mask is a segmentation mask that defines different parts within an image but may contain errors, such as misclassifying some parts of the background scene as part of the foreground object (and vice versa). Refinement refers to the process of generating a refined segmentation mask based on an existing coarse segmentation mask, which may include correcting errors and / or increasing the level of detail in the coarse segmentation mask. However, such coarse mask refinement methods are typically specific to a particular image segmentation algorithm or model and thus cannot be generalized to refine coarse masks produced by other segmentation methods. Additionally, the various types of errors that may exist in a coarse mask (e.g., errors along object boundaries, inability to capture fine-grained details in high-resolution images, and / or errors due to semantic incorrectness) can pose significant challenges during mask refinement, resulting in poor performance.
[0014] Accordingly, the present disclosure relates to image segmentation techniques, wherein an initial image segmentation mask (e.g., a rough mask output by an image segmentation model) is input into a trained diffusion model, which outputs a refined version of the image segmentation mask. For example, in a series of iterative cycles, the diffusion model can iteratively change the pixel values of the initial image segmentation mask to correct errors and gradually converge to a more accurate refined version of the initial segmentation mask. In other words, according to the techniques described herein, the segmentation mask refinement task can be represented as a data generation process, wherein refinement is achieved by applying a series of denoising diffusion steps to the initial image segmentation mask (e.g., rough mask) to generate a refined image segmentation mask with higher accuracy.
[0015] Accordingly, the techniques described herein provide various technical advantages in the field of computer image segmentation. First, they are capable of improving the accuracy of image segmentation, especially when depicting complex boundaries. This is achieved by using a discrete diffusion process, allowing for iterative refinement of the segmentation mask. The advantage of these techniques is that they are model-independent and thus versatile and applicable to various segmentation models and algorithms. In addition, the techniques described herein can improve the overall quality of segmentation, facilitating more accurate and reliable analysis in applications such as medical imaging and object recognition.
[0016] Figure 1 An example computing system 100 is schematically illustrated. The computing system 100 includes a processor 102 and a storage device 104 that stores instructions executable by the processor. As an example, the processor may include one or more central processing units (CPUs), graphics processing units (GPUs), tensor units, application specific integrated circuits (ASICs), and / or other types of processing devices. The storage device 104 may include volatile memory and / or non-volatile storage devices. In some examples, the computing system 100 is distributed across multiple physical computing devices, while in other examples, the processor 102 and the storage device 104 are included in a single physical computing device. Generally, the computing systems described herein may have any suitable capabilities, hardware configurations, and form factors, and may include any suitable number of one or more computing devices. In some examples, the computing system 100 is implemented as the computing system 600 described below with respect to Figure 6 the described computing system 600.
[0017] As shown, the computing system 100 has received an initial image segmentation mask 106 for the image 108. In other words, the initial image segmentation mask is generated for the image 108 using a suitable image segmentation technique, as discussed above. The initial image segmentation mask can be described as a "rough" image segmentation mask - for example, it may contain relatively little detail and / or contain obvious errors. The computing system can receive the image 108 from any suitable source - for example, the image can be loaded from a storage device of the computing system (e.g., storage device 104), loaded from an external storage device communicatively coupled to the computing system, received via a suitable computer network, or captured by a camera device integrated into or communicatively coupled to the computing system.
[0018] As Figure 1 shown, the initial image segmentation mask includes a plurality of mask pixels 110. The initial image segmentation mask can include any suitable number of mask pixels that correspond to the image pixels 112 of the input image 108. In some examples, the number of mask pixels in the segmentation mask is equal to the number of image pixels in the digital image - for example, the mask and the image have the same pixel resolution. Each of the mask pixels 110 of the initial image segmentation mask 106 has a pixel value. As discussed above, in some cases, the image segmentation mask is binary, where two different pixel values (e.g., 0 and 1) are used to distinguish two different segments within the image - for example, to distinguish a foreground object from a background scene. However, in general, the segmentation mask can define any suitable number of different segments within the image, which can be represented by any suitable pixel values of the mask pixels.
[0019] The initial image segmentation mask can be generated by any suitable computing device and using any suitable image segmentation technique. As a non-limiting example, the initial image segmentation mask can be generated by the computing system 100, so "receiving" the initial image segmentation mask can include generating the initial image segmentation mask. In Figure 1 the example, the initial image segmentation mask is generated by the image segmentation model 114 of the computing system 100. In other words, in some examples, the initial image segmentation mask is output by an image segmentation model that is trained to output an image segmentation mask for an input image. Such a model can take any suitable form - for example, implemented via any suitable ML and / or AI techniques. In a non-limiting example, the image segmentation model is a convolutional neural network (CNN). In other examples, the image segmentation model can use another suitable underlying architecture, such as a transformer-based architecture.
[0020] In some examples, the initial image segmentation mask can be generated by different computing devices, so the computing system 100 can "receive" the initial image segmentation mask from another suitable source in another suitable way. As an example, receiving the initial image segmentation mask can include loading the initial image segmentation mask from a storage device of the computing system (e.g., storage device 104), and / or loading the image segmentation mask from an external storage device communicatively coupled to the computing system. As another example, the initial image segmentation mask can be received via a suitable computer network, such as a local area network or a wide area network (e.g., the Internet).
[0021] In any case, in Figure 1 , the computing system inputs the initial image segmentation mask into the diffusion model 116. The diffusion model is trained to change the pixel values of a plurality of mask pixels (e.g., one or more mask pixels 110) of the initial image segmentation mask, thereby generating a refined image segmentation mask 118. In Figure 1 's example, the refined image segmentation mask includes a set of mask pixels 120 that are at least partially different from the mask pixels 110 of the initial image segmentation mask 106. In some examples, the image is input into the diffusion model together with the initial image segmentation mask, and the image and the initial image segmentation mask are processed together by the diffusion model to generate a refined image segmentation mask. This is shown in Figure 1 , where the image 108 is also input into the diffusion model 116 together with the initial image segmentation mask 106.
[0022] Generally speaking, a diffusion model can be described as a generative model that refines random noise through a learned reverse diffusion process to synthesize data, such as images or audio. The diffusion model is characterized by gradually reducing noise continuously or in multiple discrete steps to generate coherent outputs based on random or partially random inputs. Diffusion models include discrete diffusion models and continuous diffusion models. The continuous diffusion model works by transforming data through a smooth and uninterrupted process, where the changes occur in a smooth and continuous manner without distinct stages. In contrast, the discrete diffusion model operates through a series of different and independent steps. Each step in the process represents a distinct transformation, and the model adds or removes noise at quantized intervals. It should be understood that the techniques described herein can be implemented by either or both of discrete or continuous diffusion models.
[0023] In Figure 1 's example, the diffusion model is a discrete diffusion model that iteratively generates a series of intermediate image segmentation masks for the image in a series of iteration cycles. This can be done by changing the pixel values of one or more mask pixels of the previous image segmentation mask generated in the previous iteration cycle in each iteration cycle. In Figure 1In the example, the diffusion model performs multiple iterative cycles, including cycles 122A, 122B, and 122C. In each iterative cycle, corresponding intermediate image segmentation masks 124A-C are generated. In this way, through multiple iterative cycles, the initial image segmentation mask is gradually refined to converge to a refined image segmentation mask.
[0024] The diffusion model can use any suitable underlying architecture to generate intermediate image segmentation masks in each iterative cycle. In some examples, a trained neural network is used to output a series of intermediate image segmentation masks. More specifically, in some examples, a trained neural network uses a U-net architecture. It should be understood that these examples are non-limiting. As additional non-limiting examples, the diffusion model can be implemented in conjunction with a transformer-based architecture, a variational autoencoder (VAE), a generative adversarial network (GAN), a recurrent neural network, etc.
[0025] The diffusion model can be trained in a two-stage process including a forward diffusion stage and a reverse diffusion stage. In some cases, the forward diffusion stage q(x 1:T |x0) uses a Markov chain or a non-Markov chain to gradually transform the data distribution x0~q(x0) into completely noisy x T , while the reverse diffusion stage employs a gradual denoising procedure pθ(x 0:T ) to transform the random noise back to the original data distribution.
[0026] Generally speaking, continuous diffusion models follow the Gaussian assumption and define p(x T )=N(x T |0, 1). The mean and variance of the forward diffusion stage can be defined by the hyperparameter β t , while the reverse diffusion stage utilizes the mean and variance derived from the model predictions. This can be expressed as:
[0027]
[0028] p θ (x t-1 |x t )=N(x t-1 |u θ (x t , t)∑ θ (x t ,t)).
[0029] For discrete diffusion models, x T is defined as following a Bernoulli distribution B(x T |0.5). The forward diffusion stage and the reverse diffusion stage can be expressed as:
[0030] q(x t |x t-1 ) = B(x t |x t-1 (1 - β t ) + 0.5β t ),
[0031] p θ (x t-1 |x t-1 ) = B(x t-1 |f b (x t , t)).
[0032] Wherein, β t ∈ (0, 1) is a hyperparameter, and f b (x t , t) is a model for predicting the Bernoulli probability. More generally, the forward diffusion stage of the discrete diffusion model can be defined as a discrete random variable that transitions between multiple states. The process can be characterized using the state transition distribution Qt:
[0033] [Q t m,n = (x t = n|x t-1 = m).
[0034] In view of this, a diffusion model according to the techniques described herein can be applied to refine a rough mask generated via any suitable image segmentation technique. In some examples, the diffusion model can be trained in a two-stage training process including a forward diffusion stage and a reverse diffusion stage. In the forward diffusion stage, the diffusion model can adopt a discrete diffusion process (which can be expressed as a one-way random state transition) to gradually degrade the ground truth mask into a training rough segmentation mask. In other words, the forward diffusion stage can include iteratively adding noise to the ground truth image segmentation mask to generate a training rough segmentation mask. In some cases, the forward diffusion stage is a one-way process, where each mask pixel of the ground truth image segmentation mask transitions from a fine state to a rough state. In the reverse diffusion stage, the diffusion model can start from the rough segmentation mask and then gradually transform the pixels in the rough segmentation mask into a refined state, thereby correcting the mispredicted regions in the rough segmentation mask. In other words, in some examples, the reverse diffusion stage includes iteratively changing the pixel values of the rough segmentation mask during inference to generate a refined segmentation mask.
[0035] Now focusing on the forward diffusion process, the ground truth mask (denoted as m0) is transformed into a training rough segmentation mask (denoted as m T ). At any intermediate timestamp t, where t ∈ {1, 2,... T - 1}, and T represents the total number of iteration cycles, the intermediate image segmentation mask mt in the transition phase between m0 and m T . Each mask pixel in m t occupies one of two states: fine and coarse. Thus, the forward diffusion phase can be formulated as a state transition between these two states. Pixels in the fine state will retain the value of m0, and vice versa. During each iteration cycle in the forward diffusion phase, the diffusion model uses the previous intermediate image segmentation mask m t-1 , the coarse mask m T , and the state transition probability as inputs, and outputs the intermediate image segmentation mask m t for the current iteration cycle. In the context of the forward diffusion process, the state transition probability describes the probability that each pixel in m t-1 transitions to the coarse state. In some cases, this can include performing Gumbel-max sampling according to the state transition probability to obtain the transformed pixels. At this time, the transformed mask pixels will have the value of m T , while the untransformed pixels remain unchanged.
[0036] It is worth noting that, as discussed above, the transition of mask pixels from one state to another is in some cases a one-way process - for example, during the forward diffusion phase, pixels only transition from fine to coarse. This can usefully ensure that the forward diffusion phase converges to the trained coarse segmentation mask, but randomness is introduced in each iteration cycle. This is in contrast to the implementation of other diffusion models (where the forward process converges to random noise).
[0037] Using a reparameterization step, a binary random variable x can be introduced into the above process. Denoting as the one-hot vector indicating the state of the pixel (i, j) in the intermediate image segmentation mask mt. The sets and represent the fine state and the coarse state respectively. Thus, the forward process can be formulated as:
[0038] where,
[0039] where, β t ∈[0,1], and 1 - β t corresponds to the state transition probability. The form Q t can be used to embody the one-way nature of the state transition process - for example, pixels in the coarse state do not transition back to the fine state because q(x t |[0,1]) = [0,1].
[0040] The marginal distribution can be formulated as:
[0041]
[0042] Among them, In view of this, the intermediate image segmentation mask at any intermediate timestamp can be obtained without step-by-step sampling, which is beneficial to faster model training.
[0043] Now turning to the reverse diffusion stage, the training rough segmentation mask is refined to correct errors and / or improve the level of detail. However, since the fine mask and the reverse state transition probability are unknown, a neural network can be trained to predict the fine mask at each time step - for example, thus outputting the intermediate image segmentation mask at each time step. The fine mask predicted at iteration cycle t can be expressed as The confidence score of the predicted fine mask is expressed as And the neural network can be expressed as f Θ .
[0044]
[0045] Among them, I is the corresponding image to be segmented.
[0046] To obtain the reverse state transition probability, the posterior at time step t–1 can be expressed as:
[0047]
[0048] Among them, the fine state x0 is set to [1, 0] during training, representing the ground truth. During inference, x0 is unknown because the predicted may not be accurate. Since the confidence score represents the confidence of the model in the correctness of each pixel, therefore can also be interpreted as the probability that this pixel is in the fine state.
[0049] Therefore, the state of each pixel in can be obtained via thresholding, where:
[0050]
[0051] In this case, pixels with a higher confidence score will have indicating that they are in the fine state, and vice versa. However, in this one-hot form, the value of the state transition probability will only be determined by predefined hyperparameters, which may lead to a large amount of information loss.
[0052] Therefore, the soft transition can be retained by the following formula:
[0053]
[0054] This in turn allows the reverse diffusion stage to be reformulated as:
[0055]
[0056] Among them, is the reverse state transition matrix. Using the above reverse state transition probabilities, m t and as inputs, the diffusion model can convert a subset of mask pixels to a fine state at each time step, thereby correcting incorrect predictions.
[0057] During inference, given a rough mask m T and the corresponding segmented image I, all mask pixels can be initially set to the rough state. Therefore, Then, the diffusion model can iterate between: (1) forward propagation to obtain and (2) calculating the reverse state transition matrix and x t-1 ; and (3) calculating the next intermediate image segmentation mask m t-1 and based on x t-1 . This process can be iteratively repeated until a refined image segmentation mask m0 is obtained. In other words, according to this process, the pixel values of one or more mask pixels are changed at least in part based on the state transition probabilities of each mask pixel, which indicate the probability that a mask pixel changes state between the initial image segmentation mask and the refined image segmentation mask. This can occur within any suitable number of iteration cycles. In some examples, a predefined number of iteration cycles is used (e.g., a value chosen to balance accuracy and processing time), and / or the process can continue until a refined image segmentation mask with a confidence value higher than a threshold is generated.
[0058] Figure 2 Schematically illustrates the forward diffusion stage and the reverse diffusion stage. Specifically, Figure 2 Schematically represents the forward diffusion stage 200A and the reverse diffusion stage 200B. In the forward diffusion stage, the ground truth image segmentation mask 202 is transformed into a training rough segmentation mask 204 by gradually adding noise. In the reverse diffusion stage, the training rough segmentation mask 204 is used to generate a training refined segmentation mask 206.
[0059] Figure 2Shows the isolated portion 208 of the ground truth image segmentation mask, which is used to illustrate the forward diffusion phase and the reverse diffusion phase. As shown, the ground truth image segmentation mask undergoes pixel state transitions to generate an intermediate image segmentation mask, represented by the intermediate mask portion 210. This process is iterated any suitable number of times to obtain a trained rough segmentation mask, a portion of which is shown as the rough mask portion 212. Compared to the initial image segmentation mask and the intermediate image segmentation mask, the rough segmentation mask includes classification errors - for example, pixels of the foreground object are misclassified as the background scene and vice versa.
[0060] In Figure 2 the example of, the trained rough segmentation mask (represented by the rough mask portion 212) and the segmented input image (represented by the image portion 214) are both inputs to the reverse diffusion phase. In this example, in each iteration cycle, the trained neural network outputs a refined prediction and a confidence value for that prediction 218. At least in part based on the output and one or more mask pixels undergo state transitions to generate the next intermediate image segmentation mask. Similarly, this process can be iterated any suitable number of times in any suitable number of iteration cycles to output a refined mask portion 220 (which is part of the refined image segmentation mask 206).
[0061] In Figure 2 the example of, the trained neural network used to generate the intermediate image segmentation mask in each iteration cycle uses the U-net architecture 222. As a non-limiting example, the U-net architecture can be modified to accept a 4-channel input (e.g., the concatenation of the original image and the image segmentation mask from the previous iteration cycle) and output a 1-channel refined image segmentation mask. However, as discussed above, the diffusion model can use any of a variety of suitable underlying architectures to predict the segmentation mask for each iteration cycle.
[0062] In Figure 2 pixel sampling and state transitions are handled via the transformation sampling module 224. This module is used to randomly sample pixels from the current cycle mask based on the input state transition probability (represented by the state transition probability 226), thereby changing the pixel values to match the pixel values in the target mask. During training, the transformation sampling module transforms the ground truth image segmentation mask into a trained rough segmentation mask, so the "target mask" refers to the trained rough segmentation mask. During inference, the "target mask" refers to the refined image segmentation mask, and the transformation sampling module updates the pixel values in the rough mask based on the predicted refined mask and the state transition probability in each iteration cycle.
[0063] Figure 3Schematically illustrates the iterative generation of a refined image segmentation mask from an initial image segmentation mask. Specifically, Figure 3 Shows two different initial image segmentation masks 300A and 300B. These masks are input into a diffusion model, which outputs corresponding refined image segmentation masks 302A and 302B. The iterative transformation between these two image segmentation masks is shown at five different time steps ranging from t = T (corresponding to the initial rough segmentation mask) to t = 0 (corresponding to the refined image segmentation mask). At each time step, the current cycle image segmentation mask m t And the rough / fine state x of each mask pixel t Are shown. As discussed above, during inference, the x of each pixel T Is initialized to [0, 1]. It is gradually refined to obtain the refined image segmentation mask m0.
[0064] Figure 4A And Figure 4B Provide non - limiting example algorithms 400A and 400B, which can be used respectively for training and inference using a diffusion model to perform image segmentation mask refinement as described herein. Refer to Figure 4A , algorithm 400A outlines an example method for training a diffusion model, focusing on the forward diffusion stage. The method starts by inputting the total number of diffusion steps T and a dataset D. The dataset includes tuples of input images and corresponding rough and fine image segmentation masks (e.g., a training dataset). Each iteration starts by sampling a tuple from the dataset and sampling a time step t from a uniform distribution ranging from 1 to T. Initialization is performed by setting the initial image segmentation mask m0 to the fine mask (e.g., the ground - truth image segmentation mask) and setting the initial pixel state To the binary vector [1, 0].
[0065] The algorithm proceeds to define the conditional distribution Using the state transition probability matrix Sampling new values for the mask pixels from the conditional distribution Generating an intermediate image segmentation mask at least in part based on the sampled pixel state, the ground - truth image segmentation mask, and the rough mask. The last step of the iteration involves performing a gradient descent step on a loss function L, which is a function of the predicted fine mask and the ground - truth fine mask. The iteration process is repeated until convergence is achieved.
[0066] Refer to Figure 4B , algorithm 400B outlines an example method for using a diffusion model for inference to refine an image segmentation mask. The method starts from inputting the total number of diffusion steps T, an image I, and a rough mask m 粗糙Start. The algorithm proceeds to an initialization step, where x T is set to the binary vector [0, 1], and the initial image segmentation mask m 粗糙 is set to m T .
[0067] For each time step t starting from T and decreasing to 1, the algorithm uses a trained neural network parameterized by θ to compute an output intermediate image segmentation mask and a confidence value Then a state transition distribution of pixels is defined The next step involves sampling new pixel states from the state transition distribution After sampling, a new intermediate image segmentation mask m t-1 is generated. This loop iterates backward through these diffusion steps, refining the state of the image segmentation mask at each step until t = 1 is reached. Then the refined image segmentation mask m0 is output.
[0068] Figure 5 Illustrates an example method 500 for refining an image segmentation mask. Method 500 can be implemented via any suitable computing system of one or more computing devices. The computing device implementing the steps of method 500 can have any suitable capabilities, hardware configurations, and form factors. The steps of method 500 can be started, terminated, and / or looped at any appropriate time and in response to any appropriate conditions. In some examples, method 500 can be implemented by Figure 1 computing system 100 and / or Figure 6 computing system 600.
[0069] At 502, method 500 includes inputting an image into an image segmentation model, thereby generating an initial image segmentation mask. As discussed above, any suitable image segmentation technique can be used to generate the initial image segmentation model. For example, this can include an image segmentation model trained to output a segmentation mask of the input image, such as a CNN. Notably, the initial image segmentation mask may be a "rough" image segmentation mask as described above - for example, it may include misclassified pixels.
[0070] At 504, method 500 includes inputting an initial image segmentation mask into a diffusion model. As discussed above, in some examples, the diffusion model is a discrete diffusion model that iteratively refines the initial image segmentation mask over a series of iterative cycles. Thus, at 506, method 500 optionally includes iteratively generating a series of intermediate image segmentation masks over a series of iterative cycles. In this way, in each iterative cycle, the pixel values of one or more mask pixels of the previous image segmentation mask can be changed to generate a new intermediate image segmentation mask for that iterative cycle and correct errors in the original rough segmentation mask.
[0071] At 508, method 500 includes outputting a refined image segmentation mask. It should be understood that, depending on the implementation, the image segmentation mask can be "output" in various suitable ways. In some embodiments, outputting the image segmentation mask includes passing an output vector to a downstream application, transmitting the image segmentation mask to another computing device (e.g., via a suitable computer network), writing the image segmentation mask to a data file, storing the image segmentation mask in a non-volatile storage device of a computing device, and / or storing the image segmentation mask in an external storage device communicatively coupled to the computing device.
[0072] This disclosure is primarily concerned with refining the image segmentation mask of a single input image. However, it should be understood that this is not restrictive. In some examples, the techniques described herein can be used to simultaneously refine the image segmentation masks of two or more input images. For example, such input images can be different consecutive or non-consecutive video frames of a digital video. This can be achieved by adjusting the architecture of the diffusion model to accept input data with a higher dimension. As a non-limiting example, when using a U-net architecture, the U-net can be modified to a three-dimensional matrix instead of a two-dimensional matrix, which can enable the simultaneous processing of multiple image frames during the reverse diffusion phase.
[0073] In some embodiments, the methods and processes described herein can be related to the computing systems of one or more computing devices. Specifically, such methods and processes can be implemented as a computer application or service, an application programming interface (API), a library, and / or other computer program products.
[0074] Figure 6 A non-limiting embodiment of a computing system 600 that can implement one or more of the above methods and processes is schematically shown. Computing system 600 is shown in a simplified form. Computing system 600 can embody with respect to Figure 1The described computing system 100. The computing system 600 can take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones), and / or other computing devices, as well as wearable computing devices (such as smartwatches and head-mounted augmented reality devices).
[0075] The computing system 600 includes a logic processor 602, volatile memory 604, and a non-volatile storage device 606. The computing system 600 can optionally include a display subsystem 608, an input subsystem 610, a communication subsystem 612, and / or Figure 6 other components not shown.
[0076] The logic processor 602 includes one or more physical devices configured to execute instructions. For example, the logic processor can be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions can be implemented to perform a task, implement a certain data type, transform the state of one or more components, achieve a certain technical effect, or otherwise achieve an intended result.
[0077] The logic processor can include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, the logic processor can include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logic processor 602 can be single-core or multi-core, and the instructions executed thereon can be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logic processor can be distributed across two or more separate devices, which can be remotely located and / or configured to coordinate processing. Various aspects of the logic processor can be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. It will be understood that in such a case, these virtualized aspects run on different physical logic processors of various different machines.
[0078] The non-volatile storage device 606 includes one or more physical devices configured to store instructions that can be executed by the logic processor to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 606 may change, for example, to store different data.
[0079] The non-volatile storage device 606 may include removable and / or built-in physical devices. The non-volatile storage device 606 may include optical memories (such as CDs, DVDs, HD-DVDs, Blu-ray discs, etc.), semiconductor memories (such as ROMs, EPROMs, EEPROMs, FLASH memories, etc.), and / or magnetic memories (such as hard disk drives, floppy disk drives, tape drives, MRAMs, etc.) or other mass storage device technologies. The non-volatile storage device 606 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It should be understood that the non-volatile storage device 606 is configured to store instructions even when the power supply to the non-volatile storage device 606 is cut off.
[0080] The volatile memory 604 may include a physical device that includes random access memory. The volatile memory 604 is typically used by the logical processor 602 to temporarily store information during the processing of software instructions. It should be understood that when the power supply to the volatile memory 604 is cut off, the volatile memory 604 generally does not continue to store instructions.
[0081] Aspects of the logical processor 602, the volatile memory 604, and the non-volatile storage device 606 may be integrated together into one or more hardware logic components. Such hardware logic components may include field programmable gate arrays (FPGAs), program and application specific integrated circuits (PASIC / ASICs), program and application specific standard products (PSSP / ASSPs), systems on a chip (SOCs), and complex programmable logic devices (CPLDs), etc.
[0082] The terms "module", "program", and "engine" may be used to describe an aspect of the computing system 600, which is typically implemented in software by a processor to perform specific functions using portions of the volatile memory, and the functions involve transformation processing for specifically configuring the processor to perform the functions. Thus, modules, programs, or engines may be instantiated by executing instructions saved in the non-volatile storage device 606 using portions of the volatile memory 604 via the logical processor 602. It should be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" may encompass single or multiple sets of executable files, data files, libraries, drivers, scripts, database records, etc.
[0083] The display subsystem 608 (when included) can be used to present a visual representation of data stored by the non-volatile storage device 606. The visual representation can take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data stored by the non-volatile storage device and thus transform the state of the non-volatile storage device, the state of the display subsystem 608 can also be transformed accordingly to visually represent changes in the underlying data. The display subsystem 608 can include one or more display devices utilizing almost any type of technology. Such display devices can be combined with the logic processor 602, volatile memory 604, and / or non-volatile storage device 606 in a shared enclosure, or such display devices can be peripheral display devices.
[0084] The input subsystem 610 (when included) can include one or more user input devices, such as a keyboard, mouse, touch screen, or game controller, or interface with these user input devices. In some embodiments, the input subsystem can include selected natural user input (NUI) components, or interface with these components. These components can be integrated or peripheral components, and the conversion and / or processing of input actions can be performed on-board or off-board. Example NUI components can include a microphone for speech and / or voice recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; and an electric field sensing component for evaluating brain activity; and / or any other suitable sensors.
[0085] The communication subsystem 612 (when included) can be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 612 can include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via a wireless telephone network, or a wired or wireless local or wide area network, such as HDMI over Wi-Fi. In some embodiments, the communication subsystem can allow the computing system 600 to send messages to and / or receive messages from other devices via a network such as the Internet.
[0086] The following paragraphs provide additional description of the subject matter of the present disclosure. In one example, a computing system includes: a processor; and a storage device that stores instructions executable by the processor to perform the following operations: receive an initial image segmentation mask for an image; input the initial image segmentation mask into a diffusion model that is trained to change pixel values of a plurality of mask pixels of the initial image segmentation mask to generate a refined image segmentation mask for the image; and output the refined image segmentation mask. In this example or any other example, the diffusion model is a discrete diffusion model that iteratively generates a series of intermediate image segmentation masks for the image by changing pixel values of one or more mask pixels of a previous image segmentation mask generated in a previous iteration cycle in a series of iteration cycles. In this example or any other example, a trained neural network is used to output the series of intermediate image segmentation masks. In this example or any other example, the trained neural network uses a U-Net architecture. In this example or any other example, the pixel values of the one or more mask pixels are changed at least in part based on a state transition probability for each mask pixel, the state transition probability indicating the probability that the mask pixel changes state between the initial image segmentation mask and the refined image segmentation mask. In this example or any other example, the diffusion model is trained in a two-stage training process including a forward diffusion stage and a reverse diffusion stage, where the forward diffusion stage includes iteratively adding noise to a ground truth image segmentation mask to generate a training rough segmentation mask, and where the reverse diffusion stage includes iteratively changing pixel values of the rough segmentation mask during inference to generate a training refined segmentation mask. In this example or any other example, the forward diffusion stage is a one-way process in which each mask pixel of the ground truth image segmentation mask transitions from a fine state to a rough state. In this example or any other example, the image is input into the diffusion model together with the initial image segmentation mask. In this example or any other example, the initial image segmentation mask is output by an image segmentation model that is trained to output an image segmentation mask for an input image. In this example or any other example, the image segmentation model is a convolutional neural network (CNN).
[0087] In one example, a method for refining an image segmentation mask includes: at a computing system, receiving an initial image segmentation mask for an image; inputting the initial image segmentation mask into a diffusion model, the diffusion model being trained to change pixel values of a plurality of mask pixels of the initial image segmentation mask to generate a refined image segmentation mask for the image; and outputting the refined image segmentation mask. In this example or any other example, the diffusion model is a discrete diffusion model that iteratively generates a series of intermediate image segmentation masks for the image by changing pixel values of one or more mask pixels of a previous image segmentation mask generated in a previous iteration cycle in a series of iteration cycles. In this example or any other example, a trained neural network is used to output the series of intermediate image segmentation masks. In this example or any other example, the pixel values of the one or more mask pixels are changed at least in part based on a state transition probability for each mask pixel, the state transition probability indicating the probability that the mask pixel changes state between the initial image segmentation mask and the refined image segmentation mask. In this example or any other example, the diffusion model is trained in a two-stage training process including a forward diffusion stage and a reverse diffusion stage, where the forward diffusion stage includes iteratively adding noise to a ground truth image segmentation mask to generate a training rough segmentation mask, and where the reverse diffusion stage includes iteratively changing pixel values of the rough segmentation mask during inference to generate a training refined segmentation mask. In this example or any other example, the forward diffusion stage is a one-way process in which each mask pixel of the ground truth image segmentation mask transitions from a fine state to a rough state. In this example or any other example, the image is input into the diffusion model together with the initial image segmentation mask. In this example or any other example, the initial image segmentation mask is output by an image segmentation model that is trained to output an image segmentation mask for an input image. In this example or any other example, the image segmentation model is a convolutional neural network (CNN).
[0088] In one example, a computing system includes: a processor; and a storage device that stores instructions executable by the processor to perform the following operations: receiving an initial image segmentation mask for an image, the initial image segmentation mask being output by a trained image segmentation model; inputting the initial image segmentation mask into a discrete diffusion model that is trained to change pixel values of a plurality of mask pixels of the initial image segmentation mask to generate a refined image segmentation mask for the image, wherein the discrete diffusion model iteratively generates a series of intermediate image segmentation masks for the image by changing pixel values of one or more mask pixels of a previous image segmentation mask generated in a previous iteration cycle in each of a series of iteration cycles; and outputting the refined image segmentation mask.
[0089] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered restrictive as there may be various variations. The specific routines or methods described herein may represent one or more of any number of processing strategies. Accordingly, the various acts illustrated and / or described may be performed in the illustrated and / or described order, in other orders, in parallel, or omitted. Similarly, the order of the above processes may be changed.
[0090] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems, and configurations, as well as other features, functions, acts, and / or characteristics disclosed herein, and any and all equivalents thereof.
Claims
1. A computing system comprising: processor; as well as A storage device storing instructions, wherein the instructions are executable by the processor to perform the following operations: Receiving an initial image segmentation mask for the image; Inputting the initial image segmentation mask into a diffusion model, the diffusion model being trained to change pixel values of a plurality of mask pixels of the initial image segmentation mask, thereby generating a refined image segmentation mask for the image; as well as The refined image segmentation mask is output.
2. The computing system of claim 1 , wherein the diffusion model is a discrete diffusion model that iteratively generates a series of intermediate image segmentation masks for the image by changing pixel values of one or more mask pixels of a previous image segmentation mask generated in a previous iteration cycle in a series of iteration cycles.
3. The computing system of claim 2, wherein a trained neural network is used to output the series of intermediate image segmentation masks.
4. The computing system of claim 3, wherein the trained neural network employs a U-Net architecture.
5. The computing system of claim 2 , wherein the pixel values of the one or more mask pixels are changed based at least in part on a state transition probability for each mask pixel, the state transition probability indicating a probability that the mask pixel changes state between the initial image segmentation mask and the refined image segmentation mask.
6. The computing system of claim 1 , wherein the diffusion model is trained in a two-stage training process comprising a forward diffusion stage and a backward diffusion stage, wherein the forward diffusion stage comprises iteratively adding noise to a ground truth image segmentation mask to generate a training coarse segmentation mask, and wherein the backward diffusion stage comprises iteratively changing pixel values of the coarse segmentation mask during inference to generate a training refined segmentation mask.
7. The computing system of claim 6, wherein the forward diffusion stage is a one-way process in which each mask pixel of the ground truth image segmentation mask is transformed from a fine state to a coarse state.
8. The computing system of claim 1, wherein the image is input into the diffusion model along with the initial image segmentation mask.
9. The computing system of claim 1, wherein the initial image segmentation mask is output by an image segmentation model, the image segmentation model being trained to output an image segmentation mask for an input image.
10. The computing system of claim 9, wherein the image segmentation model is a convolutional neural network (CNN).
11. A method for image segmentation mask refinement, the method comprising: At a computing system, receiving an initial image segmentation mask for an image; Inputting the initial image segmentation mask into a diffusion model, the diffusion model being trained to change pixel values of a plurality of mask pixels of the initial image segmentation mask, thereby generating a refined image segmentation mask for the image; as well as The refined image segmentation mask is output.
12. The method of claim 11, wherein the diffusion model is a discrete diffusion model, which iteratively generates a series of intermediate image segmentation masks for the image by changing the pixel values of one or more mask pixels of a previous image segmentation mask generated in a previous iteration cycle in a series of iteration cycles.
13. The method of claim 12, wherein a trained neural network is used to output the series of intermediate image segmentation masks.
14. A method according to claim 12, wherein the pixel value of the one or more mask pixels is changed at least in part based on a state transition probability for each mask pixel, the state transition probability indicating the probability of the mask pixel changing state between the initial image segmentation mask and the refined image segmentation mask.
15. The method of claim 12, wherein the diffusion model is trained in a two-stage training process comprising a forward diffusion stage and a backward diffusion stage, wherein the forward diffusion stage comprises iteratively adding noise to a true image segmentation mask to generate a training coarse segmentation mask, and wherein the backward diffusion stage comprises iteratively changing pixel values of the coarse segmentation mask during inference to generate a training refined segmentation mask.
16. The method of claim 15, wherein the forward diffusion stage is a one-way process in which each mask pixel of the ground truth image segmentation mask is transformed from a fine state to a coarse state.
17. The method of claim 11, wherein the image is input into the diffusion model along with the initial image segmentation mask.
18. The method of claim 11, wherein the initial image segmentation mask is output by an image segmentation model, the image segmentation model being trained to output an image segmentation mask for an input image.
19. The method according to claim 18, wherein the image segmentation model is a convolutional neural network (CNN).
20. A computing system comprising: processor; as well as A storage device storing instructions, wherein the instructions are executable by the processor to perform the following operations: Receiving an initial image segmentation mask for an image, the initial image segmentation mask being output by a trained image segmentation model; inputting the initial image segmentation mask into a discrete diffusion model, the discrete diffusion model being trained to change pixel values of a plurality of mask pixels of the initial image segmentation mask, thereby generating a refined image segmentation mask for the image, wherein the discrete diffusion model iteratively generates a series of intermediate image segmentation masks for the image by changing the pixel values of one or more mask pixels of a previous image segmentation mask generated in a previous iteration cycle in each cycle in a series of iteration cycles; as well as The refined image segmentation mask is output.
Citation Information
Cited By
Image restoration method and device and storage medium
CN120894266A
Radio map reconstruction method and device based on double diffusion model
CN121564141A