Universal framework for panoramic segmentation of images and videos
By using the de-diffusion model in the panoramic segmentation task, the problem of customized architecture and complex loss functions in the prior art is solved, and efficient panoramic segmentation mask generation and computing resource savings are achieved.
Patent Information
- Application Number
- CN202380072622.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art requires customized architecture and complex loss functions in panoramic segmentation tasks, and it is difficult to effectively learn high-dimensional one-to-many mapping.
Using a denoising diffusion model, including an image encoder and a mask decoder, a panoramic segmentation mask is generated through generative modeling, and a simulated bit representation is used and trained through cross entropy loss and weighted loss functions.
The complex process of panoramic segmentation is simplified, the computing efficiency is improved, and high-dimensional data can be effectively processed and accurate panoramic segmentation masks are generated.
Smart Images

Figure CN120035849A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 415,619, filed on October 12, 2023. The entire contents of U.S. Provisional Patent Application No. 63 / 415,619 are hereby incorporated by reference. Technical Field
[0003] The present disclosure relates generally to machine learning and more particularly to systems and methods for performing panoptic segmentation using a denoising diffusion model. Background Art
[0004] Panoptic segmentation is a fundamental vision task that assigns semantic and instance labels to each pixel of an image. The semantic label describes the classification of each pixel (e.g., sky, car, dog, etc.), and the instance label provides a unique ID for each instance in the image (e.g., used to distinguish different instances of the same classification). This task is a combination of semantic segmentation and instance segmentation, providing rich semantic information about the scene. Since permutations of instance IDs are also valid solutions, this task requires learning a high-dimensional one-to-many mapping. Therefore, state-of-the-art methods use customized architectures and task-specific loss functions.
[0005] More specifically, while the classification categories of semantic labels are typically fixed a priori, the instance IDs assigned to objects in an image can be permuted without affecting the identified instances. For example, swapping the instance IDs of two cars will not affect the results. Therefore, a neural network trained to predict instance IDs should be able to learn a one-to-many mapping: the assignment from a single image to multiple instance IDs. Learning the one-to-many mapping is challenging, and traditional methods typically utilize a pipeline involving multiple stages of object detection, segmentation, and merging multiple predictions. Recently, end-to-end methods based on differentiable bipartite graph matching have been proposed; this effectively converts the one-to-many mapping into a one-to-one mapping based on the identified matches. However, for panoptic segmentation tasks, such methods still require customized architectures and complex loss functions with built-in inductive biases. Summary of the invention
[0006] Various aspects and advantages of the embodiments of the present disclosure will be partially set forth in the following description, or may be learned through the description, or may be learned through the practice of the embodiments.
[0007] A system of one or more computers may be configured to perform specific operations or actions by installing software, firmware, hardware, or a combination thereof on the system, which in operation causes the system to perform these actions. One or more computer programs may be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform these actions.
[0008] One general aspect includes a computer-implemented method for performing panoptic segmentation. The computer-implemented method also includes obtaining, by a computing system that may include one or more computing devices, an input image that may include a plurality of pixels. The method also includes processing, by the computing system, the input image using a denoising diffusion model to generate a panoptic segmentation mask as an output of the denoising diffusion model, wherein the panoptic segmentation mask provides a corresponding semantic identifier and a corresponding instance identifier for each of the plurality of pixels. The method also includes providing, by the computing system, the panoptic segmentation mask as an output. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each of which is configured to perform the actions of the methods.
[0009] Implementations may include one or more of the following features. In the computer-implemented method, the denoising diffusion model may include an image encoder and a mask decoder, wherein the image encoder maps an input image to a feature map, and wherein the mask decoder generates a panoramic segmentation mask from a noisy mask conditioned on the feature map. The image encoder may include a residual neural network followed by one or more transformer encoder layers. The image encoder may include convolutions with bilateral connections and upsampling operations for merging features from different resolutions. The mask decoder may include one or more transformer layers on top of a u-net, and a cross-attention layer for integrating image features from the feature map. Processing the input image using the denoising diffusion model by a computing system to generate a panoramic segmentation mask as an output of the denoising diffusion model may include: processing the input image using the denoising diffusion model by a computing system to generate an analog bit representation of the panoramic segmentation mask as an output of the denoising diffusion model; and converting the analog bit representation of the panoramic segmentation mask by a computing system into a real-valued version of the panoramic segmentation mask, wherein the corresponding semantic identifier and the corresponding instance identifier of each of the plurality of pixels may include a real value included in the real-valued version of the panoramic segmentation mask. The analog bit representation of the panoptic segmentation mask is generated according to the scaling factor, and wherein the scaling factor is equal to 0.1. The denoising diffusion model has been trained using a softmax cross entropy loss applied to the logits of the denoising diffusion model. The denoising diffusion model has been trained using a weighted loss function that assigns larger weights to mask tokens with fewer instances. The input image may include an input image frame from a video; and processing the input image by the computing system using the denoising diffusion model may include processing the input image by the computing system using the denoising diffusion model and one or more previous panoptic segmentation masks generated for one or more previous image frames before the input image frame in the video. The one or more previous panoptic segmentation masks generated for one or more previous image frames may include multiple previous panoptic segmentation masks generated for multiple previous image frames. The denoising diffusion model may include an image encoder and a mask decoder, wherein the image encoder maps the input image frame to a feature map, and wherein the mask decoder generates a panoptic segmentation mask from a noisy mask conditioned on the feature map and the one or more previous panoptic segmentation masks. Implementations of the techniques may include hardware, a method or process, or computer software on a computer-accessible medium.
[0010] One general aspect includes one or more non-transitory computer-readable media that collectively store instructions for performing panoramic segmentation. The one or more non-transitory computer-readable media also include instructions for obtaining, by a computing system, an input image that may include a plurality of pixels. The medium also includes instructions for processing, by a computing system, the input image using a denoising diffusion model to generate a panoramic segmentation mask as an output of the denoising diffusion model, wherein the panoramic segmentation mask provides a corresponding semantic identifier and a corresponding instance identifier for each of the plurality of pixels. The medium also includes instructions for providing, by the computing system, the panoramic segmentation mask as an output. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each of which is configured to perform the actions of these methods.
[0011] Implementations may include one or more of the following features. In one or more non-transitory computer-readable media, a denoising diffusion model may include an image encoder and a mask decoder, wherein the image encoder maps an input image to a feature map, and wherein the mask decoder generates a panoptic segmentation mask from a noisy mask conditioned on the feature map. Processing the input image by a computing system using the denoising diffusion model to generate a panoptic segmentation mask as an output of the denoising diffusion model may include: processing the input image by the computing system using the denoising diffusion model to generate an analog bit representation of the panoptic segmentation mask as an output of the denoising diffusion model; and converting the analog bit representation of the panoptic segmentation mask by the computing system into a real-valued version of the panoptic segmentation mask, wherein the corresponding semantic identifier and the corresponding instance identifier of each pixel in the plurality of pixels may include a real value included in the real-valued version of the panoptic segmentation mask. The input image may include an input image frame from a video; and processing the input image by the computing system using the denoising diffusion model may include processing the input image by the computing system using the denoising diffusion model and one or more previous panoptic segmentation masks generated for one or more previous image frames preceding the input image frame in the video. The one or more non-transitory computer-readable media further stores the denoising diffusion model. Implementations of the techniques may include hardware, a method or process, or computer software on a computer-accessible medium.
[0012] One general aspect includes a computing system for training a denoising diffusion model to perform panoptic segmentation. The computing system also includes instructions for obtaining, by the computing system, a training input image and a reference ground truth panoptic segmentation mask. The system also includes instructions for processing, by the computing system, the training input image using the denoising diffusion model to generate a predicted panoptic segmentation mask as an output of the denoising diffusion model, wherein the predicted panoptic segmentation mask provides a corresponding semantic identifier and a corresponding instance identifier for each of a plurality of pixels. The system also includes instructions for evaluating, by the computing system, a loss function that compares the predicted panoptic segmentation mask to a reference ground truth panoptic segmentation mask. The system also includes instructions for modifying, by the computing system, one or more parameter values of one or more parameters of the denoising diffusion model based on the loss function. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each of which is configured to perform the actions of these methods.
[0013] Implementations may include one or more of the following features. In a computing system, a loss function may include a softmax cross entropy loss applied to the logits of the denoising diffusion model. The loss function may include a weighted loss function that assigns larger weights to mask tags with fewer instances. Implementations of the technology may include hardware, a method or process, or computer software on a computer-accessible medium.
[0014] Other aspects of the disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0015] These and other features, aspects and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and appended claims.The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The specification sets forth a detailed discussion of embodiments for those of ordinary skill in the art with reference to the accompanying drawings, in which:
[0017] Figure 1 Depicted is a block diagram of an example method for performing panoptic segmentation on an input image using a denoising diffusion model, according to an example embodiment of the present disclosure.
[0018] Figure 2 Depicted is a block diagram of an example denoising diffusion model architecture for performing panoptic segmentation on an input image, according to an example embodiment of the present disclosure.
[0019] Figure 3 Depicted is a block diagram of an example method for performing panoptic segmentation on a sequence of video frames using a denoising diffusion model, according to an example embodiment of the present disclosure.
[0020] Figure 4 Depicted is a block diagram of an example denoising diffusion model architecture for performing panoptic segmentation on a sequence of video frames, according to an example embodiment of the present disclosure.
[0021] Figure 5 Depicted is a block diagram of an example method for training a diffusion model to perform panoptic segmentation according to an example embodiment of the present disclosure.
[0022] Figure 6 Depicted is a flow chart of an example method for performing panoptic segmentation inference according to an example embodiment of the present disclosure.
[0023] Figure 7 Depicted is a flow chart of an example method for training a diffusion model to perform panoptic segmentation according to an example embodiment of the present disclosure.
[0024] Fig. 8A A block diagram of an example computing system is depicted according to an example embodiment of the present disclosure.
[0025] Figure 8B A block diagram of an example computing device is depicted according to an example embodiment of the present disclosure.
[0026] Figure 8C A block diagram of an example computing device is depicted according to an example embodiment of the present disclosure.
[0027] Reference numerals repeated in multiple figures are intended to identify like features in the various implementations. DETAILED DESCRIPTION
[0028] The present disclosure provides systems and methods for performing panoptic segmentation on images and videos using a denoising diffusion model. Panoptic segmentation is a computer vision task that assigns semantic and instance labels to each pixel of an image. The semantic label describes the classification of each pixel (e.g., sky, car, dog, etc.), and the instance label provides a unique ID for each instance in the image. This task is challenging due to the need for high-dimensional one-to-many mapping, and traditional methods typically utilize complex pipelines involving object detection, segmentation, and merging multiple predictions.
[0029] In this disclosure, the panoptic segmentation task is formulated as a conditional discrete data generation problem. This is achieved by learning a generative model for panoptic masks conditioned on the input image, which are treated as arrays of discrete labels, for example. The generative model can also be applied to video data by including predictions from past frames as additional conditioning signals. This enables the model to learn to automatically track and segment objects across video frames.
[0030] Specifically, in some example implementations of the present disclosure, the generative model for predicting a panoramic segmentation mask may be a denoising diffusion model. For example, the denoising diffusion model used in the present disclosure may include an image encoder and a mask decoder. The image encoder may map raw pixel data from an input image into a high-level feature representation. The mask decoder may then generate a panoramic mask from a noisy mask conditioned on these image features. For example, given an input image, the model may start with random noise as an initial set of simulated bits and gradually refine its estimate to get closer to the value of a good panoramic mask. In some implementations, the image encoder is run only once, so the cost of multiple iterations depends only on the decoder.
[0031] Another aspect of the present disclosure relates to using analog bits to represent discrete markers in a panoramic mask. For example, a denoising diffusion model can generate analog bit representations of a panoramic mask, which can then be converted to a real-valued version of the panoramic mask. This allows the semantic identifier and instance identifier of each pixel to be represented using real values, while the model is able to operate in a space represented using analog bits.
[0032] Another aspect of the present disclosure relates to the training of denoising diffusion models. In some implementations, the model can be trained using a softmax cross entropy loss applied to the model's logits. This allows the model to directly model the underlying distribution over a set of base categories and use a weighted average of the base categories to obtain the simulated bits. Additionally or alternatively, in some implementations, the model can also be trained using a weighted loss function that assigns larger weights to mask tags associated with small objects. This can help improve the segmentation of small instances.
[0033] The systems and methods described herein can also be extended to video. For video panoptic segmentation, the model can generate a panoptic mask conditioned on one or more past mask predictions for previous image frames of the image and video. This allows the model to track and segment instances across frames without requiring explicit instance matching over time.
[0034] Thus, the present disclosure provides a general method for panoptic segmentation of images and videos. The use of a denoising diffusion model allows for modeling a large number of discrete labels simultaneously, which is difficult for other existing generative segmentation models. This method can potentially be further improved by optimizing the architecture, modeling choices, and training process described in this article.
[0035] The systems and methods of the present disclosure provide many technical effects and benefits. As an example, the present disclosure describes techniques for performing panoptic segmentation, which is a fundamental and complex visual task that assigns semantic and instance labels to each pixel of an image. The disclosed techniques address this challenge by formulating panoptic segmentation as a discrete data generation problem, such as using a denoising diffusion model to generate a panoptic segmentation mask that provides a corresponding semantic identifier and a corresponding instance identifier for each pixel of an image.
[0036] This approach has several advantages over previous techniques. Compared to previous methods that used complex, multi-stage systems, the proposed method simplifies the complex process of panoptic segmentation by using a more general framework. In particular, generative modeling of panoptic segmentation is very challenging because panoptic masks are discrete / categorized and can be very large. For example, to generate a 512×1024 panoptic mask, the model must generate over 1M discrete tokens (of both semantic and instance labels). This is expensive for autoregressive models because they are sequential in nature and do not scale well with the size of the data input. Therefore, methods that perform panoptic segmentation using autoregressive models are very computationally expensive because the forward computation of the decoder needs to be performed to predict each token. In contrast, the diffusion model described in this paper is better at handling high-dimensional data and does not operate in an inherently sequential manner, but rather predicts all tokens of the mask simultaneously. Therefore, using a diffusion model for panoptic segmentation as described in this paper represents a significant saving in computational resources such as processor cycles, memory usage, network bandwidth, etc.
[0037] The disclosed technology can be applied to a variety of fields or applications. As an example, in autonomous driving, it can help vehicles recognize and distinguish different objects and instances, such as pedestrians, other cars, and street signs, in real time, thereby improving the safety of autonomous vehicles. As another example, in the field of medical imaging, the technology can help segment different tissues, cells, or abnormalities, thereby helping to make diagnoses faster and more accurately. In addition, the technology can be used in augmented reality applications to understand and manipulate digital representations of the real world. As yet another example, in the field of robotics, it can help robots better understand and navigate their surroundings. In short, the application of this technology can potentially improve the accuracy and efficiency of any task involving image or video analysis.
[0038] Referring now to the drawings, example embodiments of the present disclosure will be discussed in further detail.
[0039] Reference now Figure 1 , depicts an exemplary process for performing panoptic segmentation on an input image using a denoising diffusion model according to an embodiment of the present disclosure. The input image 12 is obtained by a computing system including one or more computing devices. The input image 12 includes a plurality of pixels and may be a frame of any digital image or video sequence, for example, a frame from a 1080p or 4K video, or an image captured from a digital camera or mobile device.
[0040] The input image 12 may also be provided in various formats, such as JPEG, PNG, BMP or RAW. The input image 12 may be a color image or a grayscale image. The resolution of the input image 12 may vary. It may be a high-resolution image, which provides more detailed information and may potentially improve the accuracy of the panoramic segmentation. Alternatively, it may be a low-resolution image that requires less computing resources to process. The computing system may also adjust the resolution of the input image 12, for example by scaling down the high-resolution image or scaling up the low-resolution image.
[0041] In some implementations, the input image 12 may also be preprocessed before being fed to the denoising diffusion model 14. The preprocessing may include operations such as noise reduction, contrast enhancement, and normalization. These operations may help improve the quality of the input image 12 and make the panoptic segmentation task easier.
[0042] The input image 12 is processed by a denoising diffusion model 14. The denoising diffusion model 14 is a generative model that is particularly suitable for processing high-dimensional data. For example, the model can process images with thousands or even millions of pixels.
[0043] The denoising diffusion model 14 can be implemented in various computing systems, including servers, personal computers, and mobile devices. The model 14 can also be implemented in different programming languages, such as Python, Java, or C++. The specific implementation details may depend on the requirements of the panoptic segmentation task and the constraints of the computing system.
[0044] The output of the denoising diffusion model 14 is a panoptic segmentation mask 16. The panoptic segmentation mask 16 provides a corresponding semantic identifier and a corresponding instance identifier for each pixel in the input image 12. For example, the semantic identifier may classify a pixel as belonging to a class such as "sky", "car", "dog", etc., while the instance identifier assigns a unique ID to each instance in the image, thereby being able to distinguish between multiple instances of the same classification.
[0045] In some implementations, semantic identifiers may be assigned based on a predefined set of classifications. For example, for a panoramic segmentation task involving an outdoor scene, the set of classifications may include "sky", "building", "car", "pedestrian", "tree", etc. A unique semantic identifier may be assigned to each classification in the set, which is then used to mark the pixels in the mask 16. The set of classifications may be defined by the user, or it may be automatically learned by the denoising diffusion model. The semantic identifier may be represented in a variety of formats. For example, it may be represented as a binary code, a one-hot vector, or a probability distribution of the set of classifications. The specific representation may depend on the capabilities of the denoising diffusion model 14 and the requirements of the panoramic segmentation task.
[0046] The instance identifier provides a unique ID for each instance in the image. For example, if the image contains multiple cars, each car will be assigned a unique instance identifier. Depending on the specific requirements of the panoptic segmentation task, the identifier can be represented in various forms such as integers or strings. The range of the integer can be determined based on the maximum number of instances that the denoising diffusion model is expected to handle. For example, if the model is expected to handle a maximum of 1000 instances, the integer can range from 0 to 999. Integers can also be represented in various number systems such as binary, decimal, hexadecimal, etc.
[0047] Depending on the resolution of the input image and the requirements of the panoptic segmentation task, the panoptic segmentation mask 16 can be generated with various resolutions. A high-resolution mask provides more detailed information and can potentially improve the accuracy of the segmentation. On the other hand, a low-resolution mask requires less computing resources to generate and process. The resolution of the panoptic segmentation mask 16 can be adjusted by the computing system, for example, by scaling down the high-resolution mask or scaling up the low-resolution mask.
[0048] The panoptic segmentation mask 16 may be provided as an output for various applications such as object detection, instance segmentation, and image or video analysis. For example, the output may be used in an autonomous driving system, a video surveillance system, or image editing software. The panoptic segmentation mask 16 may be provided as an output in various formats such as a binary file, a text file, or an image file. The panoptic segmentation mask 16 may also be displayed on a display device or stored in a storage device.
[0049] During inference, the network generates target data in parallel, e.g., using a much smaller number of iterations than the number of pixels, which can significantly improve computational efficiency.
[0050] In some embodiments, the panoptic segmentation mask 16 is also used as a condition for generating panoptic masks for subsequent frames in the video sequence. This allows the model to track and segment instances across frames without requiring explicit instance matching over time, thereby achieving smooth and consistent instance tracking in video data.
[0051] More specifically, still referring to Figure 1 , the problem of generating a panoptic segmentation mask can be formulated as follows. The panoptic segmentation mask 16 can be represented by two channels, \(m\in\mathbb{Z}\) H×W×2 . The first channel represents the class or classification label, and the second channel represents the instance ID.
[0052] Given that the instance IDs can be permuted without changing the underlying instances, some example implementations can randomly assign integers in \([0, K]\) to instances each time an image is sampled during training, where \(K\) is the maximum number of instances allowed in any image, and 0 represents the empty label. The task of solving the panoptic segmentation problem involves learning an image-conditioned panoptic mask generation model, e.g., by maximizing \(\sum\) i \(\log P(m\) i |x i ), where \(m\) i is a random categorical variable corresponding to the panoptic mask of the image \(x\) i in the training data. Considering that the panoptic mask may consist of hundreds of thousands or even millions of discrete tokens, generative modeling can be very challenging, especially for autoregressive models.
[0053] As a solution to the above problem, some example implementations can utilize diffusion models that use simulated bits. Different from autoregressive generative models, diffusion models have been shown to be more effective for high-dimensional data. Training a diffusion model can involve learning a denoising network. During the inference phase, the network generates the target data in parallel, using significantly fewer iterations than the number of pixels. Essentially, the diffusion model learns a sequence of state transitions that transform the noise \(\epsilon\) from a known noise distribution to a data sample \(x\) in the data distribution \(p(x)\) 0 .
[0054] To learn this mapping, in some implementations, the forward transition from the data \(x\) 0 to the noisy sample \(x\) t can be defined as follows: where \(\epsilon\) is drawn from the standard normal density, \(t\) is drawn from the uniform density on \([0, 1]\), and \(\gamma(t)\) is a monotonically decreasing function from 1 to 0. During training, the neural network \(f(x\) t , \(t)\) will learn to predict \(x\) t (or \(\epsilon\)) from \(x\) 0 , typically formulated as a denoising task using the \(L\) 2 loss:
[0055]
[0056] To generate samples from the learned model, the model can start from a noise sample \(x\) TWe then perform a series of (inverse) state transitions x by iteratively applying the denoising function f with appropriate transition rules. T →x T-Δ →…→x 0 .
[0057] Traditional diffusion models assume continuous data and Gaussian noise and are not directly applicable to discrete data. To model discrete data, analog bit-based methods first convert integers representing discrete tokens into bit strings and then project the bits of the bit strings into real numbers (also called analog bits) to which the continuous diffusion model can be applied. To extract samples, analog bit-based methods use a traditional sampler in continuous diffusion and then use a final quantization step (e.g., simple thresholding) to obtain categorical variables from the generated analog bits. Examples of this approach can generally correspond to Figure 1 , where the denoising diffusion model 14 generates a panoptic segmentation mask 16 based on this principle.
[0058] Figure 2 An illustration of an exemplary denoising diffusion model architecture 200 for panoptic segmentation of an input image 12 is provided. The architecture 200 includes an image encoder 204 and a mask decoder 206. The input image 12 is an initial data point for the denoising diffusion model 200, which may have a value represented as x∈R H×W×3 Dimension.
[0059] The first step of the process involves an image encoder 204, which may be a neural network, transforming raw pixel data into a latent representation vector, thereby creating a feature map 208. For example, the image encoder 204 may be operable to transform the raw pixel data of the input image 12 into a latent representation vector, represented, for example, as R H′×W′×d A high-level feature map 208 of dimensions H′ and W′, where H′ and W′ represent the height and width of the panoramic mask 16. The size of the panoramic mask 16 can be equal to, larger than, or smaller than the original input image 12. The feature map 208 can be designed to maintain sufficient resolution and integrate features of different scales. In some implementations, the feature map 208 can be generated by the encoder 204 using a series of convolutions with bilateral connections and upsampling operations for merging features from different resolutions. For example, the encoder 204 can be a ResNet model followed by a transformer encoder layer.
[0060] Specifically, one possible implementation of the image encoder 204 may include a residual neural network followed by one or more transformer encoder layers. The residual neural network can be used to extract high-level features from the input image, and the transformer encoder layer can be used to further process these features. The specific architectures of the residual neural network and the transformer encoder layer may be different. For example, the residual neural network may include different numbers of layers, different types of activation functions, and different types of pooling operations. The transformer encoder layer may also include different numbers of layers, different types of attention mechanisms, and different types of normalization operations.
[0061] In some implementations, the image encoder 204 may also include convolutions with bilateral connections and upsampling operations for merging features from different resolutions. This allows the image encoder 204 to capture information at different scales, which may be very beneficial for panoptic segmentation tasks. Convolutions may be implemented by different types of convolutional layers, such as standard convolutional layers, dilated convolutional layers, or depthwise separable convolutional layers. Bilateral connections may be implemented by different types of connection modes, such as skip connections, residual connections, or dense connections. Upsampling operations may be implemented by different types of upsampling methods, such as nearest neighbor upsampling, bilinear upsampling, or transposed convolution upsampling.
[0062] Still refer to Figure 2 Next, the mask decoder 206 utilizes the feature map 208 along with the noisy mask 210 as its input. During the inference phase, the mask decoder 206 iteratively refines the panoramic mask, operating on the image features as conditions. More specifically, the mask decoder 206 may take as its input the concatenated image feature map from the encoder and the noisy mask (e.g., randomly initialized or from a previous iteration) and generate a refined prediction of the mask 16.
[0063] Compared to the standard U-Net architecture commonly used in image generation and image-to-image translation tasks, a notable feature of some example implementations of the mask decoder 206 is the deployment of transformer decoder layers on top of the U-Net. These layers may include a cross-attention mechanism that integrates the encoded image features 208 (e.g., before performing upsampling operations). This unique design helps to effectively refine the panoramic mask 16, thereby helping to improve the overall performance of the denoising diffusion model 200.
[0064] Therefore, one possible implementation of the mask decoder 206 may include one or more transformer layers on top of a U-Net architecture. The U-Net architecture is a convolutional neural network that is particularly effective for image segmentation tasks. It consists of a downsampling path and an upsampling path, which enables it to capture contextual and spatial information. On the other hand, transformer layers can model long-range dependencies in the data and handle variable-sized inputs, which makes them particularly useful for panoptic segmentation tasks.
[0065] In some implementations, the mask decoder 206 may also include a criss-cross attention layer for integrating the encoded image features 208. Cross-attention is a mechanism that allows the model to focus on different parts of the input when generating each part of the output. This can help the mask decoder 206 generate a more accurate panoptic segmentation mask by taking into account relevant image features.
[0066] The final output of the denoising diffusion model 200 is a panoptic segmentation mask 16 that assigns a different semantic identifier and instance identifier to each pixel present in the input image 12. This resulting mask 16 is then output, marking the completion of the panoptic segmentation process.
[0067] The denoising diffusion model 200 is particularly good at handling high-dimensional data and is a substantial improvement over traditional autoregressive generative models. The model 200 is able to model a large number of discrete labels, making it well suited for complex panoptic segmentation tasks. The architecture of the model, in particular the separation of the image encoder 204 and the mask decoder 206, enables efficient processing and iterative refinement of the panoptic mask.
[0068] Specifically, Figure 2 As shown, the architecture of the denoising diffusion model 200 is intentionally divided into two main parts: an image encoder 204 and a mask decoder 206. This separation is very important because the sampling process of the diffusion model is iterative, which means that the forward pass of the network is usually performed multiple times during inference. The image encoder 204 is responsible for transforming the raw pixel data from the input image 12 into a high-level representation vector, which may be performed only once, while the mask decoder 206 iteratively refines the panoramic mask 16 based on these image features 208.
[0069] An example inference algorithm is as follows:
[0070]
[0071] This application Figure 3An example method for performing panoptic segmentation on a sequence of video frames using a denoising diffusion model according to an example embodiment of the present disclosure is shown. In the depicted embodiment, an input image 312 is obtained from a series of video frames, which may be captured by a camera, retrieved from a digital video file, or originated from a video streaming service, for example.
[0072] The input image 312 is processed by the denoising diffusion model 14. In the context of video panoptic segmentation, Figure 3 As shown, the model can generate a panoptic segmentation mask 316 conditioned not only on the input image 312, but also on a previous panoptic segmentation mask 318 generated for a previous image frame in the video. For example, the previous image frame can be the immediately previous frame in the sequence, or it can be a frame some set number of steps ago. This approach allows the model to track and segment instances across video frames without requiring explicit instance matching over time, which might be achieved through complex object tracking algorithms or optical flow methods.
[0073] therefore, Figure 3 An example extension for video is shown. Specifically, the proposed image-conditional panoramic mask modeling using p(m|x) is directly applicable to video panoptic segmentation by considering a 3D mask of a given video (e.g., with an additional temporal dimension). To adapt to online / streaming video settings, such as Figure 3 As shown, Model 14 can model p(m t |x t , m t-1 , m t-k ), thereby generating a panoramic mask conditioned on the image and the past mask predictions. This variation can be achieved by concatenating the past panoramic mask 314 (m t-1 , m t-k ) can be easily implemented with existing noisy masks, such as Figure 3 Apart from this small change, the model can remain the same as above, which is simple and allows fine-tuning the image panorama model for videos.
[0074] By iterating the refinement process, the framework is also conveniently adapted to streaming video settings where there are strong dependencies between adjacent frames. In video settings, similar results can be achieved with fewer inference steps when there are relatively small changes in video frames. Therefore, some example implementations can adaptively set the refinement steps across video frames.
[0075] Reference now Figure 4, which illustrates a denoising diffusion model architecture 400 for implementing panoptic segmentation for a video frame sequence according to an example embodiment of the present disclosure. As shown in the figure, the denoising diffusion model 400 includes an image encoder 204 and a mask decoder 406.
[0076] The image encoder 204 can be operated to transform the raw pixel data obtained from the input image 12 into a high-level feature representation, conceptualized as a feature map 208. For example, the image encoder 204 can use a convolutional neural network or other such neural network to perform the transformation process. Further, additional components such as pooling layers and fully connected layers can also be integrated to achieve more advanced feature extraction.
[0077] The mask decoder 406 generates the panoramic mask 16 from the noisy mask 210. The generation process can be conditioned on the image features 208 obtained from the image encoder 204 and one or more previous panoramic segmentation masks such as masks 408 and 410. In some implementations, the noisy mask 210 (which can be initialized to random noise or any other suitable initialization strategy) is used as the initial simulated bits. The model 400 systematically refines these initial estimates to get closer to the optimal panoramic mask. In some implementations, the image encoder 204 is executed only once, and therefore the computational cost of multiple iterations depends primarily on the mask decoder 406.
[0078] therefore, Figure 4 Integrating prior panoptic segmentation masks 408 and 410 in a video frame processing sequence is demonstrated. For video panoptic segmentation, the model 400 can formulate a panoptic mask conditioned not only on the input image 12 but also on one or more past mask predictions corresponding to prior image frames of the video. This unique feature enables the model 400 to track and segment instances across frames without requiring explicit instance matching over time.
[0079] Finally, the output of the denoising diffusion model 400 is a panoptic segmentation mask 16, which provides a corresponding semantic identifier and a corresponding instance identifier for each pixel of the input image 12. The generation of this mask 16 symbolizes the completion of the panoptic segmentation process. The panoptic segmentation mask 16 can then be used for various applications, such as object recognition, video analysis, and autonomous navigation.
[0080] In the field of video panoptic segmentation, the denoising diffusion model 400 can be viewed as a conditional discrete data generation model that incorporates predictions from previous frames as additional conditional signals. This functionality allows the model 400 to learn to automatically track and segment objects across video frames. This approach has several advantages over previous methods, particularly in terms of handling high-dimensional data and providing substantial savings in computational resources.
[0081] refer to Figure 5, an exemplary method for training a denoising diffusion model to perform panoptic segmentation is shown. The initial stage of the training process involves acquiring a training input image 512 and a ground truth panoptic segmentation mask 518. The training input image 512 (which may be derived from a variety of databases such as, for example, ImageNet, MS-COCO, or Cityscapes) is processed by the denoising diffusion model 14 to construct a predicted panoptic segmentation mask 516.
[0082] The predicted panoptic segmentation mask 516, which is the output of the model, is then compared to the ground truth panoptic segmentation mask 518 using a loss function 520. As an example, the loss function 520 may be a softmax cross entropy loss implemented on the logits (e.g., the non-normalized output) of the denoising diffusion model 14. In particular, compared to using L 2 Different from the traditional diffusion model of denoising loss, softmax cross entropy produces better performance in panoptic segmentation tasks. Softmax cross entropy loss allows the network to directly model the underlying distribution over the base categories and use a weighted average to obtain the simulated bit.
[0083] Additionally or alternatively, the loss function 520 can be a weighted loss function that assigns larger weights to mask markers associated with small objects, thereby providing a bias for improved segmentation of smaller instances. For example, this approach can assign higher weights to mask markers associated with small objects. Loss weighting can be achieved by calculating the pixel count of each instance and assigning a weight that is inversely proportional to the 'p' power of the pixel count, where 'p' is an adjustable parameter. This approach ensures that the model gives roughly equal importance to all objects in the image, regardless of their size.
[0084] Based on the evaluation of the loss function 520, the denoising diffusion model 14 is updated. The update may include adjusting the weights and biases of the model and refining the parameters of the model via techniques such as back propagation and gradient descent. This iterative training process enables the denoising diffusion model 14 to gradually enhance its ability to perform panoptic segmentation tasks.
[0085] An example training algorithm is as follows:
[0086] def train_loss(images,masks):
[0087] ″″″images:[b,h,w,3],masks:[b,h’,w’,2].″″″
[0088] #Encode image features.
[0089] h = pixel_encoder(images)
[0090] # Discrete mask to analog bits.
[0091] m_bits=int2bit(masks).asty pe(float)
[0092] m_bits = (m_bits*2-1)*scale
[0093] #Destroy the analog bits.
[0094] t=uniform(0,1) #scalar.
[0095] eps=normal(mean=0,std=1) #Same shape as m_bits.
[0096] m_crpt=sqrt(gamma(t))*m_bits+sqrt(1-gamma(t))*eps
[0097] #Predict and calculate loss.
[0098] m_logits,=mask_decoder(m_crpt,h,t)
[0099] loss=cross_entropy(m_logits, masks)
[0100] return loss.mean()
[0101] refer to Figure 6 , the flowchart shows an illustrative method for implementing panoptic segmentation reasoning according to several embodiments currently disclosed. The method starts at step 602, where a computing system that may be composed of several computing devices obtains an input image composed of a plurality of pixels. The input image may be an independent photo or a single frame extracted from a video sequence.
[0102] Step 604 details how the computing system processes the input image using a denoising diffusion model designed to generate a panoptic segmentation mask. In some implementations, the denoising diffusion model trained to perform multiple state transitions efficiently transforms random noise in a known noise distribution into data samples that match the data distribution. This transformation can be achieved by applying a denoising function that follows a specific transformation rule.
[0103] The resulting panoptic segmentation mask assigns a unique semantic identifier and instance identifier to each pixel in the input image. The semantic identifier classifies each pixel, while the instance identifier provides a unique ID for each instance in the image, making it possible to distinguish between various instances of the same classification.
[0104] In some implementations, to create a panoptic segmentation mask, a denoising diffusion model processes an input image to produce an analog bit representation of the panoptic segmentation mask. The analog bit representation is then converted to a real-valued version of the panoptic segmentation mask, where a semantic identifier and an instance identifier for each pixel are represented as real values in the mask.
[0105] The method ends at step 606, where the computing system provides a panoptic segmentation mask as an output. The output can potentially be applied for various purposes, such as image recognition, object detection, or video analysis.
[0106] One possible implementation of step 606 may involve displaying the panoptic segmentation mask on a display device connected to or integrated with the computing system. The display device may be a monitor, a projector, a television screen, or a virtual reality headset. The panoptic segmentation mask may be displayed as an image where the color or intensity of each pixel corresponds to its semantic identifier or instance identifier. This allows a user to visually inspect the results of the panoptic segmentation.
[0107] Another possible implementation of step 606 may involve storing the panoptic segmentation mask in a storage device connected to or integrated with the computing system. The storage device may be a hard disk, a solid-state drive, a USB flash drive, a memory card, or a cloud storage service. The panoptic segmentation mask may be stored as a file in various formats, such as a binary file, a text file, or an image file.
[0108] Another possible implementation of step 606 may involve transmitting the panoptic segmentation mask to another system via a communication network. The other system may be a server, a client, a peer, or a network service. The communication network may be a local area network, a wide area network, the Internet, or a cellular network. The panoptic segmentation mask may be transmitted as a stream of packets that are then reassembled, decoded, and converted to a panoptic segmentation mask by another system. This allows the panoptic segmentation mask to be used in a distributed computing environment or incorporated into a larger data processing pipeline.
[0109] In some implementations, the denoising diffusion model used in the method can be trained using a softmax cross entropy loss applied to the model's logits, and / or by assigning a weighted loss function that assigns greater weights to mask labels associated with smaller objects. When the model is applied to a video sequence, the denoising diffusion model can create a panoramic mask conditioned on one or more past mask predictions of an image and previous image frames of the video.
[0110] This approach has several advantages over earlier panoptic segmentation methods. Specifically, the use of a denoising diffusion model allows modeling a large number of discrete labels, a task that can be challenging or even impossible for other existing generative segmentation models. In addition, the denoising diffusion model is more efficient for high-dimensional data, resulting in significant savings in computational resources.
[0111] refer to Figure 7 , a flowchart illustrating an example method for training a denoising diffusion model to perform panoptic segmentation according to an example embodiment of the present disclosure.
[0112] Step 702 involves obtaining, by a computing system, a training input image and a ground truth panoptic segmentation mask. In some implementations, in step 702, a computing system (which may be a server or server cluster) obtains training data from a data storage system, which may be a local or distributed storage system or a cloud-based storage service. The training input image may include pixel data in various formats, such as raster format, vector format, or a combination thereof. The ground truth panoptic segmentation mask (which may be manually annotated or obtained through other reliable sources) provides the correct semantic labels and instance labels for each pixel in the image.
[0113] In step 704, the computing system processes the training input image using the denoising diffusion model to generate a predicted panoptic segmentation mask as an output of the denoising diffusion model. In some implementations, the denoising diffusion model is a generative model designed to predict panoptic segmentation masks. The model can include various machine learning algorithms optimized for image processing tasks, such as deep neural networks, convolutional neural networks, and / or transformer networks.
[0114] Next, step 706 involves evaluating, by the computing system, a loss function that compares the predicted panoptic segmentation mask to the ground truth panoptic segmentation mask. In some implementations, the loss function measures the difference between the predicted panoptic segmentation mask and the ground truth panoptic segmentation mask. As an example, the loss function can be a mean squared error loss function, a cross entropy loss function, or any other suitable loss function used in machine learning tasks. The goal during the training process is to minimize the loss function, resulting in more accurate predictions from the denoised diffusion model.
[0115] Finally, step 708 involves modifying, by the computing system, one or more parameter values of one or more parameters of the denoising diffusion model based on the loss function. The parameters of the denoising diffusion model are adjusted to reduce the loss function, thereby improving the accuracy of the denoising diffusion model prediction. In some implementations, various optimization algorithms can be used to perform such adjustments, such as stochastic gradient descent (SGD), Adam, RMSProp, or other suitable optimization algorithms. The iterative process continues until the denoising diffusion model is sufficiently trained to perform accurate panoramic segmentation, which can be determined based on a predefined performance metric (such as accuracy or F1 score) reaching a predefined threshold and / or based on other stopping criteria.
[0116] Fig. 8A A block diagram of an example computing system 100 is depicted, in accordance with an example embodiment of the present disclosure. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0117] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smart phone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0118] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0119] In some implementations, the user computing device 102 may store or include one or more machine learning models 120. For example, the machine learning model 120 may be or may otherwise include various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models, including nonlinear models and / or linear models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include a multi-head self-attention model (e.g., a transformer model). Reference Figures 1 to 7 An example machine learning model 120 is discussed.
[0120] In some implementations, one or more machine learning models 120 may be received from the server computing system 130 via the network 180, stored in the memory 114 of the user computing device, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel segmentation across multiple different images).
[0121] Additionally or alternatively, one or more machine learning models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning model 140 may be implemented by the server computing system 140 as part of a web service (e.g., a segmentation service). Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.
[0122] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices by which a user can provide user input.
[0123] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0124] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by the one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0125] As described above, the server computing system 130 can store or otherwise include one or more machine learning models 140. For example, the model 140 can be or otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models can utilize an attention mechanism, such as self-attention. For example, some example machine learning models can include a multi-head self-attention model (e.g., a transformer model). Refer to Figures 1 to 7 Example models 140 are discussed.
[0126] The user computing device 102 and / or the server computing system 130 can train the model 120 and / or 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 can be separate from the server computing system 130 or can be a part of the server computing system 130.
[0127] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158, which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.
[0128] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, error back propagation. For example, a loss function may be back propagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters in multiple training iterations.
[0129] In some implementations, performing error back-propagation may include performing truncated back-propagation through time.The model trainer 160 may perform a variety of generalization techniques (eg, weight decay, dropout, etc.) to improve the generalization capabilities of the model being trained.
[0130] Specifically, the model trainer 160 can train the machine learning model 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, training pairs that can include training images and reference true segmentation masks.
[0131] In some implementations, if the user has provided consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.
[0132] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general purpose processor. For example, in some implementations, the model trainer 160 includes a program file stored on a storage device, loaded into a memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more computer executable instruction sets stored in a tangible computer readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.
[0133] Network 180 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications on network 180 may be conducted via any type of wired and / or wireless connection, using a variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0134] Fig. 8A An example computing system that can be used to implement the present disclosure is shown. Other computing systems may also be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training data set 162. In such implementations, the model 120 may be both trained and used locally at the user computing device 102. In some such implementations, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0135] Figure 8B Depicted is a block diagram of an example computing device 10 performing in accordance with an example embodiment of the present disclosure. Computing device 10 may be a user computing device or a server computing device.
[0136] Computing device 10 includes multiple applications (e.g., application 1 to application N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.
[0137] like Figure 8B As shown, each application can communicate with multiple other components of the computing device (such as, for example, one or more sensors, context managers, device state components, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to the application.
[0138] Figure 8C Depicted is a block diagram of an example computing device 50 performing in accordance with an example embodiment of the present disclosure. Computing device 50 may be a user computing device or a server computing device.
[0139] The computing device 50 includes a plurality of applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a public API across all applications).
[0140] The central intelligence layer includes many machine learning models. For example, Figure 8C As shown, a corresponding machine learning model may be provided for each application, and the corresponding machine learning model may be managed by the central intelligence layer. In other implementations, two or more applications may share a machine learning model. For example, in some implementations, the central intelligence layer may provide a single model for all applications. In some implementations, the central intelligence layer is included in the operating system of the computing device 50 or is otherwise implemented by the operating system.
[0141] The central intelligence layer may communicate with the central device data layer. The central device data layer may be a centralized repository for data of the computing device 50. Figure 8C As shown, the central device data layer can communicate with multiple other components of the computing device (e.g., such as one or more sensors, context managers, device state components, and / or additional components). In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0142] The technology discussed herein relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a variety of possible configurations, combinations, and partitions of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system, or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0143] Although the subject matter has been described in detail with respect to various specific example embodiments of the subject matter, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art may easily produce changes, modifications, and equivalents to such embodiments after understanding the foregoing. Therefore, the present disclosure does not exclude such modifications, variations, and / or additions to the subject matter that would be readily apparent to those of ordinary skill in the art. For example, a feature shown or described as part of one embodiment may be used together with another embodiment to produce yet another embodiment. Therefore, the present disclosure is intended to encompass such changes, variations, and equivalents.
Claims
1. A computer-implemented method for performing panoptic segmentation, the method include: Obtaining, by a computing system including one or more computing devices, an input image including a plurality of pixels; processing, by the computing system, the input image using a denoising diffusion model to generate a panoptic segmentation mask as an output of the denoising diffusion model, wherein the panoptic segmentation mask provides a respective semantic identifier and a respective instance identifier for each pixel in the plurality of pixels; and The panoptic segmentation mask is provided as an output by the computing system.
2. The computer-implemented method of claim 1 , wherein the denoising diffusion model comprises an image encoder and a mask decoder, wherein the image encoder maps the input image to a feature map, and wherein the mask decoder generates the panoptic segmentation mask from a noisy mask conditioned on the feature map.
3. The computer-implemented method of claim 2, wherein the image encoder comprises a residual neural network followed by one or more transformer encoder layers.
4. A computer-implemented method according to claim 2 or 3, wherein the image encoder comprises convolutions with bilateral connections and upsampling operations for merging features from different resolutions.
5. A computer-implemented method according to any one of claims 2 to 4, wherein the mask decoder comprises one or more transformer layers on top of a U-net, and a cross-attention layer for integrating image features from the feature map.
6. A computer-implemented method according to any preceding claim, wherein the input image is processed by the computing system using the denoising diffusion model to generate the panoptic segmentation mask as the output of the denoising diffusion model include: processing, by the computing system, the input image using the denoising diffusion model to generate an analog bit representation of the panoptic segmentation mask as the output of the denoising diffusion model; as well as The analog bit representation of the panoptic segmentation mask is converted, by the computing system, into a real-valued version of the panoptic segmentation mask, wherein the respective semantic identifier and the respective instance identifier for each pixel in the plurality of pixels comprise a real value included in the real-valued version of the panoptic segmentation mask.
7. The computer-implemented method of claim 6, wherein the analog bit representation of the panoptic segmentation mask is generated according to a scaling factor, and wherein the scaling factor is equal to 0.
1.
8. A computer-implemented method according to any preceding claim, wherein the denoised diffusion model has been trained using a softmax cross entropy loss applied on its logits.
9. A computer-implemented method according to any preceding claim, wherein the denoising diffusion model has been trained using a weighted loss function that assigns larger weights to mask labels with fewer instances.
10. A computer-implemented method according to any preceding claim, in: The input image comprises an input image frame from a video; and Processing, by the computing system, the input image using the denoised diffusion model includes processing, by the computing system, the input image using the denoised diffusion model and one or more previous panoptic segmentation masks generated for one or more previous image frames preceding the input image frame in the video. 11 . The computer-implemented method of claim 10 , wherein the one or more previous panoptic segmentation masks generated for the one or more previous image frames comprises a plurality of previous panoptic segmentation masks generated for a plurality of previous image frames.
12. A computer-implemented method according to claim 10 or claim 11, in: The denoising diffusion model comprises an image encoder and a mask decoder, wherein the image encoder maps the input image frame to a feature map, and wherein the mask decoder generates the panoptic segmentation mask from a noisy mask conditioned on the feature map and the one or more previous panoptic segmentation masks.
13. One or more non-transitory computer-readable media collectively storing instructions for performing panoptic segmentation, wherein execution of the instructions by a computing system causes the computing system to perform operations, the operations include: An input image including a plurality of pixels is obtained by the computing system; processing, by the computing system, the input image using a denoising diffusion model to generate a panoptic segmentation mask as an output of the denoising diffusion model, wherein the panoptic segmentation mask provides a respective semantic identifier and a respective instance identifier for each pixel in the plurality of pixels; and The panoptic segmentation mask is provided as an output by the computing system.
14. The one or more non-transitory computer-readable media of claim 13, wherein the denoising diffusion model comprises an image encoder and a mask decoder, wherein the image encoder maps the input image to a feature map, and wherein the mask decoder generates the panoptic segmentation mask from a noisy mask conditioned on the feature map.
15. The one or more non-transitory computer-readable media of claim 13 or 14, wherein the input image is processed by the computing system using the denoising diffusion model to generate the panoptic segmentation mask as the output of the denoising diffusion model include: processing, by the computing system, the input image using the denoising diffusion model to generate an analog bit representation of the panoptic segmentation mask as the output of the denoising diffusion model; as well as The analog bit representation of the panoptic segmentation mask is converted, by the computing system, into a real-valued version of the panoptic segmentation mask, wherein the respective semantic identifier and the respective instance identifier for each pixel in the plurality of pixels comprise a real value included in the real-valued version of the panoptic segmentation mask.
16. One or more non-transitory computer readable media according to claim 13, 14 or 15, in: The input image comprises an input image frame from a video; and Processing, by the computing system, the input image using the denoised diffusion model includes processing, by the computing system, the input image using the denoised diffusion model and one or more previous panoptic segmentation masks generated for one or more previous image frames preceding the input image frame in the video.
17. The one or more non-transitory computer-readable media of any one of claims 13 to 16, wherein the one or more non-transitory computer-readable media further stores the denoised diffusion model.
18. A computing system for training a denoising diffusion model to perform panoptic segmentation, the computing system comprising one or more processors and one or more non-transitory computer-readable media storing instructions for performing operations, the operations include: obtaining, by the computing system, a training input image and a reference true panoptic segmentation mask; processing, by the computing system, the training input image using the denoising diffusion model to generate a predicted panoptic segmentation mask as an output of the denoising diffusion model, wherein the predicted panoptic segmentation mask provides a respective semantic identifier and a respective instance identifier for each pixel in the plurality of pixels; evaluating, by the computing system, a loss function comparing the predicted panoptic segmentation mask to the ground truth panoptic segmentation mask; as well as One or more parameter values of one or more parameters of the denoising diffusion model are modified by the computing system based on the loss function.
19. The computing system of claim 18, wherein the loss function comprises a softmax cross entropy loss applied on the logits of the denoising diffusion model.
20. The computing system of claim 18 or 19, wherein the loss function comprises a weighted loss function that assigns larger weights to mask labels with fewer instances.