Identifying Target Objects Using a Segmentation Neural Network with Diverse Scales
By using scaled diversified segmentation neural networks and fuzzy sampling strategies in digital image editing systems, the accuracy and efficiency problems of existing systems when identifying objects depicted in digital images are solved, and more efficient and accurate object recognition is achieved.
Patent Information
- Application Number
- CN201910967936.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-24
- Filing Date
- 2019-10-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-10-12
AI Technical Summary
Existing digital image editing systems have accuracy and efficiency problems when identifying objects depicted in digital images, and conventional systems have difficulty identifying objects accurately and require a large number of user interactions.
Using a scale-diversified segmentation neural network, by generating segmentation proposals at different scales, allows users to select desired results from them, and train the neural network through fuzzy sampling strategies to improve the diversity of segmentation generation.
It improves the accuracy and efficiency of target object identification in digital images, reduces the number of user interactions, and allows the system to effectively identify objects at multiple scales.
Smart Images

Figure CN111353999B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to neural networks, and more particularly, to identifying target objects using a scale-diverse segmentation neural network. Background Art
[0002] In recent years, significant developments have been seen in hardware and software platforms for identifying and editing objects depicted in digital images. In fact, conventional digital image editing systems can select an object depicted in a digital image and then modify the digital image based on that selection (e.g., modifying the object depicted in the digital image or placing the object depicted in the digital image on a different background image). By way of illustration, conventional digital image editing systems can utilize a machine learning model trained on a large labeled digital image library to analyze a user's selection of one or more foreground pixels (e.g., via a pixel selection tool or a digital lasso tool), and then identify the object corresponding to the foreground pixels for editing.
[0003] Although conventional digital image systems can identify objects depicted in digital images, these systems still have many drawbacks in terms of accuracy and efficiency. For example, with regard to accuracy, conventional digital image editing systems often identify incorrect objects depicted in digital images. In fact, because many digital images depict a variety of different objects, there are multiple possible patterns / choices that are equally reasonable given a particular set of clicks. As a result, conventional systems often identify inaccurate objects (e.g., selecting an object that the user did not attempt to select). For example, in response to a user's indication of pixels within a logo on a person's shirt depicted in a digital image, there is ambiguity as to whether the user is attempting to select the logo, the shirt, or the person. Due to this potential ambiguity, conventional digital image editing systems often select the incorrect object.
[0004] In addition, conventional digital image editing systems also have many drawbacks in terms of efficiency. For example, conventional digital image editing systems typically require a large amount of user interaction (and a large amount of time) to select an object depicted in a digital image. In fact, conventional digital image editing systems may require a large number of different foreground and / or background pixel inputs to accurately identify the pixels corresponding to an object depicted in a digital image. By way of illustration, to isolate and select a shirt worn by a person depicted in a digital image, a conventional digital image editing system may require a large amount of user input to distinguish the foreground pixels of the desired shirt from the background pixels. This problem is only exacerbated when the desired object has visual features and characteristics similar to those of background objects (e.g., a digital image of a tree in front of background shrubs).
[0005] In addition, as described above, some digital image editing systems utilize machine learning models trained on large repositories of training digital images to identify objects depicted in digital images. Constructing and managing a training digital image repository with corresponding ground truth masks requires significant computational resources and time, which further reduces the efficiency of conventional systems. Some digital image editing systems attempt to avoid these computational costs by using encoding rules or heuristics to select models for objects. However, these non-machine learning methods introduce additional problems in terms of efficiency and accuracy. In fact, such systems are limited to handcrafted low-level features, which results in inefficient selection of different objects and excessive user interaction.
[0006] These and other problems exist with respect to identifying objects in digital visual media. SUMMARY OF THE INVENTION
[0007] Embodiments of the present disclosure utilize systems, non-transitory computer-readable media, and methods for training and utilizing neural networks to identify multiple potential objects depicted in digital media at different scales to provide benefits and / or solve one or more of the foregoing or other problems in the art. In particular, the disclosed systems can utilize a neural network to generate a set of scale-varying segmentation proposals based on user input. Specifically, given an image and user interaction, the disclosed systems can generate a diverse set of segmentations at different scales from which the user can select a desired result.
[0008] To train and evaluate such models, the disclosed systems can employ a training pipeline that synthesizes diverse training samples without the need to collect or generate new training data sets. In particular, the disclosed systems can utilize a training input sampling strategy that mimics ambiguous user input where multiple possible segmentations are equally reasonable. In this way, the disclosed systems can clearly encourage the model to more accurately learn the diversity in segmentation generation. Thus, the disclosed systems can utilize a fuzzy sampling strategy to generate training data, thereby effectively training a neural network to generate multiple semantically valid segmentation outputs (at different scale variations).
[0009] Additional features and advantages of one or more embodiments of the present disclosure are outlined in the following description, and will in part become apparent from the description, or can be learned by practice of these example embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] By using the drawings described briefly below, the detailed description provides one or more embodiments with additional features and details.
[0011] FIG. 1A illustrates an overview of a conventional segmentation method.
[0012] Figure 1BIllustrated is an overview of generating multiple object segmentation outputs from a digital image by a segmentation neural network utilizing scale diversification according to one or more embodiments.
[0013] Figures 2A - 2C Illustrated are the digital input, layers, and output of a scale-diversified segmentation neural network according to one or more embodiments, the scale-diversified segmentation neural network generating multiple object segmentation outputs corresponding to multiple scales using multiple channels.
[0014] Figure 3 Illustrated is a schematic diagram for training a scale-diversified segmentation neural network according to one or more embodiments.
[0015] Figure 4 Illustrated is a representation of explicit sampling and fuzzy sampling according to one or more embodiments.
[0016] Figures 5A - 5C Illustrated are generating explicit positive samples, explicit negative samples, explicit ground truth segmentations, fuzzy positive samples, fuzzy negative samples, and fuzzy ground truth segmentations according to one or more embodiments.
[0017] Figure 6 Illustrated is a schematic diagram for identifying a ground truth scale for ground truth segmentation according to one or more embodiments.
[0018] Figure 7 Illustrated is a schematic diagram of a multi-stage scale-diversified segmentation neural network according to one or more embodiments.
[0019] Figure 8 Illustrated is a schematic diagram of a scale-diversified segmentation neural network having a scale proposal neural network for generating input scales according to one or more embodiments.
[0020] Figures 9A - 9C Illustrated is a computing device having a graphical user interface according to one or more embodiments, the graphical user interface including user interface elements for identifying a user indicator and inputs at different scales and providing object segmentation outputs corresponding to different scales for display.
[0021] Figures 10A - 10D Illustrated is a computing device having a graphical user interface according to one or more embodiments, the graphical user interface including user interface elements for identifying a user indicator and providing object segmentation outputs corresponding to different scales for display.
[0022] Figure 11 Illustrated is a schematic diagram of an example environment in which a digital object selection system is implemented according to one or more embodiments.
[0023] Figure 12 FIG. illustrates a schematic diagram of a digital object selection system according to one or more embodiments.
[0024] Figure 13 FIG. illustrates a flowchart of a series of operations for generating an object segmentation output by using a trained scale-diverse segmentation neural network according to one or more embodiments.
[0025] Figure 14 FIG. illustrates a flowchart of a series of operations for training a scale-diverse segmentation neural network to generate an object segmentation output according to one or more embodiments.
[0026] Figure 15 FIG. illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0027] The present disclosure describes one or more embodiments of a digital object selection system that trains and utilizes a scale-diverse segmentation neural network to analyze digital images at different scales and identify different target objects depicted in the digital images. In particular, the digital object selection system can utilize a single-stage or multi-stage scale-diverse segmentation neural network to suggest multiple object segmentation outputs at different scales based on minimal user input. The digital object selection system can improve target object selection by allowing the user to select a suggested option from semantically meaningful alternatives defined with respect to scale, which results in improved interpretation of each output and identification of the target object after just a few user interactions.
[0028] Moreover, the digital object selection system can effectively train a scale-diverse segmentation neural network by clearly encouraging segmentation diversity by using explicit sampling and fuzzy sampling methods. In this way, the digital object selection system can simulate the ambiguity that occurs in user indicators / user input and learn diversity in segmentation generation. Thus, the object selection system can effectively train and utilize a scale-diverse segmentation neural network to resolve ambiguity and accurately select target objects within a digital image with minimal user input.
[0029] For illustration, in one or more embodiments, a digital object selection system receives a digital image (which depicts a target object) and user indicators (e.g., a foreground click, a background click, and / or an edge click corresponding to the target object). In response, the digital object selection system can utilize a scale-diverse segmentation neural network to generate multiple object segmentation outputs. Specifically, the digital object selection system can generate a first object segmentation output at a first scale based on the digital image and the user indicators using a scale-diverse segmentation neural network. Moreover, the digital object selection system can generate a second object segmentation output at a second scale based on the digital image and the user indicators using a scale-diverse segmentation neural network. Optionally, the digital object selection system can generate third, fourth, fifth, or more object segmentation outputs. As described above, the digital object selection system can provide object segmentation outputs at varying scales for display, thereby allowing a client device to select an object segmentation output that aligns with one or more target objects or other desired outputs.
[0030] As just described, the digital object selection system can generate a segmentation based on user indicators corresponding to a target object depicted in a digital image. In particular, the digital object selection system can analyze various user inputs indicating how one or more pixels are related to the target object depicted in the digital image. For example, the digital object selection system can analyze a foreground indicator (e.g., a foreground click), a background indicator, an edge indicator, a boundary region indicator (e.g., a bounding box), and / or a language indicator provided via a client device. Then, the digital object selection system can generate an object segmentation selection based on these multiple user input patterns and the digital image.
[0031] As discussed above, user indicators / inputs are typically ambiguous. The digital object selection system can address this ambiguity by generating scale-defined diverse object segmentation outputs. For example, the digital object selection system can define different scales in terms of size and aspect ratio. The digital object selection system can train and utilize a scale-diverse segmentation neural network that generates multiple segmentation outputs corresponding to (e.g., suitable for) different scale anchor boxes of different sizes and aspect ratios. For example, the scale-diverse segmentation neural network can generate segmentation masks and / or segmentation boundaries that indicate different objects (or object groupings) depicted in the digital image associated with different scales.
[0032] In one or more embodiments, the digital object selection system may also generate a semantically more meaningful object segmentation output by applying an object verification model that is part of a scale-diverse segmentation neural network. For example, the digital object selection system may incorporate a trained object classifier into the scale-diverse segmentation neural network architecture to determine (via an object score) that the proposed object segmentation output reflects the object depicted in the digital image or a semantically meaningful result.
[0033] When generating an object segmentation output, the digital object selection system may provide the object segmentation output for display via a client device. For example, the digital object selection system may provide different object segmentation outputs for display via a client device, thereby allowing a user to identify the object segmentation output that aligns with the target object or other desired output. Based on the user selection of the object segmentation output, the digital object selection system may select the corresponding target object (and modify the digital image based on the target object).
[0034] The digital object selection system may utilize a single-stage or multi-stage scale-diverse segmentation neural network. For example, in one or more embodiments, the digital object selection system utilizes a single-stage scale-diverse segmentation neural network that includes multiple output channels corresponding to multiple (predefined) scales. By utilizing different output channels that are trained to identify object segmentation outputs at different scales, the single-stage scale-diverse segmentation neural network can generate multiple object segmentation outputs in a single pass.
[0035] In other embodiments, the digital object selection system may utilize a multi-stage scale-diverse segmentation neural network. In particular, the digital object selection system may utilize a multi-stage scale-diverse segmentation neural network that is trained to analyze a continuous range of input scales (e.g., rather than predefined input scales via different scale channels). For example, the digital object selection system may utilize a multi-stage scale-diverse segmentation neural network with an additional scale input plane to analyze the scale input, and the multi-stage scale-diverse segmentation neural network generates an object segmentation output specific to the scale input. The digital object selection system may generate various different object segmentation outputs based on different scale inputs.
[0036] The digital object selection system may identify different scale inputs and generate different object segmentation outputs based on user input and / or based on a scale proposal neural network. For example, in one or more embodiments, the digital object selection system provides user interface elements for receiving scale input from a user (e.g., via a scale input slider or a timer input element that expands the scale based on user input time). In other embodiments, the digital object selection system may utilize a trained scale proposal neural network that analyzes the digital image and user indicators to generate an input scale.
[0037] As described above, the digital object selection system can also effectively train a scale-diverse segmentation neural network. In particular, the digital object selection system can utilize a supervised training method based on ground truth segmentation corresponding to a specific scale and training indicators within a training digital image to train the scale-diverse segmentation neural network. Additionally, the digital object selection system can generate training data from an existing training repository. For example, the digital object selection system can generate positive and negative samples from existing training images. Moreover, the digital object selection system can generate clear samples and ambiguous samples. For example, the digital object selection system can generate clear samples by collecting training indicators from foreground pixels and background pixels that define a single ground truth segmentation. The digital object selection system can generate ambiguous samples by collecting training indicators from common foreground pixels and / or common background pixels corresponding to multiple ground truth segmentations.
[0038] Compared with conventional systems and methods, the digital object selection system offers various advantages and benefits. For example, by generating multiple object segmentation outputs at different scale levels, the digital object selection system can improve the accuracy of identifying target objects in digital images. In fact, since user indicators / inputs regarding different combinations of objects in digital images are usually ambiguous, the digital object selection system can generate multiple object segmentation outputs to identify segmentations that are accurately aligned with the target object. In fact, the digital object selection system allows the user to select the segmentation that most closely approximates the desired output and provide additional refinement if necessary. Additionally, in one or more embodiments, compared with handcrafted low-level features, by leveraging a scale-diverse segmentation neural network, the digital object selection system learns to better understand the deep representation of the semantic content of the image.
[0039] Furthermore, the digital object selection system can also improve efficiency relative to conventional systems. In fact, the digital object selection system can utilize a scale-diverse segmentation neural network to analyze user indicators corresponding to digital images at different scales to generate a set of object segmentation selections. By providing a set of object segmentation selections for user interaction, the digital object selection system can allow for the efficient selection of an object segmentation corresponding to a specific target object depicted in the digital image with minimal user input. In fact, the digital object selection system can simplify the user selection process by allowing the user to select from a set of suggested selections after only a few clicks (or even a single click).
[0040] Moreover, the digital object selection system can provide additional efficiency in training scale-diverse segmentation neural networks. As described above, the digital object selection system can utilize existing training data to train scale-diverse segmentation neural networks, which reduces the processing power and time required to construct a labeled training dataset. Additionally, by using explicit and / or ambiguous training samples, the digital object selection system can improve efficiency while also improving the performance of generating distinguishable, diverse, and semantically relevant segmentations with respect to different scales.
[0041] As illustrated by the previous discussion, the present disclosure uses various terms to describe the features and advantages of the digital object selection system. Additional details regarding the meaning of these terms are now provided. For example, as used herein, the term "neural network" refers to a machine learning model that can be tuned (e.g., trained) based on inputs to approximate an unknown function. In particular, a neural network can include a model of interconnected artificial neurons (in different layers) that communicate and learn to approximate complex functions and generate outputs based on multiple inputs provided to the model. For example, a neural network can include a deep convolutional neural network (i.e., "CNN"), a fully convolutional neural network (i.e., "FCN"), or a recurrent neural network (i.e., "RNN"). In other words, a neural network is an algorithm that implements deep learning techniques, i.e., machine learning that uses a set of algorithms to attempt to model high-level abstractions in data.
[0042] Moreover, as used herein, a "scale-diverse segmentation neural network" refers to a neural network that generates object segmentation outputs for digital images based on scales. In particular, a scale-diverse segmentation neural network includes a fully convolutional neural network that analyzes user indicators (e.g., in the form of distance map input channels) and digital images at different scales (e.g., anchor regions such as anchor boxes) (e.g., in the form of RGB input channels) to generate object segmentation outputs (e.g., segmentation boundaries and segmentation masks).
[0043] As used herein, the term "scale proposal neural network" refers to a neural network that generates different scales. In particular, a scale proposal neural network includes a neural network that analyzes an input digital image and user indicators and generates multiple proposed scales. For example, the digital object selection system can utilize a scale proposal neural network to generate one or more scales that are utilized by the scale-diverse segmentation neural network to analyze the digital image.
[0044] As used herein, the term "object verification model" refers to a computer-implemented algorithm that determines an indication that a scale corresponds to one or more objects. In particular, the object verification model includes layers of a neural network that predicts an object score, the object score indicating whether a particular scale configuration includes an object. For example, the object verification model can include an object classifier neural network that determines an object score, the object score indicating whether an object segmentation output at a particular scale includes a complete object or a partial object.
[0045] As used herein, the term "digital image" refers to any digital visual representation (e.g., digital symbol, picture, icon, or illustration). For example, the term "digital image" includes digital files having the following file extensions: JPG, TIFF, BMP, PNG, RAW, or PDF. A digital image can include a part or portion of other digital visual media. For example, a digital image can include one or more frames of a digital video. Thus, a digital image can also include digital files having the following file extensions: FLV, GIF, MOV, QT, AVI, WMV, MP4, MPG, MPEG, or M4V. In fact, although many example embodiments are described with respect to digital images, the digital object selection system can also select objects in digital video frames.
[0046] As used herein, the term "object" refers to an item, design, person, or thing. In particular, the term object includes a person or thing depicted (represented) in a digital image. An object can include other objects. For example, a person (i.e., an object) in a digital image can include a shirt, pants, shoes, a face, etc. Similarly, a group of animals in a digital image can include multiple individual animals. Also, as used herein, the term "target object" refers to an object depicted in a digital image that is being attempted to be identified or selected. For example, the term "target object" includes an object reflected in a digital image that a user is attempting to identify or select.
[0047] As used herein, the term "user indicator" refers to user input related to a target object of a digital image (e.g., a user's selection of one or more pixels). In particular, the term user indicator includes user input indicating one or more pixels of a digital image and an indication of how the one or more indicated pixels correspond to the target object depicted in the digital image. For example, a user indicator can include a positive indicator (also referred to as a foreground indicator, such as a click or swipe on a foreground pixel of a target object), a negative indicator (also referred to as a background indicator, such as a click or swipe on a background pixel not included as part of the target object), an edge indicator (e.g., a click along the boundary or edge between a target object and the background), a bounding region indicator (e.g., user input of a bounding box or other shape enclosing the target object), or a language indicator (e.g., a language input, such as a text input or voice input indicating the pixels of a target object).
[0048] As used herein, the term "object segmentation output" (or "segmentation" or "object segmentation") refers to an indication of multiple pixels depicting one or more objects. For example, an object segmentation output can include a segmentation boundary (e.g., a boundary line or curve indicating the edges of one or more objects) or a segmentation mask (e.g., a binary mask identifying the pixels corresponding to an object).
[0049] As used herein, the term "scale" refers to an indication of relative portion, size, extent, or degree. In particular, scale includes an indication of a portion, size, extent, or degree of a digital image. For example, the term scale can include an anchor region of a particular size, shape, and / or dimension (e.g., an anchor box or an anchor circle). By way of illustration, the scale can include an anchor box having a particular size (e.g., area or dimension) and aspect ratio that defines a portion of a digital image. Similarly, the scale can include an anchor circle (or other shape) having a particular radius that defines a portion of a digital image.
[0050] As used herein, the term "training" is used as a modifier to describe information, data, or objects used to train a neural network. For example, a training digital image depicting a training object refers to a digital image depicting an object (e.g., an object corresponding to a ground truth mask or a collection of individual objects) used to train a neural network. Similarly, a training indicator refers to a user indicator (or a sample approximating a user indicator) used to train a neural network. As described below, a training indicator can include an explicit indicator (sometimes referred to as an explicit sample, which refers to a training indicator indicating a specific object segmentation in a digital image) and / or an ambiguous indicator (sometimes referred to as an ambiguous sample, which refers to a training indicator indicating multiple possible object segmentations in a digital image). Similarly, as used herein, the term "ground truth segmentation" refers to a segmentation of the pixels indicating a training object (e.g., a ground truth boundary or a ground truth mask).
[0051] Additional details regarding various embodiments of the digital object selection system will now be provided in conjunction with the illustrative drawings. For example, as discussed above, by generating multiple scale-varied object segmentation outputs, the digital object selection system can improve efficiency and accuracy relative to conventional systems. FIGS. 1A- Figure 1B A comparison is made between applying conventional methods to identify target objects in digital images and one or more embodiments of the digital object selection system.
[0052] Specifically, FIG. 1A illustrates a digital image 100 and a corresponding user indicator 101 (i.e., a foreground (or positive) indicator). As shown, the conventional system provides the digital image 100 and the user indicator to a model 102, and the model 102 identifies a segmentation 104 of three dogs depicted in the digital image 100. However, as shown in FIG. 1A, the digital image 100 contains multiple different objects, and the user indicator 101 is ambiguous as to which combination of different objects is desired as the target object. In fact, the digital image 100 depicts three different dogs on a blanket lying on a bed. Thus, the user indicator 101 may indicate a desire to select one dog, two dogs, three dogs, three dogs and the blanket, or three dogs, the blanket, and the bed. Despite this ambiguity, the model 102 generates a segmentation 104 of three dogs.
[0053] This method requires various additional user inputs to select a specific target object. For example, to select a single dog, the conventional system of FIG. 1A would require many user indicators to distinguish the desired dog from the other objects depicted in the digital image 100. For example, the client device would need to capture negative user indicators around the desired dog to exclude the blanket, the bed, and the other dogs from the resulting segmentation.
[0054] In contrast, Figure 1B illustrates the use of a scale-diverse segmentation neural network 106 according to one or more embodiments of the digital object selection system. As shown, the digital object selection system utilizes the scale-diverse segmentation neural network 106 to analyze the digital image 100 and the user indicator 101 to generate multiple scale-varied object segmentation outputs 108-112. In fact, as Figure 1B shown, the object segmentation output 108 identifies a single dog, the object segmentation output 110 identifies three dogs, and the object segmentation output 112 identifies three dogs and the blanket on which the dogs are sitting. The digital object selection system provides the object segmentation outputs 108-112 for display via the client device. Moreover, if the user attempts to select a single dog, the user can interact with the object segmentation output 108 via the client device. Thus, by providing a single user indicator, the user can identify the appropriate segmentation from among the multiple segmentations generated by the digital object selection system.
[0055] As Figure 1B shown, the digital object selection system generates multiple object segmentation outputs 108 - 112 based on multiple scales. In fact, the digital object selection system can apply a first (small) scale to generate object segmentation output 108, a second (medium) scale to generate object segmentation output 110, and a third (large) scale to generate object segmentation output 112. As illustrated, the digital object selection system can thus generate multiple semantically meaningful segmentations (e.g., segmentations depicting meaningful complete objects) in a logical progression (e.g., based on scale) to allow for fast and accurate target object selection.
[0056] Although Figure 1B three object segmentation outputs are illustrated, the digital object selection system can generate additional (or fewer) object segmentation outputs. For example, in some embodiments, the digital object selection system generates 12 object segmentation outputs of different scales (e.g., segmentations including a bed, two dogs, etc.). Also, although Figure 1B a specific user indicator (e.g., a positive indicator) is illustrated, the digital object selection system can analyze various different inputs.
[0057] In fact, as described above, the digital object selection system can analyze various combinations of user inputs via a scale - diverse segmentation neural network to generate various different object segmentations. For example, Figures 2A - 2C illustrates the input of a scale - diverse segmentation neural network 201, the architecture of the scale - diverse segmentation neural network 201, and the output of the scale - diverse segmentation neural network 201 according to one or more embodiments.
[0058] Specifically, Figure 2A illustrates a digital image 200 with user indicators, where the user indicators include a positive indicator 204 (e.g., a positive click on the pixels of the target object) and a negative indicator 206 (e.g., a negative click on the background pixels outside the target object). The digital object selection system can recognize various types of user inputs as positive and negative indicators. For example, in one or more embodiments, the digital object selection system recognizes a left - mouse - button click, a single - tap touch gesture, a circle, or other types of user inputs as an indication of a positive user indicator. Similarly, the digital object selection system can recognize a right - mouse - button click, a double - tap touch gesture, an "x" as an indication of a negative user indicator.
[0059] As Figure 2A shown, the digital object selection system utilizes the digital image and user indicators to generate a distance map. For example, as Figure 2AAs shown, the digital object selection system generates distance maps 210, 212 based on digital image 200 and user indicators 204, 206. In particular, the digital object selection system generates a positive distance map 210 based on positive user indicator 204. Moreover, the digital object selection system generates a negative distance map 212 based on negative user indicator 206.
[0060] As used herein, a "distance map" refers to a digital item that reflects the distance between a pixel in a digital image and a selected pixel. For example, a distance map can include a database or digital file that includes the distance between a pixel in a digital image and a selected pixel (i.e., a positive or negative user indicator). For example, a positive distance map includes digital items that reflect the distance between a pixel in a digital image and a selected pixel that is part of a target object. Similarly, a negative distance map includes digital items that reflect the distance between a pixel and a selected pixel that is not part of a target object.
[0061] For example, with respect to Figure 2A , the positive distance map 210 includes a two-dimensional matrix having entries for each pixel in the digital image 200. Specifically, the positive distance map 210 includes a matrix having entries for the pixels in the digital image 200, where each entry reflects the distance between the pixel corresponding to the entry and the positive user indicator 204. Thus, as illustrated, entry 214 in the positive distance map 210 reflects the distance (i.e., 80 pixels) between the pixel corresponding to entry 214 and the pixel corresponding to the positive user indicator 204.
[0062] Similarly, the negative distance map 212 includes a two-dimensional matrix having entries for the pixels in the digital image 200. Specifically, each entry in the negative distance map 212 reflects the distance between the pixel corresponding to the entry and the negative user indicator 206. Thus, as illustrated, entry 216 reflects the distance (i.e., 255 pixels) between the pixel corresponding to entry 216 and the pixel corresponding to the negative user indicator 206.
[0063] As Figure 2AAs shown, the digital object selection system can also provide (alternatively or) additional maps 213 as inputs to the scale-variant segmentation neural network 201. For example, with respect to an edge indicator (e.g., a click indicating the edge of a target object), the digital object selection system can provide an edge distance map reflecting the distance between the selected edge pixels and other pixels of the digital image. Similarly, for a bounding box indicator, the digital object selection system can provide a bounding distance map reflecting the distance between any pixel of the digital image and the pixels of the bounding box. The digital object selection system can provide each distance map via a specific channel (e.g., an edge channel for the edge distance map), which is trained to analyze a specific type of user input.
[0064] Although Figure 2A FIG. illustrates a single positive user indicator and a single negative user indicator, it should be understood that the digital object selection system can also generate distance maps based on additional (or fewer) user indicators. For example, in the case where the digital object selection system receives multiple positive user indicators (or multiple edge indicators), the digital object selection system generates a distance map reflecting the distance between the pixels and the closest user indicator. Similarly, in the case where the digital object selection system receives multiple negative user indicators, the digital object selection system generates a negative distance map reflecting the distance between the pixels and the closest negative user indicator. In other embodiments, the digital object selection system generates a separate distance map for each user indicator.
[0065] In addition to the distance maps, the digital object selection system also utilizes one or more color channels. For example, with respect to Figure 2A , the digital object selection system utilizes channels for three colors, an R channel 218 (corresponding to red), a G channel 220 (corresponding to green), and a B channel 222 (corresponding to blue). In particular, in one or more embodiments, each color channel 218 - 222 includes a two-dimensional matrix (e.g., a color map) having an entry for each pixel in the digital image 200. Specifically, as shown, the B channel 222 includes a matrix (e.g., a blue map) having an entry for each pixel in the digital image 200, where each entry (e.g., entry 224) reflects the amount of blue corresponding to each pixel. Thus, entry 224 corresponding to a pixel having very little blue reflects a low value (i.e., one) in the B channel 222.
[0066] Although illustrated as three separate channels, the digital object selection system can utilize fewer or more channels. For example, the digital object selection system can utilize four color channels in conjunction with a CMYK image. Similarly, the digital object selection system can utilize a single color channel with respect to a grayscale image. Moreover, although with respect to Figure 2AShown as the R, G, and B channels, it will be understood that the digital object selection system can utilize various other colors or color spaces for the color channels. For example, in one or more embodiments, the digital object selection system utilizes the LAB color space and LAB color channels instead of the RGB color space and RGB color channels.
[0067] In one or more embodiments, the digital object selection system generates an image / user interaction pair (e.g., a combination of a distance map and color channels). For example, Figure 2A An image / user interaction pair is generated by combining the user interaction data reflected in the positive distance map 210 and the negative distance map 212 and the image data reflected in the color channels 218 - 222.
[0068] In one or more embodiments, the digital object selection system utilizes a series of equations and algorithms to generate an image / user interaction pair. For example, in one or more embodiments, the digital object selection system defines a sequence of user indicators which includes a set of positive user indicators (e.g., positive user indicator 204) and a set of negative user indicators (e.g., negative user indicator 206). In one or more embodiments, the digital object selection system utilizes the Euclidean distance transform (or some other distance metric, such as a truncated distance map or a non - linear Gaussian distribution) to transform and into separate channels U 1 (e.g., positive distance map 210) and U 0 (e.g., negative distance map 212). Each channel U 1 and U 0 reflects a two - dimensional matrix having the same height and width as the digital image (e.g., digital image 200). More specifically, the number of rows in the matrix is equal to the number of pixel rows in the digital image, and the number of columns in the matrix is equal to the number of pixel columns in the digital image.
[0069] To calculate the distance value at position (i, j) (e.g., entry 214 in the positive distance map 210 or entry 216 in the negative distance map 212), t ∈ {0, 1}, in one or more embodiments, the digital object selection system defines an operator f that calculates the minimum Euclidean distance (or other distance) between a point (e.g., a pixel in the digital image 200) and a set (e.g., the set of positive user indicators including positive user indicator 204). In other words, the digital object selection system defines an operator f such that for a given set of points where (i, j) is the point position, and then for any point, Moreover, the digital object selection system can be defined by the following formula (e.g., for each entry in the distance map):
[0070]
[0071] In one or more embodiments, for data storage efficiency, the digital object selection system employs the unsigned integer part of and truncates it at 255.
[0072] Thus, with respect to Figure 2A , the digital object selection system utilizes channels U 1 and U 0 to generate a positive distance map 210 and a negative distance map 212. For example, channel U 1 provides the matrix illustrated for the positive distance map 210. Moreover, the digital object selection system combines color channels 218 - 222 with the distance maps reflecting U 1 and U 0 to generate an image / user interaction pair.
[0073] In other words, before coupling with the RGB input image, the digital object selection system can convert sparse binary positive and negative clicks into two truncated Euclidean distance maps u = (u + ; u - ) of the union of the user's positive clicks and the union of the user's negative clicks to form a 5 - channel input (x, u).
[0074] As Figure 2A shown, the digital object selection system can also provide scale 226 (or additional scales) as an input to the scale - diversified segmentation neural network 201. For example, scale 226 can include size (e.g., the vertical dimension or the horizontal dimension of the anchor box) and aspect ratio.
[0075] As described above, in some embodiments, the digital object selection system utilizes a multi - stage scale - diversified segmentation neural network, which can consider various scales (e.g., any scale entry along a continuous range suitable for the digital image) as inputs. In such an embodiment, the digital object selection system can utilize scale 226 as an input to generate an object segmentation output. Additional details regarding providing scales as inputs to a neural network (e.g., a multi - stage scale - diversified segmentation neural network) are provided below (e.g., with respect to Figure 7 ).
[0076] In other embodiments, the digital object selection system may utilize a network architecture that includes channels for different scales and generate object segmentation outputs according to different scales. For example, the digital object selection system may define a set of scales and then include output channels for each scale in a scale-diverse segmentation neural network. As described above, using this single-stage method, the digital object selection system can generate multiple object segmentation maps in a single pass. In Figures 2B - 2C Additional details regarding such single-stage scale-diverse segmentation neural networks are discussed in the remainder of
[0077] For example, in one or more embodiments, the digital object selection system defines different scales for different combinations of aspect ratio a and size s (e.g., scale diversity). Mathematically, given P sizes and Q aspect ratios, there are M = PQ possible scale combinations, S = {(s p , a q ) | p = 1,..., P, q = 1,..., Q}. Given an input image and some user input the digital object selection system can formulate the task of synthesizing a diverse set of segmentations as a learning mapping function f(; θ, S), which is parameterized by θ and conditioned on a set of predefined scales S:
[0078]
[0079] where is a set of scale-diverse segmentation outputs, where each segmentation output o i corresponds to a 2D scale in S.
[0080] For illustration, in one or more embodiments, the digital object selection system resizes the digital image 200 to 512 × 512. Then, the digital object selection system uses 3 aspect ratios (1:1, 1:2, and 2:1) and 3 scales (64, 128, 256). In addition, the digital object selection system includes 3 anchors with sizes 16, 32, and 512 and an aspect ratio of 1:1, resulting in a total of 12 proposals. Although the foregoing example utilizes 12 anchor boxes with specific sizes and aspect ratios, the digital object selection system can utilize various different anchors (e.g., circular anchors), anchors of various different sizes and / or shapes, and different numbers of anchors (e.g., 5 or 20).
[0081] When generating Figure 2A the input illustrated in, the digital object selection system can utilize the scale-diverse segmentation neural network 201 to analyze the input. For example, Figure 2BFIG. illustrates an example architecture of a scale-diverse segmentation neural network 201 (e.g., a single-stage scale-diverse segmentation neural network) according to one or more embodiments. Specifically, Figure 2B FIG. illustrates a scale-diverse segmentation neural network 201 that includes an input layer 232 of 512×512×5 (e.g., for the Figure 2A 5-channel input discussed in Figure 2B which can be modified for different input indicators), an encoder 234, a decoder 238, and an output layer 240 of 512×512 with M output channels 240a-240m. As
[0082] shown in
[0083] In one or more embodiments, the scale-diverse segmentation neural network 201 includes a fully convolutional neural network. For example, the digital object selection system can utilize a ResNet-101 variant of DeepLabv3+ that is equipped with (1) dilated convolutional kernels (i.e., increasing the output resolution while maintaining the same amount of network parameters), (2) an atrous spatial pyramid pooling (ASPP) encoder (encoder 234) for encoding rich multi-scale context information, and (3) a decoder (decoder 238) for recovering object boundaries.
[0084] Using Figure 2B the architecture illustrated in Figure 2AThe described input. For example, a digital object selection system can analyze an encoded 512×512 color map and a distance map via an encoder 234. Specifically, the encoder 234 can utilize tunable parameters (e.g., internal weighting parameters that can be modified during training, such as via backpropagation) to generate one or more latent feature maps that reflect the features of the digital image and the input indicators. Then, the digital object selection system can utilize a decoder 238 and an output layer 240 to analyze the latent feature maps for M varying scales. As shown, the output layer 240 includes M output channels 240a - 240m for each of the M different scales.
[0085] As shown, the digital object selection system can also utilize an object verification model 242 of the scale - diverse segmentation neural network 201. In fact, as discussed above, not all scales necessarily correspond to meaningful selections. The digital object selection system can utilize the object verification model 242 to filter and / or remove scales (i.e., segmentation outputs) that do not include meaningful object selections. For illustration, the digital object selection system can utilize the object verification model 242 to remove segmentations that include partial or incomplete objects or other non - semantically meaningful outputs.
[0086] As illustrated, a global average pooling layer 243 and a fully - connected layer 244 can analyze the latent feature maps generated by the encoder 234 to output M object scores (e.g., confidence scores) for M scale - diverse segmentations. The digital object selection system can analyze the object scores to determine the scales that depict semantically meaningful results. For example, the digital object selection system can filter / remove segmentations with low object scores (e.g., below a threshold object score). Similarly, the digital object selection system can provide segmentations with high object scores (e.g., above a threshold object score) for display.
[0087] For illustration, in one or more embodiments, the scale - diverse segmentation neural network 201 applies one or more post - processing mechanisms to remove segmentation proposals that depict incomplete objects. Generally, meaningful segmentations include confident predictions at most of the pixel locations (from the confidence scores generated by the output layer 240 or the object scores generated via the object verification model 242). On the other hand, in meaningless proposals (e.g., proposals that do not include an object or include a partial object), there are a large number of uncertain predictions. Thus, in some embodiments, the digital object selection system applies a threshold to each indicator to obtain a binary mask.
[0088] For example, if the object score / confidence score for a pixel is higher than a threshold, the digital object selection system can use 1 in the binary mask for that pixel. Similarly, if the object score / confidence score for a pixel is lower than the threshold, the digital object selection system can apply 0 in the binary mask for that pixel. Then, the digital object selection system can determine the IoU (Intersection over Union) score between the prediction of the scale-diverse segmentation neural network 201 and the thresholded binary mask. The calculated IoU score is used as a validation score to decide whether a proposal should be presented to the user. Then, the digital object selection system can present proposals with high validation scores to the user.
[0089] In fact, as Figure 2C shown, the digital object selection system utilizes a scale-diverse segmentation neural network 201 to generate and display multiple object segmentation outputs (in one or more formats). For example, as shown, the digital object selection system can generate a first object segmentation output 250 (corresponding to a first scale) and a second object segmentation output 252 (corresponding to a second scale). Although Figure 2C two object segmentation outputs are illustrated, the digital object selection system can generate additional object segmentation outputs (e.g., M segmentation outputs).
[0090] As described above, the digital object selection system can generate object segmentation outputs including segmentation boundaries and / or segmentation masks. In fact, as Figure 2C shown, the first object segmentation output 250 includes a segmentation boundary 254 and a segmentation mask 256. Similarly, the second object segmentation output 252 includes a segmentation boundary 258 and a segmentation mask 260.
[0091] As shown, the segmentation boundaries 254, 258 illustrate the boundaries or edges corresponding to one or more target objects depicted in the digital image 200. For example, the segmentation boundaries 254, 258 can include a probability map indicating the probability that each pixel in the digital image corresponds to the boundary or edge of a target object in the digital image. Such segmentation boundaries can be used in various post-processing algorithms (such as a graph cut algorithm) to precisely cut or isolate a specific object from the digital image. Thus, the digital object selection system can utilize the segmentation boundary 254 as part of a graph cut algorithm to isolate the objects 202, 205 from the digital image 200.
[0092] Similarly, Figure 2CThe figure illustrates that the segmentation masks 256, 260 identify foreground and background pixels corresponding to different segmentations. For example, the segmentation masks 256, 260 can include a probability map indicating the probability that each pixel in the digital image is part of the target object. Such segmentation masks can also be used in various post-processing algorithms. For example, the digital object selection system can select and edit all pixels in the segmentation mask 256 that meet a threshold confidence level to modify the objects 256, 260 in the digital image 200.
[0093] As described above, the digital object selection system can also train a scale-diverse segmentation neural network. Figure 3 The figure illustrates training a scale-diverse segmentation neural network according to one or more embodiments (e.g., Figure 2B the scale-diverse segmentation neural network 201 illustrated in the figure). Specifically, Figure 3 the figure illustrates using a training digital image 300 (with positive training indicators) and ground truth segmentations 304, 306 at different scales to train the scale-diverse segmentation neural network 201.
[0094] As described above, the digital object selection system provides the training digital image 300 and the training indicators to the scale-diverse segmentation neural network 201. In particular, as Figure 2A described, the digital object selection system can generate an RGB channel and a distance map (e.g., an image / user interaction pair) and provide the RGB channel and the distance map as training inputs.
[0095] As illustrated, the scale-diverse segmentation neural network 201 analyzes the training inputs and generates predicted segmentations 302a - 302m at different scales. For example, the digital object selection system can generate a first predicted segmentation 302a at a first scale (e.g., a first size and aspect ratio), and a second predicted segmentation 302b at a second scale (e.g., a second size and aspect ratio).
[0096] Then, the digital object selection system can compare the predicted segmentations 302a - 302m with the ground truth segmentations. In particular, the digital object selection system can determine a measure of loss by applying a loss function to each predicted segmentation and its corresponding ground truth segmentation. Then, the digital object selection system can modify the parameters of the scale-diverse segmentation neural network 201 based on this comparison (e.g., based on backpropagation of the measure of loss).
[0097] For example, the digital object selection system may perform an action 308 of comparing the predicted segmentation 302b (corresponding to scale 2) with the ground truth segmentation 304 (corresponding to scale 2). Based on this comparison, the digital object selection system may determine a loss metric between the predicted segmentation 302b and the ground truth segmentation 304. Then, the digital object selection system may perform an action 312 of modifying the internal parameters of the scale-varied segmentation neural network 201 (e.g., modifying the weighted parameters of the encoder, decoder, output layer, and other layers to reduce the loss metric). For illustration, the digital object selection system may modify the internal parameters of the channels corresponding to scale 2 via backpropagation to train the digital object selection system to more accurately identify the segmentation at scale 2.
[0098] In some cases, the channels of a particular scale (e.g., the predicted segmentation) will not correspond to the ground truth. For example, as Figure 3 shown, the predicted segmentation 302a does not have a corresponding ground truth segmentation at that scale (e.g., no object falls within a particular anchor box size and aspect ratio). Thus, as Figure 3 shown, the digital object selection system may identify the channels of the scale-varied segmentation neural network 201 that have corresponding ground truth segmentations, and the digital object selection system backpropagates only at the matching scales (e.g., leaving the other channels unaffected).
[0099] For example, as Figure 3 shown, the digital object selection system identifies the ground truth segmentations 304, 306 corresponding to scale 2 and scale 3. The digital object selection system performs actions 308, 310 of comparing the predicted segmentations 302b, 302c with the corresponding ground truth segmentations 304, 306, and also performs actions 312, 314 of backpropagating based on this comparison to modify the scale-varied segmentation neural network 201. As shown, the digital object selection system does not compare the predicted segmentations 302a, 302m with the corresponding ground truth segmentations, or backpropagate along these channels.
[0100] In one or more embodiments, the digital object selection system identifies which scales have corresponding ground truth segmentations by comparing the ground truth segmentations with multiple scales. The digital object selection system identifies those scales that have corresponding ground truth segmentations (e.g., ground truth segmentations that fill a threshold portion of a particular scale). Specifically, given a ground truth segmentation mask y, the digital object selection system may calculate its size s y and aspect ratio a y . Then the digital object selection system finds the set where IoU is the intersection over union, and box(s p , a q ) is a box of size s p and aspect ratio aq a bounding box, where the center is the center of the bounding box enclosing the ground truth y. Then, the digital object selection system can backpropagate the loss only through these branches. Although the digital object selection system can utilize various different loss functions, in one or more embodiments, the digital object selection system utilizes:
[0101]
[0102] where l is the standard sigmoid cross-entropy loss.
[0103] By repeatedly analyzing different training images and training indicators, generating prediction segments of different scales, and comparing the prediction segments with the ground truth segments specific to a particular scale, the digital object selection system can train a scale-diverse segmentation neural network to accurately generate segmentations across different scales.
[0104] In one or more embodiments, the digital object selection system can also train the object verification model 242 of the scale-diverse segmentation neural network 201. Specifically, as shown, the object verification model 242 generates predicted object scores (e.g., a vector of M-dimensional scores corresponding to each scale). Then, the digital object selection system compares the predicted object scores with the ground truth object verification. Specifically, the digital object selection system can identify those scales that actually include an object (e.g., a complete object), and compare the predicted object scores with the ground truth object verification (e.g., using a loss function). Then, the digital object selection system can train the object verification model 242 by modifying the internal parameters of the object verification model 242 to reduce the loss function.
[0105] Although the digital object selection system can utilize various different loss functions during training, in some embodiments, the digital object selection system utilizes a class-balanced sigmoid cross-entropy loss to train the object verification model 242. In fact, the digital object selection system can use this loss function because the distribution of positive / negative samples may be unbalanced (e.g., only a small group of scales contain an object).
[0106] As just discussed, the digital object selection system can utilize training images, training indicators, and the corresponding ground truth segments of different scales for the training images and training indicators to train a scale-diverse segmentation neural network. The digital object selection system can effectively and accurately generate this training data. Moreover, as previously discussed, the digital object selection system can generate both explicit training indicators and ambiguous training indicators to more effectively and accurately train a scale-diverse segmentation neural network. Figure 4 , Figures 5A - 5CAdditional details are provided regarding the generation of training samples including explicit indicators and fuzzy indicators. Also, Figure 6 Additional details are provided regarding identifying an appropriate scale corresponding to ground truth segmentation for training.
[0107] Figure 4 Illustrated are a set of explicit training indicators 402a - 402c, 404 for training image 300 and a set of fuzzy training indicators 408, 410 for training image 300. Specifically, the explicit training indicators include negative explicit training indicators 402a - 402c and positive explicit training indicator 404. Additionally, the fuzzy training indicators include positive fuzzy training indicator 408 and negative fuzzy training indicator 410.
[0108] As illustrated, the explicit training indicators 402a - 402c, 404 together indicate a single ground truth segmentation 406 within the training digital image 300. In fact, the explicit training indicators 402a - 402c, 404 exclude ground truth segmentations including other dogs, blankets, or beds, but rather correspond only to the ground truth segmentation 406 depicting a dog. In contrast, the fuzzy training indicators 408, 410 indicate multiple ground truth segmentations 412, 414. In fact, the positive training indicator 408 and the negative training indicator 410 can equally indicate ground truth segmentation 412 (indicating a single dog) or ground truth segmentation 414 (indicating all three dogs). By generating and training a segmentation neural network with scale - diverse definite training samples and fuzzy training samples, the digital object selection system can improve the diversity and accuracy of the resulting object segmentations generated by the scale - diverse segmentation neural network.
[0109] As described above, the digital object selection system can generate training samples (including explicit training data and fuzzy training data) from an existing training data repository. For example, Figures 5A - 5C Additional details are provided regarding generating explicit training indicators (corresponding to explicit ground truth segmentations) and fuzzy training indicators (corresponding to fuzzy ground truth segmentations) from existing training data. Specifically, Figure 5A Illustrated is that the digital object selection system performs an action 502 of identifying an object depicted in the training digital image 300. For example, the digital object selection system can perform action 502 by accessing an existing training data repository of labeled digital images. In fact, the digital object selection system can access a digital repository of digital images with objects (e.g., pixels of the objects) identified in the digital images. The existing training repository typically includes digital images with segmentations of the objects depicted in the digital images.
[0110] However, traditional training repositories typically do not include training indicators, or corresponding to different scales as described above regarding Figure 3Diverse training splits (utilized). As Figure 5A As shown, the digital object selection system can perform object-based combinations to generate actions 504 for different splits. In particular, the digital object selection system can identify the objects depicted in the digital image (according to action 502) and combine the objects to generate different splits. For example, the digital object selection system generates splits 504a - 504d, which are combinations of different objects within the training image 300.
[0111] In one or more embodiments, the digital object selection system identifies splits 504a - 504d based on proximity or distance within the digital image. For example, the digital object selection system can identify an object (for the first split) and adjacent objects (for the second split). Then, the digital object selection system can generate a hierarchical list of splits based on different combinations of the adjacent objects. Specifically, for each instance in the digital image (e.g., multiple dogs), the digital object selection system can find all adjacent instances (e.g., all adjacent dogs). Then, the digital object selection system can build a hierarchical list of splits based on different combinations of the instances (e.g., expanding the split depicting multiple dogs).
[0112] In some embodiments, the digital object selection system combines adjacent instances in a class-agnostic manner. In particular, the digital object selection system does not consider the object class when generating a diverse set of ground truth splits (e.g., the digital object selection system combines a dog and a blanket, not just dogs). In other embodiments, the digital object selection system can generate ground truth splits based on the class.
[0113] Moreover, in one or more embodiments, the digital object selection system uses other factors (in addition to or instead of proximity or distance) when generating a set of ground truth splits. For example, the digital object selection system can consider depth. In particular, the digital object selection system can combine the objects in the digital image depicted at similar depths (and exclude object combinations if the objects are at different depths beyond a specific depth difference threshold).
[0114] As Figure 5A As shown, the digital object selection system can then generate explicit samples and / or fuzzy samples from the identified masks. Regarding explicit sampling, the digital object selection system can perform action 506 of identifying a single mask (e.g., a split) from splits 504a - 504d. Then, the digital object selection system can perform action 508 of explicitly sampling from the identified mask. In this way, the digital object selection system can generate training data that includes negative explicit training indicators 510 and positive explicit training indicators 512 corresponding to the explicit ground truth split (i.e., the identified mask). Regarding Figure 5BAdditional details regarding explicit sampling are provided.
[0115] Similarly, a digital object selection system can generate a blurred sample by performing an action 516 that identifies multiple masks. For example, the digital object selection system can select two or more segments from segmentations 504a - 504d. Then, the digital object selection system can perform an action 518 of performing blurred sampling from the multiple masks to generate training data that includes a negative blurred training indicator 522, a positive blurred training indicator 524, and a blurred ground truth segmentation 520 (e.g., multiple masks). Regarding Figure 5C Additional details regarding blurred sampling are provided.
[0116] Figure 5B Additional details regarding explicit sampling according to one or more embodiments are provided. As Figure 5B shown, the digital object selection system performs an action 506 by identifying a mask of a single dog depicted in the training image 300. The action 508 of explicit sampling is performed by performing an action 530 of sampling positive training indicators from the foreground based on the identified mask. Specifically, the digital object selection system samples pixels within the mask identified at action 506. Moreover, the digital object selection system performs an action 532 of sampling negative samples from the background based on the identified mask. Specifically, the digital object selection system samples pixels outside the mask identified at action 506.
[0117] The digital object selection system can utilize various methods to generate positive and negative training samples. For example, in one or more embodiments, the digital object selection system utilizes random sampling techniques (inside or outside the mask). Moreover, in other embodiments, the digital object selection system utilizes random sampling techniques within non - target objects.
[0118] However, random sampling may not provide sufficient information about the boundaries, shape, or features of the target object in training a neural network. Thus, in one or more embodiments, the digital object selection system samples training indicators based on the position (or distance from) of other training indicators. More specifically, in one or more embodiments, the digital object selection system samples positive training indicators to utilize the positive training indicators to cover the target object (e.g., such that samples extending across the target object fall within a threshold distance of the boundary and / or exceed a threshold distance from another sample). Similarly, in one or more embodiments, the digital object selection system samples negative training indicators to utilize the negative training indicators to surround the target object (e.g., fall within a threshold distance of the target object).
[0119] Figure 5C Additional details regarding blurred sampling are provided. As Figure 5CAs shown, the digital object selection system identifies multiple masks at operation 516, such as a mask of a single dog and a mask of three dogs. The digital object selection system can select multiple masks from a set of segmentations (at operation 504) in various ways. For example, the digital object selection system can select multiple masks by random sampling. In other embodiments, the digital object selection system can select multiple masks based on proximity (e.g., distance within the digital image) or depth.
[0120] When performing operation 518, the digital object selection system performs operation 540 of identifying a common foreground region and / or background region from the multiple masks. In fact, as illustrated, the digital object selection system performs operation 540 by identifying a common foreground 540a that indicates pixels of a dog common to two masks. Moreover, the digital object selection system performs operation 540 by identifying a common background 540b that indicates pixels not included in the set of three dogs (e.g., background pixels common to the two masks).
[0121] After identifying the common foreground region and / or background region, the digital object selection system then performs operation 542 of sampling positive blur training indicators from the common foreground. For example, as Figure 5C shown, the digital object selection system can sample within the common foreground 540a to generate samples within the dog.
[0122] Moreover, the digital object selection system can also perform operation 544 of sampling negative blur samples from the common background. For example, as Figure 5C shown, the digital object selection system samples from the common background 540b to generate samples outside the region depicting all three dogs.
[0123] It is noted that each of the positive training indicators and negative training indicators sampled in operations 542 and 544 is ambiguous because they do not distinguish between the multiple masks identified at operation 516. In fact, both the positive blur training indicator and the negative blur training indicator will be consistent with identifying a single dog or multiple dogs in the training image 300.
[0124] As Figure 5C shown, the digital object selection system can also perform operation 546 of identifying other reasonable ground truth segmentations (in addition to the multiple masks identified at operation 516). The digital object selection system performs operation 546 by analyzing the segmentations identified at operation 504 to determine whether there are any additional segmentations that will satisfy the positive training indicators and negative training indicators identified at operations 542, 544. As Figure 5CAs shown, the digital object selection system determines that the segmentation 504d satisfies the positive training indicator and the negative training indicator. Thus, the segmentation 504d can also be used as an additional ground truth segmentation for the positive and negative fuzzy training indicators.
[0125] As described above, in addition to generating training indicators, the digital object selection system can also determine the ground truth scale corresponding to the ground truth segmentation (e.g., to align the ground truth with an appropriate scale when training a scale-diverse segmentation neural network). Figure 6 Illustrated is identifying the ground truth scale corresponding to the ground truth segmentation according to one or more embodiments. Specifically, Figure 6 Illustrated is the ground truth segmentation 602 for the training image 600. The digital object selection system performs the action 604 of identifying a set of scales. As Figure 6 shown, the digital object selection system identifies scales 604a - 604e that include anchor boxes of different sizes and aspect ratios. In one or more embodiments, the digital object selection system identifies scales 604a - 604e based on the channels of the scale-diverse segmentation neural network. For example, the first scale 604a can reflect the corresponding scale of the first channel of the scale-diverse segmentation neural network 201.
[0126] After identifying a set of scales, the digital object selection system performs the action 606 of identifying the scale (e.g., anchor box) corresponding to the ground truth segmentation. Specifically, the digital object selection system can find the closest matching anchor box to train the selection model. For example, in one or more embodiments, the digital object selection system determines the center of the bounding box that encloses the ground truth segmentation. Next, the digital object selection system aligns a set of anchors (from action 604) conditioned on this center. Then, the digital object selection system determines the similarity between B and each anchor box based on the intersection over union (IoU). The anchor box with the maximum IoU is considered the scale corresponding to this particular selection.
[0127] As Figure 6 shown, the digital object selection system can identify the scale corresponding to the ground truth segmentation as the ground truth scale. The digital object selection system can use this matching method to find the ground truth scale for each possible ground truth mask. Moreover, as Figure 3 described above, in one or more embodiments, the digital object selection system only backpropagates the gradients of the scale-diverse segmentation neural network on the matching anchors, while keeping other anchors unaffected.
[0128] Many of the foregoing examples and illustrations have been discussed with respect to a scale-diverse segmentation neural network 201 (e.g., a single-stage scale-diverse segmentation neural network). As discussed above, the digital object selection system may also utilize a multi-stage scale-diverse segmentation neural network that takes various scales as inputs to the neural network. In fact, as discussed above, a scale-diverse segmentation neural network that utilizes channels without a predetermined scale may allow additional flexibility in generating segmentations that reflect any scale on a continuous range. For example, a possible drawback of a one-stage method is that due to discretization, some intermediate scales corresponding to semantically meaningful selections may be missing. An alternative is to define a continuous scale variation so that all possible selections can be obtained.
[0129] For example, Figure 7 illustrates the utilization of a multi-stage scale-diverse segmentation neural network according to one or more embodiments. Compared to the scale-diverse segmentation neural network 201 illustrated in Figure 2B , the scale-diverse segmentation neural network 706 does not include multiple channels for each scale. Instead, the scale-diverse segmentation neural network 706 receives a scale input and then generates an object segmentation output based on the scale input. The scale-diverse segmentation neural network 706 may generate multiple object segmentation outputs in response to multiple input scales.
[0130] For example, as shown in Figure 7 , the digital object selection system provides a digital image 700 with a user indicator 702 to the scale-diverse segmentation neural network 706. Additionally, the digital object selection system provides a first (small) scale 704. The scale-diverse segmentation neural network 706 analyzes the digital image 700, the user indicator 702, and the first scale 704, and generates an object segmentation output 708 corresponding to the first scale.
[0131] The digital object selection system also provides the digital image 700, the user indicator 702, and a second (larger) scale 705 to the scale-diverse segmentation neural network 706. The scale-diverse segmentation neural network 706 analyzes the digital image 700, the user indicator 702, and the second scale 705, and generates an object segmentation output 710 corresponding to the second scale.
[0132] As described above, the architecture of the scale-diverse segmentation neural network 706 is different from Figure 2BThe architecture of the scale-diverse segmentation neural network 201 illustrated therein. For example, the digital object selection system appends the scale as an additional channel to form a 6D input (image, user input, scale), which will be passed forward to the scale-diverse segmentation neural network 706. For illustration, for the scale channel, the digital object selection system can generate a scale map that repeats the scale value (scalar) at each pixel location. Thus, given the same (image, user input), this formulation forces the model to learn to produce different selections conditioned on the given scale.
[0133] In some embodiments, the digital object selection system can input the scale value in a different way rather than using the scale input plane. For example, the digital object selection system can utilize a single scale value (instead of the entire scale plane). The scale-diverse segmentation neural network 201 can analyze the scale value that is a numerical input to generate an object segmentation output corresponding to the scale value.
[0134] Moreover, when generating the scale-diverse segmentation neural network 706, the digital object selection system replaces multiple scale output channels from the scale-diverse segmentation neural network 201 with scale output channels corresponding to the input scale. Thus, for a specific discretized scale input, the digital object selection system can generate a specific object segmentation output via the scale-diverse segmentation neural network 706. Additionally, in some embodiments, the scale-diverse segmentation neural network 706 does not include the object verification model 242.
[0135] The digital object selection system can train the scale-diverse segmentation neural network 706 in a manner similar to the scale-diverse segmentation neural network 201. The digital object selection system can identify training images, generate ground truth segmentations and corresponding training segmentations, and train the scale-diverse segmentation neural network by comparing the predicted segmentation with the ground truth segmentation.
[0136] Because the scale-diverse segmentation neural network 706 takes the input scale into account, the digital object selection system can also utilize the training scale to train the scale-diverse segmentation neural network 706. For example, the digital object selection system can provide the scale corresponding to the ground truth segmentation (e.g., the ground truth scale from Figure 6 ) as the training input scale. Then, the scale-diverse segmentation neural network 706 can generate a predicted segmentation corresponding to the training input scale and compare the predicted segmentation with the ground truth segmentation corresponding to the training scale. Then, the scale-diverse segmentation neural network can backpropagate based on this comparison and modify the tunable parameters of the scale-diverse segmentation neural network 706.
[0137] When training the scale-diverse segmentation neural network 706, the digital object selection system can generate training indicators and ground truth segmentations as described above (e.g., with respect to Figures 5A - 5C ). Moreover, the digital object selection system can determine the training scale corresponding to the ground truth segmentation. For example, in one or more embodiments, the digital object selection system utilizes the methods described above (e.g., with respect to Figure 6 ). In some embodiments, the digital object selection system determines the training scale by determining the size and aspect ratio of the ground truth segmentation and using that size and aspect ratio as the training scale.
[0138] Although Figure 7 describes a scale-diverse segmentation neural network that considers user input with only a single output channel, the digital object selection system can utilize a scale-diverse segmentation neural network that considers input scales while retaining multiple scale output channels. For example, in one or more embodiments, the digital object selection system can utilize a scale-diverse segmentation neural network that considers scale input and has an architecture similar to that of the scale-diverse segmentation neural network 201 (with additional input channels) of Figure 2B . For example, the scale-diverse segmentation neural network can receive one or more input scales and then generate an object segmentation output using only those channels corresponding to the input scale (e.g., using the channel closest to the input scale). In this way, the digital object selection system can receive multiple scale inputs and generate multiple object segmentation outputs corresponding to the scale inputs in a single pass.
[0139] As just described, in one or more embodiments, the digital object selection system can identify scale inputs. The digital object selection system can identify scale inputs in various ways. For example, in one or more embodiments, the digital object selection system receives user input at different scales. Additional details regarding user interfaces and user interface elements for receiving user input scales are provided below (e.g., with respect to Figures 9A - 9C ).
[0140] In other embodiments, the digital object selection system can utilize a scale proposal neural network to generate scale inputs. For example, Figure 8 illustrates generating and utilizing scales via a scale proposal neural network 806. As shown in Figure 8 , the scale-diverse segmentation neural network provides a digital image 802 and a user indicator 804 to the scale proposal neural network 806. The scale proposal neural network 806 generates one or more scales, and then the one or more scales are analyzed by the scale-diverse segmentation neural network 808 as inputs to generate one or more object segmentation outputs 810.
[0141] The digital object selection system can train a scale proposal neural network 806 to generate scales corresponding to objects and user indicators depicted in a digital image. For example, the digital object selection system can provide training images and training indicators to the scale proposal neural network 806 to generate one or more predicted scales. The digital object selection system can compare the one or more predicted scales with the ground truth scale.
[0142] For example, the digital object selection system can identify a training object in a training image and identify the ground truth scale corresponding to the training object (e.g., the ground truth scale that includes the training object). Then, the digital object selection system can use the identified ground truth scale to compare the predicted scales generated via the scale proposal neural network 806. Then, the digital object selection system can modify the parameters of the scale proposal neural network 806 based on this comparison. In this way, the digital object selection system can identify diverse scales suitable for a digital image and then identify diverse object segmentation outputs corresponding to the diverse scales.
[0143] As previously mentioned, the digital object selection system can provide various graphical user interfaces and interface elements via a computing device for providing digital images, receiving user indicators, and providing object segmentation outputs. For example, Figure 9A FIG. illustrates a computing device 900 depicting a user interface 902 generated via a digital object selection system. As shown, the user interface 902 includes a digital image 904, user indicator elements 908-912, and a scale input slider element 914.
[0144] Specifically, the user interface 902 includes a foreground user indicator element 908, a background user indicator element 910, and an edge user indicator element 912. Based on the user's interaction with the foreground user indicator element 908, the background user indicator element 910, and / or the edge user indicator element 912, the digital object selection system can identify and receive different types of user indicators. For example, as Figure 9A shown, the foreground user indicator element 908 is activated and the user has selected pixels of the digital image 904. In response, the digital object selection system identifies the positive user indicator 906.
[0145] Although the user interface 902 illustrates three user indicator elements 908-912, the digital object selection system can generate a user interface with additional user indicator elements. For example, as previously mentioned, the digital object selection system can generate a user interface with a bounding box indicator element and / or a voice user indicator element.
[0146] As described above, the user interface 902 also includes a scale input slider element 914. Based on a user's interaction with the scale input slider element 914, the digital object selection system can identify a user input for a certain scale used to generate an object segmentation output. For example, Figure 9A The figure shows the scale input slider element 914 in a first position 916 corresponding to a first scale.
[0147] The digital object selection system can identify various scales based on a user's interaction with the scale input slider element 914. For example, as Figure 9B shown, the digital object selection system identifies a user input via the scale input slider element 914 at a second position 920 corresponding to a second scale. Based on the positive user indicator 906 and the second scale, the digital object selection system can generate an object segmentation output.
[0148] For example, Figure 9B The figure shows a user interface 902 including an object segmentation output 922. In particular, the digital object selection system analyzes the positive user indicator 906 and the second scale via a scale-diverse segmentation neural network to generate the object segmentation output 922. Specifically, the digital object selection system utilizes a multi-stage scale-diverse segmentation neural network (as described with respect to Figure 7 ) to analyze the second scale, the digital image 904, and the positive user indicator 906 as inputs to generate the object segmentation output 922.
[0149] The digital object selection system can generate additional object segmentation outputs of different scales based on user inputs of different scales. For example, Figure 9C The figure shows the user interface 902 after receiving an additional user input at a third position 930 corresponding to a third scale via the scale input slider element 914. The digital object selection system analyzes the digital image 904, the positive user indicator 906, and the third scale via a scale-diverse segmentation neural network and generates a second segmentation output 932. Moreover, the digital object selection system provides the second segmentation output 932 for display via the user interface 902.
[0150] Thus, a user can modify the scale input slider element 914 and dynamically generate different object segmentation outputs. After identifying an object segmentation output corresponding to a target object (e.g., the head of the mushroom or the whole mushroom as shown in Figure 9C ), the user can select the object segmentation output via the computing device 900. For example, the user can interact with an editing element to modify the object segmentation output corresponding to the target object.
[0151] Although Figures 9A - 9CIllustrated is a particular type of user interface element (e.g., a slider element) for providing a scale input, but the digital object selection system can utilize various elements to identify the scale input. For example, in one or more embodiments, the digital object selection system utilizes a timing element that modifies the scale input based on the amount of time of user interaction. For example, if the user presses the timing element, the digital object selection system can generate different object segmentation outputs based on the amount of time the user presses the timing element. Thus, for example, the digital object selection system can generate a dynamically increasing object segmentation output based on a single press-and-hold event of the timing element via computing device 900.
[0152] Similarly, in one or more embodiments, the digital object selection system utilizes a pressure element that modifies the scale based on the amount of pressure corresponding to user interaction. For example, if computing device 900 includes a touch screen, the amount of pressure of the user input can determine the corresponding scale (e.g., the digital object selection system can dynamically modify the segmentation based on the identified amount of pressure).
[0153] In one or more embodiments, the digital object selection system can identify different scale values based on a scroll event (e.g., from a mouse wheel) or based on a pinch event (e.g., a two-finger movement on a tablet). For illustration, the digital object selection system can detect a vertical pinch to modify the vertical scale size and detect a horizontal pinch to modify the horizontal scale size. Moreover, in some embodiments, the digital object selection system utilizes two slider elements (e.g., one slider element for modifying the vertical dimension and another slider element for modifying the horizontal dimension).
[0154] Similarly, although Figures 9A - 9C the slider element can select a continuous scale range, in some embodiments, the digital object selection system utilizes sticky slider elements corresponding to a set of scales (e.g., predefined scales or those corresponding to semantically meaningful segments). For example, the slider knob can stick at a particular scale or position until the knob is moved close enough to the next scale corresponding to a semantically meaningful output. In this case, intermediate results (selections with scales that do not correspond to semantically meaningful outputs) will not be visible, and only a set of high-quality proposals will be shown to the user.
[0155] In other embodiments, the digital object selection system generates a histogram or plot of all recommended scales at the upper part of the slider when the user has full control of the slider. The user can obtain all intermediate results and visualize the "growing" process of the selection as he / she moves the slider. The plot serves as a guide to show the user where the possible good proposals might be.
[0156] As described above, the digital object selection system can also generate multiple object segmentation outputs and provide the object segmentation outputs for simultaneous display. For example, Figure 10A FIG. illustrates a computing device 1000 displaying a user interface 1002 generated by a digital object selection system according to one or more embodiments. The user interface 1002 includes a digital image 1004 and user indicator elements 908-912. As Figure 10A shown, the digital object selection system identifies a positive user indicator 1006 within the digital image 1004 (e.g., a click on a hat when the foreground user indicator element 908 is active). Based on the positive user indicator 1006, the digital object selection system generates multiple object segmentation outputs 1010a-1010c corresponding to multiple scales in a segmentation output region 1008.
[0157] In contrast to Figures 9A - 9C this, the digital object selection system generates multiple object segmentation outputs 1010a-1010c simultaneously (or almost simultaneously) without user input for a certain scale. The digital object selection system utilizes a single-stage scale-diverse segmentation neural network to analyze the digital image 1004 and the user indicator 1006 to generate the object segmentation outputs 1010a-1010c.
[0158] As described above, the digital object selection system can determine multiple different scales for generating the object segmentation outputs 1010a-1010c. In some embodiments, the digital object selection system utilizes different scales corresponding to different channels of a scale-diverse segmentation neural network (e.g., Figure 2B the channels described in Figure 8 ) to generate the object segmentation outputs. In other embodiments, the scale-diverse segmentation neural network can utilize a scale proposal neural network (as
[0159] described in Figure 10B ) to generate scales. Regardless of the method, the digital object selection system can utilize different scales to generate the object segmentation outputs 1010a-1010c, and then the user can interact with the object segmentation outputs 1010a-1010c.
[0160] For example, as Figure 10C shown, the digital object selection system identifies the user's selection of the first object segmentation output 1010a. In response, the digital object selection system also provides a corresponding object segmentation selection 1020 in the digital image 1004. In this way, the user can quickly and effectively review multiple object segmentation outputs and select a specific object segmentation output corresponding to the target object.
[0160] As Figure 10C shown, the user can select different object segmentation outputs, and the digital object selection system can provide corresponding object segmentation selections. For example, in Figure 10CIn [the system], the digital object selection system identifies the user's interaction with the third object segmentation element 1010c. In response, the digital object selection system generates a corresponding object segmentation selection 1030 within the digital image 1004.
[0161] The digital object selection system can also refine the object segmentation based on additional user selections. For example, regarding Figure 10D , after selecting the third object segmentation element 1010c, the user provides an additional user selection. Specifically, the third object segmentation element 1010c omits a portion of the shirt depicted in the digital image 1004. The user activates the edge user indicator element 912 and provides an edge indicator 1042 (e.g., a click at or near the edge of the shirt shown in the digital image 1004). In response, the digital object selection system modifies the object segmentation selection 1030 to generate a new object segmentation 1032 that includes the portion of the shirt that was initially omitted. Thus, the digital object selection system can generate multiple object segmentation selections and also consider additional user indicators to identify the segmentation that aligns with the target object.
[0162] Although Figure 10D no additional object segmentation output proposals are included, in one or more embodiments, after receiving an additional user indicator, the digital object selection system generates an additional set of object segmentation output proposals. Thus, if the user indicator remains ambiguous, the digital object selection system can provide an additional set of object segmentation outputs to reduce the time and user interaction required to identify the target object.
[0163] In addition, although Figures 10A - 10D a specific number of object segmentation outputs (i.e., three) are illustrated, the digital object selection system can generate various different object segmentation outputs. For example, as described above, in one or more embodiments, the digital object selection system generates and provides 12 segmentations. By way of illustration, in some embodiments, the digital object selection system generates and provides 12 segmentations, but emphasizes (e.g., outlines with additional boundaries) those segmentations with the highest quality (e.g., the highest confidence score or object score). In other embodiments, the digital object selection system filters out segmentations with low confidence scores or low object scores.
[0164] In addition, although Figures 10A - 10D multiple object segmentation proposals are provided as separate visual elements, the digital object selection system can display the object segmentation proposals as different overlays on a single digital image. For example, the digital object selection system can overlay all proposals on the digital image 1004 with different color codes (e.g., different colors corresponding to different scales), where the user can simply drag the cursor to select or deselect proposals.
[0165] As discussed above, the digital object selection system can improve efficiency and accuracy. In fact, researchers have conducted experiments to illustrate the improvements offered by the digital object selection system over conventional systems. The common practice for evaluating the performance of a single-output interactive image segmentation system is as follows: Given an initial positive click at the center of the object of interest, the model for evaluation outputs an initial prediction. Subsequent clicks are iteratively added to the center of the region with the maximum error label, and this step is repeated until the maximum number of clicks (fixed at 20) is reached. The intersection over union (IoU) is recorded at each click. The average number of clicks required to achieve a specific IoU on a specific dataset is reported.
[0166] However, because the digital object selection system can produce multiple segmentations, the study also considered the amount of interaction required when selecting one of the predictions in the selection. This is because to add a new click to the center of the region with the maximum error, the study needs to pick one of the M segmentations as the output of the model to calculate the segmentation error. To achieve this goal, the researchers maintain a "default" segmentation branch and increment the number of changes when the user needs to change from the "default" segmentation mask to another segmentation mask.
[0167] The researchers compared the digital object selection system with multiple image segmentation models based on published benchmarks with instance-level annotations, including the PASCAL VOC validation set and the Berkeley dataset. The researchers evaluated the digital object selection system with respect to "Deep Interactive Object Selection (DISO)" by N. Xu et al., "Regional Interactive Image Segmentation Networks (RIS-Net)" by J. H. Liew et al., "Iteratively Trained Interactive Segmentation (ITIS)" by S. Mahadevan et al., "Deep Extreme Cut: From extreme points to object segmentation (DEXTR)" by K. Maninis et al., "Interactive Image Segmentation With Latent Diversity (LDN)" by Z. Li et al., and "A Fully Convolutional Two-Stream Fusion Network For Interactive Image Segmentation (FCFSFN)" by Y. Hu et al. Results showing the improvements on clicks obtained by the digital object selection system are provided in Table 1. As shown, the digital object selection system produces the fewest number of clicks across all systems.
[0168] Table 1
[0169]
[0170] As described above, the digital object selection system can be implemented in conjunction with one or more computing devices. Figure 11 A diagram illustrating an environment 1100 in which the digital object selection system can operate. As Figure 11 shown, the environment 1100 includes server device(s) 1102 and client devices 1104a - 1104n. Moreover, each of the devices within the environment 1100 can communicate with each other via a network 1106 (e.g., the Internet). Although Figure 11Illustrates a particular arrangement of components, but various additional arrangements are possible. For example, the server device(s) 1102 may communicate directly with the client devices 1104a - 1104n instead of via the network 1106. Also, although Figure 11 illustrates three client devices 1104a - 1104n, in alternative embodiments, the environment 1100 includes any number of user client devices.
[0171] As Figure 11 shown, the environment 1100 may include client devices 1104a - 1104n. The client devices 1104a - 1104n may include various computing devices, such as one or more personal computers, laptop computers, mobile devices, mobile phones, tablets, dedicated computers, including the computing devices described below with respect to Figure 14 description.
[0172] Also, as Figure 11 shown, the client devices 1204a - 1204n and the server device(s) 1102 may communicate via the network 1106. The network 1106 may represent a network or a collection of networks (such as the Internet, a corporate intranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a cellular network, a wide area network (WAN), a metropolitan area network (MAN), or a combination of two or more such networks). Thus, the network 1106 can be any suitable network through which the client devices 1104a - 1104n can access the server device(s) 1102, and vice versa. Additional details about the network 1106 are provided below (e.g., with respect to Figure 14 ).
[0173] In addition, as Figure 11 shown, the environment 1100 may also include the server device(s) 1102. The server device(s) 1102 may generate, store, analyze, receive, and transmit various types of data. For example, the server device(s) 1102 may receive data from a client device such as client device 1104a and send the data to another client device such as client device 1104b. The server device(s) 1102 may also transmit electronic messages between one or more users in the environment 1100. In some embodiments, the server device(s) 1102 are data servers. The server device(s) 1102 may also include communication servers or web hosting servers. Additional details about the server device(s) 1102 will be discussed below (e.g., with respect to Figure 14 ).
[0174] As shown, the server device(s) 1102 includes a digital media management system 1108 that can manage the storage, selection, editing, modification, and distribution of digital media such as digital images or digital videos. For example, the digital media management system 1108 can collect digital images (and / or digital videos) from the client device 1104a, edit the digital images, and provide the edited digital images to the client device 1104a.
[0175] As Figure 11 shown, the digital media management system 1108 includes a digital object selection system 1110. The digital object selection system 1110 can identify one or more target objects in a digital image. For example, the server device(s) 1102 can receive an indication of a pixel in the digital image from the client device 1104a via the client device 1104a. The digital object selection system 1110 can utilize a scale-diverse segmentation neural network to generate multiple object segmentations and provide the multiple object segmentations for display via the client device 1104a.
[0176] In addition, the digital object selection system 1110 can also train one or more scale-diverse segmentation neural networks. In fact, as discussed above, the digital object selection system 1110 can generate training data (e.g., training images, explicit training indicators, fuzzy training indicators, and ground truth segmentations) and utilize the training data to train the scale-diverse segmentation neural network. In one or more embodiments, a first server device (e.g., a third-party server) trains the scale-diverse segmentation neural network, and a second server device (or client device) applies the scale-diverse segmentation neural network.
[0177] Although Figure 11 illustrated as implemented via the server device(s) 1102, the digital object selection system 1110 can be implemented wholly or in part by separate devices 1102-1104n of the environment 1100. For example, in one or more embodiments, the digital object selection system 1110 is implemented on the client device 1102a. Similarly, in one or more embodiments, the digital object selection system 1110 can be implemented on the server device(s) 1102. Moreover, different components and functions of the digital object selection system 1110 can be implemented separately among the client devices 1204a-1204n, the server device(s) 1102, and the network 1106.
[0178] Now referring Figure 12 , additional details regarding the capabilities and components of the digital object selection system 1110 according to one or more embodiments will be provided. In particular, Figure 12A schematic diagram showing an example architecture of a digital object selection system 1110 of a digital media management system 1108 implemented on a computing device 1200.
[0179] As shown, the digital object selection system 1110 is implemented via the computing device 1200. Generally, the computing device 1200 can represent various types of computing devices (e.g., (multiple) server devices 1102 or client devices 1104a - 1104n). As Figure 7 shown, the digital object selection system 1110 includes various components for performing the processes and features described herein. For example, the digital object selection system 1110 includes a training data manager 1202, a scale - diverse segmentation neural network training engine 1204, a digital image manager 1206, a scale - diverse segmentation neural network application engine 1208, a user input manager 1210, a user interface facility 1212, and a storage manager 1214. Each of these components is described in turn below.
[0180] As Figure 12 shown, the digital object selection system 1110 includes a training data manager 1202. The training data manager 1202 can receive, manage, identify, generate, create, modify, and / or provide training data for the digital object selection system 1110. For example, as described above, the digital object selection system can access a training repository, generate training indicators (e.g., positive training indicators, negative training indicators, explicit training indicators, and / or fuzzy training indicators), identify the ground - truth segmentation corresponding to the training indicators, and identify the ground - truth scale corresponding to the ground - truth segmentation.
[0181] Additionally, as Figure 12 shown, the digital object selection system 1110 further includes a scale - diverse segmentation neural network training engine 1204. The scale - diverse segmentation neural network training engine 1204 can tune, teach, and / or train a scale - diverse segmentation neural network. As described above, the scale - diverse segmentation neural network training engine 1204 can utilize the training data generated by the training data manager 1202 to train a single - stage and / or multi - stage scale - diverse segmentation neural network.
[0182] As Figure 12 shown, the digital object selection system 1110 further includes a digital image manager 1205. The digital image manager 1205 can identify, receive, manage, edit, modify, and provide digital images. For example, the digital image manager 1205 can identify digital images (from client devices or image repositories), provide the digital images to the scale - diverse segmentation neural network to identify target objects, and modify the digital images based on the identified target objects.
[0183] Moreover, as Figure 12 shown in, the digital object selection system 1110 further includes a scale-diverse segmentation neural network application engine 1208. The scale-diverse segmentation neural network application engine 1208 can generate, create, and / or provide object selection outputs based on scale. For example, as discussed above, the scale-diverse segmentation neural network application engine 1208 can analyze digital images and user indicators via a trained scale-diverse segmentation neural network to create, generate, and / or provide one or more scale-based object selection outputs.
[0184] In addition, as Figure 12 shown in, the digital object selection system 1110 further includes a user input manager 1210. The user input manager 1210 can obtain, identify, receive, monitor, capture, and / or detect user input. For example, in one or more embodiments, the user input manager 1210 identifies one or more user interactions with respect to the user interface. The user input manager 1210 can detect user input of one or more user indicators. In particular, the user input manager 1210 can detect user input of one or more user indicators with respect to one or more pixels in a digital image. For example, in one or more embodiments, the user input manager 1210 detects user input of a point or pixel in a digital image (e.g., a mouse click event or a touch event on a touch screen). Similarly, in one or more embodiments, the user input manager 1210 detects user input of a stroke (e.g., a mouse click, drag, and release event). In one or more embodiments, the user input manager 1210 detects user input of a bounded region (e.g., a mouse click, drag, and release event). Additionally, in one or more embodiments, the user input manager 1210 detects user input of an edge (e.g., a mouse click and / or drag event) or voice input.
[0185] As Figure 12 shown in, the digital object selection system 1110 further includes a user interface facility 1212. The user interface facility 1212 can generate, create, and / or provide one or more user interfaces with corresponding user interface elements. For example, the user interface facility 1212 can generate user interfaces 902 and 1002 and corresponding elements (e.g., slider elements, timer elements, image display elements, and / or segmentation output regions).
[0186] The digital object selection system 1110 also includes a storage manager 1214. The storage manager 1214 maintains data for the digital object selection system 1110. The storage manager 1214 can maintain any type, size, or kind of data as needed to perform the functions of the digital object selection system 1110. As illustrated, the storage manager 1214 can include digital images 1216, object segmentation outputs 1218, scale-diverse segmentation neural networks 1220, and training data 1222 (e.g., training images depicting training objects, training indicators corresponding to the training objects, training scales, and ground truth segmentations of the training images and training indicators corresponding to different scales).
[0187] Each of the components 1202 - 1214 of the digital object selection system 1110 can include software, hardware, or both. For example, the components 1202 - 1214 can include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (such as a client device or a server device). When executed by one or more processors, the computer-executable instructions of the digital object selection system 1110 can cause the (multiple) computing devices to perform the feature learning methods described herein. Alternatively, the components 1202 - 1214 can include hardware, such as a dedicated processing device for performing a particular function or group of functions. Alternatively, the components 1202 - 1214 of the digital object selection system 1110 can include a combination of computer-executable instructions and hardware.
[0188] In addition, the components 1202 - 1214 of the digital object selection system 1110 can be implemented, for example, as one or more operating systems, one or more standalone applications, one or more modules of an application, one or more plug-ins, one or more library functions or functions that can be called by other applications, and / or a cloud computing model. Thus, the components 1202 - 1214 can be implemented as a standalone application, such as a desktop application or a mobile application. In addition, the components 1202 - 1214 can be implemented as one or more web-based applications hosted on a remote server. The components 1202 - 1214 can also be implemented in a set of mobile device applications or "apps". By way of illustration, the components 1202 - 1214 can be implemented in an application that includes, but is not limited to, Creative After and Sensei. "ADOBE", "CREATIVE CLOUD", "PHOTOSHOP", "INDESIGN", "LIGHTROOM", "ILLUSTRATOR", "AFTER EFFECTS", and "SENSEI" are registered trademarks or trademarks of Adobe Systems Incorporated in the United States and / or other countries.
[0189] In Figures 1B - 12 , corresponding text and examples provide non-transitory computer-readable media for a variety of different methods, systems, devices, and digital object selection systems 1110. In addition to the foregoing, one or more embodiments may also be described according to a flowchart including actions for achieving a particular result, as Figures 13 - 14 shown. More or fewer actions may be performed to execute Figures 13 - 14 the series of actions illustrated. Additionally, the actions may be performed in a different order. Additionally, the described actions may be repeated or performed in parallel with each other, or repeated or performed in parallel with different instances of the same or similar actions.
[0190] Figure 13 FIG. illustrates a flowchart of a series of actions 1300 for a segmentation neural network that utilizes scale diversification to generate object segmentation outputs based on diverse scales according to one or more embodiments. Although Figure 13 illustrates actions according to one or more embodiments, alternative embodiments may omit, add, reorder, and / or modify Figure 13 any of the actions shown. Figure 13 The actions of can be performed as part of a method. Alternatively, the non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause a computing device to perform Figure 13 the actions of. In some embodiments, the system may perform Figure 13 the actions of.
[0191] As Figure 13 shown, a series of actions 1300 includes an action 1310 of identifying a user indicator. In particular, action 1310 may include identifying a user indicator that includes one or more pixels of a digital image. The digital image may depict one or more target objects. More particularly, action 1320 may include identifying one or more of a positive user indicator, a negative user indicator, or a boundary user indicator regarding one or more expected target objects. Additionally, action 1310 may also include receiving (or identifying) the digital image and the user indicator that includes one or more pixels of the digital image.
[0192] As Figure 13As shown in, a series of operations 1300 further includes an operation 1320 of generating a first object segmentation output by using a scale-diverse segmentation neural network. In particular, operation 1320 may include using a scale-diverse segmentation neural network to generate a first object segmentation output of a first scale based on a digital image and a user indicator. For illustration, in one or more embodiments, the scale-diverse segmentation neural network includes a plurality of output channels corresponding to a plurality of scales. Thus, operation 1320 may include using a first output channel corresponding to the first scale to generate the first object segmentation output.
[0193] More particularly, operation 1320 may include generating one or more distance maps. Generating one or more distance maps may include generating one or more of a positive distance map, a negative distance map, or a boundary map. Operation 1320 may include generating a positive distance map reflecting the distance of a pixel from a positive user indicator. Operation 1320 may include generating a negative distance map reflecting the distance of a pixel from a negative user indicator. Operation 1320 may further include generating an edge distance map reflecting the distance of a pixel from an edge user indicator.
[0194] Operation 1320 may further include generating one or more color maps. For example, generating one or more color maps may include: generating a red map reflecting the amount of red corresponding to each pixel, a green map reflecting the amount of green corresponding to each pixel, and a blue map reflecting the amount of blue corresponding to each pixel.
[0195] Operation 1320 may further include generating one or more feature maps from one or more color maps and one or more distance maps. In particular, operation 1320 may include using a neural network encoder to generate one or more feature maps from one or more color maps and one or more distance maps.
[0196] A series of operations 1300 may further include generating a plurality of object segmentation outputs of different scales. In particular, a series of operations 1300 may include generating a plurality of object segmentation outputs (a first object segmentation output of a first scale, a second object segmentation output of a second scale, etc.) by processing one or more feature maps using a neural network decoder. In one or more embodiments, the first scale includes a first size and a first aspect ratio, and the second scale includes a second size and a second aspect ratio.
[0197] Therefore, as Figure 13As shown in [description], a series of operations 1300 further includes an operation 1330 of generating a second object segmentation output using a scale-diverse segmentation neural network. In particular, operation 1330 may include using a scale-diverse segmentation neural network to generate a second object segmentation output of a second scale based on a digital image and a user indicator. By way of illustration, operation 1330 may include generating a second object segmentation output using a second output channel corresponding to the second scale. Operation 1330 may include the steps described above with respect to operation 1320 and may be performed in parallel with operation 1320 described above.
[0198] A series of operations 1300 may further include processing one or more feature maps generated by an object verification model to generate a plurality of object scores. For example, a series of operations 1300 may include generating object scores for each of a plurality of scales by processing one or more feature maps through a global pooling layer and a fully connected layer.
[0199] A series of operations 1300 may further include selecting an object segmentation output having a high object score for display. For example, a series of operations 1300 may include filtering / removing object segmentation outputs having low object scores such that only object segmentation outputs having high object scores are provided for display. Thus, a series of operations 1300 may include identifying that the first object segmentation output and the second object segmentation output have high object scores and selecting the first object segmentation output and the second object segmentation output for display based on the high object scores.
[0200] Alternatively, operation 1320 may include identifying a first input scale. For example, operation 1320 may include identifying a selection of the first input scale based on user input using a slider. Then, operation 1320 may include providing one or more distance maps, one or more color maps, and the first input scale (e.g., the first scale) to a scale-diverse segmentation neural network. The scale-diverse segmentation neural network may use the one or more distance maps, the one or more color maps, and the input scale (e.g., the first scale) to generate a first object segmentation output of the first scale. Then, a series of operations 1300 may include identifying a second input scale (e.g., the user may determine that the first object segmentation output is too small). In such an implementation, operation 1330 may include providing one or more distance maps, one or more color maps, and the second input scale (e.g., the second scale) to a scale-diverse segmentation neural network. The scale-diverse segmentation neural network may use the one or more distance maps, the one or more color maps, and the second input scale to generate a second object segmentation output of the second scale.
[0201] In addition, as Figure 13As shown, a series of operations 1300 also includes an operation 1340 of providing a first object segmentation output and a second object segmentation output for display (e.g., providing a plurality of object segmentation outputs for display). For example, in one or more embodiments, operation 1340 includes providing a scale slider user interface element for display; in response to a user input identifying a first position corresponding to a first scale via the scale slider user interface element, providing the first object segmentation output for display; and in response to a user input identifying a second position corresponding to a second scale via the scale slider user interface element, providing the second object segmentation output for display. In one or more embodiments, the first object segmentation output includes at least one of a segmentation mask or a segmentation boundary.
[0202] In one or more embodiments, a series of operations 1300 also includes (at least one of the following) using a scale proposal neural network to analyze a digital image and a user indicator to generate a first scale and a second scale; or determining a first scale based on the amount of time of user interaction. For example, a series of operations 1300 may include determining a first scale based on a first amount of time of user interaction (e.g., the amount of time of a click and hold), and determining a second scale based on a second amount of time of user interaction (e.g., an additional amount of time after the click and hold until a release event).
[0203] Moreover, a series of operations 1300 may also include applying an object verification model of a scale-diverse segmentation neural network to determine an object score corresponding to the first scale; and providing the first object segmentation output for display based on the object score. For example, a series of operations 1300 may include applying an object verification model of a scale-diverse segmentation neural network to determine a first object score corresponding to the first scale and a second object score corresponding to the second scale; and providing the first object segmentation output and the second object segmentation output for display based on the first object score and the second object score. Additionally, a series of operations 1300 may also include identifying a user's selection of the first object segmentation output; and based on the user's interaction with the first object segmentation output, selecting pixels of the digital image corresponding to one or more target objects.
[0204] In addition to (or instead of) the above operations, in some embodiments, a series of operations 1300 includes a step of using a scale-diverse segmentation neural network to generate a plurality of object segmentation outputs corresponding to a plurality of scales based on a digital image and a user indicator. In particular, the algorithms and operations described above with respect to Figures 2A - 2C and Figure 7 may include corresponding operations (or structures) for the step of using a scale-diverse segmentation neural network to generate a plurality of object segmentation outputs corresponding to a plurality of scales based on a digital image and a user indicator.
[0205] Figure 14 FIG. illustrates a flowchart of a series of operations 1400 for training a scale-diverse segmentation neural network to generate an object segmentation output based on scale diversity. Although Figure 14 illustrates operations in accordance with one or more embodiments, alternative embodiments may omit, add, reorder, and / or modify Figure 14 any of the operations shown in Figure 14 The operations of may be performed as part of a method. Alternatively, a non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause a computing device to perform Figure 14 the operations of. In some embodiments, a system may perform Figure 14 the operations of.
[0206] As Figure 14 shown in, a series of operations 1400 includes an operation 1410 of identifying a training digital image depicting a training object, one or more training indicators, and a ground truth segmentation at a first scale (e.g., training data stored in at least one non-transitory computer-readable storage medium). For example, operation 1410 may include identifying a training digital image depicting a training object; one or more training indicators corresponding to the training object; and a first ground truth segmentation corresponding to the first scale, the training object, and one or more training indicators. In one or more embodiments, operation 1410 further includes identifying a second ground truth segmentation corresponding to a second scale, the training object, and one or more training indicators.
[0207] Additionally, in one or more embodiments, the training object includes a first object and a second object, and the one or more training indicators include an ambiguous training indicator associated with the training object and the first object. Moreover, operation 1410 may include generating the ambiguous training indicator by: identifying a common foreground for the training object and the first object; and sampling the ambiguous training indicator from the common foreground that conveys the training object and the first object. Further, in some embodiments, the one or more training indicators include an ambiguous training indicator and an explicit training indicator. Operation 1410 may also include generating the explicit training indicator by sampling a positive explicit training indicator from a region of the digital image corresponding to the first ground truth segmentation. Additionally, operation 1410 may also include comparing the first ground truth segmentation with multiple scales to determine that the first scale corresponds to the first ground truth segmentation.
[0208] Additionally, as Figure 14As shown, a series of operations 1400 includes an operation 1420 of generating a first predicted object segmentation output of a first scale by using a scale-diverse segmentation neural network. For example, operation 1420 may include using a scale-diverse segmentation neural network to analyze a training digital image and one or more training indicators of the first scale to generate a first predicted object segmentation output. In one or more embodiments, the scale-diverse segmentation neural network includes a plurality of output channels corresponding to a plurality of scales. Moreover, operation 1420 may include using a first output channel corresponding to the first scale to generate a first predicted object segmentation output. Additionally, operation 1420 may further include using a scale-diverse segmentation neural network to analyze a training digital image and one or more training indicators of a second scale to generate a second predicted object segmentation output.
[0209] Moreover, as Figure 14 shown, a series of operations 1400 includes an operation 1430 of comparing the first predicted object segmentation output with a first ground truth segmentation. For example, operation 1430 may include modifying adjustable parameters of the scale-diverse segmentation neural network based on a comparison between the first predicted object segmentation output and the first ground truth segmentation corresponding to the first scale, a training object, and one or more training indicators. Additionally, operation 1430 may further include comparing the second predicted object segmentation output with a second ground truth segmentation.
[0210] In addition to (or instead of) the above operations, in some embodiments, a series of operations 1400 includes steps for training a scale-diverse segmentation neural network to analyze training indicators corresponding to a training digital image and generate object segmentation outputs corresponding to different scales. In particular, the algorithms and operations described above with respect to Figure 3 and Figure 7 may include corresponding operations for the steps of training a scale-diverse segmentation neural network to analyze training indicators corresponding to a training digital image and generate object segmentation outputs corresponding to different scales.
[0211] Embodiments of the present disclosure may include or utilize a special-purpose or general-purpose computer including computer hardware (such as, for example, one or more processors and system memory), as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be at least partially implemented as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (such as any of the media content access devices described herein). Generally, a processor (such as a microprocessor) receives instructions from a non-transitory computer-readable medium (such as memory) and executes those instructions, thereby performing one or more processes (including one or more of the processes described herein).
[0212] A computer-readable medium can be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the present disclosure can include at least two distinct types of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0213] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives (“SSDs”) (e.g., RAM-based), flash memory, phase change memory (“PCM”), other types of memory, other optical disk storage devices, magnetic disk storage devices, or other magnetic storage devices, or any other medium that can be used to store desired program code components in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.
[0214] “Network” is defined as one or more data links that support the transportation of electronic data between computer systems and / or modules and / or other electronic devices. When information is passed or provided to a computer via a network or another communication connection (wired, wireless, or a combination of wired or wireless), the computer properly views that connection as a transmission medium. Transmission media can include networks and / or data links that can be used to carry desired program code components in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0215] In addition, when program code components in the form of computer-executable instructions or data structures reach various computer system components, they can be automatically transferred from the transmission medium to the non-transitory computer-readable storage medium (device) (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., “NIC”) and then ultimately transferred to the computer system RAM and / or less volatile computer storage media (devices) at the computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0216] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a particular function or group of functions. In some embodiments, the computer-executable instructions are executed by a general-purpose computer to transform the general-purpose computer into a special-purpose computer implementing elements of the present disclosure. The computer-executable instructions can be, for example, binary files, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims need not be limited to the above-described features or acts. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0217] Those skilled in the art will understand that the present disclosure can be practiced in a network computing environment having many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, and the like. The present disclosure can also be implemented in a distributed system environment where local and remote computer systems that are network-linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0218] Embodiments of the present disclosure can also be implemented in a cloud computing environment. As used herein, the term "cloud computing" refers to a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the marketplace to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or low service provider interaction, and then scaled accordingly.
[0219] The cloud computing model can include various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and the like. The cloud computing model can also expose various service models such as, for example, software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). Different deployment models (such as private cloud, community cloud, public cloud, hybrid cloud, etc.) can also be used to deploy the cloud computing model. Additionally, as used herein, the term "cloud computing environment" refers to an environment that employs cloud computing.
[0220] Figure 15FIG. illustrates a block diagram of an example computing device 1500 that may be configured to perform one or more of the processes described above. It will be appreciated that one or more computing devices such as computing device 1500 may represent the computing devices described above (e.g., computing device 900, computing device 1000, server devices 1102, client devices 1104a - 1104n, and / or computing device 1200). In one or more embodiments, computing device 1500 may be a mobile device (e.g., a mobile phone, smartphone, PDA, tablet computer, laptop computer, camera, tracker, watch, wearable device, etc.). In some embodiments, computing device 1500 may be a non - mobile device (e.g., a desktop computer or other type of client device). Additionally, computing device 1500 may be a server device that includes cloud - based processing and storage capabilities.
[0221] As Figure 15 shown, computing device 1500 may include one or more processors 1502, a memory 1504, a storage device 1506, an input / output interface 1508 (or "I / O interface 1508"), and a communication interface 1510, which may be communicatively coupled by means of a communication infrastructure (e.g., a bus 1512). Although computing device 1500 is shown in Figure 15 FIG., Figure 15 the components illustrated in Figure 15 FIG. are not intended to be limiting. Additional or alternative components may be used in other embodiments. Additionally, in certain embodiments, computing device 1500 includes fewer components than Figure 15 shown in Figure 15 FIG. The components of computing device 1500 shown in
[0222] FIG. will now be described in additional detail.
[0222] In a particular embodiment, the processor(s) 1502 includes hardware for executing instructions, such as instructions that make up a computer program. By way of example and not limitation, to execute instructions, the processor(s) 1502 may retrieve (or fetch) instructions from an internal register, an internal cache, the memory 1504, or the storage device 1506 and decode and execute them.
[0223] Computing device 1500 includes a memory 1504 that is coupled to the processor(s) 1502. The memory 1504 may be used to store data, metadata, and programs for execution by the processor(s). The memory 1504 may include one or more of volatile and non - volatile memory, such as random access memory ("RAM"), read - only memory ("ROM"), solid - state drive ("SSD"), flash memory, phase - change memory ("PCM"), or other types of data storage devices. The memory 1504 may be internal memory or distributed memory.
[0224] The computing device 1500 includes a storage device 1506 that includes storage means for storing data or instructions. By way of example and not limitation, the storage device 1506 may include the non-transitory storage media described above. The storage device 1506 may include a hard disk drive (HDD), flash memory, a universal serial bus (USB) drive, or a combination of these or other storage devices.
[0225] As shown, the computing device 1500 includes one or more I / O interfaces 1508 that provide the one or more I / O interfaces 1508 to allow a user to provide input (e.g., user strokes) to the computing device 1500, receive output therefrom, and otherwise transfer data to and from it. These I / O interfaces 1508 may include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or a combination of these I / O interfaces 1508. The touch screen may be activated using a stylus or a finger.
[0226] The I / O interface 1508 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, the I / O interface 1508 is configured to provide graphic data to a display for presentation to a user. The graphic data may represent one or more graphical user interfaces and / or may be any other graphic content for a particular implementation.
[0227] The computing device 1500 may also include a communication interface 1510. The communication interface 1510 may include hardware, software, or both. The communication interface 1510 provides one or more interfaces for communication (such as, for example, packet-based communication) between the computing device and one or more other computing devices or one or more networks. By way of example and not limitation, the communication interface 1510 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as WI-FI. The computing device 1500 may also include a bus 1512. The bus 1512 may include hardware, software, or both that connect the components of the computing device 1500 to each other.
[0228] In the foregoing specification, the present invention has been described with reference to specific example embodiments of the present invention. The various embodiments and aspects of the present invention(s) have been described with reference to the details discussed herein, and the accompanying drawings illustrate the various embodiments. The above description and drawings are illustrative of the present invention and should not be construed as limiting the present invention. Numerous specific details have been described to provide a thorough understanding of the various embodiments of the present invention.
[0229] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. Additionally, the steps / actions described herein may be repeated or performed in parallel with each other, or repeated or performed in parallel with different instances of the same or similar steps / actions. Accordingly, the scope of the present invention is indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computer system to: Identify a user indicator including one or more pixels of a digital image depicting one or more target objects; Based on the digital image and the user indicator, use a scale-diverse segmentation neural network to generate a first object segmentation output at a first scale corresponding to a first anchor region containing a first portion of the digital image; Based on the digital image and the user indicator, use the scale-diverse segmentation neural network to generate a second object segmentation output at a second scale corresponding to a second anchor region containing the first portion and a second portion of the digital image; And In response to a first input via a graphical user interface element, provide the first object segmentation output for display, and in response to a second input via the graphical user interface element, provide the second object segmentation output for display.
2. The non-transitory computer-readable medium according to claim 1, wherein the scale-diverse segmentation neural network includes a plurality of output channels corresponding to a plurality of scales, and further includes instructions that, when executed by the at least one processor, cause the computer system to: Use a first output channel corresponding to the first scale to generate the first object segmentation output; and Use a second output channel corresponding to the second scale to generate the second object segmentation output.
3. The non-transitory computer-readable medium according to claim 1, wherein the first scale includes a first size and a first aspect ratio, and the second scale includes a second size and a second aspect ratio.
4. The non-transitory computer-readable medium according to claim 1, further including instructions that, when executed by the at least one processor, cause the computer system to provide the first object segmentation output and the second object segmentation output for display by: Providing the graphical user interface element for display by presenting a scale slider user interface element; In response to a user input identifying a first position corresponding to the first scale via the scale slider user interface element, providing the first object segmentation output for display; And In response to a user input identifying a second position corresponding to the second scale via the scale slider user interface element, providing the second object segmentation output for display.
5. The non-transitory computer-readable medium according to claim 1, further including instructions that, when executed by the at least one processor, cause the computer system to perform at least one of the following: Use a scale proposal neural network to analyze the digital image and the user indicator to generate the first scale and the second scale; or Determine the first scale based on the amount of time of user interaction.
6. The non-transitory computer-readable medium according to claim 1, further including instructions that, when executed by the at least one processor, cause the computer system to: Apply the object verification model of the scale-diversified segmentation neural network to determine a first object score corresponding to the first scale and a second object score corresponding to the second scale; and Based on the first object score and the second object score, provide the first object segmentation output and the second object segmentation output for display.
7. The non-transitory computer-readable medium according to claim 1, wherein the first object segmentation output includes at least one of the following: a segmentation mask or a segmentation boundary.
8. The non-transitory computer-readable medium according to claim 1, further comprising instructions that, when executed by the at least one processor, cause the computer system to: Identify a user input that selects the first object segmentation output; and Based on the user input that selects the first object segmentation output, select the pixels of the digital image corresponding to the one or more target objects.
9. A computer-implemented method for using scale-variant deep learning to identify digital objects depicted within digital visual media in a digital media environment for editing digital visual media, the method comprises: Steps for training a scale-diversified segmentation neural network to analyze training indicators corresponding to training digital images and generate object segmentation outputs corresponding to different scales; Receive a digital image and a user indicator, the user indicator including one or more pixels of the digital image; Steps for generating multiple object segmentation outputs corresponding to multiple scales by using the scale-diversified segmentation neural network based on the digital image and the user indicator; and Provide the multiple object segmentation outputs for display by: Providing a graphical user interface element for selecting an object segmentation output to display; and In response to multiple different inputs via the graphical user interface element, provide the multiple object segmentation outputs for display.
10. The computer-implemented method according to claim 9, wherein: The multiple scales include a first scale having a first size and a first aspect ratio and a second scale having a second size and a second aspect ratio.
11. The computer-implemented method according to claim 10, wherein: The multiple object segmentation outputs include a first object segmentation output and a second object segmentation output, the first object segmentation output includes a first object depicted in the digital image, the second object segmentation output includes the first object and a second object depicted in the digital image, The first object corresponds to the first scale, and The first object and the second object together correspond to the second scale.
12. The computer-implemented method according to claim 11, wherein providing the multiple object segmentation outputs for display includes: In response to a user input identifying the first scale, provide the first object segmentation output for display; and In response to a user input identifying the second scale, provide the second object segmentation output for display.
13. The computer-implemented method according to claim 9, wherein the training indicator includes a set of explicit training indicators and a set of fuzzy training indicators.
14. A system for displaying an output, the system comprising: one or more memory devices including digital images depicting one or more target objects; and one or more computing devices configured to cause the system to: identify a user indicator including one or more pixels of the digital image; generate a first object segmentation output at a first scale using a scale-diverse segmentation neural network according to the user indicator, the first scale corresponding to a first anchor region including a first portion of the digital image; generate a second object segmentation output at a second scale using the scale-diverse segmentation neural network according to the user indicator, the second scale corresponding to a second anchor region including the first portion and a second portion of the digital image; and in response to a first input via a graphical user interface element, provide the first object segmentation output for display, and in response to a second input via the graphical user interface element, provide the second object segmentation output for display.
15. The system according to claim 14, wherein the scale-diverse segmentation neural network includes a plurality of output channels corresponding to a plurality of scales, and wherein the one or more computing devices are further configured to cause the system to generate the first object segmentation output using a first output channel corresponding to the first scale and generate the second object segmentation output using a second output channel corresponding to the second scale.
16. The system according to claim 14, wherein the one or more computing devices are further configured to cause the system to: generate the first object segmentation output including a first object within the first anchor region; and generate the second object segmentation output including the first object and a second object within the second anchor region.
17. The system according to claim 16, wherein the one or more computing devices are further configured to cause the system to provide the first object segmentation output and the second object segmentation output for display by: providing the graphical user interface element for display by presenting a scale slider user interface element; in response to a user input identifying a first position of the scale slider user interface element, provide the first object segmentation output indicating the first object for display; and in response to a user input identifying a second position of the scale slider user interface element, provide the second object segmentation output indicating the first object and the second object for display.
18. The system according to claim 14, wherein the one or more computing devices are further configured to cause the system to: determine an input time associated with an interaction via the graphical user interface element; in response to determining the first input based on the input time associated with the interaction via the graphical user interface element, provide the first object segmentation output for display; and In response to determining the second input based on the input time associated with the interaction via the graphical user interface element, provide the second object segmentation output for display.
19. The system according to claim 14, wherein the one or more computing devices are further configured to cause the system to: Determine a vertical scale size and a horizontal scale size in response to a user input; and Provide the first object segmentation output or the second object segmentation output for display based on the vertical scale size and the horizontal scale size.
20. The system according to claim 14, wherein the one or more computing devices are further configured to cause the system to: Generate an object score corresponding to the first scale using the scale-diversified segmentation neural network; and Emphasize the first object segmentation output within the display based on the object score.
Citation Information
Patent Citations
System and method for displaying results of an image processing system that has multiple results to allow selection for subsequent image processing
US20090252429A1