Implementation of natural language processing and multiple object detection models for automatic selection of objects in images.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-12-10
- Publication Date
- 2026-04-09
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
background
[0001] In recent years, there has been a significant increase in digital image manipulation. Advances in both hardware and software have enhanced individuals' ability to capture, create, and edit digital images. For example, the hardware in most modern computing devices (such as smartphones, tablets, servers, desktops, and laptops) allows for digital image manipulation without noticeable lag or processing delays. Similarly, software improvements enable individuals to modify, combine, filter, or otherwise manipulate digital images. Examples of digital image manipulation include detecting an object, copying an object from an image to a new background, or removing an object from an image.
[0002] Despite these improvements in digital image processing, conventional systems suffer from numerous problems regarding flexibility, accuracy, and operational efficiency, particularly in detecting and selecting objects within digital images. For example, many conventional systems have limited functionality in their ability to select objects based on natural language sentence input. If a user's input for selecting an object includes a colloquial term, many conventional systems are unable to recognize it. Many systems ignore the colloquial term as an unrecognized object and, as a result, fail to select an object within the digital image.Similarly, many conventional systems are unable to recognize specialized or specific terms that involve a high degree of granularity. Again, conventional systems are too inflexible to produce adequate results.
[0003] In another example of inflexibility, many conventional systems fail to identify or process relationships between objects contained in a natural language-based object selection request. For instance, if the natural language-based object selection request calls for the selection of an object that relates to another object, many conventional systems cannot correctly identify which object should be selected. Although a small number of conventional systems have addressed this issue, they remain limited to a generic relationship operator that can only handle simplified relationships.
[0004] Furthermore, conventional systems are not accurate. As mentioned above, many conventional systems fail, for example, to recognize one or more object terms in a natural language-based object selection request. Consequently, these conventional systems fail to provide results to the user. Alternatively, some conventional systems fail to correctly recognize an object term and return an incorrect object. In either case, conventional systems provide incorrect, imprecise, and inaccurate results to the user. Similarly, if conventional systems fail to recognize relationships between multiple objects in a natural language-based object selection request, they cannot provide accurate results to the user.
[0005] Furthermore, conventional systems are inefficient. For example, they have significant drawbacks in the automatic detection and selection of objects. If a conventional system provides an inaccurate result, it wastes computing resources and real-time memory. In another example, many conventional object detection systems are end-to-end neural networks. If an error occurs or the desired result is not achieved, users or even the system's designers are unable to pinpoint which component is malfunctioning. Instead, the entire system must be retrained and adjusted until the desired result is achieved—a process that can consume considerable time and computing resources.
[0006] Furthermore, many conventional systems provide inefficient mouse-based tools that also require manual selection of a desired object. These tools are often inaccurate and impractical for many selection tasks. As a result, considerable time and user interaction with various selection tools waste significant computing resources when detecting, displaying, selecting, and correcting object selections in digital images. Systems for modifying or identifying digital images using voice or text input are known from DE 10 2018 007 937 A1, DE 10 2016 010 909 A1, DE 10 2017 010 210 A1, and DE 10 2018 010 162 A1. Correcting object selections in digital images.
[0007] Systems for modifying or identifying digital images using voice or text input are known from DE 10 2018 007 937 A1, DE 10 2016 010 909 A1, DE 10 2017 010 210 A1 and DE 10 2018 010 162 A1.
[0008] These problems and tasks, along with other problems and tasks, exist in image processing systems with regard to detecting and selecting objects in digital images. Brief summary
[0009] Implementations of the present disclosure offer advantages and / or solve one or more of the aforementioned problems or other prior art problems in systems, non-temporary computer-readable media, and methods for automatically selecting detected objects in a digital image based on natural language input. For example, the disclosed systems can employ natural language processing tools to detect objects and their corresponding relationships within natural language object selection queries. The disclosed systems can, for example, determine alternative object terms for unrecognized objects in a natural language object selection query.In another example, the disclosed systems identify several types of object relationships in a natural language-based object selection query and employ different object relationship models to select the requested query object. In yet another example, the disclosed systems can employ an object selection pipeline consisting of interchangeable object-detecting neural networks and models to accurately detect and automatically select the query object identified in the natural language-based object selection query.
[0010] The following description shows additional features and advantages of one or more implementations of the disclosed systems, computer media and methods. Brief description of the drawing
[0011] The detailed description provides one or more implementations with additional specificity and detail based on the accompanying drawing, which is briefly described below. Fig. Figure 1 shows a schematic diagram of an environment in which an object selection system can operate according to one or more implementations. Fig. Figure 2 shows a schematic diagram of the automatic detection and selection of a query object in an image based on natural language user input according to one or more implementations. Fig. Figures 3A to 3D show a graphical user interface for representing a process of automatically detecting and selecting a query object in an image according to one or more implementations. Fig. Figure 4 shows a schematic diagram of an object selection pipeline according to one or more implementations. Fig. Figure 5 shows a flowchart of identifying and selecting a query object based on an alternative object concept according to one or more implementations. Fig. 6A and Fig. Figure 6B shows several mapping tables used to identify alternative mapping terms of a query object, corresponding to one or more implementations. Fig. Figures 7A to 7D show a graphical user interface for representing a process of selecting a query object based on one or more alternative terms according to one or more implementations. Fig. Figure 8 shows a flowchart of the deployment of an object relationship model to detect a query object among multiple objects associated with a natural language-based user input, according to one or more implementations. Fig. Figure 9A shows a graphical user interface of a digital image that includes a natural language-based object selection query, according to one or more implementations. Fig. 9B shows a component graph of the natural language-based object selection query of Fig. 9A. Fig. Figure 10 shows a block diagram of several object relationship models corresponding to one or more implementations. Fig. Figure 11 shows a schematic diagram illustrating the use of a relative object position model for selecting an object according to one or more implementations. Fig. Figures 12A to 12D show a graphical user interface for representing a process of using an object relationship model to select a query object from among several objects associated with a natural language-based user input, according to one or more implementations. Fig. Figure 13 shows a table for evaluating different implementations of the object selection system according to one or more implementations. Fig. Figure 14 shows a schematic diagram of the object selection system according to one or more implementations. Fig. Figure 15 shows a flowchart of a sequence of operations for deploying object relationship models to detect a query object according to one or more implementations. Fig. Figure 16 shows a flowchart of a sequence of operations for inserting alternative object terms and multiple object detection models for detecting a query object according to one or more implementations. Fig. Figure 17 shows a block diagram of an exemplary computing device for implementing one or more implementations of the present disclosure. Detailed description
[0012] This disclosure describes one or more implementations of an object selection system for accurately detecting and automatically selecting user-requested objects (e.g., query objects) in a digital image. In particular, in one or more implementations, the object selection system employs one or more natural language-based processing tools to identify object terms and relationships within a natural language-based object selection query. Additionally, the object selection system can establish and update an object selection pipeline to select a requested query object from among one or more query objects specified in the natural language-based object selection query.Furthermore, the object selection system can add, update, or replace sections of the object selection pipeline to improve the overall accuracy and efficiency of automatic object selection within an image.
[0013] For clarity, the object selection system can generate and deploy an object selection pipeline to detect objects within a query string (that is, within a natural language-based object selection query). Generally, the object selection pipeline includes multiple object detection models (such as neural networks) designed to identify different classes of objects. For example, the object selection pipeline might include models that detect specific class objects, models that detect known class objects, models that detect concepts, and / or models that detect general objects. This allows the object selection system to employ the object detection model that most accurately and efficiently detects an object identified from the query string.
[0014] The object selection system can employ natural language processing tools to detect a query object among multiple objects specified in a natural language object selection query. For illustrative purposes, in one or more implementations, the object selection system generates a component graph to identify multiple objects (and / or object classes) from a query string (i.e., a natural language object selection query). Additionally, the object selection system can identify a relationship type from the component graph and, based on this relationship type, identify one object relationship model among several. The object selection system can generate object masks for each of the identified objects and analyze these masks to determine an object that satisfies the object relationship model (for example, the query object).Furthermore, the object selection system can provide the query object selected within a digital image in response to the natural language-based object selection query.
[0015] The object selection system can employ one or more object relationship models to detect a query object from a query string. In various implementations, the object selection system generates a component graph or a tree to specify each of the objects (or object classes) contained in a query string, as well as the relationship type between objects. For each object, the object selection system can use the object selection pipeline to identify a target object-detecting model that accurately detects the object and / or generates an object mask for each detected instance of the object.
[0016] In additional implementations, the object selection system can determine an object relationship model based on the relationship type specified in the component graph. Based on the relationship type, the object selection system selects, for example, a target object relationship model from among several object relationship models. For illustrative purposes, the object selection system can determine an object relationship model from an object contact model, a relative object position model, a background / foreground object model, or another object relationship model. When selecting a target object relationship model, the object selection system can determine which of the detected objects fulfills the target object relationship model, which the object selection system can then select as the request object for provision within the digital image.
[0017] In additional implementations or alternative approaches, the object selection system can employ natural language-based processing tools to determine alternative object terms in order to detect a query object. For illustrative purposes, in one or more implementations, the object selection system identifies a query object from a natural language-based object selection query. Based on the identified query object, the object selection system determines that the query object does not correspond to any known object class recognized by a model that detects known object classes. Accordingly, the object selection system can use a mapping table to identify one or more alternative object terms for the query object.Based on one of the alternative object concepts, the object selection system can use a model that detects a known object class to detect and generate an object mask for the query object. Additionally, the object selection system can provide the selected query object within a digital image in response to the natural language-based object selection query.
[0018] The object selection system can therefore use a mapping table to identify one or more alternative object terms for the query object or other objects detected in a query string. In various implementations, the object selection system uses the mapping table to identify synonyms of an object that has been determined to correspond to an unknown object class (for example, an unrecognized object). In some implementations, the mapping table provides hypernyms of an unrecognized object.
[0019] By providing synonyms and / or hypernyms as alternative object terms for an object of an unrecognized class, the object selection system can detect the object more efficiently and accurately. For example, if the object selection system does not recognize an object in the query string, it attempts to detect the object using a generic detector that detects objects of unknown object class. Conversely, if the object selection system recognizes a synonym of the object as belonging to a known object class, it can employ a neural network that detects known object classes to detect the object accurately and efficiently. Similarly, the object selection system can more accurately detect hypernyms of an unrecognized object by using a target-object-detecting model in the object selection pipeline.
[0020] Depending on whether the object selection system detects a synonym or a hypernym, it can further modify the object selection pipeline. For example, if the object selection system uses a hypernym for an unrecognizable object, it can exclude various paths and / or object detection models from the object selection pipeline, as described below. The object selection system can also add additional stages to the object selection pipeline, such as a verification stage, depending on whether the object is a synonym or a hypernym.
[0021] As mentioned above, the object selection system offers numerous advantages and useful features over conventional systems due to its practical application (for example, the automatic selection of objects within images using one or more natural language processing tools). The object selection system can, for instance, automatically detect and select objects across a wide range of object types and classes based on natural language user input (that is, a natural language object selection query). The object selection system can employ various natural language processing tools and techniques to detect objects that conventional systems would otherwise fail to detect or would only detect with the use of unnecessary computing resources.Accordingly, the object selection system offers increased flexibility, improved efficiency and enhanced functionality compared to conventional systems.
[0022] For illustrative purposes, the object selection system offers increased flexibility by detecting objects that are not otherwise recognized or that do not belong to a known object class. For example, the object selection system can recognize objects in a natural language-based object selection query that uses colloquial terms not recognized by conventional systems. Similarly, the object selection system can recognize objects in a natural language-based object selection query that uses highly granular or specific terms, also not recognized by conventional systems.In this way, the object selection system can expand the breadth and range of objects that can be detected by using more accurate and efficient object-detecting models (for example, by using specialized object-detecting models and known object-class detecting models).
[0023] Additionally, the object selection system offers improved accuracy compared to conventional systems. For example, the object selection system improves object detection accuracy by better identifying objects specified in a natural language-based object selection query. If a conventional system does not recognize an object term, it is largely unable to detect the object. In cases where the conventional system can use a generic object-detecting network to detect the object, it often returns the wrong object or an inaccurate selection.In contrast, the object selection system can use specialized object-detecting models or known object-class-detecting models to accurately identify the object by utilizing the natural language-based processing tools and techniques described here.
[0024] In another example, the object selection system can more accurately identify a query object by using a target object relationship model. Each of the object relationship models is tailored to specific types of object relationships. Therefore, by using an object relationship model that matches the relationship type between objects, the object selection system can achieve more accurate results than conventional systems that use a generic, one-size-fits-all relationship model (universal relationship model). In many cases, the generic relationship model is unable to recognize or correctly apply the relationship between objects.
[0025] Furthermore, the object selection system offers improved efficiency compared to conventional systems by employing target object detection models and an object selection pipeline. For example, using a specialized object detection model or a model that detects a known object class to detect an object is more efficient and accurate than using a generic object detection network, which involves many additional steps to identify an object. Similarly, using a target object relationship model is more efficient and requires fewer computations and steps than using a generic relationship model.
[0026] Furthermore, unlike closed, conventional end-to-end systems, the object selection system can isolate the faulty component in the object selection pipeline and repair it when an error occurs. Additionally, the object selection system can add further components to improve accuracy. For example, it can add specialized neural networks to the object selection pipeline that detect frequently requested objects. Similarly, the object selection system can replace components within the object selection pipeline with more efficient versions and / or update them.
[0027] Additionally, the object selection system significantly reduces the number of steps many traditional systems require a user to perform when selecting an object within an image. Instead of relying on inefficient mouse-based tools to manually select an object, the user "tells" the object selection system which object to select (for example, by providing verbal cues in a natural language-based object selection query or query string), and the system automatically detects and accurately selects the object. Furthermore, the object selection system greatly simplifies the object selection process to one or two simple user actions to achieve accurate results, rather than the numerous steps previously required to achieve only mediocre results.
[0028] Additional advantages and useful properties of the object selection system will become apparent from the following description. As explained in the preceding discussion, the present disclosure also employs a variety of terms to describe features and advantages of the object selection system. Before describing the object selection system with reference to the figures below, further details regarding the meaning of these terms are given.
[0029] For the purposes of this document, the term "digital image" (or simply "image") refers to a digital graphics file that displays one or more objects when played back. Specifically, an image can contain one or more objects belonging to any suitable object type or object class. In various implementations, the image processing system displays an image on a computing device, such as a client device. Additional implementations allow a user to modify or change an image, as well as generate new images. For example, the image processing system allows a user to copy an object selected in one image to the background of a second image. Furthermore, a digital image can include one or more frames in a video or animation, along with other digital images.
[0030] For the purposes of this text, the term "object" refers to a visual representation of a subject, concept, or sub-concept in an image. Specifically, an object is a set of pixels in an image that are combined to form a visual representation of an object, item, sub-object, component, or element. An object can correspond to a wide range of classes and concepts. For example, objects include specialized objects, object categories (such as concept objects), object classes, objects from known classes, and unknown object classes (such as objects that were not used when training any of the object-detecting neural networks). In some implementations, an object may contain multiple instances of the object.For example, an image of a rosebush contains multiple instances of roses; alternatively, the object term "furniture" can contain subgroups of a chair, a desk, and a sofa. In one or more implementations, an object contains sub-objects, parts, or sections. For example, a person's face or leg can be an object that is part of another object (such as the person's body). In another example, an object is a shirt that can be part of another object (such as a person).
[0031] As mentioned above, the object selection system can accurately detect and automatically select an object within an image based on a query string. For the purposes of this document, the term "natural language-based object selection query" or equivalently "query string" refers to a text string containing one or more object terms (i.e., words) that specify a target object. A query string can be a natural language-based user input containing a noun representing a query object. Additionally, a query string can include object terms for other objects that are related to the query object. Generally, the object selection system receives a query string when a user requests that the system automatically select an object in an image. In some implementations, the query string is provided as a text string.In alternative implementations, the object selection system detects an alternative user input, such as speech data, and converts the alternative user input into text to produce the query string.
[0032] As mentioned earlier, a query string can contain a query object. The term "query object" refers to the object in the query string that the user requests for detection and selection. A noun in the query string, for example, specifies the query object. If a query string contains multiple objects (for example, several nouns), the query object is generally the first one listed. In additional implementations, the query string includes additional words, such as adjectives and adverbials, that specify attributes of the query object. As mentioned above, the query string can also contain other nouns (and corresponding attributes) that indicate a relationship to the query object. In the context of this document, the term "object attribute" refers to a descriptive word used to further identify the query object.Examples of object attributes include color, size, length, shape, position, location, pattern, composition, expression, feel, strength and / or flexibility.
[0033] For the purposes of this document, the term "alternative object term" refers to a replacement term for an object or object term in a query string. In one or more implementations, an alternative object term is a synonym of an object term in a query string. In some implementations, an alternative object term is a hypernym of the object term in a query string. In several implementations, an alternative object term is a hyponym of the object term in a query string. As described here, the object selection system can use one or more mapping tables to identify an alternative object term for an object term in a query string.
[0034] In the present context, the term "mapping table" refers to a data structure that links related words. A mapping table is, for example, a database, table, or list containing terms (such as object terms) and their corresponding alternative object terms. A mapping table can provide synonyms, hypernyms, hyponyms, and / or other terms for a given term. In some implementations, the object selection system uses multiple mapping tables, such as a synonym mapping table and a hypernym mapping table. The mapping table can be updated at regular intervals or whenever the object selection system detects new alternative object terms for a given object term.
[0035] For the purposes of this document, the terms "object mask," "segmentation mask," or "object segmentation" refer to the specification of multiple pixels that represent an object. An object mask can, for example, include a segmentation boundary (such as a limiting line or curve to specify an edge of one or more objects) or a segmentation mask (such as a binary mask to identify pixels corresponding to an object). Generating an object mask is sometimes referred to as "selecting" a target object (that is, identifying pixels that represent the target object).
[0036] For the purposes of this document, the term "approximate boundary" refers to the specification of an area containing an object that is larger and / or less precise than an object mask. In one or more implementations, an approximate boundary can include at least a portion of a query object and portions of the image that do not include the query object. An approximate boundary can include any shape, such as a square, rectangle, circle, oval, or any other outline that surrounds an object. In one or more implementations, an approximate boundary includes a bounding box.
[0037] The term "object selection pipeline" refers to a collection of components and actions used to detect and select a query object in an image. In various implementations, the object selection system uses a subset of the components and actions in the object selection pipeline to detect and select a query object in an image, with the output of one component serving as input for another. The components and actions can include neural networks, machine learning models, heuristic models, and / or functions. Furthermore, the components and actions in the object selection pipeline can be interchangeable, removable, replaceable, or updatable, which will be described in more detail below.
[0038] For the purposes of this document, the term "component graph" refers to a data structure that characterizes objects and relationships derived from a query string. A component graph is, for example, a data tree containing multiple objects to be detected and a relationship type between those objects. In one or more implementations, the component graph includes the components "locate," "reference," and "intersect," where "locate" specifies one or more objects to be detected, "reference" specifies a relationship type, and "intersect" specifies an object relationship between objects based on the relationship type. Further details regarding component graphs and an example of a component graph will be provided in the following section. Fig. 9B indicated.
[0039] The term "relationship type" refers to a characteristic property of how objects in a query string (for example, in a digital image) are related to each other. A relationship type, for instance, specifies the object relationship (such as a spatial relationship) between two or more objects. Examples of relationship types include spatial relationships, proximity relationships, depth relationships, relative position relationships, absolute position relationships, and exclusion relationships.
[0040] As mentioned above, the object selection system can use object relationship models to identify a query object in an image. For the purposes of this document, the term "object relationship model" refers to a relationship operator that can select an input object (i.e., the query object) based on its relationship to one or more input objects. In various implementations, the object relationship model employs heuristics and / or rules to select the query object. Some implementations use a machine learning approach. Examples of object relationship models include, among others, an object touch model, a relative object position model, and a background / foreground object model.
[0041] For the purposes of this discussion, the object selection system can employ machine learning and various neural networks in different implementations. The term "machine learning model" refers to a computer representation that can be tuned (e.g., trained) based on inputs to approximate unknown functions. Specifically, a machine learning model can include a model that uses algorithms to learn from known data and make predictions about it by analyzing the known data to learn and generate outputs that reflect patterns and attributes of the known data. The term "machine learning model" can encompass linear regression models, logistic regression models, random forest models, SVG (Support Vector Machines), neural networks, or decision tree models.A machine learning model can therefore perform high-level abstractions on data by generating data-driven predictions or decisions from the known input data.
[0042] Machine learning can include neural networks (for example, a natural language processing neural network, a specialized object-detecting neural network, a concept-based object-detecting neural network, a known object class-detecting neural network, an object-suggesting neural network, an unknown object class-detecting neural network, a domain-suggesting neural network, a concept-embedding neural network, an object-masking neural network, an object-classifying neural network, an object category-detecting neural network, and / or a selected object attribute-detecting neural network), data-based models (for example, a natural language processing model, a general object-detecting model,a model that detects an unknown object class, a model that recognizes an object, a filter model and / or a selection object attribute model) or a combination of networks and models.
[0043] For the purposes of this document, the term "neural network" refers to a machine learning model comprising interconnected artificial neurons that communicate and learn to approximate complex functions and generate outputs based on multiple inputs provided to the model. The term "neural network" includes, for example, an algorithm (or set of algorithms) that implements deep learning techniques, employing a set of algorithms to model high-level abstractions in data using monitoring data to fine-tune parameters of the neural network.Examples of neural networks include a convolutional neural network (CNN), a residual learning neural network (Residual Learning Neural Network), a recurrent neural network (RNN), a graph neural network, a generative-adversarial neural network (GAN), a region-based CNN (R-CNN), a faster R-CNN, a mask R-CNN, and single-shot detect SSD networks.
[0044] The figures show Fig. 1 A schematic diagram of an environment 100 in which the object selection system 106 can operate according to one or more implementations. As in Fig. As shown in Figure 1, environment 100 includes a client device 102 and a server device 110, which are connected via a network 108. Additional details relating to computing devices (for example, the client device 102 and the server device 110) are given below in conjunction with Fig. 17 is indicated. Furthermore, it shows Fig. 17 additional details relating to networks, such as the network 108 shown.
[0045] Although Fig. While environment 100 represents a specific number, type, and arrangement of components, various additional configurations of the environment are possible. For example, environment 100 can contain any number of client devices. In another example, server device 110 can represent a set of connected server devices. In yet another example, client device 102 can communicate directly with server device 110, bypassing network 108 or using a separate and / or additional network.
[0046] As shown, environment 100 includes client device 102. In various implementations, client device 102 is assigned to a user (for example, a user client device), such as a user requesting automatic selection of an object in an image. Client device 102 may include an image processing system 104 and an object selection system 106. In some implementations, image processing system 104 implements object selection system 106. In alternative implementations, object selection system 106 is separate from image processing system 104. Although image processing system 104 and object selection system 106 are shown on client device 102, in some implementations, image processing system 104 and object selection system 106 are located remotely from client device 102 (for example, on server device 110), as further explained below.
[0047] The Image Processing System 104 generally facilitates the creation, modification, sharing, and / or detection of digital images. For example, the Image Processing System 104 provides a variety of tools related to image creation and editing (such as photo editing). The Image Processing System 104 provides, for example, selection tools, color correction tools, and image manipulation tools. The Image Processing System 104 can also work in conjunction with one or more applications for generating or modifying images. In one or more implementations, the Image Processing System 104 works, for example, in conjunction with digital design applications or other image editing applications.
[0048] In some implementations, the Image Processing System 104 provides an intelligent image processing assistant that performs one or more automatic image processing operations for the user. For example, the Image Processing System 104 receives a natural language-based object selection request (or request string) such as "Make the red dress yellow," "Make the background blurred and gray," or "Increase the contrast on the water." As part of fulfilling the request, the Image Processing System 104 uses the Object Selection System 106 to automatically select the appropriate request object identified in the request string. The Image Processing System 104 can then employ additional system components (such as a color replacement tool, a blur filter, or an image adjustment tool) to perform the requested operation in relation to the detected request object.
[0049] As mentioned above, the image processing system 104 includes the object selection system 106. As described in detail below, the object selection system 106 accurately detects and automatically selects objects in an image based on a user request (for example, based on a user-provided query string). In many implementations, the object selection system 106 employs natural language processing tools and an object selection pipeline to determine which object-detecting neural networks to use based on the query object, and which additional neural networks and / or models to use to select the specific requested query object.
[0050] As shown, environment 100 also includes server device 110. Server device 110 includes an object selection server system 112. In one or more implementations, object selection server system 112 provides and / or renders similar functionality to that described here in connection with the object selection system. In some implementations, object selection server system 112 supports object selection system 106 on client device 102.
[0051] In one or more implementations, the server device 110 can contain the complete object selection system 106 or a part thereof. In particular, the object selection system 106 on the client device 102 can download an application from the server device 110 (for example, an image editing application from the object selection server system 112) or a part of a software application.
[0052] In some implementations, the object selection server system 112 may include a web hosting application that allows the client device 102 to interact with content and services hosted on the server device 110. For illustrative purposes, in one or more implementations, the client device 102 accesses a web page supported by the server device 110, which hosts the models that enable automatic object selection in images based on a request string provided by the user via the client device 102.In another example, client device 102 includes an image processing application that provides the image and the request string to the object selection server system 112 on server device 110. Server device 110 then detects the request object using one or more natural language processing tools and one or more object detection models, and in turn provides an object mask of the detected request object to client device 102. Using this object mask, the image processing application on client device 102 then selects the detected request object.
[0053] The next character is Fig. 2, which provides an overview of the use of the object selection system for automatically selecting an object in an image. In particular, it shows Fig. 2. A sequence of operations 200 for the automatic detection and selection of a query object in an image based on natural language user input, according to one or more implementations. In various implementations, the object selection system 106 performs the sequence of operations 200. In some implementations, an image processing system and / or an image processing application performs one or more of the operations included in the sequence of operations 200.
[0054] The procedure is carried out as described in Fig. Figure 2 shows an operation (202) of the object selection system 106: identifying a query string that specifies an object to be selected in an image. For example, a user uses an image editing program to edit an image. While editing the image, the user wants to select a specific object within the image. Accordingly, the object selection system 106 provides the user with a graphical interface that allows the user to enter a query string requesting automatic selection of the object. The object selection system 106 can allow the user to provide the query string as typed text or spoken words, which the object selection system 106 then converts into text. As shown in Fig. As shown in Figure 2 in conjunction with process 202, the object selection system 106 can receive the query string "boy's cap on the right".
[0055] In response to the detection of a query string by the object selection system 106, process 204 is performed: inserting a mapping table to identify alternative object terms for unknown object terms in the query string. For example, the object selection system 106 analyzes the query string and determines that the object term "cap" corresponds to a known object class. However, the object selection system 106 also determines that the object term "boy" is not recognized as belonging to a known object class. Accordingly, the object selection system 106 can use a mapping table to look up alternative object terms, such as synonyms or hypernyms. As shown, the object selection system 106 identifies the alternative object terms "man," "boy," and "male person" for the unrecognized object term "boy."Additional details relating to the use of figure tables and alternative object terms are explained below. Fig. 5 to 7D indicated.
[0056] As shown, the sequence of operations 200 includes operation 206, in which the object selection system 106 generates object masks for each of the objects in the query string with the alternative object terms. For example, if objects are identified either directly or based on an alternative object term in the query string, the object selection system 106 can use the object selection pipeline to generate masks for each of the objects. As shown in connection with operation 206, the object selection system 106 generated an object mask for each man (for example, the alternative object term for "boy") and both caps. Additional details related to the use of the object selection pipeline are given below in conjunction with Fig. 4 provided. In one example, the object selection system 106 can use an object detection model trained to select people based on a mapping of the query term to "human" using the mapping table.
[0057] In some cases, the object selection system 106 detects multiple objects in the query string. For example, the selection of the query object is predicated on another object specified in the query string. Accordingly, the sequence of operations 200, as shown, includes operation 208, in which the object selection system 106 uses a component graph and a corresponding object relationship model to identify the query object from the object masks.
[0058] In various implementations, the object selection system 106, as part of operation 208, can generate a component graph from the query string that locates objects within the query string and specifies relationships between them. As described below, the object selection system 106 can use relationship types to select a target object relationship model to determine the query object from the objects in the query string. In this example, the object selection system 106 can use the object relationship model to identify the relationship between the two men, in order to identify the man on the right. Additional details regarding the use of a component graph and an object relationship model to identify the query object are described below. Fig. 8 to 12D are specified.
[0059] As in Fig. As shown in Figure 2, the sequence of operations 200 includes operation 210, in which the object selection system 106 provides the selected query object within the image. The object selection system 106 provides the image, for example, on a computing device, with the query object being automatically selected in response to receiving the query string. As shown, the cap of the boy on the left is selected in the image. In additional implementations, the object selection system 106 can automatically perform additional steps with the selected query object based on instructions detected in the query string, such as "Make the cap of the boy on the right blue".
[0060] The object selection system 106 can perform operations 202 to 210 in a variety of sequences. For example, the object selection system 106 can perform operation 208, inserting a component graph, before performing operation 206, generating object masks. In some implementations, the object selection system 106 omits operation 208, inserting a component graph.
[0061] Fig. Figures 3A to 3D show a client device 300 having a graphical user interface 302, where a process of selecting a query object in a figure 304 based on an object detection request (i.e., based on a natural language-based object selection request) is depicted, according to one or more embodiments. The client device 300 in Fig. 3A and Fig. 3B can represent the client device 102 as shown above. Fig. 1 has been introduced. The client device 300 includes, for example, an image processing application that implements the image processing system 104 with the object selection system 106. The graphical user interface 302 in Fig. 3A and Fig. 3B can be generated by the image editing application, for example.
[0062] As in Fig. As shown in Figure 3A, the graphical user interface 302 includes image 304 within an image editing application. Image 304 shows, on the left, a woman holding a surfboard and entering the water at the beach, and on the right, a man walking along the beach. For clarity, image 304 has been simplified and does not include any additional objects or object classes.
[0063] The image processing system and / or the object selection system 106 can provide an object selection interface 306 in response to the detection that a user has selected an option to have an object automatically selected. For example, the object selection system 106 provides the object selection interface 306 as a selection tool within the image processing application. As shown, the object selection interface 306 can include a text field into which a user can enter a natural language-based object selection request in the form of a request string (i.e., "Lady holding a surfboard"). The selection interface 306 also includes selectable options. For example, the object selection interface includes a selectable element to confirm (i.e., "OK") or cancel (i.e., "Cancel") the object detection request.In some implementations, the object selection interface 306 includes additional elements, such as a selectable option to record audio input from a user dictating the request string.
[0064] Based on the detection of the query string from the object detection request, the object selection system 106 can initiate the detection and selection of the query object. As described in more detail below, the object selection system 106 uses, for example, natural language-based processing tools to identify objects in the query string and also to classify the objects as known objects, specialized objects, object categories, or unknown objects.
[0065] For illustrative purposes, object selection system 106 can generate a component graph to determine that the query string contains multiple objects (i.e., "queen" and "surfboard") as well as a relationship between the two objects (i.e., "holds"). Additionally, object selection system 106 can classify each of the objects identified in the query string, for example, classifying the term "queen" as belonging to an unknown object class and the term "surfboard" as belonging to a known object class. When determining that the object term "queen" is not recognized or is unknown, object selection system 106 can further employ a mapping table to identify alternative object terms, for example, by identifying the alternative object term "person," which corresponds to a known object class.
[0066] Using the identified objects in the query string, the object selection system 106 can employ an object selection pipeline. Based on the classification of an object, the object selection system 106 can, for example, employ an object detection model that detects the corresponding object extremely efficiently and accurately. Using the term "person," the object selection system 106 can, for example, select a specialized object-detecting model to detect each of the people included in Figure 304. The specialized object-detecting model and / or an additional object-masking model can also create object masks of the two people, as shown in Fig. 3B is shown, generate.
[0067] Similarly, object selection system 106 can use the object selection pipeline to select a model detecting a known object class in order to detect the surfboard in Figure 304, and in some cases can also generate an object mask of the surfboard. For illustrative purposes, Fig. 3C a mask of the surfboard in picture 304.
[0068] As mentioned above, the object selection system 106 can isolate the query object from the detected objects. In various implementations, the object selection system 106 can, for example, use the relationship type specified in the component graph to select an object relationship model. As noted above, the component graph can specify the relationship "hold" between the lady (that is, the person) and the surfboard. Based on the relationship type, the object selection system 106 can therefore select an object relationship model that is tailored to detect the intersection of the first object "hold" with a second object.
[0069] In one or more implementations, the object relationship model uses a machine learning model to determine the overlap between the two objects. In alternative implementations, the object relationship model uses heuristic models. In any case, the object selection system 106 uses the selected object relationship model to determine a query object by identifying an overlapping object mask that satisfies the object relationship model.
[0070] For illustrative purposes, it shows Fig. 3D shows the result of the query string. In particular, it shows Fig. 3D, that an object mask selects the woman holding the surfboard (and not the surfboard or the running man). In this way, the user is able to easily modify the woman as desired within image 304 in the image editing application. In various implementations, the object selection system 106 can additionally provide the automatically selected woman in response to the query string, without the intermediate steps that Fig. 3B and Fig. 3D are assigned, to show. In response to the detection of the request string in Fig. 3A can also automatically use the object selection system 106 to provide the in Fig. The 3D result will jump while the intermediate steps are performed in the background.
[0071] In Fig. Sections 4 to 12D show additional details relating to how the object selection system generates and uses 106 natural language-based processing tools and the object selection pipeline to automatically select and accurately detect objects requested in an object detection request. In particular, it shows Fig. 4. An exemplary implementation of the object selection pipeline. Fig. Figures 5 to 7D show the use of mapping tables and alternative object terms to select object detection models within the object selection pipeline. Fig. Figures 8 to 12D show the use of a component graph and an object relationship model to identify a query object among several detected objects specified in the query string.
[0072] As mentioned, shows Fig. Figure 4 shows a schematic diagram of an object selection pipeline 400 according to one or more implementations. In different implementations, the object selection system performs 106 operations included in the object selection pipeline 400. In alternative implementations, the image processing system and / or the image processing application can perform one or more of the included operations.
[0073] As shown, the object selection pipeline 400 includes an operation 402 in which the object selection system 106 retrieves a query string corresponding to an image (that is, a digital image), for example, an image in an image editing application, as described above. Generally, the image contains one or more objects. The image can contain background objects (that is, a landscape), foreground objects (that is, image subjects), and / or other types of objects.
[0074] Additionally, the 402 operation may include the fact that object selection system 106 is receiving a query string. Object selection system 106, for example, provides an object selection interface (which, for example, is located in...). Fig. (as shown in Figure 3A) is ready, into which a user can enter one or more words (for example, user input) that specify the query object they want the object selection system to select automatically. As described above, in some implementations, the object selection system 106 can allow alternative forms of user input, such as audio input that tells the object selection system 106 to select an object in the image.
[0075] Next, the object selection pipeline 400 includes an operation 404 in which the object selection system 106 generates a component graph of the query string to identify one or more objects and / or relationship types. The object selection system 106 can, for example, use one or more natural language processing tools, as described below, to generate a component graph. If the query string contains multiple objects, the component graph can also specify a relationship type between the objects.
[0076] As shown, the object selection pipeline 400 includes an operation 406 in which the object selection system 106 determines whether any of the objects (for example, any of the object terms) corresponds to an unknown object class. In one or more implementations, the object selection system 106 can compare each of the objects with known object classes (for example, known objects, specialized objects, object categories) to determine whether the object is known or unknown. The object selection system 106 can compare an object with a list or a lookup table to determine whether an object detection model has been trained to specifically detect the object.
[0077] If one of the objects in the query string is determined not to correspond to an unknown object class (for example, if the object is not recognized), the object selection system 106 can proceed to operation 408, in which the object selection pipeline 400 of the object selection system 106 uses a mapping table to identify an alternative object term. As described below in conjunction with Fig. If the object is specified in detail in sections 5 to 7D, the object selection system 106 can often identify a synonym or hypernym for the unrecognized object term, which the object selection system 106 can then use in conjunction with the object selection pipeline 400 to detect the object. If the object selection system 106 cannot identify an alternative object term, it can jump to operation 424 of the object selection pipeline 400, as described in more detail below. Furthermore, as described below, the object selection system 106 can use the mapping table to determine alternative object terms for terms in a query string that are initially not recognized by the object selection system 106 within the SOP or the object selection pipeline 400.
[0078] As shown, object selection pipeline 400 includes an operation 410 in which object selection system 106 determines whether the alternative object term is a synonym or a hypernym. If the alternative object term is a synonym, object selection system 106 can proceed to operation 412 in object selection pipeline 400, as described below. A synonym of an unrecognized object term often includes at least one alternative object term that object selection system 106 recognizes as belonging to a known object class. In this way, for the purpose of object detection using an object detection model, object selection system 106 treats the synonymous alternative term as if it were the object term (i.e., as a substitute object term).
[0079] If the alternative object term is a hypernym, the object selection system 106 can, in one or more implementations, proceed to operation 416 of the object selection pipeline 400, as described below. The object selection system 106 can skip operation 412 for efficiency if a hypernym of an unrecognized object would likely not recognize the object because the hypernym's scope is broader than that of the unknown object. As with synonyms, hypernyms of an unrecognized object term often include at least one alternative object term that is recognized by the object selection system 106. In alternative implementations, the object selection system 106 provides the hypernym of the unrecognized object for operation 412 instead of operation 416.
[0080] As in Fig. As shown in Figure 4, if the object selection system 106 recognizes any of the objects (see, for example, operation 406) or recognizes a synonym for any unknown object (see, for example, operation 410), the object selection system 106 can proceed to operation 412. As also shown by the dashed box, the object selection system 106 can iterate through operations 412 to 426 of the object selection pipeline 400 for each object included in the query string.
[0081] As shown, operation 412 of object selection pipeline 400 involves object selection system 106 determining whether the object term (for example, the object term or the alternative object term) corresponds to a specialized network. If a specialized network exists for the query object, object selection system 106 can identify a specific specialized network based on the query object. For example, object selection system 106 can compare the object term with several specialized object-detecting neural networks to identify the specialized object-detecting neural network that best matches the object. For the object term "sky," object selection system 106 can, for example, identify that a specialized object-detecting neural network "sky" is best suited to identify and select the query object.
[0082] As shown in Procedure 414, the object selection system 106 can detect the object (for example, based on the object concept or an alternative object concept) using the identified specialized network. Specifically, the object selection system 106 can employ the identified specialized object-detecting neural network to locate and detect the object within the image. For example, the object selection system 106 can use the specialized object-detecting neural network to generate a bounding box around the detected object in the image. In some implementations, if multiple instances of the object are present in the image, the object selection system 106 can employ the specialized object-detecting neural network to identify each instance separately.
[0083] In one or more implementations, a specialized object network may include a specialized object-detecting neural network, such as body parts, and / or a specialized clothing-detecting neural network. Additional details regarding the use of specialized object-detecting neural networks can be found in U.S. Patent Application No. 16 / 518,880 (Attorney File No. 20030.257.1), filed on July 19, 2019, entitled "Utilizing Object Attribute Detection Models To Automatically Select Instances Of Detected Objects in Images," which is hereby incorporated in its entirety by reference.
[0084] As shown, object selection pipeline 400 includes operation 426, which receives the output of operation 414 along with the outputs of operations 418, 422, and 424. Operation 426 involves object selection system 106 generating an object mask for the detected object. In some examples, operation 426 involves object selection system 106 employing an object-masking neural network. In various implementations, object selection system 106 can provide a bounding box for an object-masking neural network that generates a mask for the detected query object. If multiple bounding boxes are provided, object selection system 106 can utilize the object-masking neural network to generate multiple object masks under the multiple bounding boxes (for example, one object mask for each instance of the detected query object).
[0085] When generating an object mask for a detected object (or for any detected object instance), the object-masking neural network can segment the pixels in the detected object from the other pixels in the image. For example, the object-masking neural network can create a separate image layer that makes pixels corresponding to the detected object positive (e.g., binary 1), while making the remaining pixels in the image neutral or negative (binary 0). When the object mask layer is combined with the image, only the pixels of the detected object are visible. The generated object mask can provide segmentation that allows for the selection of the detected object within the image.
[0086] The object-masking neural network can correspond to one or more deep neural networks or models that select an object based on bounding box parameters corresponding to the object within an image. In one or more implementations, the object-masking neural network employs techniques and solutions found in the paper "Deep GrabCut for Object Selection" by Ning Xu et al., published on July 14, 2017, which are hereby incorporated in their entirety by reference. For example, the object-masking neural network can employ a deep-grad-cut solution instead of a salience mask transfer. In another example, the object-masking neural network can employ techniques and solutions found in the paper published on July 31, 2017, by Ning Xu et al., which are hereby incorporated in their entirety by reference.The US patent application filed in October 2017 with publication number 2019 / 0130229 “Deep Salient Content Neural Networks for Efficient Digital Object Segmentation”, the US patent application filed on July 13, 2018, No. 16 / 035,410 “Automatic Trimap Generation and Image Segmentation”, and the US patent application filed on November 18, 2015, No. 10,192,129 “Utilizing Interactive Deep Learning To Select Objects In Digital Visual Media”, which are hereby incorporated in their entirety by reference.
[0087] As in Fig. As shown in Figure 4, if the object selection system 106 determines during operation 412 that the object does not correspond to any specialized network, it can perform an additional determination related to the object. As shown, the object selection pipeline 400 includes operation 416, in which the object selection system 106 determines whether the object concept (for example, the object concept or the alternative object concept) corresponds to a known object class. In various implementations, the object selection system 106 employs, for example, an object-detecting neural network trained to detect objects belonging to a number of known object classes.Accordingly, object selection system 106 can compare an object class of the object (for example, based on the object concept or an alternative object concept) with the known object classes to determine whether the object is part of the known object classes. If so, object selection system 106 can proceed to operation 418 of object selection pipeline 400. Otherwise, object selection system 106 can proceed to operation 420 of object selection pipeline 400, as described below.
[0088] As mentioned earlier, object selection pipeline 400 includes operation 418, in which object selection system 106 detects the object using a network pertaining to a known object class. Known object classes can include object classes that have been labeled in training images and used to train an object-detecting neural network. Accordingly, based on the detection that the object belongs to a known object class, object selection system 106 can employ a known-object-class-detecting neural network to detect the object optimally with regard to accuracy and efficiency. Furthermore, object selection system 106 can provide the detected object to the object-masking neural network to generate an object mask (see, for example, operation 426), as described above.Additional details relating to the 418 procedure can be found in US patent application no. 16 / 518,880 (attorney file no. 20030.257.1) “Utilizing Object Attribute Detection Models To Automatically Select Instances Of Detected Objects In Images”, filed on July 19, 2019, which is hereby incorporated in its entirety by reference.
[0089] If the object selection system 106 determines that the object does not correspond to a specialized network (see, for example, operation 412) or a known object class (see, for example, operation 416), the object selection system 106 can perform an additional determination. For illustrative purposes, the object selection pipeline 400 includes operation 420, in which the object selection system 106 determines whether the object concept (for example, the object concept or an alternative object concept) corresponds to an object category (for example, uncountable objects such as water, road, and ceiling). If the object concept corresponds to an object category, the object selection system 106 determines that concept-based object-detecting techniques are used to detect the object, as described below.
[0090] For illustrative purposes, object selection pipeline 400 includes an operation 422 in which object selection system 106 detects the object using a concept-detecting network (that is, using a concept-based object-detecting neural network and / or a panoptic semantic segmenting neural network). In general, a concept-detecting network can include an object-detecting neural network trained to detect objects based on concepts, a background landscape, and other high-level descriptions of objects (e.g., semantics). Additional details relating to operation 418 are provided in U.S. Patent Application No. 16 / 518,880 (Attorney File No. 20030.257), filed on July 19, 2019.1) “Utilizing Object Attribute Detection Models To Automatically Select Instances Of Detected Objects In Images” is specified, which is hereby included in its entirety by reference.
[0091] As in Fig. As shown in Figure 4, the object selection system 106 provides the object detected by the concept-detecting network to the object-masking neural network in order to generate an object mask of the detected object (i.e., operation 426). For example, the object selection system 106 represents the detected semantic domain of an object concept within the image. As mentioned above, the object-masking neural network can generate a segmentation of the detected object, which the object selection system 106 uses as a selection of the detected object.
[0092] Up to this point in object selection pipeline 400, object selection system 106 has been able to detect objects according to known object classes. Using the object concept or an alternative object concept, object selection system 106 has been able to map the object concept to an object-detecting model trained to detect that object concept. Although the list of known object classes often runs into the tens of thousands, object selection system 106 sometimes fails to recognize an object. Nevertheless, object selection system 106 can extend its object detection capabilities by detecting objects of unknown categories. In this way, object selection system 106 can add additional layers to object selection pipeline 400 to facilitate the detection of unknown objects.
[0093] For illustrative purposes, if object selection system 106 determines during operation 420 that the object is not part of an object category, it can detect the object using a generic object-detecting network, as shown in operation 424 of the object selection pipeline. In one or more implementations, the generic object-detecting model involves the use of a concept-masking model and / or an auto-labeling model.
[0094] As mentioned above, the object selection pipeline 400 can include different object-detecting models. For example, in one or more implementations, the object selection pipeline 400 includes a panoptic segmenting detecting model. In one or more implementations, a panoptic segmenting detecting model detects instances of objects using a segmentation-based approach. The object selection system 106 can employ additional (or fewer) object-detecting models that can be easily added to (or removed from) the object selection pipeline 400.
[0095] As in Fig. As shown in Figure 4, the object selection pipeline 400 can include an operation 428 in which the object selection system 106 determines whether a relationship exists between multiple objects. For illustrative purposes, if the object string contains only a single object (i.e., the query object), the object selection system 106 can skip to operation 432 of providing the selected query object within the image. Otherwise, if the query string contains multiple objects for which the object selection system 106 has generated multiple object masks, the object selection system 106 can identify one query object among several detected objects.
[0096] In particular, the object selection system 106 can determine the relationship type between the detected objects in various implementations. For example, the object selection system 106 analyzes the component graph (generated, for instance, during operation 404) to determine the relationship type. Based on this relationship type, the object selection system 106 can isolate the query object from the multiple detected objects.
[0097] For illustrative purposes, object selection pipeline 400 includes an operation 430 in which object selection system 106 uses an object relationship model to detect the query object. Specifically, object selection system 106 selects an object relationship model that has been partially trained on objects exhibiting the given relationship type specified in the component graph. Next, object selection system 106 can isolate the query object from among the multiple objects based on identifying the object that satisfies the object relationship model. Additional details related to identifying a query object based on a suitable object relationship model are provided below in conjunction with Fig. 8 to 12D are specified.
[0098] As shown in operation 432 of object selection pipeline 400, once object selection system 106 isolates an object mask of the request object, object selection system 106 provides the request object within the image. Object selection system 106 can provide the selected object (or the selected instance of the object) to a client device assigned to a user, for example. Object selection system 106 can, for example, automatically select the object within the image for the user within the aforementioned image editing application.
[0099] Fig. Figure 4 and the identified corresponding figures describe different implementations of selecting objects in an image. Accordingly, the processes and algorithms associated with Fig. 4 and in the subsequently identified figures (for example) Fig. 5 to 12D) describe an exemplary structure and architecture for performing a step of object detection using natural language-based processing tools and an object-detecting neural network selected from several object-detecting neural networks. The [document / section] in conjunction with Fig. 4, Fig. 5 and Fig. The flowcharts described in section 8 provide a structure for one or more algorithms according to the object selection system 106.
[0100] As described above, the object selection pipeline 400 comprises various components that the object selection system 106 uses to detect a query object. Many of these components are interchangeable with updated versions as well as with entirely new components. Accordingly, if errors occur, the object selection system 106 can identify the source of the error and perform an update. Additionally, the object selection system 106 can add further components to the object selection pipeline to improve its performance in detecting objects in images. Further details regarding the modification and updating of the object selection pipeline 400 with interchangeable modules can be found in U.S. Patent Application No. 16 / 518,880 (Attorney File No. 20030.257), filed on July 19, 2019.1) “Utilizing Object Attribute Detection Models To Automatically Select Instances Of Detected Objects In Images”, which is hereby included in its entirety by reference.
[0101] The next character is Fig. Figure 5 shows a flowchart with a sequence of operations (500) for identifying and selecting a query object based on an alternative object concept according to one or more implementations. As mentioned above, this represents Fig. 5 additional details related to operations 408 to 410 of object selection pipeline 400, which are described above in conjunction with Fig. As described in section 4, the sequence of processes is ready. 500 of Fig. Figure 5 shows how the object selection system 106 can modify the object selection pipeline based on a mapping table.
[0102] For simplicity, the sequence of operations is described using a query string containing a single object to be selected (i.e., the query object). As shown above in connection with the object selection pipeline, the query string can contain multiple objects, which the object selection system 106 detects using one or more object-detecting models, and from which the object selection system 106 can isolate a query object using an object-relationship model.
[0103] As shown, the sequence of operations 500 includes operation 502, in which the object selection system 106 analyzes a query string to identify an object term that represents a query object. As mentioned above, the object selection system 106 can identify an object in the query string in several ways. For example, the object selection system 106 parses the query string to determine nouns within it. In other implementations, the object selection system 106 generates a component graph that specifies the objects within the query string.
[0104] Additionally, the sequence of operations 500 includes an operation 504 in which the object selection system 106 determines that the object term of the query object does not correspond to any known object class. As mentioned above, the object selection system 106 can, for example, search a list or a database of known objects and / or known object classes for the object term. If the object selection system 106 does not recognize the object term, it can determine that the object term corresponds to an unknown object class.
[0105] The sequence of operations 500 includes an operation 506 in which the object selection system 106 uses a mapping table to identify one or more alternative object terms. For example, the object selection system 106 accesses one or more mapping tables to identify synonyms, hypernyms, hyponyms, or other alternative object terms for the unrecognized object term. As described above, synonyms include terms that are interchangeable with the unknown object term. Hypernyms include terms that are broader than an unknown object term. Hyponyms include terms that are narrower than an unknown object term (for example, subcategories of the unknown object term). Examples of a synonym mapping table and a hypernym mapping table are given below in conjunction with Fig. 6A and Fig. 6B is indicated.
[0106] In many implementations, the alternative object terms for an unknown object term correspond to known object terms. For example, in one or more implementations, the mapping table contains only alternative object terms that are known. In other implementations, the mapping table may contain one or more alternative object terms that are themselves unknown object terms for an unknown object term. In these latter implementations, the object selection system can choose not to select the alternative object term and search for further or different alternative object terms.
[0107] As shown, the sequence of operations 500 includes an operation 508 in which the object selection system 106 determines whether the alternative object term is a synonym or a hypernym. If the alternative object term for the query object is a synonym, the object selection system 106 can use the object selection pipeline as described above in conjunction with Fig. The object selection system 106, for example, replaces the unknown object concept with the known alternative object concept, then identifies the object-detecting model in the object selection pipeline that detects the alternative object concept most accurately and efficiently, and uses it.
[0108] For illustrative purposes, it shows Fig. 5, that operation 508 can proceed to multiple operations if the alternative object term is a synonym or a highly reliable match. If, as shown in operation 510, the object selection system 106 determines that the synonymous alternative object term is a known object, then the object selection system 106 detects the alternative object term using the network relating to a known object class, which was described above in connection with operation 418. Fig. 4 has been described. If, as shown in process 512, the object selection system 106 determines that the synonymous alternative object concept is a specialized object, then the object selection system 106 detects the alternative object concept using the network relating to a specialized object, which was described above in connection with process 414. Fig. 4 has been described.
[0109] Additionally, as shown in process 514, if the object selection system 106 determines that the synonymous alternative object concept is a known object concept or category, the object selection system 106 detects the alternative object concept using the concept-detecting network described above in connection with process 422. Fig. 4 has been described. As shown in process 516, if the object selection system 106 does not recognize the synonymous alternative object concept, the object selection system 106 can detect the alternative object concept using the general object-detecting model described above in connection with process 424. Fig. 4 has been described.
[0110] As described above, each of the selected object-detecting networks can detect the alternative object concept. In additional implementations, the selected object-detecting network can generate one or more object masks for one or more instances of the detected object. In alternative implementations, the object selection system 106 uses an object-masking neural network to generate one or more object masks of the query object.
[0111] Again, in operation 508 of the sequence of operations 500, if the alternative object term is a hypernym or a less reliable match, the object selection system 106 can modify the object selection pipeline. For illustrative purposes, if the alternative object term is a hypernym, the object selection system 106 can restrict the available operations from the sequence of operations to operation 510, which detects the alternative object term using a network relating to a known object class. In some implementations, the object selection system 106 can exclude the option of proceeding to operations 512 through 516 if the alternative object term is a hypernym. In other implementations, the object selection system 106 can provide the hypernym for one of operations 512 through 516.
[0112] This means, in particular, that operation 510 can handle both highly reliable matches (i.e., synonyms) and less reliable matches (i.e., hypernyms). Matches with a highly reliable alternative object term do not require additional object verification by object selection system 106 to ensure that the query object is the object detected by the network pertaining to a known object class. However, if the alternative object term is a hypernym, object selection system 106 may perform one or more additional operations to verify that the detected objects correctly include the query object specified in the query string.
[0113] Is the alternative object concept, for example, a hypernym, as in Fig. As represented by the dashed arrows in Figure 5, the sequence of operations 500 can include an operation 518 in which the object selection system 106 uses labeling to filter out detected instances that do not match the query object. In one or more implementations, for example, the object selection system 106 uses an auto-labeling model to label each of the detected objects, after which the object selection system 106 deselects instances of the object that do not match the query object. In alternative implementations, the object selection system 106 uses a different labeling approach to label and deselect detected instances of the query object that are different from the query object.
[0114] For illustrative purposes, if the query string contains the words "the badminton player," the object selection system 106 can detect the hypernymous alternative object terms "person," "man," and "woman," which are hypernyms to the object term "badminton player." In response, the known object class can identify multiple people in an image. However, since "person" is a much broader term than the query object term "badminton player," the known object class can identify more people in an image than just the badminton player. Accordingly, the object selection system 106 can employ an auto-labeling model to verify the masks, and / or bounding boxes for each detected person in the image correctly contain the badminton player or another type of person.
[0115] As mentioned above, Object Selection System 106 largely does not employ other object detection models (that is, operations 512 through 516) to detect a hypernymous alternative object concept for a query object. In detail, this means that because hypernyms have a broad object scope, most specialized object-detecting models are incompatible, as they correspond to specific objects and object classes. Similarly, most concept-detecting networks are trained to detect specific object categories. With regard to a general-purpose object-detecting network, Object Selection System 106, because it employs an auto-labeling model as part of filtering out non-query object selections, is unable to utilize the auto-labeling model as an additional verification / filtering step.
[0116] In one or more implementations, the object selection system 106 identifies a hyponym or subgroup for a query object with an unknown object class. In many implementations, the object selection system 106 treats hyponyms in the same way it treats synonymous alternative object terms. However, because a hyponym can include subcategories of a query object, the object selection system 106 can perform multiple iterations of the sequence of operations 500 to detect each instance of the query object. For example, if the query object is "Furniture," hyponyms can include the subcategory objects "Table," "Chair," "Lamp," and "Desk." Accordingly, the object selection system 106 can repeat operations 510 through 516 for each subcategory instance of the query object.
[0117] As in Fig. As shown in Figure 5, the sequence of operations can include an operation 520 in which the object selection system 106 provides the selected object mask within the image. As described above, the object selection system 106 can provide an image with the requested object by automatic selection in response to the query string. Therefore, even if the requested object in the query string initially did not correspond to any known object class, the object selection system 106 can use a mapping table and one or more alternative object terms to detect it more accurately and efficiently.By transferring the object detection of unknown objects in a query string to target object-detecting models, such as the specialized object-detecting model and the known object class-detecting model, instead of using the general object-detecting model as the default, the object selection system 106 can drastically improve the performance of the object selection system 106 in terms of accuracy and efficiency.
[0118] Although Fig. Regarding the use of a mapping table to determine alternative object terms for nouns in a query string, the object selection system 106 can also employ one or more mapping tables at different stages of the object selection pipeline. For example, if the object selection system 106 determines that a term (e.g., an attribute) in the query string is not congruent with one or more models in the object selection pipeline, the object selection system 106 can use the mapping table to identify an alternative object term (e.g., synonyms) that the object selection system 106 recognizes. In another example, the object selection system 106 uses a mapping table to determine alternative terms for labels that are generated for a potential object.In this way, the object selection system 106 can use intelligent matching to verify that the terms in the user's request match object classes and models used by the object selection system 106.
[0119] Fig. 6A and Fig. Section 6B shows examples of tables of figures. In particular, they show Fig. 6A and Fig. 6B several mapping tables used to identify alternative mapping terms of a query object, according to one or more implementations. Fig. Figure 6A, for example, shows a synonym mapping table 610, which includes a first column with object terms (i.e., nouns) and a second column with corresponding synonyms. Similarly, Figure 6A shows... Fig. 6B a hypernym mapping table 620, which includes a first column with object terms (that is, nouns) and a second column with corresponding hypernyms.
[0120] In the synonym mapping table 610, the object selection system 106 can look up an object term for an unrecognized query object (for example, a query object with an object term that does not correspond to any known object class) and identify one or more synonyms that correspond to the object term. For illustrative purposes, the object selection system 106 can identify the synonymous alternative object term "woman" for the object term "lady".
[0121] As mentioned above, in various implementations, each of the synonyms listed in synonym mapping table 610 can correspond to a known object class. In one or more implementations, for example, each object term in synonym mapping table 610 is mapped to at least one synonym that corresponds to a known object class. In some implementations, synonym mapping table 610 contains only alternative object terms that correspond to known object classes. In alternative implementations, synonym mapping table 610 may contain additional alternative object terms, such as object terms that do not correspond to any known object class.
[0122] In some implementations, the synonyms are ordered based on detectability. In one example, the first synonym listed corresponds to an object detectable by a specialized object-detecting network, while the second synonym listed corresponds to an object detectable by a network pertaining to a known object class. In another example, the first synonym listed corresponds to an object detectable by a network pertaining to a known object class, while the next synonyms listed correspond to less efficient and effective object-detecting models (for example, a general object-detecting model).
[0123] As mentioned above, shows Fig. 6B the hypernym mapping table 620, which is similar to the synonym mapping table 610 of Fig. 6A is shown. As demonstrated, the hypernym mapping table 620 contains hypernyms that are mapped to the object concepts. For the sake of illustration, the hypernym mapping table 620 maps the object concept "dog" to the hypernyms "pet", "animal", and "mammal".
[0124] In various implementations, the object selection system 106 generates a combined mapping table that includes an object concept mapping for both synonyms and hypernyms. In some implementations, the mapping table can also include hyponyms. Furthermore, in one or more implementations, hyponyms can be determined by looking up object concepts in the second hypernym column backwards and identifying hyponyms from the first object concept column. For example, for the object concept "human," the object selection system 106 can identify the hyponyms "lady" and "buddy," among other alternative object concepts.
[0125] Additionally, the object selection system 106 can modify the figure table or tables. For example, if new colloquial terms or more granular terms are identified, the object selection system 106 can add entries for these terms to the figure table. In some implementations, the object selection system 106 retrieves new terms from an external source, such as an online dictionary database or a crowdsourced website.
[0126] Fig. Figures 7A to 7D show a graphical user interface (720) for selecting a query object based on one or more alternative terms, according to one or more implementations. For the sake of simplicity, they include Fig. 7A to 7D the client device 300 introduced above. The client device 300 includes, for example, an image processing application and the object selection system 106 within a graphical user interface 702. As in Fig. As shown in Figure 7A, the graphical user interface 702 includes an image 704 of a room with various pieces of furniture. Additionally, the graphical user interface 702 includes an object selection interface 706, as described above in conjunction with Fig. 3A describes where the user provides the request string "chair".
[0127] Upon detecting that the user is requesting automatic selection of the chair from image 704 based on the query string, the object selection system 106 can use natural language processing tools and the object selection pipeline 400 to determine how best to fulfill the request. For example, the object selection system 106 can determine that the query object in the query string is "chair". Furthermore, the object selection system 106 can determine that the object term "chair" does not correspond to any known object class.
[0128] Based on the determination that the query object does not correspond to any known object class, the object selection system 106 can access a mapping table to identify an alternative object term for the query object. In this example, the object selection system 106 is unable to identify a synonym for the query object. However, the object selection system 106 identifies the hypernym "furniture".
[0129] Using the hypernymous alternative object term (that is, "furniture"), the object selection system 106 can employ the network relating to a known object class to identify each instance of furniture in the image. For illustrative purposes, it shows Fig. 7B, that the object selection system 106 generates an object mask for the lamp 708a, the sofa 708b and the chair 708c, since each of these objects falls under the hypernym "furniture".
[0130] As mentioned above, the object selection system 106 can modify the object selection pipeline described above to add a verification step when a hypernym is used as an alternative object term. For illustrative purposes, the following is shown. Fig. 7C shows that the object selection system 106 uses an auto-labeling model to generate labels for each of the detected instances of the query object. In particular, Figure 704 in Fig. 7C a lamp label 710a, a sofa label 710b and a chair label 710c in conjunction with the corresponding detected furniture.
[0131] In one or more implementations, the object selection system 106 uses labels to select the query object. In some implementations, the object selection system 106 compares the unrecognized query object term with each of the labels to determine which label has the strongest correspondence to the query object (for example, a correspondence score above a correspondence threshold). In other implementations, the object selection system 106 does not find a strong correspondence between the unrecognized query object term (i.e., "chair") and the labels. However, for another detected object in the image (i.e., the lamp), the object selection system 106 determines a weak correspondence score for the label "chair" and a much stronger correspondence score for the label "lamp".Accordingly, the object selection system 106 can cancel the selection of the lamp within the image based on the fact that the correspondence value for the label "lamp" is stronger than the correspondence value for the label "chair".
[0132] In many implementations, the object selection system 106 can generate multiple labels for each of the detected instances of the query object. Furthermore, the object selection system 106 can assign a reliability score to each generated label. In this way, the object selection system 106 can use the reliability score of a label as a weighting factor when determining a correspondence score between the label and the detected query object instance, or when comparing labels to determine which object an unrecognized object corresponds to.
[0133] As in Fig. As shown in Figure 7D, the object selection system 106 can select chair 708c as the query object and display it within image 704 in response to the query string. In some implementations, the object selection system 106 can display multiple selected objects within image 704 if, due to a lack of reliability, it is unable to deselect detected instances. For example, if the object selection system 106 is unable to identify and label the detected instances of "furniture," it will display all three selected objects within image 704 to the user. This allows the image processing application to enable a user to manually deselect objects that are not the query object, which is a simpler task compared to manually selecting an object.
[0134] Fig. Figure 8 shows a flowchart of a sequence of operations 800 for deploying an object-relationship model to detect a query object among multiple objects associated with a natural language-based user input, according to one or more implementations. The sequence of operations 800 may involve the object selection system 106 deploying a computation graph that provides the structure of a natural language-based input (for example, a query string) to implement handling of the selection requirements that specify an object or domain (for example, an object class) using relationships between objects or object parts. As mentioned above, Figure 8 shows Fig. 8 additional details regarding process 404 and process 430 of object selection pipeline 400, which were mentioned above in conjunction with Fig. 4 has been described.
[0135] As shown, the sequence of operations 800 includes an operation 802 in which the object selection system 106 receives a query string specifying a query object to be selected within an image. In one or more implementations, the query string contains multiple objects, such as the query string "woman in dress". In these implementations, the detection and selection of the query object is performed based on another object included in the query string. Furthermore, the first noun (that is, "woman") is generally the query object, while the second noun (that is, "dress") is a secondary object used to identify the query object among multiple instances of the first noun in an image.
[0136] Additionally, the sequence of operations 800 includes an operation 804, in which the object selection system 106 generates a component graph from the query string. As described above, a component graph identifies each of the objects identified in the query string, as well as the relationship type between objects if multiple objects exist. In one or more implementations, the component graph is structured as a tree. In alternative implementations, the component graph is structured in the format of a table, a list, or a database. Fig. 9A and Fig. 9B shows a visual example of a component graph based on a query string and also provides further details related to generating a component graph.
[0137] As shown, the sequence of operations 800 includes an operation 806 in which the object selection system 106 identifies several object classes and a relationship from the component graph. For example, the object selection system 106 analyzes the component graph to analyze each of the objects included in the query string. In many implementations, one or more of the objects have an object class that corresponds to a class of known objects. For the sake of illustration, the object selection system 106 generates a component graph for the query string "Woman in Dress" that identifies the objects "Woman" and "Dress". In some implementations, at least one of the objects in the query string has an object class that corresponds to an object category (for example, "Sky", "Water", or "Background").
[0138] As mentioned above, the object selection system 106, in various implementations, can identify the relationship between objects from the component graph. Based on an analysis of the component graph, the object selection system 106 identifies, for example, the relationship between two or more objects. For the sake of illustration, the object selection system 106 generates a component graph for the query string "woman in dress" that identifies the relationship term "in" within the query string.
[0139] In one or more implementations, the relationship between two objects is indicated by a preposition, such as "in," "with," "of," "on," "over," and "at / near." In some implementations, the relationship between two objects is indicated by a verb (for example, an action), such as "hold," "carry," "catch," "jump," "drive," and "touch." In additional implementations, the relationship between two objects is indicated by a phrase, such as "above," "in the background," and "next to."
[0140] In one or more implementations, the object selection system 106 uses a lookup table, a list, or a database to map relationship terms in the component graph to relationship types. For example, the object selection system 106 can determine that the terms "on" and "carry" correspond to a first relationship type, the terms "behind" and "on" to a second relationship type, and the terms "in the background" and "in the foreground" to a third relationship type. In alternative implementations, the object selection system 106 determines relationship types for the component graph using a machine learning model that analyzes the query string and the component graph and predicts the relationship type.
[0141] As shown, the sequence of operations 800 includes an operation 808 in which the object selection system 106 detects an initial set of objects for the first identified object class. The image, for example, can contain one or more instances of the first identified object class. For the sake of illustration, the object selection system 106 detects each instance of a woman found in the image for the first object "woman" in the query string "woman in dress". As mentioned above, when a query string contains multiple objects, the first identified object in the query string often includes the query object, with the user requesting that the object selection system 106 automatically select a specific instance of the first object as the query object based on its relationship to another object (for example, a secondary object included in the query string).If the image contains only one instance of the request object, the request string must only contain the request object without reference to any additional related objects.
[0142] In various implementations, the object selection system 106 uses the object selection pipeline, as described above, to select the optimal object-detecting model for the first object class to detect each instance of that first object class. For example, if the first object class corresponds to a specialized object class, the object selection system 106 detects each instance of the first object class using a specialized object-detecting network focused on detecting that object class. If the first object class corresponds to a known object class, the object selection system 106 similarly detects each instance of the first object class using the network pertaining to the known object class.
[0143] In one or more implementations, the object selection system 106 generates an object mask for each detected instance of the first object class. In alternative implementations, the object selection system 106 generates a bounding box for each detected instance of the first object class without generating object masks. In these implementations, the object selection system 106 can generate the object mask after isolating one or more specific instances of the query object. In this way, the object selection system 106 can reduce computational load by not generating object masks for instances of an object that were not requested for selection by the user.
[0144] As shown, the sequence of operations 800 includes an operation 810 in which the object selection system 106 detects a second set of objects for the second identified object class. As with the first object class, the object selection system 106 can use the object selection pipeline to detect one or more instances of the second object class. For the second object class "dress" in the query string "woman in dress," the object selection system 106 generates, for illustrative purposes, an object mask for each dress detected in the image. Again, as described above, in some implementations the object selection system 106 can generate a bounding box and / or an object mask for each detected instance of the second object class (i.e., dress).
[0145] As shown, the sequence of operations 800 includes an operation 812 in which the object selection system 106 selects an object relationship model based on the identified relationship type. In one or more implementations, the object selection system 106 selects an object relationship model from a group of object relationship models based on the identified relationship type. In some implementations, the object selection system 106 uses a canonical list of relationship operators (that is, object relationship models) that map relationship-based words and phrases (for example, relationship types) in a query string to a specific object relationship model.Based on the fact that each relationship type specifies a paired relationship between objects, the object selection system 106 identifies a corresponding object relationship model (for example, a paired relationship operator) that is tailored to the paired relationship (for example, the relationship type). Examples and additional disclosure regarding object relationship models are given below in conjunction with . Fig. 10 is given.
[0146] Additionally, the sequence of operations 800 includes an operation 814 in which the object selection system 106 uses the selected object relationship model to determine an overlap between the first set of object classes and the second set of object classes. For example, the object selection system 106 determines an overlap between the first set of detected objects and the second set of detected objects that satisfy the relationship boundary conditions of the selected object relationship model. The object selection system 106 can determine whether one or more detected instances of the first object class satisfy a relationship threshold with one or more detected instances of the second object class as specified by the selected object relationship model.
[0147] In various implementations, the object selection system 106 tests each instance of the first detected set of objects against each instance of the second detected set of objects to determine whether pairs satisfy the selected object relationship model. For the sake of illustration, the object selection system 106, for the query string "woman in dress" and an object relationship model that determines whether a first object is inside a second object, can test each detected instance of a woman in the image against each detected instance of a dress to determine a pair that satisfies the object relationship model.
[0148] In one or more implementations, the object-relationship model uses heuristics to determine whether a pair of detected objects satisfies the model. For example, the object-relationship model is defined by a set of rules that specify when a first detected object overlaps with a second detected object. The rules for each object-relationship model can vary based on the type of relationship applied between the objects. Additional examples of the use of different object-relationship models are given below in conjunction with Fig. 10 is given.
[0149] In alternative implementations, the object relationship model employs training and / or deployment of a machine learning model to determine an overlap of detected objects that satisfy the object relationship model. For example, the object selection system 106 uses surveillance training to teach an object relationship neural network to predict when a target overlap between two objects will occur, based on the objects, the relationship type, and / or the image.
[0150] As shown, the sequence of operations 800 includes an operation 816 in which the object selection system 106 selects the query object based on the output of the selected object relationship model. The selected object relationship model outputs one or more pairs of objects that satisfy the object relationship model. Based on this output, the object selection system 106 can select the query object. For example, the object selection system 106 selects the first object instance contained in a pair. For the sake of illustration, the selected object relationship model outputs a pair of objects for the query string "woman in dress" in which a woman is represented by a dress. From this output, the object selection system 106 can select the instance "woman" contained in the pair as the query object.
[0151] In one or more implementations, the object selection system 106 can generate an object mask for the selected query object. The object selection system 106 uses the neural network pertaining to an object mask, which was described above in connection with the object selection pipeline. In alternative implementations, the object selection system 106 can select the object mask for the selected query object if the object has been generated beforehand.
[0152] As shown, the sequence of operations 800 includes an operation 818 in which the object selection system 106 provides the selected query object within the image. The object selection system 106 automatically selects the query object in the image in response to the query string (that is, the natural language-based object selection query), as described above.
[0153] As mentioned above, show Fig. 9A and Fig. Section 9B provides a visual example of a component graph based on a query string and offers further details related to generating a component graph. In this context, it states Fig. 9A represents a graphical user interface of a digital image that includes a natural language-based object selection query, according to one or more implementations. As shown, it includes Fig. 9A the client device 300 introduced above, which includes an image processing application and the object selection system 106 within a graphical user interface 902.
[0154] As shown, the graphical user interface 902 includes an image 904 of a man on a bicycle, a woman standing next to a bicycle, and a background landscape. Additionally, the graphical user interface 902 includes an object selection interface 906, which is described above in conjunction with Fig. 3A has been described, and in which the user provides the request string "Person on a bike".
[0155] In Fig. 9B shows a component graph. In particular, it shows Fig. 9B a component graph 910, which was generated based on the query string 912. As described above, the object selection system 106 can determine the structure of the query string based on the component graph 910, including which objects are involved in a relationship and which relationship is being requested. For illustrative purposes, the component graph 910 is structured as a binary tree with leaf nodes corresponding to the words in the query string 912. In alternative implementations, the component graph 910 can be structured in alternative arrangements, as described above.
[0156] As shown, component graph 910 contains several node types. For example, component graph 910 includes localization nodes (for example, a first localization node 916a and a second localization node 916b) that correspond to objects detected in query string 912. As described above, the object selection system 106 can use the object selection pipeline for each localization node to detect all instances of the corresponding object. Additionally, component graph 910 includes a relationship node 918, which corresponds to a relationship between objects in query string 912. As shown, relationship node 918 has two child nodes and describes the predicate of the relationship (for example, "on") and the second localization node 916b (for example, "wheel").Additionally, as described above, the object selection system 106 can select an object relationship model based on the relationship node 918.
[0157] Additionally, component graph 910 includes an intersection node 914. As shown, intersection node 914 contains the two child nodes of the first localization node 916a and the relationship node 918. In various implementations, intersection node 914 provides an indication of the object relationship from relationship node 918. As described above, object selection system 106 can use the selected object relationship model to disambiguate, locate, and select the correct instance of the object associated with the first localization node 916a. For example, object selection system 106 executes a set of heuristic rules corresponding to the object relationship model to locate the instance of the object associated with the first localization node 916a, which object selection system 106 then selects as the query object.
[0158] As described above, the object selection system 106 can generate a component graph. In various implementations, the object selection system 106 generates the component graph, for example, using a selection parsing model (such as a natural language-based decomposition model). In one or more implementations, the object selection system 106 employs, for example, the techniques and solutions found in "Using Syntax to Ground Referring Expressions in Natural Languages" by Cirik et al., published in 2018. Additionally or alternatively, the object selection system 106 can use other methods to generate a structured decomposition of the query string for grounding the language input.
[0159] As mentioned above, includes Fig. 10 examples of using different object relationship models. For illustrative purposes, shown below. Fig. 10. A block diagram with multiple object relationship models corresponding to one or more implementations. In particular, it extends Fig. 10 the process 812 of selecting an object relationship model based on the identified relationship type, as described above Fig. 8 has been described. As in Fig. As shown in Figure 10, the process includes 812 different object relationship models that the object selection system 106 can select, including an object touch model 1002, a relative object position model 1004, a background / foreground object model 1006, and other object relationship models 1008.
[0160] In one or more implementations, the object selection system 106 selects the object contact model 1002 as the selected object relationship model. For example, the object selection system 106 selects the object contact model 1002 if the relationship type (e.g., relationship node) in the component graph corresponds to objects that are close to each other in the image or touch each other there. The object selection system 106 also selects the object contact model 1002 if the relationship node contains an action, such as "hold," "carry," or "touch," or a preposition, such as "in," "with," or "of."
[0161] In various implementations, the object selection system 106 implements the touch operator TOUCHING (Object_1, Object_2) to check for intersection or overlap within the image in a first object (for example, Object_1) and a second object (for example, Object_2). In various implementations, the object touch model 1002, to achieve greater accuracy, checks the object areas within the object masks of the first and second objects for overlap. In one or more implementations, the object touch model 1002 determines, for example, whether the overlap between the two object masks is greater than 0. In alternative implementations, the touch threshold is a distance greater than 0, such as 1, 2, 5, 10, or more pixels.
[0162] In some implementations, even when the overlap threshold is 0, two adjacent object masks may still fail to satisfy the object touch model 1002. For example, the object mask neural network used to generate the object masks might be unable to produce an exact object mask, where the edges of two adjacent object masks are separated by a few pixels, resulting in a non-overlap scenario. In these implementations, the object selection system 106 and / or the object touch model 1002 can perform image dilation for one or more iterations on one or more of the object masks to expand their area before checking for an overlap.
[0163] In some implementations, the object selection system 106 selects the relative object position model 1004 as the object relationship model. In many implementations, the relative object position model 1004 determines a relative position of a first object (that is, an object instance) with respect to a second object. Examples of relative position types include "left," "right," "above," "over," "below," "under," "below," and "below" of. In several implementations, the object selection system 106 selects the relative object position model 1004 based on the identification of a relationship type that corresponds to one of the relative positions listed above.
[0164] In many implementations, the Relative Object Position Model 1004 includes a relative position table that maps a position to a canonical form. For example, for the position "above," the relative position table might include the relative positions "on," "over," "above," "above," and "toward above." Additionally, for the position "right," the relative position table might include the relative positions "to the right," "to the right of," and "to the right."
[0165] In various implementations, the Relative Object Position Model 1004 implements the relative position operator RELATIVE_POSITION(Bounding_Box_1, Bounding_Box_2, Relative_Position) to check whether the first bounding box (for example, Bounding_Box_1) is in a relative position (Relative_Position) relative to the second bounding box (for example, Bounding_Box_2), where Relative_Position is one of the canonical forms of "top", "bottom", "left", or "right". The following is described below. Fig. 11 provides an example of how to determine whether a first object has a relative position to a second object.
[0166] In various implementations, the object selection system 106 selects the background / foreground object model 1006 as the selected object relationship model. For example, if the object selection system 106 detects a relationship type corresponding to the background or foreground of an image or different depths between objects, it can select the background / foreground object model 1006. The object selection system 106 detects the "background" or "foreground" relationship type from the component graph.
[0167] In various implementations, the background / foreground object model 1006 implements a foreground operator and / or a background operator. For example, in one or more implementations, the background / foreground object model 1006 implements the IN_BACKGROUND(Object_1) operation to determine that the first object (for example, the object mask of the first object) is in the background of the image. In these implementations, the background / foreground object model 1006 can employ a specialized model from the object selection pipeline to detect the foreground and / or background in the image. Specifically, the background / foreground object model 1006 can receive a background mask from the specialized model and / or an object mask-related neural network.
[0168] Furthermore, the background / foreground object model 1006 can compare the object mask belonging to the first image with the background mask to determine whether there is an overlap that satisfies the background / foreground object model 1006. For example, the background / foreground object model 1006 can determine whether the majority of the first object lies within the background. For illustrative purposes, the background / foreground object model 1006 can apply the following heuristic in one or more implementations: intersection(first object mask, background mask) / area(first object mask)≥0.5
[0169] This ensures that most of the first object mask overlaps with the background mask.
[0170] Similarly, the background / foreground object model 1006 can implement the IN_FOREGROUND(Object_1) operation. The background / foreground object model 1006 can also employ other operations, such as the IN_FRONT(Object_1, Object_2) operation or the BEHIND(Object_1, Object_2) operation, which uses a depth mapping to determine the relative depths of the first and second objects.
[0171] In some implementations, the object selection system 106 selects one or more other object relationship models 1008 as the selected object relationship model. The object selection system 106 selects an object relationship model that determines whether an object has a specific color (for example, HAS_COLOR(Object_1, Color)). In another example, the object selection system 106 selects an object relationship model that compares the size between objects (for example, IS_LARGER(Object_1, Object_2)). Furthermore, the object selection system 106 can use an object relationship model that considers other object attributes, such as size, length, shape, position, location, pattern, composition, expression, feel, strength, and / or flexibility.
[0172] As mentioned above, shows Fig. 11. Determining whether a first object has a relative position to a second object (for example, based on the relationship type "on"). In particular, shows Fig. 11. The use of the relative object position model 1004 to select an object, according to one or more implementations. For simplicity, this includes Fig. 11 the bounding box of a first object 1102 and the bounding box of a second object 1104.
[0173] Additionally includes Fig. 11. In addition, there is a first height 1106 corresponding to the height of the bounding box of a first object 1102 and a second height 1108 that is proportional to the first height 1106. As shown, the second height 1108 is one-quarter of the height of the first height 1106. In alternative implementations, the second height 1108 is larger or smaller (for example, half or one-eighth of the height of the first height 1106). The size and shape of the relative position threshold zone 1110 can vary based on the relationship type and / or the object relationship model used. The relationship type of the above, for example, can extend from above the first object 1102 to the top edge of the image.
[0174] As also shown, the second height 1108 can be used to define a relative position threshold zone 1110 based on the second height 1108 and the width of the bounding box of a first object 1102. If the bounding box of the second object 1104 has one or more pixels within the relative position threshold zone 1110, the relative object position model 1004 can determine that the relative position is satisfied. In some implementations, at least a minimal fraction of the bounding box zone of the second object 1104 (for example, 15% or more) must be within the relative position threshold zone 1110 before the relative object position model 1004 is satisfied.
[0175] In one or more implementations, the relative object position model 1004 can generate a separate relative position threshold zone corresponding to each canonical position. For example, for the canonical position "right," the relative object position model 1004 can generate a right relative position zone threshold using the same or different proportions for the relative position threshold zone 1110.
[0176] In alternative implementations, the relative object position model 1004 can use the same separate relative position threshold zone and rotate the image before determining whether the relative position is satisfied. For example, for the relative object position operator RELATIVE_POSITION(Object_1, Object_2, Right), the relative object position model 1004 can rotate the image 90° clockwise and the relative position threshold zone 1110, which is defined in Fig. As shown in Figure 11, apply. In this way, the relative object position model 1004 simplifies the relative object position operator to check any relative position based on a single canonical position (for example, above) and the relative position threshold zone 1110.
[0177] Fig. Figures 12A to 12D show a graphical user interface 1202 for implementing an object relationship model for selecting a query object from among several objects associated with a natural language-based user input, according to one or more implementations. For the sake of simplicity, they include Fig. 12A to 12D the client device 300 introduced above. The client device 300 includes, for example, an image processing application that implements the image processing system 104 and the object selection system 106. As in Fig. As shown in Figure 12A, the graphical user interface 1202 includes an image 1204 of a man on a bicycle, a woman standing next to a bicycle, and a background landscape. Additionally, the graphical user interface 1202 includes an object selection interface 1206, as described above in conjunction with Fig. 3A has been described and in which the user provides the request string "Person on a bike".
[0178] To automatically select the query object in the query string, as described above, the object selection system 106 can generate a component graph to identify the objects in the query string as well as a relationship type. For example, the object selection system 106 generates a component graph for the query string, such as the one shown in Fig. The component graph shown in Figure 9B, which has been described above, can identify the two object classes (for example, "Person" and "Wheel") as well as the relationship type "on".
[0179] As described above, the object selection system 106 can use the object selection pipeline to select the object detection model best suited to detecting people (for example, a specialized object-detecting neural network), to detect each instance of a person in Figure 1204 using the selected object-detecting model, and in some cases, to generate object masks for each of the detected instances. For illustrative purposes, Fig. 12B, that the object selection system 106 detects a first person instance 1208a and a second person instance 1208b.
[0180] Additionally, the object selection system 106 can use the object selection pipeline to select an object detection model that best detects bicycles (for example, a neural network relating to a known object class), to detect one or more instances of the wheel in Figure 1204 using the selected object detection model, and in some cases, to generate object masks for each of the detected wheel instances. For illustrative purposes, Fig. 12C, that the object selection system 106 detects a first radix instance 1210a and a second radix instance 1210b.
[0181] In one or more implementations, the object selection system 106 initially fails to identify an object-detecting neural network corresponding to the object term "wheel" (for example, the term "wheel" does not correspond to any known object class). As described above, the object selection system 106 can use a mapping table to identify a synonym, such as "bicycle," which the object selection system 106 recognizes as a known object class. Using the alternative object term "bicycle," the object selection system 106 can then use the neural network relating to a known object class to detect each instance of the wheel.
[0182] As described above, the object selection system 106 can select an object relationship model based on the relationship type to determine an overlap between instances of people and instances of bicycles. For example, in one or more implementations, the object selection system 106 selects the object contact model based on the relationship type "on" to check for overlaps between the two sets of objects. Specifically, the object contact model performs the following tests: TOUCHING (first person instance 1208a, first radius instance 1210a), TOUCHING (first person instance 1208a, second radius instance 1210b), TOUCHING (second person instance 1208b, first radius instance 1210a) and TOUCHING (second person instance 1208b, second radius instance 1210b).
[0183] When performing each of the aforementioned tests, the object selection system 106 can determine that the intersection between the first person instance 1208a and the first wheel instance 1210a satisfies the object contact model, while the other tests fail to satisfy the object contact model. Accordingly, the object contact model can generate an output indicating that the first person instance 1208a is "on" a wheel.
[0184] Based on the output of the object contact model (for example, the selected object relationship model), the object selection system 106 can automatically select the first person instance 1208a. For illustrative purposes, it shows Fig. 12D is the first person instance 1208a selected within image 1204. The object selection system 106 can automatically select the first person instance 1208a in image 1204 in response to the query string.
[0185] Note that in many implementations, the object selection system 106 does not display any intermediate actions to the user. Rather, the object selection system 106 appears to automatically detect and precisely select the query object in response to the user's query string request. In other words, the graphical user interface 1202 jumps from Fig. 12A to Fig. 12D. In alternative implementations, the object selection system 106 displays one or more of the intermediate actions to the user. For example, the object selection system 106 displays the graphical user interfaces and each selected object instance (and / or bounding boxes of each detected object instance), as in Fig. 12B and Fig. 12C is shown.
[0186] Fig. Figure 13 shows an evaluation table 1310 for evaluating different implementations of the object selection system according to one or more implementations. In this context, evaluators tested various implementations of the object selection system 106 to test whether the object selection system 106 offers improvements over a basic selection model.
[0187] For the evaluations, the evaluators used the quality measure "Intersection over Union" (IoU) of an output mask compared to a ground-truth mask for a natural language-based object selection query (i.e., a query string). Specifically, the evaluators processed a test dataset of approximately 1000 images and 2000 query strings.
[0188] In the implementations of the object selection system 106 described here, the evaluators found considerable improvements compared to basic models. As shown in evaluation table 1310, the mean loU value, for example, ranges from 0.3036 for the basic selection model (see, for example, the first row of evaluation table 1310) to 0.4972 (see, for example, the last row of evaluation table 1310), which is due to added improvements to the natural language-based processing tools described here (for example, the use of a component graph and a mapping table). Evaluation table 1310 confirms that the object selection system 106 empirically improves the accuracy of object selection models.
[0189] Fig. Section 14 provides additional details regarding the capacity and components of the object selection system 106 according to one or more implementations. In particular, it shows Fig. 14 A schematic diagram of an exemplary architecture of the object selection system 106, which can be implemented in the image processing system 104 and hosted on a computing device 1400. The image processing system 104 can correspond to the image processing system 104 described above in conjunction with Fig. 1 has been described.
[0190] As shown, the object selection system 106 resides on a computing device 1400 within an image processing system 104. The computing device 1400 can generally represent various types of client devices. In some implementations, for example, the client is a mobile device, such as a laptop, tablet, mobile phone, smartphone, and the like. In other implementations, the computing device 1400 is a non-mobile device, such as a desktop computer, a server, or another type of client device. Additional details relating to the computing device 1400 are given below and also with reference to Fig. 17 described.
[0191] As in Fig. As shown in Figure 14, the object selection system 106 comprises various components for performing the processes and features described here. For example, the object selection system 106 includes a digital image manager 1410, a user input detector 1412, an object concept mapping manager 1414, an object relationship model manager 1416, an object detection model manager 1418, an object mask generator 1420, and a memory manager 1422. As shown, the memory manager 1422 includes digital images 1424, object concept mapping tables 1426, component graphs 1428, object relationship models 1430, object detection models 1432, and an object mask model 1434. Each of the components listed above is described below.
[0192] As mentioned above, the object selection system 106 includes the digital image manager 1410. In general, the digital image manager 1410 facilitates the identification, access, receipt, acquisition, generation, import, export, copying, modification, removal, and organization of images. In one or more implementations, the digital image manager 1410 operates in conjunction with an image processing system 104 (for example, an image processing application) to access and process images as described above. In some implementations, the digital image manager 1410 communicates with the storage manager 1422 to store and retrieve digital images 1424, for example, within a digital image database managed by the storage manager 1422.
[0193] As shown, the object selection system 106 includes the user input detector 1412. In various implementations, the user input detector 1412 can detect, receive, and / or facilitate user input on the computing device 1400 in any suitable manner. In some cases, the user input detector 1412 detects one or more user interactions (for example, a single interaction or a combination of interactions) with respect to a user interface. For example, the user input detector 1412 detects user interaction from a keyboard, a mouse, a touch-sensitive page, a touch-sensitive screen, and / or any other input device in conjunction with the computing device 1400.The User Input Detector 1412 detects, for example, user input of a query string (that is, a natural language-based object selection query) presented by an object selection interface that requests the automatic selection of an object within an image. Additionally, the User Input Detector 1412 detects further user input from a mouse selection and / or touch input to specify an object location within the image, as described above.
[0194] As shown, the object selection system 106 includes the object concept mapping manager 1414. In one or more implementations, the object concept mapping manager 1414 creates, generates, accesses, modifies, updates, removes, and / or otherwise manages object concept mapping tables 1426 (that is, mapping tables). As described in detail above, the object concept mapping manager 1414 can, for example, generate, insert, and update one or more object concept mapping tables 1426 that map object concepts to synonymous object concepts, hypernymic object concepts, and / or hyponymous object concepts (that is, alternative object concepts).In general, the object concept mapping manager 1414 maps object concepts for objects that do not correspond to any known object class to one or more alternative object concepts, as explained above.
[0195] In one or more implementations, the object selection system 106 may include a component graph manager. In various implementations, the component graph manager performs the creation, generation, access, modification, updating, removal, and / or other management of component graphs 1428. In one or more implementations, the component graph manager generates a component graph. In alternative implementations, the component graphs 1428 are generated by an external party or source.
[0196] As shown, the object selection system 106 includes the object relationship model manager 1416. In various implementations, the object relationship model manager 1416 analyzes component graphs 1428 and uses them to identify objects according to query strings and relationship types between their objects, as described above. Additionally, in some implementations, the object relationship model manager 1416 selects an object relationship model from among the object relationship models 1430 based on the relationship type identified in the component graph. Furthermore, the object relationship model manager 1416 can determine an overlap of objects identified in the query string / component graph that satisfy the selected object relationship model.
[0197] As shown, the object selection system 106 includes the object detection model manager 1418. In various implementations, the object detection model manager 1418 performs pre-storage, creation, generation, training, updating, accessing, and / or deployment of the object-detecting neural networks disclosed herein. As described above, the object detection model manager 1418 detects one or more objects within an image (for example, a query object) and generates a boundary (for example, a bounding box) to indicate the detected object.
[0198] Additionally, in a number of implementations, the object detection model manager 1418 can communicate with the memory manager 1422 to store, access, and insert the object detection models 1432. In various implementations, the object detection models 1432 include one or more specialized object-detecting models 1434 (for example, a sky-detecting neural network, a face-detecting neural network, a body or body parts-detecting neural network, a skin-detecting neural network, a clothing-detecting neural network, and a waterfall-detecting neural network), known-class object-detecting neural networks 1436 (for example, for detecting objects with classes learned from training data),Category-based object-detecting neural networks 1438 (for example, for detecting uncountable objects such as soil, water, and sand) and general-purpose object-detecting neural networks 1440 (for example, for detecting objects from unknown object classes), each of which has been described above. Additionally, the object detection model manager 1418 may include one or more neural networks in conjunction with the aforementioned object-detecting neural networks to detect objects within an image, such as auto-labeling neural networks 1442, object-suggesting neural networks 1444, area-suggesting neural networks 1446, and concept-embedding neural networks 1448.The object detection model manager 1418 can employ various object-detecting neural networks within the object selection pipeline to detect objects within a query string, as described above.
[0199] Additionally, as shown, the object selection system 106 includes the object mask generator 1420. In one or more implementations, the object mask generator 1420 creates, builds, and / or generates accurate object masks from detected objects. For example, the object detection model manager 1418 provides a boundary of an object (such as a detected query object) to the object mask generator 1420, which uses one or more object mask models 1432 to generate an object mask of the detected object, as described above. As also explained above, in various implementations, the object mask generator 1420 generates multiple object masks if multiple instances of the query object are detected.
[0200] Each of the components 1410 to 1450 of the object selection system 106 can include software, hardware, or both. For example, components 1410 to 1450 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices, such as a client device (e.g., a mobile client device) or a server device. When executed by the one or more processors, the computer-executable instructions of the object selection system 106 can cause a computing device to perform the feature learning procedures described herein. Alternatively, components 1410 to 1450 can include hardware, such as a special-purpose processing device for performing a particular function or group of functions.Additionally, components 1410 to 1450 of the object selection system 106 can include a combination of computer-executable instructions and hardware.
[0201] Furthermore, components 1410 to 1450 of the object selection system 106 can be implemented as one or more operating systems, as one or more standalone applications, as one or more modules of an application, as one or more plug-ins, as one or more library functions or functions that can be called by other applications, and / or as a cloud computing model. Therefore, components 1410 to 1450 can be implemented as a standalone application, such as a desktop or mobile application. Additionally, components 1410 to 1450 can be implemented as one or more web-based applications hosted on a remote server. Components 1410 to 1450 can also be implemented in a suite of mobile device applications, or "apps." For illustrative purposes, components 1410 to 1450 can be implemented in an application that includes, among other things, Adobe software. ® INDEG® , ADOBE ACROBAT ® , ADOBE ® ILLUSTRATOR ® , ADOBE PHOTOSHOP ® , ADOBE ® CREATIVE CLOUD ® “ADOBE”, “INDESIGN”, “ACROBAT”, “ILLUSTRATOR”, “PHOTOSHOP” and “CREATIVE CLOUD” are either registered trademarks or trademarks of Adobe Inc. in the United States and / or other countries.
[0202] Fig. Sections 1 to 14, the corresponding text, and the examples provide a number of different procedures, systems, devices, and non-temporary computer-readable media of the object selection system 106. Additionally, the processes described here can also be performed repeatedly or in parallel with each other or with other instances of the same or other processes.
[0203] As mentioned, show Fig. 15 and Fig. 16 each a flowchart with a sequence of processes corresponding to one or more implementations. Although Fig. 15 and Fig. While 16 processes each show an implementation, alternative implementations can omit, add, rearrange, and / or modify any of the shown processes. The processes of Fig. 15 and Fig. 16 can be carried out as part of a procedure. Alternatively, a non-temporary computer-readable medium can contain instructions which, when executed by one or more processors, cause a computing device to perform the operations of Fig. 15 and Fig. 16. In some implementations, a system can perform the operations of Fig. 15 and Fig. 16.
[0204] For illustrative purposes, it shows Fig. Figure 15 shows a flowchart with a sequence of operations (1500) for employing object relationship models to detect a query object according to one or more implementations. In various implementations, the sequence of operations (1500) is implemented on one or more computing devices, such as client device (102), server device (110), client device (300), or computing device (1400). Additionally, in some implementations, the sequence of operations (1500) is implemented in a digital environment for creating or editing digital content (e.g., digital images).The sequence of operations 1500, for example, is implemented on a computing device with a memory containing a digital image, a component graph of a query string to specify a first object to be selected based on a second object within a digital image, and one or more object-detecting neural networks.
[0205] The sequence of operations 1500 includes an operation 1510 of generating a component graph of a query string. Operation 1510 can specifically imply generating a component graph of a query string to identify multiple object classes and a relationship type between those multiple object classes. In one or more implementations, operation 1510 involves parsing the component graph of the query string to identify an identifier of the relationship type between the first object and the second object. In various implementations, operation 1520 involves employing a natural language-based decomposition model that generates the component graph from the query string.
[0206] As shown, the sequence of operations 1500 includes an operation 1520 of generating object masks for each object instance identified in the component graph. Operation 1520 can specifically imply generating one or more object masks for each of the multiple object classes using one or more object detection models. In one or more implementations, operation 1520 involves generating one or more object mask instances for each of the first and second objects using one or more object-detecting neural networks. In some implementations, the multiple object classes include a specialized object, a known object, an object category, a concept object, or an unknown object.
[0207] As in Fig. As shown in Figure 15, the sequence of operations 1500 further includes an operation 1530 for identifying an object relationship model from the component graph. Operation 1530 can specifically involve identifying one object relationship model from among several object relationship models based on the relationship type. In one or more implementations, operation 1530 involves identifying an object relationship model based on the relationship type by matching the relationship type with an object relationship model. In different implementations, the multiple object relationship models include an object contact model, a relative object position model, and a background / foreground object model.
[0208] As shown, the sequence of operations 1500 includes an operation 1540 of analyzing the object masks to identify a query object that satisfies the object-relationship model. Operation 1540 may specifically include analyzing the one or more object masks generated for each of the multiple object classes to identify a query object that satisfies the object-relationship model. In one or more implementations, operation 1540 involves determining a query object by identifying an overlap between the one or more object mask instances for the first object and the one or more object mask instances for the second object that satisfies the object-relationship model.
[0209] In some implementations, operation 1540 involves determining the query object based on the object-relationship model using heuristic rules. In alternative implementations, operation 1540 involves determining the query object based on the object-relationship model using machine learning. Operation 1540 may also include, in various implementations, identifying a query object that satisfies the object-relationship model by identifying an overlap between one or more object mask instances for the first object and one or more object mask instances for the second object that satisfies the object-relationship model.
[0210] The sequence of operations 1500 also includes, as shown, an operation 1550 for providing a digital image with the selected query object. Operation 1550 can specifically imply providing a digital image with the query object that was selected in response to receiving the query string. In one or more implementations, operation 1550 includes providing the digital image with the object mask in response to receiving the query string.
[0211] The sequence of operations 1500 can also include a number of additional operations. In additional implementations, the sequence of operations 1500 includes an operation for inserting a mapping table to identify one of several object classes based on a determination of one or more alternative object terms for the object class. In one or more implementations, the sequence of operations 1500 includes additional operations for identifying an object attribute that is assigned to an object of one or more object classes, and for detecting a target instance of the object based on the object attribute and one or more object detection models.
[0212] As mentioned above, shows Fig. Figure 16 shows a flowchart of a sequence of operations (1600) for the application of object relationship models to detect a query object according to one or more implementations. In different implementations, the sequence of operations (1600) is implemented on one or more computing devices, such as client device (102), server device (110), client device (300), or computing device (1400).
[0213] The sequence of operations 1600 includes an operation 1610 of identifying a query string containing a query object. Operation 1610 can specifically imply parsing a query string that specifies a query object to be selected in a digital image. In some implementations, operation 1610 also includes parsing the query string to identify a noun that specifies the query object. In various implementations, operation 1610 further includes parsing the noun to determine an object class type of the query object. In exemplary implementations, operation 1610 involves receiving text input from the user associated with a client device and identifying the text input as a query string (that is, as a natural language-based object selection query).
[0214] As shown, the sequence of operations 1600 also includes an operation 1620 of determining that the query object is not a recognizable object. Operation 1620 can specifically imply determining that the query object does not correspond to any known object class. In one or more implementations, operation 1620 involves not identifying the query object in a list or database of known object classes.
[0215] As in Fig. As shown in Figure 16, the sequence of operations 1600 further includes an operation 1630 for inserting a mapping table to identify an alternative object term for the query object. Operation 1630 may, in particular, include inserting a mapping table to identify one or more alternative object terms for the query object based on the fact that the query object does not correspond to any known object class. In one or more implementations, the one or more alternative object terms include a synonym of the query object that corresponds to a known object class. In some implementations, operation 1630 includes updating the mapping table to modify the one or more alternative object terms of the query object.
[0216] As shown, the sequence of operations 1600 also includes an operation 1640 of selecting a known object-detecting neural network. Operation 1640 can, in particular, include determining, based on at least one of the one or more alternative object concepts of the query object, that a known object-detecting neural network is selected. In one or more implementations, the known object-detecting neural network includes a specialized object-detecting neural network corresponding to one or more alternative object concepts of the query object. In alternative implementations, the known object-detecting neural network includes an object-class-detecting neural network corresponding to one or more alternative object concepts of the query object.
[0217] In various implementations, one or more alternative object terms include a hypernym of the query object that corresponds to a known object class. In additional implementations, operation 1640 may include the use of a labeling model to identify the query object among multiple instances of the query object according to the query object's hypernym.In one or more implementations, process 1640 may also include, for example, inserting the labeling model to label the multiple instances of the query object according to a hypernym of the query object, inserting the mapping table to identify one or more additional alternative object terms for a labeled instance of the query object among the multiple instances of the query object, and filtering out the labeled instance of the query object based on the one or more additional alternative object terms that do not correspond to the query object.
[0218] As shown, the sequence of operations 1600 also includes an operation 1650 for generating an object mask for the query object. Operation 1650 can specifically imply generating an object mask for the query object using a neural network that detects a known object. In various implementations, the neural network responsible for generating an object mask uses a boundary (for example, a bounding box) to identify the detected query object and generate an accurate object mask for it.
[0219] As shown, the sequence of operations 1600 also includes an operation 1660 for providing the digital image with the object mask. Operation 1660 can specifically imply providing the digital image with the object mask for the request object in response to receiving the request string. Operation 1660 can specifically imply providing the image with an object mask of the request object to a client device assigned to a user. In some implementations, operation 1660 includes automatically selecting the detected request object within an image processing application using the object mask of the request object.
[0220] For the purposes of this document, the term "digital environment" generally refers to an environment that is implemented, for example, as a standalone application (such as a PC or mobile application running on a computing device), as an element of an application, as a plug-in for an application, as library functions, as a computing device, and / or as a cloud computing system. A digital media environment enables the object selection system to create, execute, and / or modify the object selection pipeline described herein.
[0221] Implementations of this disclosure may include or employ a special-purpose or general-purpose computer that incorporates computer hardware, such as one or more processors and system memory, as described in more detail below. Implementations within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented, at least in part, as instructions embodied on a non-temporary computer-readable medium and executable by one or more computing devices (for example, any of the media content access devices described herein).In general, a processor (for example, a microprocessor) receives instructions from a non-temporary, computer-readable medium (for example, memory) and executes these instructions, thereby carrying out one or more processes, including one or more of the processes described here.
[0222] Computer-readable media can be any available media accessible to a general-purpose or special-purpose computer system. Computer-readable media that store computer-executable instructions are non-temporary computer-readable storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. By way of example, and not by limitation, implementations of the disclosure may include at least two clearly distinct types of computer-readable media, namely non-temporary computer-readable storage media (devices) and transmission media.
[0223] Non-temporary computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, SSDs (Solid State Drives) (for example, based on RAM), flash memory, phase-change memory (PCM), other types of memory or storage, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code resources in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.
[0224] A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted to or made available to a computer via a network or other communication link (either wired, wireless, or a combination of both), the computer treats the link as a transmission medium. Transmission media may include a network and / or data links that can be used to carry desired program code resources in the form of computer-executable instructions or data structures, and which a general-purpose or special-purpose computer can access. Combinations of the foregoing are to be included within the scope of computer-readable media.
[0225] In the implementation of various computer system components, program code resources in the form of computer-executable instructions or data structures can be automatically transferred from transmission media to non-temporary, computer-readable storage media (devices) (or vice versa). Computer-executable instructions or data structures received over a network or data link can, for example, be buffered in RAM within a network interface module (such as a "NIC") and then, if necessary, transferred to the computer system RAM and / or to less volatile computer storage media (devices) on a computer system. It should therefore be clear that non-temporary, computer-readable storage media (devices) can be included in computer system components that also (or even primarily) use transmission media.
[0226] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a certain function or group of functions. In some implementations, computer-executable instructions are executed by a general-purpose computer to transform the general-purpose computer into a special-purpose computer that implements elements of the disclosure. The computer-executable instructions may be, for example, binaries, instructions in an intermediate format such as assembly language, or even source code. Although the subject matter of the invention has been described in a language specific to structural features and / or methodological processes, it should be clear that the subject matter of the invention defined in the appended claims is not necessarily limited to the features or processes described above.Rather, the described features and processes are revealed as exemplary forms of implementing the claims.
[0227] It is obvious to a person skilled in the art that the disclosure can be practically implemented in network computing environments with many types of computer system configurations, including PCs, desktop computers, laptop computers, information processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure can also be practically implemented in distributed system environments where local and remote computer systems connected via a network (either by wired data links, wireless data links, or a combination of both) perform tasks alike.In a distributed system environment, program modules can reside in both local and remote memory storage devices.
[0228] Implementations of this disclosure may also be implemented in cloud computing environments. For the purposes of this disclosure, the term "cloud computing" refers to a model that enables on-demand network access to a shared pool of configurable computing resources. Cloud computing can be used, for example, in a marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with minimal administrative overhead or interaction from a service provider, and then scaled accordingly.
[0229] A cloud computing model can be composed of various characteristics or properties, such as on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and the like. A cloud computing model can also offer different service models, such as Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). Furthermore, a cloud computing model can be deployed using various deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, and the like. For the purposes of this document, the term "cloud computing environment" refers to an environment in which cloud computing is used.
[0230] Fig. Figure 17 shows a block diagram of a computing device 1700, which can be configured to perform one or more of the processes described above. It should be clear that one or more computing devices, such as computing device 1700, can represent the computing devices described above (for example, client device 102, server device 110, client device 300, or computing device 1400). In one or more implementations, computing device 1700 can be a mobile device (for example, a laptop, tablet, smartphone, mobile phone, camera, tracker, wristwatch, wearable device, and the like).In some implementations, the Computing Device 1700 can be a non-mobile device (for example, a desktop computer, a server device, a web server, a file server, a social networking system, a program server, an app store, or a content provider). Furthermore, the Computing Device 1700 can be a server device that includes cloud-based processing and storage capabilities.
[0231] As in Fig. As shown in Figure 17, the computing device 1700 can include one or more processors 1702, a memory 1704, a storage device 1706, input / output interfaces (“I / O”) 1708, and a communication interface 1710, which can be coupled by means of a communication infrastructure (for example, by means of a bus 1712). Although in Fig. 17 the calculating device 1700 is shown, are the in Fig. The components shown in Figure 17 are not meant to be limiting. Additional or alternative components may be used in other implementations. Furthermore, in certain implementations, the computing device 1700 includes fewer components than those shown in Figure 1700. Fig. 17 components shown. Fig. The 1700 calculating device shown in section 17 will now be described in more detail.
[0232] In certain implementations, the 1702 processor(s) includes hardware for executing instructions, such as those that constitute a computer program. For example, and not as a limitation, the 1702 processor(s) can retrieve instructions from an internal register, an internal cache, the 1704 memory, or the 1706 storage device, and then decode and execute them.
[0233] The 1700 computing device includes the 1704 memory, which is coupled to the 1702 processor(s). The 1704 memory can be used to store data, metadata, and programs for execution by the processor(s). The 1704 memory can include one or more volatile and non-volatile memory types, such as random access memory (RAM), read-only memory (ROM), a solid-state drive (SSD), flash memory, phase-change memory (PCM), or other types of data storage. The 1704 memory can be internal or distributed.
[0234] The computing device 1700 includes a storage device 1706 with memory for storing data or instructions. For example, and not as a limitation, the storage device 1706 may include a non-temporary storage medium as described above. The storage device 1706 may include a hard disk drive (HDD), flash memory, a USB (Universal Serial Bus USB) drive, or a combination of these or other storage devices.
[0235] The computing device 1700 includes, as shown, one or more I / O interfaces 1708 designed to allow a user to provide input (such as user strokes) to the computing device 1700, receive output from it, and otherwise transfer data to and from it. The I / O interfaces 1708 may include a mouse, a keypad or keyboard, a touchscreen, a camera, an optical scanner, a network interface, a modem, other known I / O devices, or a combination of such I / O interfaces 1708. The touchscreen may be activated by a stylus or a finger.
[0236] The 1708 I / O interfaces can include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (such as a display screen), one or more output drivers (such as display drivers), one or more audio speakers, and one or more audio drivers. In certain implementations, the 1708 I / O interfaces are configured to provide graphical data for presentation to a user. This graphical data can represent one or more graphical user interfaces and / or any other graphical content, provided it is useful for a particular implementation.
[0237] The computing device 1700 may also include a communication interface 1710. The communication interface 1710 may include hardware, software, or both. The communication interface 1710 provides one or more interfaces for communication (such as packet-based communication) between the computing device and one or more other computing devices or one or more networks. By way of example, and not as a limitation, the communication interface 1710 may include a network interface controller (NIC) or a network adapter for communication with an Ethernet or other wired network, or a wireless NIC (WNIC) or a wireless adapter for communication with a wireless network, such as Wi-Fi. The computing device 1700 may also include a bus 1712.The 1712 bus can contain hardware, software, or both that connect the components of the 1700 computing device.
[0238] In the foregoing description, the invention has been described with reference to specific exemplary implementations. Various implementations and aspects of the invention(s) are described with reference to the details explained herein, with the accompanying drawing illustrating the different implementations. The foregoing description and the drawing are for illustrative purposes only and should not be interpreted as limiting the invention. Numerous specific details have been described to facilitate a thorough understanding of the various implementations of the present invention.
[0239] The present invention may be embodied in other specific forms without departing from its essence or essential characteristics. The implementations described are to be considered in every respect merely illustrative and not restrictive. For example, the methods described herein may be carried out with fewer or more steps / operations, or the steps / operations may be carried out in different sequences. In addition, the steps / operations described herein may be repeated, carried out in parallel, or carried out in parallel with other instances of the same or similar steps / operations. The scope of the present invention is therefore defined by the appended claims and not by the preceding description. All modifications that are in accordance with the meaning and scope of the claims shall be included in their scope.
Citation Information
Patent Citations
Utilizing interactive deep learning to select objects in digital visual media
US10192129B2
Deep salient content neural networks for efficient digital object segmentation
US20190130229A1
structured modelling, extraction and localization of knowledge from images
DE102016010909A1
Image Matting via Deep Learning
DE102017010210A1
Removing and replacing objects in images according to a guided user dialog.
DE102018007937A1