Open Vocabulary Object Detection in Images
The system addresses the limitations of conventional object detection by using a neural network architecture to process images and query embeddings, enabling open-vocabulary detection and improving efficiency through independent processing and contrastive pre-training.
Patent Information
- Application Number
- JP2024565183
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-06
- Filing Date
- 2023-05-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-05-05
AI Technical Summary
Conventional object detection systems are limited to detecting objects within a small, fixed set of categories, and they struggle with efficiently processing images and text embeddings for open-vocabulary object detection.
The system employs a neural network architecture that includes an image encoding sub-network, a localization sub-network, and a classification sub-network to process images and query embeddings, enabling open-vocabulary object detection by generating classification score distributions for each object embedding across a set of query embeddings.
This approach allows for the detection of objects across any category, improving inference efficiency by processing images independently of query embeddings, and pre-training the networks contrastively to enhance downstream performance on object detection tasks.
Smart Images

Figure 2025516346000001_ABST
Abstract
Description
Background Art
[0001] This specification relates to the processing of data using a machine learning model.
[0002] A machine learning model receives an input and generates an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models that generate an output based on the received input and the values of the model's parameters.
[0003] Some machine learning models are deep models that use multiple layers of the model to generate an output for the received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers, and each layer applies a non-linear transformation to the received input to generate an output.
Summary of the Invention
[0004] This specification describes a system for performing object detection in an image, implemented as a computer program on one or more computers located in one or more locations.
[0005] Throughout this specification, "embedding" may refer to an ordered set of numbers, e.g., a vector, matrix, or other tensor of numbers.
[0006] According to one aspect, a method executed by one or more computers is provided. The method includes: (i) obtaining a set of (i) an image and (ii) one or more query embeddings, where each query embedding represents a respective category of an object; processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image, including processing the image using an image encoding sub-network of the object detection neural network to generate a set of object embeddings, processing each object embedding using a localization sub-network of the object detection neural network to generate localization data defining a corresponding region of the image, and processing (i) the set of object embeddings and (ii) the set of query embeddings using a classification sub-network of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings; and generating a set of the image and the query embeddings, where each respective classification score distribution for each object embedding defines the likelihood that the region of the image corresponding to the object embedding depicts an object included in the category represented by the query embedding.
[0007] In some embodiments, for one or more of the query embeddings, obtaining the query embedding includes obtaining a text sequence describing the category of the object and processing the text sequence using a text encoding sub-network of the object detection neural network to generate the query embedding.
[0008] In some embodiments, the image encoding sub-network and the text encoding sub-network are pre-trained, and the pre-training involves: (i) obtaining training images, (ii) positive text sequences that characterize the training images, and (iii) one or more negative text sequences that do not characterize the training images; using the image encoding sub-network to generate an embedding of the training images; using the text encoding sub-network to generate respective embeddings of each of the positive text sequences and the negative text sequences; and repeatedly performing operations including co-training the image encoding sub-network and the text encoding sub-network to promote (i) increasing the similarity between the embedding of the training images and the embedding of the positive text sequences, and (ii) decreasing the similarity between the embedding of the training images and the embedding of the negative text sequences.
[0009] In some embodiments, using the image encoding sub-network to generate an embedding of the training images includes using the image encoding sub-network to process the training images to generate a set of object embeddings for the training images, and using an embedding neural network to process the object embeddings to generate an embedding of the training images.
[0010] In some embodiments, the embedding neural network is co-trained with the image encoding sub-network and the text encoding sub-network.
[0011] In some embodiments, co-training the image encoding sub-network and the text encoding sub-network includes co-training the image encoding sub-network and the text encoding sub-network to optimize an objective function that includes a contrastive loss term.
[0012] In some embodiments, after pre-training the image encoding sub-network and the text encoding sub-network, the object detection neural network is trained to optimize an objective function that measures the performance of the object detection neural network for the task of object detection within an image.
[0013] In some embodiments, the objective function that measures the performance of the object detection neural network for the task of object detection within an image includes a bipartite matching loss term.
[0014] In some embodiments, using the classification sub-network of the object detection neural network to process (i) a set of object embeddings and (ii) a set of query embeddings to generate, for each object embedding, a respective classification score distribution over the set of query embeddings includes using one or more neural network layers of the classification neural network to process each object embedding to generate a corresponding classification embedding, and for each object embedding, using (i) the classification embedding corresponding to the object embedding and (ii) the query embeddings to generate a classification score distribution over the set of query embeddings, where the measurement of the similarity between the classification embedding and each query embedding defines the likelihood that the region of the image corresponding to the object embedding depicts an object included in the category represented by the query embedding, and includes generating a measurement of the similarity between the classification embedding and the query embedding.
[0015] In some embodiments, processing each object embedding using one or more neural network layers of a classification neural network to generate a corresponding classification embedding includes generating each classification embedding by projecting the corresponding object embedding into a latent space that includes a query embedding.
[0016] In some embodiments, generating each measurement of similarity between a classification embedding and each query embedding includes, for each query embedding, calculating an inner product between the classification embedding and the query embedding.
[0017] In some embodiments, for each object embedding, processing the object embedding using a localization subnetwork to generate localization data that defines a corresponding region of an image includes processing the object embedding using the localization subnetwork to generate localization data that defines a bounding box within the image.
[0018] In some embodiments, processing an image using an image encoding subnetwork to generate a set of object embeddings includes generating a set of initial object embeddings by an embedding layer of the image encoding subnetwork, where each initial object embedding is derived at least in part from a corresponding patch within the image, and processing the set of initial object embeddings by a plurality of neural network layers including one or more self-attention neural network layers to generate a set of final object embeddings.
[0019] In some embodiments, processing an object embedding using a localization subnetwork to generate localization data that defines a corresponding region of an image includes generating a set of offset coordinates that define an offset of the corresponding region of the image from a location of an image patch corresponding to the object embedding.
[0020] In some embodiments, the text encoding subnetwork includes one or more self-attention neural network layers.
[0021] In some embodiments, the method includes determining, for one or more of the object embeddings, that the region of the image corresponding to the object embedding depicts an object included in the category represented by the query embedding based on the classification score distribution for the object embedding.
[0022] In an embodiment, for one or more of the query embeddings, obtaining the query embedding includes obtaining one or more query images, each query image including a respective target region depicting an example of the target object, generating a respective embedding of the target object for each query image, and generating the query embedding by combining the embeddings of the target object within the query images.
[0023] In some embodiments, for each query image, generating an embedding of the target object within the query image includes processing the query image using an image encoding subnetwork to generate a set of object embeddings for the query image, processing each object embedding for the query image using a localization subnetwork to generate localization data defining the corresponding region of the query image, determining, for each object embedding for the query image, a respective measure of overlap between (i) the target region of the query image depicting an example of the target object and (ii) the region of the query image corresponding to the object embedding for the query image, and selecting, from the set of object embeddings for the query image, an object embedding as the embedding of the target object within the query image based on the measure of overlap.
[0024] In some embodiments, generating a query embedding by combining embeddings of target objects in a query image includes averaging the embeddings of the target objects in the query image.
[0025] According to another aspect, there is provided a system including one or more computers and one or more storage devices communicatively coupled to the one or more computers, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the methods described herein.
[0026] According to another aspect, there is provided one or more non-transitory computer-readable media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the methods described herein.
[0027] The subject matter described herein can be implemented in certain embodiments so as to realize one or more of the following advantages.
[0028] The system described in this specification can perform "open-vocabulary" object detection. That is, the system can detect objects of any object category. In contrast, conventional object detection systems may be limited to detecting objects within a small, fixed set of object categories. This system enables open-vocabulary object detection by making it possible to specify object categories in the form of query embeddings. Each query embedding represents a respective object category and can be generated in any suitable way, for example, from a text sequence that describes the object category (e.g., "a bird perched on a tree") or from an image that shows an example of an object within the object category. In particular, this system enables image-conditioned object detection. That is, a query embedding is derived from an image that shows an example of an object, thereby enabling the detection of objects that are difficult to describe in text but easy to capture in an image (e.g., special technical parts).
[0029] In the case of a text sequence, the system can perform object detection by comparing a set of query embeddings generated by processing the text sequence using a text-encoding neural network with a set of object embeddings generated by processing an image using an image-encoding neural network. The text-encoding neural network and the image-encoding neural network can operate independently, which can lead to a dramatic improvement in inference efficiency. For example, in a system where the text-encoding neural network and the image-encoding neural network are fused, a forward pass through the image-encoding neural network is required to encode a query, and this needs to be repeated for each image-query combination. In contrast, the system described herein can process an image once using the image-encoding neural network and then generate any number of query embeddings without reprocessing the image. The same applies to other query modalities. The image-encoding neural network and the query embedding generator are independent entities. The query embeddings may be generated separately (and may be generated on different systems) and provided to the system to perform object detection (and vice versa).
[0030] The system can pre-train a text-encoding neural network and an image-encoding neural network "contrastively" to learn representations of images and text in a shared embedding space, such that, for example, embeddings of semantically similar images and text tend to be closer in the embedding space with respect to object categories depicted in the images and defined in the text. Training data for the contrastive pre-training of the text-encoding neural network and the image-encoding neural network is abundantly available. By pre-training the text-encoding neural network and the image-encoding neural network contrastively, the downstream performance of the text-encoding neural network and the image-encoding neural network on the task of object detection can be significantly improved, for which relatively little training data may be available.
[0031] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0032]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Best Mode for Carrying Out the Invention
[0033] Like reference numerals and symbols in the various drawings refer to like elements.
[0034] FIG. 1 shows an exemplary object detection system 100. The object detection system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations, where the systems, components, and techniques described below are implemented.
[0035] The object detection system 100 is configured to process a set of an image 105 and one or more query embeddings 110 to generate an object detection output 145.
[0036] The image 105 can be generated using any suitable imaging device, such as a camera, a microscope imaging device, a telescope imaging device, a medical imaging device (e.g., a computed tomography (CT) imaging device, an X-ray imaging device, or an ultrasonic (US) imaging device), etc. The image 105 can show (depict) one or more objects. For example, an image captured by a camera can show objects such as road signs, vehicles, pedestrians, buildings, etc. As another example, an image captured by a medical imaging device can show anatomical structures (e.g., organs, tumors, etc.), tissues, regions of the human or animal body, etc. The image 105 can include pixel intensity values.
[0037] Each query embedding in the set of one or more query embeddings 110 represents a respective object category and can be generated in any suitable way, for example, from a text sequence that describes the object category (e.g., "a bird perched on a tree") or from an image that shows an example of an object in the object category. In particular, the object detection system 100 enables image-conditioned object detection. That is, a query embedding is derived from an image showing an example of an object, thereby enabling the detection of objects that are difficult to describe in text but easy to capture in an image (e.g., special technical parts). For example, exemplary techniques for generating query embeddings 110 representing object categories from text sequences or images will be described in more detail below with reference to FIGS. 2-4.
[0038] The object detection output can include one or more object detection instances. Each object detection instance can be associated with (i) location-specific data that identifies a region of the image and (ii) object identification data that characterizes the type of object shown in the region of the image. In some cases, the object identification data can include a classification score distribution that defines, for each query embedding, the likelihood that the region of the image shows an object within the object category represented by the query embedding. In some cases, the object identification data defines that the region of the image shows an object included in an object category associated with a particular query embedding from the set of query embeddings. Thus, the object detection output can characterize the location and type of objects shown in the image.
[0039] The object detection system 100 can receive an image 105 and a query embedding 110 from a user or another source, for example, via an application programming interface (API) available to the system. The object detection system 100 can provide an object detection output 145, for example, for transmission via a data communication network (e.g., the Internet), or for display on an interface (e.g., a graphical user interface), or for storage in memory.
[0040] The object detection system 100 can be implemented in any suitable location, such as within a data center, or on a user device (e.g., a smartphone, tablet, or laptop), or in a distributed manner, for example, where certain parts of the system are implemented in a data center and other parts of the system are implemented on a user device.
[0041] The object detection system 100 can include an object detection neural network 150, and the object detection neural network 150 itself includes an image encoding subnet 115, a localization subnet 125, and a classification subnet 135, which are each described in more detail throughout this specification.
[0042] The image encoding subnet 115 processes the image 105 to generate a set of object embeddings 120. One or more of the object embeddings can each correspond to a respective object shown in the image.
[0043] The set of object embeddings 120 can include any suitable number of object embeddings, e.g., a static (pre - defined) number of object embeddings, or a number of object embeddings that depends on the image 105. The image encoding sub - network 115 can have any suitable neural network architecture, e.g., include any suitable type of neural network layer (e.g., attention layer, convolutional layer, fully - connected layer, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and be connected in any suitable configuration (e.g., as a directed graph of layers). A specific exemplary architecture of the image encoding sub - network 115 will be described in more detail below with reference to FIG. 3.
[0044] The localization sub - network 125 processes each object embedding within the set of object embeddings 120 to generate localization data 130 that defines the corresponding region of the image. The localization sub - network can have any suitable neural network architecture, e.g., include any suitable type of neural network layer (e.g., attention layer, convolutional layer, fully - connected layer, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and be connected in any suitable configuration (e.g., as a directed graph of layers).
[0045] In some examples, the localization data can define a bounding box within the image corresponding to each object embedding. For example, if the image 105 depicts several different objects, the localization sub - network 125 can define a corresponding bounding box for each object within the image.
[0046] The classification subnet 135 processes a set of object embeddings 120 and a set of query embeddings 110 to generate, for each object embedding, a respective classification score distribution 140 across the set of query embeddings. The classification score distribution 140 for an object embedding 120 defines, for each query embedding 110, the likelihood that the region of the image corresponding to the object embedding (i.e., according to the localization data) depicts an object included in the category represented by the query embedding.
[0047] The classification subnet 135 can have any suitable neural network architecture, for example, include any suitable type of neural network layer (e.g., attention layer, convolutional layer, fully connected layer, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and be connected in any suitable configuration (e.g., as a directed graph of layers). A specific exemplary architecture of the classification subnet will be described in more detail below with reference to FIG. 3.
[0048] The object detection system 100 can process the localization data 130 and the classification score distribution 140 to generate an object detection output 145. The object detection output includes one or more object detection instances, and each object detection instance can be associated with (i) localization data identifying a region of the image and (ii) object identification data characterizing the type of object shown in the region of the image (as described above).
[0049] In particular, for each of one or more of the object embeddings, the object detection system 100 can generate a respective object detection instance. For example, for each of the one or more object embeddings, the object detection system 100 can generate an object detection instance associated with (i) location data generated by a location subnet for the object embedding and (ii) object identification data based on a classification score distribution generated by a classification subnet for the object embedding. In some cases, the object identification data includes the classification score distribution. In some cases, the object identification data identifies a query embedding associated with the highest score under the classification score distribution.
[0050] In some examples, the object detection system 100 can refrain from generating an object detection instance for an object embedding, for example, based on the classification score distribution for the object embedding. For example, the object detection system 100 can refrain from generating an object detection instance for an object embedding if the maximum score included in the classification score distribution for the object embedding is less than a threshold value, for example, a threshold value of 0.001, or 0.01, or 0.1.
[0051] The training of the object detection neural network will be described in more detail below with reference to FIGS. 4-5.
[0052] FIG. 2 is a flowchart of an exemplary process 200 for detecting and locating an object. For convenience, process 200 is described as being executed by a system of one or more computers located in one or more locations. For example, an object detection system, such as object detection system 100 of FIG. 1, can be appropriately programmed in accordance with this specification and can execute process 200.
[0053] The system obtains a set of images and query embeddings (step 202). Each query embedding represents a respective category of an object.
[0054] An image can depict one or more objects. For example, an image captured by a camera can show objects such as road signs, vehicles, pedestrians, buildings, etc. As another example, an image captured by a medical imaging device can show anatomical structures (e.g., organs, tumors, etc.), tissues of the human or animal body, regions, etc.
[0055] The system can obtain query embeddings in any of a variety of ways. Next, some exemplary techniques by which the system can generate query embeddings will be described.
[0056] In some examples, the system generates one or more of the query embeddings by obtaining a text sequence that describes the category of an object. Next, the system processes the text sequence using a text encoding subnetwork to generate a query embedding for the given text sequence. For example, the given text sequence can be a short description of the type of object (e.g., "big brown butterfly") or "full moon".
[0057] In some examples, the system obtains one or more of the query embeddings by first obtaining one or more query images. Each query image includes a target region that shows an example of the target object. For example, the query image can depict a bird on a tree branch in front of a blue background. If the target object is a bird, the target region can be the region that depicts the bird. The system generates an embedding for each target object in each query image and generates a query embedding by combining the embeddings of the target objects in the query images. In some examples, the system generates a query embedding by averaging the embeddings of the target objects in the query images to combine the embeddings of the target objects in the query images.
[0058] The system can generate an embedding of the target object for each query image by using an image encoding subnetwork to process the query image to generate a set of object embeddings for the query image. The system can use a localization subnetwork to process each object embedding for the query image to generate localization data that defines the corresponding region of the query image. For each object embedding for the query image, the system can determine a respective measure of overlap using, for example, an intersection over union (IoU) measure between the target region of the query image that depicts an example of the target object and the region of the query image corresponding to the object embedding for the query image. The system can select an object embedding from the set of object embeddings for the query image as the embedding of the target object within the query image based on the measure of overlap. For example, the system can select the object embedding with the highest measure of overlap.
[0059] For example, a query image can depict a tennis ball and a racket. A target region can depict a tennis ball. A set of object embeddings can include an object embedding corresponding to the tennis ball and an object embedding corresponding to the racket. The system can determine two overlap measurements, one between the object embedding corresponding to the tennis ball and the target region that depicts the tennis ball, and the other between the object embedding corresponding to the racket and the target region that depicts the tennis ball. The system can determine that there is a higher overlap measurement between the object embedding corresponding to the tennis ball and the target region that depicts the tennis ball, and can select this object embedding as the target object embedding. The overlap measurement can be a percentage or score on a given scale, or an intersection over union (IoU) measurement.
[0060] The image encoding subnetworks, text encoding subnetworks, and location determination subnetworks can have any suitable neural network architecture, for example, can include any suitable type of neural network layer (e.g., self-attention layer, convolutional layer, fully connected layer, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and be connected in any suitable configuration (e.g., as a directed graph of layers).
[0061] Training of the image encoding subnetworks, text encoding subnetworks, and location determination subnetworks will be described in detail below with reference to FIGS. 4-5.
[0062] The system processes an image using an image encoding sub-network to generate a set of object embeddings (step 204). For example, the system can generate an initial set of object embeddings by an embedding layer of the image encoding sub-network. In particular, the system can divide the image into a grid of patches. Each initial object embedding can be derived from a corresponding patch within the image. For example, each initial object embedding can be generated by concatenating the pixels within the corresponding patch in the image into a vector. The system can process the initial set of object embeddings by one or more neural network layers, such as including one or more self-attention neural network layers, to generate a final set of object embeddings. The final set of object embeddings can include one or more embeddings for each of one or more objects depicted within the image.
[0063] In some embodiments, the system can generate a predefined number of object embeddings, such as 500, for each input image. The image can include fewer objects than the predefined number of object embeddings. For example, the image can include only two objects. In this case, the system may generate two object embeddings corresponding to the objects within the input image and 498 object embeddings that do not correspond to any object within the input image. For example, these may correspond to embeddings generated from the background of the image. Alternatively, there may be multiple object embeddings generated for each object. In particular, when the object embeddings are derived from image patches, the object may span multiple patches. When the object embeddings are derived from image patches, they can be considered as image region embeddings.
[0064] The system processes each object embedding to generate location-specific data that defines a corresponding region of the image using a location-specific sub-network (step 206).
[0065] The system can use a localization subnetwork to process each object embedding and generate localization data that defines regions of the image, such as bounding boxes within the image. For example, if the image depicts several different objects, the localization data can define corresponding bounding boxes for each object within the image.
[0066] In some examples, the system can generate a set of offset coordinates as the localization data. The offset coordinates can be represented as (x,y) displacements. The offset coordinates can define the offset of a corresponding region of the image from an image patch corresponding to the object embedding.
[0067] The system uses a classification subnetwork to process a set of object embeddings and a set of query embeddings to generate, for each object embedding, a classification score distribution over the set of query embeddings (step 208). The classification score distribution for an object embedding defines, for each query embedding, the likelihood that the region of the image corresponding to the object embedding (i.e., according to the localization data) depicts an object included in the category represented by the query embedding.
[0068] The system can generate a set of classification score distributions by first processing each object embedding using one or more neural network layers of a classification neural network to generate a corresponding classification embedding. Next, the system can generate, for each object embedding, a classification score distribution over the set of query embeddings using the classification embeddings corresponding to the object embedding and the query embeddings.
[0069] In some examples, the system generates each classification embedding by projecting the corresponding object embedding into a latent space that includes the query embeddings.
[0070] In some examples, the system determines respective measurements of similarity between each classification embedding and each query embedding (e.g., L 1 norm, L 2 norm, cosine similarity, inner product, etc.). A measurement of similarity between a classification embedding and a query embedding can define the likelihood that a region of an image corresponding to the object embedding depicts an object included in the category represented by the query embedding.
[0071] The system generates an object detection output from the localization data and the classification score distribution (step 210). The object detection output includes one or more object detection instances, and each object detection instance can be associated with (i) localization data that identifies a region of the image, and (ii) object identification data that characterizes the type of object shown in the region of the image. In particular, the system can generate a respective object detection instance for each of one or more of the object embeddings. For example, for each of one or more object embeddings, the system can generate an object detection instance associated with (i) localization data generated by a localization subnetwork for the object embedding, and (ii) object identification data based on a classification score distribution generated by a classification subnetwork for the object embedding. In some cases, the object identification data includes the classification score distribution. In some cases, the object identification data identifies a query embedding associated with the highest score under the classification score distribution.
[0072] In some examples, the system can refrain from generating an object detection instance for an object embedding, for example, based on the classification score distribution for the object embedding. For example, the system can refrain from generating an object detection instance for an object embedding if the maximum score included in the classification score distribution for the object embedding is less than a threshold value, such as a threshold value of 0.001, or 0.01, or 0.1.
[0073] For example, the system can receive a query embedding representing the category "duck" and a query embedding representing the category "dog". The system can receive an input image depicting a duck and a cat. The set of object embeddings for the image includes an embedding representing a duck and an embedding representing a cat. For the object embedding representing a duck, the classification score distribution can include a score of 0.93 for the category "duck" and a score of 0.01 for the category "dog". For the object embedding representing a cat, the classification score distribution can include a score of 0.02 for the category "duck" and a score of 0.01 for the category "dog". The system can determine not to generate an object detection instance if the maximum score included in the classification score distribution for the object embedding is less than a threshold value of 0.1. Thus, the system does not generate an object detection instance for the object embedding representing a cat, and the system generates an object detection instance for the object embedding representing a duck.
[0074] In some cases, the set of query embeddings includes query embeddings designated as "background" query embeddings. Background query embeddings can represent the background portion of an image, e.g., the portion of the image that does not show a particular object. In response to determining that a background embedding is associated with the maximum score under the classification score distribution for object embeddings, the system can refrain from generating an object detection instance for the object embedding.
[0075] The object detection output includes one object detection instance associated with localization data identifying the region of the image depicting the duck and object identification data characterizing the duck. For example, the localization data can be a bounding box surrounding the region of the image depicting the duck. The object identification data can identify that, for the object embedding representing the duck, the query embedding representing the category "duck" is associated with the highest score under the classification score distribution across the set of query embeddings.
[0076] In some examples, the system can be implemented on a robot. The robot can receive natural language instructions to perform tasks. The natural language instructions can request the robot to identify a particular object (e.g., "Clean the kitchen using X"). The software installed on the robot can use the open vocabulary object detection capabilities of an object detection neural network to identify the object specified by the natural language instructions.
[0077] In some examples, the system can be implemented on a robot that has previously seen one or more objects. The robot can receive instructions to identify objects similar to the previously seen objects. The object detection neural network can generate a set of query embeddings using the previously seen objects. The robot can use the object detection neural network to detect similar objects in the future.
[0078] Figure 3 shows an exemplary architecture 300 of the object detection system of FIG. 1. This architecture shows a text encoding neural network 310 and an image encoding neural network 330.
[0079] The image encoding sub-network and the text encoding sub-network can have any suitable neural network architecture, for example, can include any suitable type of neural network layer (e.g., self-attention layer, convolutional layer, fully connected layer, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and be connected in any suitable configuration (e.g., as a directed graph of layers). Exemplary architectures include vision transform encoders.
[0080] The text encoding sub-network 310 receives a text sequence 305 that describes an object category, such as "giraffe", "tree", "car". The text sequence 305 can be the name of the category or other text description. The text encoding sub-network processes the text sequence 305 to generate a set of query embeddings 315 corresponding to the text sequence. Each query embedding consists of a distinct token sequence representing an individual object description and is processed individually by the text encoder.
[0081] The image encoding neural network 330 receives an input image 325 that is divided into patches. The image encoding neural network 330 generates a set of object embeddings 355. The set of object embeddings includes object-specific embeddings for each object shown in the input image 325. The input image 325 shows two giraffes and one tree. In this example, there can be object embeddings for each giraffe and the tree.
[0082] The location-specific sub-network 335 processes the object embeddings 355 to generate location-specific data that defines a set of bounding boxes 340 within the image. The bounding boxes are represented as box coordinates. For each object embedding, the location-specific sub-network generates a prediction box representing the location of the object.
[0083] The classification sub-network 350 processes a set of object embeddings and a set of query embeddings to generate, for each object embedding, a classification score distribution 360 across the set of query embeddings. Thereby, for each object embedding, a predicted class 320 can be generated.
[0084] The query embeddings 315 are, from left to right, "giraffe", "tree", and "car". The classification score distribution corresponding to the topmost object embedding indicates probabilities of 0.9 for giraffe, 0.1 for tree, and 0.1 for car. Thus, the classification sub-network 350 predicts that the object embedding belongs to the class "giraffe". The classification score distribution corresponding to the bottommost object embedding indicates probabilities of 0.1 for giraffe, 0 for tree, and 0.1 for car. Thus, the classification sub-network 350 predicts that the object embedding does not belong to one of the query classes.
[0085] The architecture 300 does not include a fusion between the image encoder 330 and the text encoder 310. That is, the image encoding neural network 330 and the text encoding neural network 310 are separate neural networks that execute their processing independently of each other. Although early fusion may seem intuitively beneficial, encoding the query 315 requires a forward pass through the entire image model and needs to be repeated for each image / query combination, resulting in a decrease in inference efficiency. In this architecture, the system 100 can compute query embeddings independently of the images, enabling the use of thousands of queries per image and allowing for the use of far more queries than would be possible with early fusion.
[0086] The object detection system 100 processes the predicted classes 320 and the predicted boxes 340 to generate respective object detection instances for each of one or more of the object embeddings. An object detection instance can define, for example, a bounding box and a classification score distribution for the bounding box.
[0087] FIG. 4 is a flow diagram of an exemplary process 400 for pre-training and fine-tuning an image encoding sub-network and a text encoding sub-network included in an object detection neural network, for example, described with reference to FIG. 1. For convenience, process 400 is described as being executed by one or more computer systems located in one or more locations, for example, an object detection system described with reference to FIG. 1.
[0088] The system pre-trains an image encoding sub-network and a text encoding sub-network (step 405). The pre-training includes repeatedly performing a series of operations on a set of training images. For example, for each training image in the set of training images, the system can obtain a positive text sequence and one or more negative text sequences. The positive text sequence characterizes at least one of the objects depicted in the training image, and the negative text sequence does not characterize any of the objects depicted in the training image. For example, the training image can be a great purple emperor. The positive text sequence may be "great purple emperor", and the negative text sequences may be "peacock butterfly", "moth", "caterpillar", and "starfish". The system can obtain positive text sequences and negative text sequences for training images from various possible sources. For example, the system can scrape positive text sequences for training images as captions of the images from an image database. The negative text sequences can be, for example, randomly sampled text sequences.
[0089] The system can generate an embedding of a training image using an image encoding subnetwork. For example, the system can use the image encoding subnetwork to process a training image to generate a set of object embeddings for the training image (as described with reference to FIGS. 1 and 2, for example). Next, the system can use an embedding neural network to process the object embeddings to generate an embedding of the training image. The embedding neural network can have any suitable neural network architecture, for example, include any suitable type of neural network layer (such as a pooling layer, a fully connected layer, an attention layer, etc.) in any suitable number (such as 1 layer, 5 layers, or 10 layers) and be connected in any suitable configuration (such as as a directed graph of layers). A specific exemplary architecture of the embedding neural network is described with reference to FIG. 5. In some embodiments, the embedding neural network is co-trained with an image encoding subnetwork and a text encoding subnetwork as described in more detail below.
[0090] The system uses a text encoding subnetwork to generate individual embeddings of a positive text sequence and a negative text sequence, respectively. That is, for each of the positive text sequence and the negative text sequence, the system uses the text encoding subnetwork to process the text sequence and generate an embedding of the text sequence.
[0091] The system co-trains an image encoding sub-network, a text encoding sub-network, and optionally an embedding neural network to promote (i) increasing the similarity between the embedding of the training image and the embedding of the positive text sequence, and (ii) decreasing the similarity between the embedding of the training image and the embedding of the negative text sequence. For example, the system can co-train the image encoding sub-network, the text encoding sub-network, and the embedding neural network to optimize an objective function that includes a contrastive loss term or a triplet loss term. The objective function can measure the similarity between the embeddings using, for example, the inner product, or an L 1 similarity measure, or an L 2 similarity measure, or other suitable similarity measure. For each training image, the system can determine the gradient of the objective function with respect to the parameters of the image encoding sub-network, the text encoding sub-network, and optionally the embedding neural network. The system can then use the gradient to adjust the parameter values of the image encoding sub-network, the text encoding sub-network, and optionally the embedding neural network, for example, using the update rules of a suitable gradient descent optimization algorithm (e.g., RMSprop or Adam).
[0092] That is, the system can pre-train the text encoding neural network and the image encoding neural network "contrastively" to learn the representations of images and text in a shared embedding space, such that, for example, the embeddings of images and text having a common object tend to be closer in the embedding space. By pre-training the text encoding neural network and the image encoding neural network contrastively, the downstream performance of the text encoding neural network and the image encoding neural network on the task of object detection can be significantly improved.
[0093] In some examples, after the pre-training is completed, the embedded neural network is discarded.
[0094] After pre-training the image encoding sub-network and the text encoding sub-network, the system then fine-tunes the image encoding sub-network and the text encoding sub-network with respect to the task of object detection (step 410). The system fine-tunes both the image encoding sub-network and the text encoding sub-network from start to finish.
[0095] When fine-tuning, the system repeatedly performs a series of operations on a set of training examples. For example, each training example can include a training image, one or more positive text sequences, each bounding box associated with each positive text sequence, and one or more negative text sequences. Each positive text sequence can characterize each object shown within the corresponding bounding box in the training image. The negative text sequences do not characterize the objects shown in the training image.
[0096] The system can use the image encoding sub-network to process the training image to generate a set of object embeddings for the training image (e.g., as described with reference to FIGS. 1 and 2).
[0097] The system uses the text encoding sub-network to generate individual embeddings for each of the positive text sequences and the negative text sequences. That is, for each of the positive text sequences and the negative text sequences, the system uses the text encoding sub-network to process the text sequence and generate an embedding of the text sequence.
[0098] The system uses a localization subnetwork to generate localization data that defines a set of bounding boxes within an image. For each object embedding, the localization subnetwork generates a predicted box representing the location of the object.
[0099] The system trains an object detection neural network to optimize an objective function that measures the performance of the object detection neural network for the task of object detection within an image. In some examples, the objective function includes a loss term that encourages the object detection neural network to produce accurate object classifications. For example, the objective function can include a bipartite matching loss term. In some examples, the objective function includes a loss term that encourages the localization subnetwork to perform accurate localization. For example, the objective function can include a loss term that measures the error between (i) a bounding box associated with a positive text sequence of a training example and (ii) a bounding box generated by the object detection neural network for the corresponding object embedding, for each positive text sequence of the training example.
[0100] FIG. 5 shows, for example, the pre-training of an image encoding subnetwork 525 and a text encoding subnetwork 510 included in the object detection neural network described with reference to FIG. 1.
[0101] The text encoding subnetwork 525 receives a positive text sequence 505 such as, for example, "a bird perched on a tree". The image encoding subnetwork receives a training image 520 that matches the positive text sequence, such as, for example, an image of a bird perched on a tree.
[0102] The text encoding subnetwork 510 processes the positive text sequence 505 to generate a text embedding 515 corresponding to the positive text sequence. The image encoding subnetwork 525 processes the input image 520 to generate an embedding of the training image 535.
[0103] The input image 520 can be divided into patches. The image encoding subnetwork 525 generates a set of object embeddings 545.
[0104] The embedding neural network 530 processes the object embeddings 545 to generate an embedding of the training image 535. In this example, the embedding neural network includes a pooling layer (e.g., an average pooling layer or a max pooling layer) and a projection (e.g., fully connected) layer.
[0105] The object detection system can train the text encoding subnetwork and the image encoding subnetwork as a single batch at a time. The batch can include one or more training images along with their respective positive and negative text sequences. For example, the batch can include five training images depicting a bird, a cat, a dog, a squirrel, and a giraffe, along with appropriate descriptions. For the image of the bird, the positive text sequence could be "a bird perched on a tree" and the negative text sequence could be "a giraffe standing on the grass".
[0106] The text embedding 515 is added to a set of text embeddings 550 corresponding to other text descriptions in the batch. The embedding of the training image 535 is added to a set of image embeddings corresponding to other training images in the batch.
[0107] The object detection system 100 optimizes an objective function that includes a contrastive loss term over all the images within the batch 540. The objective function promotes increasing the similarity between the embedding of the training image 535 and the embedding of the positive text sequence 515, and decreasing the similarity between the embedding of the training image and the embedding of the negative text sequence. The illustration of FIG. 5 shows a grid having text embeddings 550 on one axis and embeddings of training images 555 on the other axis. The intersection of each text embedding and each image embedding is represented by either a “+” or a “−”, where “+” indicates that the text embedding and the image embedding are promoted to be similar, and “−” indicates that the text embedding and the image embedding are promoted not to be similar.
[0108] The term “configured to” is used herein in relation to systems and computer program components. A system of one or more computers being configured to perform a particular operation or action means that the system has installed therein software, firmware, hardware, or a combination thereof that causes the operating system to perform the operation or action. One or more computer programs configured to perform a particular operation or action means that one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.
[0109] The subject matter and embodiments of the functional operations described in this specification can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware, such as the structures disclosed in this specification and their structural equivalents, or combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs encoded on a tangible non-transitory storage medium, i.e., can be implemented as one or more modules of computer program instructions, executed by a data processing apparatus or controlling the operation of a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively, or in addition, the program instructions can be encoded in an artificially generated transmitted signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver device for execution by a data processing apparatus.
[0110] The term "data processing apparatus" refers to data processing hardware and includes, by way of example, any kind of apparatus, device, and machine for data processing, including programmable processors, computers, or multiple processors or multiple computers. The apparatus may be, or further include, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The apparatus optionally includes, in addition to hardware, code to create an execution environment for a computer program (e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them).
[0111] A computer program, which may be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or in the form of a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program can be stored in a part of a file that holds one or more scripts stored in a document of a markup language, in a single file dedicated to the program of interest, or in multiple related files, such as files that store one or more modules, subprograms, or parts of code. A computer program can be deployed to be executed on one computer or can be distributed over one or more locations and executed on multiple computers interconnected by a data communication network.
[0112] As used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine is implemented as one or more software modules or components and installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and executed on the same one or more computers.
[0113] The processes and logical flows described in this specification can be implemented by one or more programmable computers executing one or more computer programs to act on input data and generate output. The processes and logical flows can also be implemented by, or in combination with, special-purpose logic circuitry, e.g., an FPGA or ASIC.
[0114] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors, or both, or any other kind of central processing unit. In general, a central processing unit receives instructions and data from a read only memory, a random access memory, or both. Basic components of a computer are a central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry. In general, a computer also includes, or is operatively coupled to, one or more mass storage devices for storing data, such as, by way of example only, magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Further, a computer can be embedded in another device, such as, by way of example only, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, a universal serial bus (USB) flash drive, etc.
[0115] Computer-readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and any form of non-volatile memory, media, and memory devices, such as CD-ROM and DVD-ROM disks.
[0116] To interact with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other kinds of devices can also be used to interact with a user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input received from the user can be in any form including acoustic, speech, or tactile input. Further, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user (e.g., by sending a web page to a web browser on the user's device in response to a request received from the web browser). Also, the computer can interact with the user by sending a text message or other form of message to a personal device (e.g., a smartphone running a messaging application) and then receiving a response message from the user.
[0117] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing the general and numerical calculations portions of machine learning training or production (i.e., inference, workload).
[0118] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework or the Jax framework.
[0119] Embodiments of the subject matter described herein can be implemented in a computing system that includes, for example, a backend component as a data server, or a middleware component, such as an application server, or a frontend component, such as a client computer having a graphical user interface, a web browser, or an app through which a user can interact with embodiments of the subject matter described herein, or a computing system that includes any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium, such as by a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), such as the Internet.
[0120] The computing system can include clients and servers. Clients and servers are generally in a remote state from each other and typically interact via a communication network. The relationship between a client and a server results from computer programs that are executed on respective computers and have the relationship of client and server with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device for the purpose of, for example, displaying the data to a user who interacts with a device functioning as a client and receiving user input from the user. Data generated on the user device (e.g., the result of a user interaction) can be received at the server from the device.
[0121] Although this specification contains many details of specific embodiments, these should not be construed as limiting the scope of any invention or the scope of what can be claimed, but rather as an explanation of features that may be specific to a particular embodiment of a particular invention. Specific features described herein as separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described as a single embodiment may also be implemented separately or in any suitable sub-combination in a plurality of embodiments. Furthermore, even if a plurality of features are described above as functioning in a particular combination and are first described as such in the claims, one or more features may, in some cases, be deleted from the combination described in the claims, and furthermore, the combination described in the claims may be directed to a sub-combination or a variation of a sub-combination.
[0122] Similarly, although operations are shown in the drawings and described in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order or sequential order shown in order to obtain a desirable result, or that all operations shown be performed. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and the described program components and systems may generally be integrated into a single software product or packaged into a plurality of software products.
[0123] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims may be performed in a different order and still result in a desirable outcome. As one example, the processes shown in the accompanying figures do not necessarily require that the particular order or series of orders shown be followed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
Claim 1 A method executed by one or more computers, comprising: obtaining a set of (i) an image and (ii) one or more query embeddings, each query embedding representing a respective category of an object; processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image, processing the image using an image encoding subnet of the object detection neural network to generate a set of object embeddings, processing each object embedding using a localization subnet of the object detection neural network to generate localization data defining a corresponding region of the image, processing (i) the set of object embeddings and (ii) the set of query embeddings using a classification subnet of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings; wherein each respective classification score distribution for each object embedding defines, for each query embedding, the likelihood that the region of the image corresponding to the object embedding depicts an object included in the category represented by the query embedding; processing the image and the set of query embeddings; the method as claimed in claim 1. Claim 2 For one or more of the query embeddings, obtaining the query embedding comprises: obtaining a text sequence describing a category of an object; processing the text sequence using a text encoding subnet of the object detection neural network to generate the query embedding; the method as claimed in claim 1. Claim 3 The image encoding subnet and the text encoding subnet are pre-trained, and the pre-training comprises (i) a training image, (ii) a positive text sequence that characterizes the training image, and (iii) one or more negative text sequences that do not characterize the training image; generating an embedding of the training image using the image encoding sub-network; generating respective embeddings of the positive text sequence and the negative text sequences using the text encoding sub-network; co-training the image encoding sub-network and the text encoding sub-network to promote (i) increasing the similarity between the embedding of the training image and the embedding of the positive text sequence and (ii) decreasing the similarity between the embedding of the training image and the embedding of the negative text sequence; The method according to claim 2, comprising repeatedly performing an operation including the above.
4. Generating the embedding of the training image using the image encoding sub-network comprises processing the training image using the image encoding sub-network to generate a set of object embeddings for the training image; processing the object embeddings using an embedding neural network to generate the embedding of the training image; The method according to claim 3, comprising the above.
5. The method according to claim 4, wherein the embedding neural network is co-trained with the image encoding sub-network and the text encoding sub-network.
6. Co-training the image encoding sub-network and the text encoding sub-network comprises co-training the image encoding sub-network and the text encoding sub-network to optimize an objective function including a contrastive loss term The method according to any one of claims 3 to 5, comprising the above.
7. After the pre-training of the image encoding sub-network and the text encoding sub-network, the object detection neural network is trained to optimize an objective function that measures the performance of the object detection neural network for the task of object detection in an image. The method according to any one of claims 3 to 6.
8. The objective function for measuring the performance of the object detection neural network for the task of object detection in an image includes a bipartite matching loss term. The method according to claim 7.
9. Using the classification sub-network of the object detection neural network to process (i) the set of object embeddings and (ii) the set of query embeddings, for each object embedding, generating a respective classification score distribution over the set of query embeddings, Processing each object embedding using one or more neural network layers of the classification neural network to generate a corresponding classification embedding, For each object embedding, generating the classification score distribution over the set of query embeddings by using (i) the classification embedding corresponding to the object embedding and (ii) the query embedding, Generating a measure of similarity between the classification embedding and each query embedding, which defines the likelihood that the region of the image corresponding to the object embedding depicts an object included in the category represented by the query embedding. Including generating the classification score distribution as described above. The method according to any one of claims 1 to 8.
10. Processing each object embedding using one or more neural network layers of the classification neural network to generate a corresponding classification embedding, Generating each classification embedding by projecting the corresponding object embedding into a latent space that includes the query embedding. The method according to claim 9.
11. Generating the respective measure of similarity between the classification embedding and each query embedding is for each query embedding. Calculating the inner product between the classification embedding and the query embedding The method according to claim 9, comprising: **Claim 12** For each object embedding, using the localization sub-network to process the object embedding to generate localization data that defines the corresponding region of the image Using the localization sub-network to process the object embedding to generate localization data that defines a bounding box within the image The method according to any one of claims 1 to 11, comprising: **Claim 13** Using the image encoding sub-network to process the image to generate a set of object embeddings Generating a set of initial object embeddings by an embedding layer of the image encoding sub-network, wherein each initial object embedding is at least partially derived from the corresponding patch in the image Processing the set of initial object embeddings by a plurality of neural network layers including one or more self-attention neural network layers to generate a set of final object embeddings The method according to any one of claims 1 to 12, comprising: **Claim 14** Using the localization sub-network to process the object embedding to generate localization data that defines the corresponding region of the image The method according to claim 13, comprising generating a set of offset coordinates that define an offset of the corresponding region of the image from the location of the image patch corresponding to the object embedding **Claim 15** The method according to any one of claims 1 to 14, when dependent on claim 2, wherein the text encoding sub-network includes one or more self-attention neural network layers **Claim 16** For one or more of the object embeddings Determining that the region of the image corresponding to the object embedding depicts an object included in the category represented by the query embedding based on the classification score distribution for the object embedding The method according to any one of claims 1 to 15, further comprising: **Claim 17** For one or more of the query embeddings, obtaining the query embedding comprises: obtaining one or more query images, each query image including a respective target region depicting an example of a target object, and generating, for each query image, a respective embedding of the target object in the query image, and generating the query embedding by combining the embeddings of the target object within the query images, The method according to any one of claims 1 to 16, comprising: **Claim 18** For each query image, generating the embedding of the target object in the query image comprises: processing the query image using the image encoding sub-network to generate a set of object embeddings for the query image, and processing each object embedding for the query image using the localization sub-network to generate localization data defining a corresponding region of the query image, and for each object embedding for the query image, determining a respective measure of overlap between (i) the target region of the query image depicting the example of the target object and (ii) the region of the query image corresponding to the object embedding for the query image, and selecting, based on the measure of overlap, an object embedding from the set of object embeddings for the query image as the embedding of the target object in the query image, The method according to claim 17, comprising: **Claim 19** Generating the query embedding by combining the embeddings of the target object in the query image comprises: averaging the embeddings of the target object in the query image The method according to claim 18, comprising: **Claim 20** one or more computers, and one or more storage devices communicatively coupled to the one or more computers, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the respective methods according to any one of claims 1 to 19, A system comprising: **Claim 21** One or more non-transitory computer storage media that, when executed by one or more computers, store instructions that cause the one or more computers to perform the operations of each of the methods of any one of claims 1 to 19.
Citation Information
Patent Citations
Medical visual question answering
US20220130499A1