Instant level scene recognition using visual language model
By integrating object recognition models and visual language models to generate detailed descriptions of images, the solution addresses the limitations of current technologies in searching and retrieving information from images, achieving improved accuracy and effectiveness.
Patent Information
- Application Number
- JP2024074824
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-27
- Filing Date
- 2024-05-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current technologies face challenges in effectively searching for and retrieving detailed information about specific objects or scenes from images, due to limitations in text search capabilities and the inability to express complex concepts accurately.
The proposed solution involves using a combination of object recognition models and visual language models to process images, generating fine-grained object recognition outputs and language outputs, which are then integrated to produce extended language outputs that provide detailed descriptions of the objects and scenes.
This approach enables the generation of detailed and accurate descriptions of images, improving the effectiveness of search results and content generation by leveraging instance-level object recognition and scene understanding.
Smart Images

Figure 2025073967000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates generally to visual language model output enhancement based on instance-level object recognition. More particularly, the present disclosure relates to leveraging visual language model processing in conjunction with object recognition processing to generate detailed output that can be utilized for search result determination and / or generative model content generation. [Background technology]
[0002] Understanding the world at large can be difficult. Whether an individual is trying to understand what an object in front of them is, determining where else that object may be found, and / or determining where an image on the Internet was captured from, text search alone can be difficult. In particular, users may struggle to determine what words to use. Additionally, words may not be descriptive and / or rich enough to generate the desired results.
[0003] Additionally, the content being requested by the user may not be readily available to the user based on the user not knowing where to search, based on the content's storage location, and / or based on the content not existing. A user may be requesting search results based on imagined concepts without an obvious way to express the imagined concepts. Summary of the Invention [Means for solving the problem]
[0004] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned through practice of the embodiments.
[0005] One exemplary aspect of the present disclosure is directed to a computer-implemented method. The method may include obtaining, by a computing system including one or more processors, image data. The image data may include an input image. The method may include processing, by the computing system, the input image with an object recognition model to generate a fine-grained object recognition output. The fine-grained object recognition output may describe identity details regarding objects shown in the input image. The method may include processing, by the computing system, the input image with a visual language model to generate a linguistic output. The linguistic output may include a set of predicted words predicted to describe the input image. In some implementations, the set of predicted words may include coarse-grained terms that describe predicted identities of objects shown in the input image. The method may include processing, by the computing system, the fine-grained object recognition output and the linguistic output to generate an augmented linguistic output. The augmented linguistic output may include a set of predicted words in which the coarse-grained terms are replaced with the fine-grained object recognition output.
[0006] In some implementations, processing an input image with an object recognition model by a computing system to generate fine-grained object recognition output may include detecting an object in the input image, generating an object embedding, determining an image cluster associated with the object embedding, and processing a web resource associated with the image cluster to determine identifying details about the object. Generating the object embedding may include generating a bounding box associated with a location of the object in the input image, generating image segments based on the bounding box, and processing the image segments with the embedding model to generate the object embedding.
[0007] In some implementations, processing the fine-grained object recognition output and the linguistic output by the computing system to generate the extended linguistic output may include processing the linguistic output by the computing system to determine a plurality of text tokens associated with features in the input image, determining by the computing system that a particular token of the plurality of text tokens is associated with the object, and replacing by the computing system the particular token with the fine-grained object recognition output. Determining by the computing system that a particular token of the plurality of text tokens is associated with the object may include processing the fine-grained object recognition output with an embedding model to generate an instance-level embedding, processing by the computing system the plurality of text tokens with the embedding model to generate a plurality of token embeddings, and determining by the computing system that the instance-level embedding is associated with the particular embedding associated with the particular token.
[0008] In some implementations, the method may include processing, by the computing system, the augmented language output with a second language model to generate a natural language response to the augmented language output. The natural language response may include additional information associated with the augmented language output. The coarse-grained term may include an object type. The fine-grained object recognition output may include detailed identification information of the object. In some implementations, the method may include providing, by the computing system, the augmented language output in an augmented reality experience. The augmented reality experience may include the augmented language output overlaid on a live video feed of an environment.
[0009] Another exemplary aspect of the present disclosure is directed to a computing system for image captioning. The system may include one or more processors and one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include obtaining image data. The image data may include an input image. The operations may include processing the input image with an object recognition model to generate an object recognition output. The object recognition output may describe identity details regarding an object shown in the input image. The operations may include processing the input image with a visual language model to generate a linguistic output. The linguistic output may include a set of words predicted to describe the input image. In some implementations, the set of words may include terms that describe predicted identities of objects shown in the input image. The operations may include processing the object recognition output and the linguistic output with the visual language model to generate an augmented linguistic output. The augmented linguistic output may include a set of words with the terms replaced with the object recognition output.
[0010] In some implementations, the input image may describe an object in an environment with one or more additional objects. The object recognition output may be associated with the object. The language output may be associated with the object and the environment with one or more additional objects. The visual language model may be trained on a training dataset including a plurality of image caption pairs. The plurality of image caption pairs may include a plurality of training images and a plurality of respective captions associated with the plurality of training images. The input image may be processed in parallel with the object recognition model and the visual language model to perform parallel determination of the object recognition output and the language output. The visual language model may include one or more text encoders, one or more image encoders, and one or more decoders. In some implementations, the object recognition model may include one or more classification models. The object recognition output may include instance level object recognition associated with the object. The language output may include scene understanding associated with the input image.
[0011] Another exemplary aspect of the present disclosure is directed to one or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations may include obtaining image data. The image data may include an input image. The operations may include processing the input image to determine an object recognition output. The object recognition output may describe identity details regarding an object shown in the input image. The operations may include processing the input image with a visual language model to generate a linguistic output. The linguistic output may include a set of words predicted to describe the input image. In some implementations, the set of words may include terms that describe predicted identities of objects shown in the input image. The operations may include processing the object recognition output and the linguistic output to generate an augmented linguistic output. The augmented linguistic output may include a set of words with the terms replaced with the object recognition output. The operations may include determining one or more search results associated with the augmented linguistic output. The one or more search results may be associated with one or more web resources.
[0012] In some implementations, processing the input image to determine the object recognition output may include processing the input image with a search engine to determine text data describing object identification. Processing the input image to determine the object recognition output may include processing the input image with an embedding model to generate image embeddings and determining one or more object labels based on the image embeddings. The one or more object labels may include identification details about the object shown in the input image. In some implementations, determining one or more search results associated with the augmented language output may include determining that a plurality of search results are responsive to a search query that comprises the augmented language output. The operations may further include providing the plurality of search results for display.
[0013] Other aspects of the present disclosure are directed to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.
[0014] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain related principles.
[0015] Detailed descriptions of embodiments directed to persons skilled in the art are set forth herein and refer to the accompanying figures. [Brief description of the drawings]
[0016] [Figure 1] FIG. 1 is a block diagram of an exemplary detailed image captioning system, according to an exemplary embodiment of the present disclosure. [Diagram 2] FIG. 1 is a block diagram of an exemplary generative model-enabled search system, according to an exemplary embodiment of the present disclosure. [Diagram 3] FIG. 2 is a flowchart diagram of an example method for performing augmented language output generation, according to an example embodiment of the present disclosure. [Figure 4A] FIG. 1 is a diagram of an exemplary fine-grained object recognition using scene understanding, according to an exemplary embodiment of the present disclosure. [Figure 4B] FIG. 1 is a diagram of an exemplary fine-grained object recognition using scene understanding, according to an exemplary embodiment of the present disclosure. [Figure 4C] FIG. 1 is a diagram of an exemplary fine-grained object recognition using scene understanding, according to an exemplary embodiment of the present disclosure. [Diagram 5] FIG. 1 is a diagram of an exemplary generative model response system, according to an exemplary embodiment of the present disclosure. [Figure 6A] 1 is a diagram of an exemplary scene description generation template, according to an exemplary embodiment of the present disclosure. [Figure 6B] 1 is a diagram of an exemplary scene description generation template, according to an exemplary embodiment of the present disclosure. [Figure 7] FIG. 2 is a flowchart diagram of an exemplary method for performing object-aware scene recognition, according to an exemplary embodiment of the present disclosure. [Figure 8] FIG. 1 is a flowchart diagram of an exemplary method for performing a generative model-assisted search, according to an exemplary embodiment of the present disclosure. [Figure 9A] FIG. 1 is a block diagram of an exemplary computing system for performing instance-level scene recognition, according to an exemplary embodiment of the present disclosure. [Figure 9B] FIG. 1 is a block diagram of an exemplary computing system for performing instance-level scene recognition, according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] Reference numbers that are repeated among the figures identify like features in various implementations.
[0018] In general, the present disclosure is directed to systems and methods for detailed instance-level scene recognition. In particular, the systems and methods disclosed herein may leverage an object recognition system and a visual language model to generate detailed captions, queries, and / or prompts associated with an input image. For example, an object recognition system (e.g., a system with one or more object recognition models) may process an input image to generate an object recognition output that describes the recognition of a particular object of a particular object class. The object recognition output may describe detailed identification information of unique object instances. Additionally, a visual language model may process an input image to generate a linguistic output that describes the scene recognition for a scene shown in the input image. The linguistic output may include details that describe an environment and one or more objects in the environment. The linguistic output may not include the granularity and / or specificity of the object recognition output. The object recognition model and the visual language model may process the input image in parallel to reduce latency. The object recognition output and the linguistic output may then be processed to generate an extended linguistic output that describes the scene recognition of the linguistic output along with the specificity and / or particularity of the object recognition output. For example, the language output may include a particular object class identity for an object shown in the input image, while the extended language output may include specific indications of the instance-level identity of the shown object (e.g., brand and model names for products, names of shown people, names for works of art, and / or species and subspecies identities for plants or animals).
[0019] The augmented language output may then be utilized as a query and / or prompt to obtain additional information associated with the scene and / or object shown in the input image. In some implementations, an input text may be received along with the input image, and the language output and / or the augmented language output may be generated based in part on the input text. Thus, a user may ask a question about the shown scene, and detailed scene recognition may be generated, and detailed queries and / or prompts may be generated that include the semantic intent of the question and the recognition information of the augmented language output. The augmented language output may be processed using a search engine and / or a generative model (e.g., a large-scale language model, a visual language model, an image generation model, etc.) to generate additional information that may be responsive to the input question.
[0020] The visual language model can leverage learned image and language associations to generate natural language captions for images; however, the visual language model may struggle with details including object specificity. Lack of specificity may lead to the generation of generalized queries and / or prompts, thereby failing to provide results that are specific and / or applicable to the features shown in the image. For example, a user may provide an image with the question "how do I take care of this?" The visual language model may process the image and determine that the image shows a plant, which may be leveraged to generate an improved query of "what do plants need to stay alive and grow?" The improved query may be processed to determine search results that may be associated with general care instructions for plants, which may include watering twice a week, half a day of direct sunlight, and loamy soil. However, generalized care instructions may not be suitable for the specific plant shown in the image (e.g., succulents (e.g., Agave plants) may need less water and different soil, and Celastrus revoluta may prefer shade over direct sunlight). Thus, the use of generalized information for object classes may be detrimental to care and counter to the original purpose of the input.
[0021] The systems and methods disclosed herein can process images in parallel using a visual language model and a fine-grained object recognition model to generate output that is scene-aware and object-aware while being formatted in a natural language format. The parallel processing can be separate and independent such that the scene-aware output and the object-aware output are determined separately and without the influence of the other. Token replacement can be utilized to replace the coarse-grained object recognition of the visual language model (e.g., object class recognition (e.g., plants, humans, cars, buildings, etc.)) with the fine-grained recognition of the instance-level object recognition system (e.g., unique object recognition indicating specific object identities (e.g., Tiger Lily, George Washington, Model T Soft Top Convertible with 5L engine, Monticello, etc.)). For example, the systems and methods can include processing an input image using the object recognition system to generate an object recognition output that describes identity details for a specific object shown in the input image. The identity details can include instance-level identity information describing unique and detailed identity information for the object. The systems and methods may also use the visual language model to process the input image to generate a linguistic output that describes the scene recognition for the entire scene shown in the input image. The scene recognition may be less specific than the object recognition output. Thus, the systems and methods may process the object recognition output and the linguistic output to generate an augmented linguistic output that leverages the specificity of the scene recognition and object recognition outputs of the linguistic output.
[0022] Pairing instance-level object recognition with visual language model processing can be utilized to generate detailed captions, queries, and / or prompts. Combining scene understanding with instance understanding can be leveraged for image retrieval, image indexing, automatic content generation and / or understanding, and / or other image understanding tasks. For example, the augmented language output can be leveraged as and / or to generate detailed queries and / or detailed prompts to obtain and / or generate additional information. The specificity can lead to improved adaptation of search results and / or generated prompts.
[0023] Different objects within the same object class may have different properties for maintenance, use, assembly, and / or repair, which may cause a generalized search query to produce search results that may not be relevant for that particular object. Thus, by leveraging scene understanding along with object understanding, output can be generated that can be processed with a search engine and / or machine learning models to generate object-aware information.
[0024] A multimodal large-scale language model (e.g., a large-scale visual language model) can be tuned and / or trained to have a rough understanding of an image. For example, by processing an image with a large-scale language model, it may be possible to output "This is a black dog sitting on a beach". However, an object recognition system can be trained and / or configured to recognize objects at instance-level granularity. In the same image, the object recognition system can recognize the breed of dog as an Australian Kelpie and the beach as Bondi Beach. When combining these two systems, the system and method can teach and / or condition the large-scale language model that the scene includes an Australian Kelpie sitting on Bondi Beach. The large-scale language model can then learn and / or be prompted to describe the scene at instance-level granularity. The systems and methods disclosed herein can be utilized to recognize every product in an aisle as a user passes by the products, and can then help the user find products that meet their dietary restrictions and / or other preferences and criteria.
[0025] In some implementations, object recognition and / or scene recognition techniques (including visual search) may be utilized to tune and / or train a visual language model for instance-level recognition, which may include training the visual language model for attribute specificity based on output from the visual search.
[0026] The systems and methods disclosed herein may be utilized to process a number of different data types (e.g., image data, text data, video data, audio data, statistical data, graph data, latent coding data, and / or multimodal data) to generate output (e.g., image data, text data, video data, audio data, statistical data, graph data, latent coding data, and / or multimodal data) that may be in a number of different data formats. For example, input data may include video, which may be processed to generate a video summary, which may include a natural language summary, a timeline, a flowchart, an audio file in the form of a podcast, and / or a comic book. An object recognition system may be utilized for object-specific details, while a scene understanding model (e.g., a visual language model) may be utilized for scene recognition and / or frame group understanding. In some implementations, one or more additional models may be utilized for context understanding. For example, a hierarchical video encoder may be utilized for frame understanding, frame sequence understanding, and / or full video understanding. Audio input processing may include utilization of a text-to-speech model, which may be implemented as part of a language model.
[0027] The systems and methods of the present disclosure provide several technical effects and benefits. As an example, the systems and methods may be utilized to generate instance-level scene recognition output. In particular, the systems and methods disclosed herein may leverage a visual language model in parallel with an object recognition system to generate natural language output that is both scene-aware and object-specific. The augmented language output may then be utilized as a query for search and / or a prompt for generative model content generation.
[0028] Another exemplary technical effect and benefit relates to improved computational efficiency and increased capabilities of a computing system. For example, a technical benefit of the disclosed systems and methods is the ability to reduce computational resources required for detailed query and / or detailed prompt generation. In particular, training and / or tuning a language model for instance-level object recognition may be computationally expensive and may require large training datasets. Additionally, training a language model for such specificity may be computationally expensive for model inference. The process disclosed herein may reduce training time and resource costs for detailed image captioning to generate instance-level scene recognition output. In some implementations, to reduce latency, an input image may be processed in parallel with an object recognition model and a visual language model. Alternatively and / or additionally, when performed by a computing device with limited processing power, an input image may be processed at different times with an object recognition model and a visual language model.
[0029] Exemplary embodiments of the present disclosure will now be described in further detail with reference to the figures.
[0030] 1 illustrates a block diagram of an exemplary detailed image captioning system 10 according to an exemplary embodiment of the present disclosure. In some implementations, the detailed image captioning system 10 is configured to receive and / or obtain a set of input data including image data 12 describing an environment with one or more objects, and generate, determine, and / or provide an augmented language output 22 describing object instance level scene recognition as a result of receiving the image data 12. Thus, in some implementations, the detailed image captioning system 10 may include a visual language model 18 operable to perform scene recognition and an object recognition block 14 operable to perform object recognition.
[0031] In particular, the detailed image captioning system 10 can acquire input data, which can include image data 12 that describes one or more input images. The one or more input images can describe an environment and one or more objects. The environment can include a room, a landscape, a city, a town, a sky, and / or other environment. In some implementations, the environment describes a user environment generated with one or more image sensors of a user computing device. The one or more objects can include products, people, plants, animals, works of art, structures, landmarks, and / or other objects.
[0032] The image data 12 may be processed using parallel processing pipelines. A first pipeline may process the image data 12 to generate an object recognition output (e.g., instance-level object recognition). A second pipeline may process the image data 12 to generate a linguistic output 20 that describes the scene recognition. The detailed image captioning system 10 may then process the output of the pipelines to generate an extended linguistic output 22 that describes a detailed image caption.
[0033] For example, object recognition block 14 (e.g., an object recognition system) may process image data 12 to generate object recognition output 16 that describes fine-grained object recognition. Object recognition output 16 may include identification details for one or more objects shown in one or more input images. The identification details may describe instance-level recognition associated with a particular object in a particular object class, which may include product model names, plant species and / or subspecies, particular people's names, names of artworks, location names, and / or other instance-level identifiers.
[0034] The object recognition block 14 may include an object recognition model, which may include one or more machine-learned models. The object recognition model may be trained and / or configured to process an image, detect an object, segment a portion of the image that includes the object, and then process the image segments to generate a recognition output. The object recognition model may include a detection model that processes an input image to generate a bounding box that indicates the location of the detected object. A segmentation model of the object recognition model may then segment the detected object based on the bounding box to generate an image segment for the detected object. The image segments may then be processed with a classification model of the object recognition model to generate an object classification. The object classification may then be processed to generate the object recognition output 16.
[0035] Alternatively and / or additionally, the object recognition block 14 may include one or more embedding models. The one or more embedding models may process the image data 12 and / or image segments to generate one or more image embeddings. The one or more image embeddings may be utilized to query the embedding space for similar embeddings, nearby embeddings, embedding clusters, and / or embedding labels (e.g., labels that describe learned properties for the learned distribution in the embedding space). The similar embeddings, nearby embeddings, embedding clusters, and / or embedding labels may be utilized to obtain multiple web resources determined to be associated with an object shown in one or more input images. The multiple web resources may be processed to determine details associated with the object, which may include product name, object origin, object listing, object location, other instances of the object, other identifiers, and / or other details. The details may then be utilized to generate the object recognition output 16. The multiple web resources may be sources of content items that are embedded to generate similar embeddings, nearby embeddings, and / or other embeddings in the embedding clusters.
[0036] The object recognition block 14 may generate an object recognition output 16 for each object shown in the input image. Alternatively and / or additionally, the object recognition block 14 may determine a focus object and / or an object of interest based on object location, object size, image semantics, image focus, occurrence in the sequence of input images, and / or other contextual attributes.
[0037] The visual language model 18 can process the image data 12 to generate linguistic output 20. The linguistic output 20 can be a natural language text string that describes scene recognition for a scene (e.g., an environment and one or more objects) shown in one or more input images. The linguistic output 20 can include coarse-grained recognition output associated with a location and / or one or more objects, which can include class identification information for the location and one or more objects.
[0038] The visual language model 18 may include a language model trained, configured, and / or tuned to process multimodal data, which may include tuning for an image understanding task. For example, the visual language model 18 may be trained on a training dataset that includes image caption pairs. The image caption pairs may include training images and respective training captions for the particular training images. The training and / or tuning may include processing the training images with the visual language model 18 to generate predicted text strings. The predicted text strings and respective training captions may be processed to evaluate a loss function to generate a gradient descent. The gradient descent may then be backpropagated to adjust one or more parameters of the visual language model.
[0039] Alternatively and / or additionally, the visual language model may include a text encoder and an image encoder, which may be jointly trained and / or jointly tuned, to encode input data, which may then be processed with a decoder to generate the visual language model output. In some implementations, an image embedding model may be trained to process images and generate image embeddings, which may then be processed with the large-scale language model. The image embeddings may describe representations associated with image features.
[0040] The object recognition output 16 and the language output 20 may then be processed to generate an augmented language output 22. The augmented language output 22 may include the object recognition granularity of the object recognition output 16 along with the scene understanding of the language output 20. In some implementations, the augmented language output 22 may describe detailed image captions for one or more input images.
[0041] 2 illustrates a block diagram of an exemplary generative model-enabled search system 200, in accordance with an exemplary embodiment of the present disclosure. The generative model-enabled search system 200 is similar to the detailed image captioning system 10 of FIG. 1, except that the generative model-enabled search system 200 further includes a search result 226 determination and a generative model 228 for generating a generative response 230.
[0042] In particular, the generative model-enabled search system 200 can obtain input data, which may include image data 212 describing one or more input images (e.g., one or more images of beef wellington on a platter with kale on a red tablecloth) and text data 232 describing a request for specific information (e.g., a request for a recipe for the depicted beef wellington). The one or more input images can describe an environment and one or more objects. The environment may include a room, a landscape, a city, a town, a sky, and / or other environment. In some implementations, the environment describes a user environment generated with one or more image sensors of a user computing device. The one or more objects may include products, people, plants, animals, works of art, structures, landmarks, and / or other objects.
[0043] The image data 212 and / or text data 232 may be processed using one or more image processing pipelines. A first pipeline may process the image data 212 to generate object recognition output (e.g., instance-level object recognition). A second pipeline may process the image data 212 and / or text data to generate linguistic output 220 describing the scene recognition. These pipelines may be executed in parallel, serially, and / or in a self-attention loop. The generative model-assisted search system 200 may then process the output of the pipelines and / or the text data 232 to generate an augmented linguistic output 222 describing a detailed image caption. The augmented linguistic output 222 may then be processed using a search engine and / or a generative model 228 to obtain and / or generate additional data (e.g., one or more search results 226 and / or one or more model-generated responses 230).
[0044] For example, the object recognition block 214 (i.e., object recognition system) may process the image data 212 to generate an object recognition output 216 describing fine-grained object recognition. The object recognition output 216 may include identity details about one or more objects shown in one or more input images (e.g., "beef wellington," Jane Doe, Mona Lisa, Sixteenth Chapel, Washington Monument, Brand X Model YZ Smartphone, etc.). The identity details may describe instance-level recognition (e.g., recognition for that unique object shown) associated with a particular object in a particular object class, which may include product model names, plant species and / or subspecies, names of particular people, names of works of art, location names, and / or other instance-level identifiers.
[0045] The object recognition block 214 may include an object recognition model, which may include one or more machine-learned models (e.g., one or more embedding models, one or more detection models, one or more segmentation models, one or more classification models, one or more semantic understanding models, one or more feature extractors, and / or one or more other models). The object recognition model may be trained and / or configured to process an image, detect an object, segment a portion of the image that includes the object (e.g., segment an image portion within the object and / or segment the object from the image), and then process the image segments to generate a recognition output. The object recognition model may include a detection model that processes an input image to generate a bounding box that indicates the location of the detected object. A segmentation model of the object recognition model may then segment the detected object based on the bounding box to generate an image segment for the detected object. The image segments may then be processed with a classification model of the object recognition model to generate an object classification. The object classification may then be processed to generate the object recognition output 216.
[0046] Alternatively and / or additionally, the object recognition block 214 may include one or more embedding models. The one or more embedding models may process the image data 212 and / or image segments to generate one or more image embeddings. The one or more image embeddings may be utilized to query the embedding space for similar embeddings, nearby embeddings, embedding clusters, and / or embedding labels (e.g., labels that describe learned properties for the learned distributions in the embedding space). The similar embeddings, nearby embeddings, embedding clusters, and / or embedding labels may be utilized to obtain multiple web resources determined to be associated with the object shown in the one or more input images. The multiple web resources may be processed to determine details associated with the object, which may include product name, object origin, object listing, object location, other instances of the object, other identifiers, and / or other details. The details may then be utilized to generate the object recognition output 216. Multiple web resources may be sources of content items that are embedded to generate similar embeddings, nearby embeddings, and / or other embeddings in an embedding cluster.
[0047] The object recognition block 214 may generate an object recognition output 216 for each object shown in the input image. Alternatively and / or additionally, the object recognition block 214 may determine a focus object and / or an object of interest based on object location, object size, image semantics, image focus, occurrence in the sequence of input images, and / or other contextual attributes. In some implementations, the particular object selected for processing may be based on the text data 232 (e.g., "what are recipes for this item?" would cause a food item to be processed, while "what is that on the right?" would cause the object on the right of the input image to be processed).
[0048] The visual language model 218 can process the image data 212 and / or the textual data 232 to generate the linguistic output 220. The linguistic output 220 can be a natural language text string describing a scene recognition for a scene (e.g., an environment and one or more objects) shown in one or more input images. The linguistic output 220 can include a coarse-grained recognition output associated with a location and / or one or more objects, which can include class identification information for the location and one or more objects. In some implementations, the linguistic output 220 can be a scene recognition output including structure, format, and / or additional language based on processing the textual data 232. For example, the textual data 232 can include "What is the origin of this food item?" and the linguistic output can include "The image depicts a formal dinner, in which the food item is a pastry, which is presented on a plate with a vegetable on a red table." Object recognition outputs for an example image may include "beef wellington," "kale," and / or "maroon tablecloth."
[0049] The visual language model 218 may include a language model trained, configured, and / or tuned to process multimodal data, which may include tuning for an image understanding task. For example, the visual language model 218 may be trained on a training dataset that includes image caption pairs. The image caption pairs may include training images and respective training captions for the particular training images. The training and / or tuning may include processing the training images with the visual language model 218 to generate predicted text strings. The predicted text strings and respective training captions may be processed to evaluate a loss function to generate a gradient descent. The gradient descent may then be backpropagated to adjust one or more parameters of the visual language model.
[0050] Alternatively and / or additionally, the visual language model may include a text encoder and an image encoder, which may be jointly trained and / or jointly tuned, to encode input data, which may then be processed with a decoder to generate the visual language model output. In some implementations, an image embedding model may be trained to process images and generate image embeddings, which may then be processed with the large-scale language model. The image embeddings may describe representations associated with image features.
[0051] The object recognition output 216, the language output 220, and / or the text data 232 may then be processed with an augmentation model 224 to generate an augmented language output 222. The augmentation model 224 may include a visual language model, other language models, and / or other generative models. The augmentation model 224 may be trained and / or tuned to identify text tokens associated with the same object and to augment the language output to replace coarse-grained terms of the language output 220 with fine-grained terms of the object recognition output 216. The augmented language output 222 may include the object recognition granularity of the object recognition output 216 as well as the scene understanding of the language output 220. In some implementations, the augmented language output 222 may describe detailed image captions for one or more input images. For example, augmented language output 222 may include, "The image depicts a formal dinner, in which the food item is a beef wellington, which is presented on a plate with kale on a maroon tablecloth."
[0052] In some implementations, the augmented linguistic output 222 and / or the textual data 232 may be processed with a search engine to determine one or more search results 226. The augmented linguistic output 222 and / or the textual data 232 may include and / or be formatted as a query (e.g., "recipe for a beef wellington"). The one or more search results 226 may be responsive to the query posed by the augmented linguistic output 222 and / or the textual data 232 (e.g., a web page, a video, and / or a book with a beef wellington recipe).
[0053] Alternatively and / or additionally, the augmented linguistic output 222 and / or the text data 232 may be processed with a generative model 228 (e.g., a large-scale language model, a text-to-image model, and / or other generative model) to generate one or more model-generated responses 230. The augmented linguistic output 222 and / or the text data 232 may include and / or be formatted as a prompt (e.g., “generate step by step instructions for a beef wellington”). The one or more model-generated responses 230 may be responsive to the prompt in the augmented linguistic output 222 and / or the text data 232 (e.g., the model-generated response may include step-by-step instructions for making beef wellington based on processing one or more beef wellington recipes and may include images with the steps outlined in the text (e.g., images pulled from a web resource and / or images generated using a text-to-image generation model)).
[0054] 3 shows a flow chart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. Although FIG. 3 shows steps performed in a particular order for purposes of illustration and explanation, the method of the present disclosure is not limited to the specifically shown order or arrangement. Various steps of method 300 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0055] At 302, the computing system may acquire image data. The image data may include an input image. The input image may include one or more objects in an environment. The environment may include a room, a landscape, and / or other environment. The one or more objects may include people, structures, animals, plants, monuments, art, products, and / or other objects. The computing system may acquire and / or generate the image data using a computing device, which may include a mobile computing device, a smart wearable, a smart appliance, and / or other computing device.
[0056] At 304, the computing system can process the input image with the object recognition model to generate a fine-grained object recognition output. The fine-grained object recognition output can describe identity details about the object shown in the input image. The object can include a product, and the fine-grained object recognition output can include a unique product label, and the unique product label can include a model name, a registration number, a model number, a product specific name, and / or identity details. In some implementations, the object can include a person, and the fine-grained object recognition output can include a name of the person. Alternatively and / or additionally, the object can include a work of art (e.g., a sculpture, a painting, a photograph, etc.), and the fine-grained object recognition output can include a name for the work of art. The object recognition model can include one or more embedding models, one or more detection models, one or more feature extractors, one or more classification models, one or more segmentation models, one or more search engines, and / or one or more other models.
[0057] In some implementations, processing the input image with the object recognition model to generate a fine-grained object recognition output may include detecting an object in the input image, generating an object embedding, determining an image cluster associated with the object embedding, and processing a web resource associated with the image cluster to determine identifying details about the object. Generating the object embedding may include generating a bounding box associated with a location of the object in the input image, generating an image segment based on the bounding box, and processing the image segment with the embedding model to generate the object embedding. Additionally and / or alternatively, the object embedding may be utilized to search the embedding space for nearby embeddings and / or similar embeddings. In some implementations, one or more embedding clusters may be determined to be associated with the object embedding. The nearby embeddings, similar embeddings, and / or embedding clusters may be associated with image embeddings, text embeddings, document embeddings, multimodal embeddings, and / or other embeddings. Content items and / or web resources associated with nearby embeddings, similar embeddings, and / or embedding clusters may be retrieved and / or processed to determine identity details. Identification details may include precise names (and / or classifications) for particular objects. In some implementations, an input image may show multiple objects and multiple object recognition outputs may be generated based on the multiple objects.
[0058] At 306, the computing system can process the input image with the visual language model to generate a linguistic output. The linguistic output can include a set of predicted words predicted to describe the input image. The set of predicted words can include coarse-grained terms describing a predicted identity of an object shown in the input image. In some implementations, the coarse-grained terms can include an object type. The fine-grained object recognition output can include a detailed identity of the object. The set of predicted words can include one or more sentences including the coarse-grained terms. The coarse-grained terms can include a generalized description of the fine-grained object recognition output (e.g., the coarse-grained terms can include "vacuum," while the fine-grained object recognition output can include "an RF-600 cordless Dilred XL Vacuum"). The visual language model may include one or more transformer models, one or more autoregressive language models, one or more image encoder models, one or more text encoder models, one or more decoder models, one or more diffusion models, and / or one or more other models. The visual language model may be trained on image caption pairs, trained with contrastive learning, may include a prefix language model, may include masked language modeling, may include image text matching, may include learned sequence representations, may be trained with black-box optimization, and / or may include multimodal fusion with cross-attention. In some implementations, the visual language model may be trained on multiple templates, multiple scene types, multiple object classes, and / or multiple natural language examples.
[0059] At 308, the computing system can process the fine-grained object recognition output and the linguistic output to generate an augmented linguistic output. The augmented linguistic output may include a set of predicted words in which the coarse-grained terms are replaced with the fine-grained object recognition output. The augmented linguistic output may be generated by processing the fine-grained object recognition output and the linguistic output with a visual language model, an augmentation model, a natural language processing model, and / or one or more other models. The replacement may be determined by determining that the coarse-grained terms are associated with objects described by the fine-grained object recognition output. Alternatively and / or additionally, the augmented linguistic output may include a different structure, syntax, phrasing, and / or instructions than the linguistic output based on the processing of the fine-grained object recognition output.
[0060] In some implementations, processing the fine-grained object recognition output and the language output to generate an extended language output may include processing the language output to determine a plurality of text tokens associated with features in the input image, determining that a particular token of the plurality of text tokens is associated with the object, and replacing the particular token with the fine-grained object recognition output. In some implementations, determining that a particular token of the plurality of text tokens is associated with the object may include processing the fine-grained object recognition output with an embedding model to generate an instance-level embedding, processing the plurality of text tokens with the embedding model to generate a plurality of token embeddings, and determining that the instance-level embedding is associated with the particular embedding associated with the particular token.
[0061] In some implementations, the computing system can process the augmented language output with a second language model to generate a natural language response to the augmented language output. The natural language response may include additional information associated with the augmented language output. In some implementations, the augmented language output can be processed with a visual language model to generate a response. The response may include text data, image data, audio data, latent coding data, and / or multimodal data. The language output, object recognition output, augmented language output, and / or the response may be conditional and / or based on input text, which may be obtained with the input image.
[0062] Additionally and / or alternatively, the computing system can provide augmented language output in an augmented reality experience. The augmented reality experience may include augmented language output overlaid on top of a live video feed of the environment. The augmented reality experience may be provided via a viewfinder of a mobile computing device, via a smart wearable, and / or via other computing devices. The augmented reality experience may be leveraged to ask about and receive additional information about the user's environment via the augmented reality interface.
[0063] 4A-4C show diagrams of an exemplary fine-grained object recognition with scene understanding according to an exemplary embodiment of the present disclosure. In particular, FIG. 4A shows processing input text 402 and an input image 404 to generate an object-aware prompt to generate a model-generated response 410.
[0064] The input text 402 may include "How much water does this need?" The input image 404 may include a plant in a flower pot on a window ledge. The input text 402 and input image 404 may be obtained via a mobile computing device, a smart wearable, and / or other computing device. The input may be obtained via one or more user interfaces, which may include a viewfinder interface, a search interface, an augmented reality interface, and / or an assistant interface.
[0065] The input image 404 may be processed using an object recognition system and a visual language model to generate recognition data 406 describing the object recognition and scene recognition. For example, the recognition data 406 may include an instance-level object recognition (e.g., Fasciated haworthia) generated using the object recognition system. Additionally, the recognition data 406 may include a predicted image caption (e.g., "aloe vera plant in a black pot on a window sill") generated using the visual language model, which may be associated with the scene recognition. In some implementations, the recognition data 406 may include web resources identified as associated with the object in the input image 404. For example, a web page with the title "How to care for a zebra succulent aka haworthia an easy step by step guide to everything you need to know to save a dying zebra succulent sundaze saltair" may be determined to be associated with an object in the input image 404 based on a visual search, which may include a reverse image search and / or an embedding-based search. The web resources may be processed to determine that the web page is associated with the plant "Fasciated haworthia." Entity recognition may then be utilized to determine and / or confirm fine-grained object recognition.
[0066] The recognition data 406 and the input text 402 may be processed to generate an enhanced prompt 408. The enhanced prompt may leverage the recognition data 406 and the input text 402 to generate a prompt that includes the intent of the input text 402 along with more information determined based on the input image 404. The enhanced prompt 408 may include, "How much water does Fasciated haworthia need?"
[0067] The improved prompt 408 may be processed with one or more generative models to generate a model-generated response 410. The model-generated response 410 may be responsive to the input text 402 and the input image 404. The model-generated response 410 may be generated by obtaining and processing web data associated with the search results for the improved prompt. The generative models may then leverage information in the web data to generate a summary of responses to the improved prompt 408. Improved prompt 408 might include, "Haworthia plants don't need to be watered often because they store water efficiently. You should only water them when the soil has been completely dry for a few days. This could be every two weeks, but in warmer months or warmer climates, it could be more often. In more humid environments, it may not be as frequent."
[0068] 4B illustrates processing input text 432 and input image 434 to generate an image-aware model-generated content item 444. Input text 432 can describe a prompt for generative model generation. For example, input text 432 can include "Write a haiku about this place." Input image 434 can show a location with houses and a landscape with grass, bushes, and trees.
[0069] The visual language model 436 may process the input text 432 and the input image 434 to generate a preliminary content item 438. The visual language model 436 may process the input image 434 to generate a scene recognition, which may be processed with the input text 432 to generate a preliminary content item 438. The preliminary content item 438 may be responsive to the prompt of the text input 432, including details from the scene recognition. The preliminary content item 438 may include "A quaint village in the countryside, with stone houses and pub. A peaceful place to stay."
[0070] The instance level recognition block 440 may process the input image 434 to generate an instance level recognition 442 of the indicated location. The instance level recognition block 440 may include a visual search system for fine-grained recognition of objects and / or locations. The instance level recognition 442 may include "Palant, Denmark."
[0071] The preliminary content items 438 and the instance level recognition 442 may then be processed to generate an image aware model generated content item 444. The image aware model generated content item 444 may be generated by processing the preliminary content items 438 and the instance level recognition 442 with a generative model (e.g., visual language model 436). In some implementations, the image aware model generated content item 444 may be generated based on the text token identifications and replacements. The image aware model generated content item 444 may include "Palant in the Denmark countryside, with stone houses and pub. A peaceful place to stay."
[0072] 4C illustrates processing an input image 452 and input text to generate an improved query 458. The input image 452 may show a particular humidifier. The input text may include a question about the product shown. For example, the input text may include "what's its use?"
[0073] The input image 452 may be processed using an object recognition system and a visual language model to generate recognition data 454. For example, the recognition data 454 may include instance-level object recognition (e.g., QLX Drip XA-Q2 and QLX Drip Humidifier) generated using the object recognition system. Additionally, the recognition data 454 may include predicted image captions (e.g., "white humidifier on a white background," which may be associated with scene recognition) generated using the visual language model. In some implementations, the recognition data 454 may include web resources identified as associated with the object in the input image 452. For example, a web page with the title "Qax Cool-Mist Humidifier, 1 Gal. - Clear & White" may be determined to be associated with the object in the input image 452 based on a visual search, which may include a reverse image search and / or an embedding-based search. The web resource may be processed to determine that the web page is associated with the product "QLX Drip Humidifier." Entity recognition may then be utilized to determine and / or confirm fine-grained object recognition. In some implementations, the recognition data 454 may include optical character recognition generated by performing optical character recognition on the input image 452 to determine that the input image 452 includes the text "QLX."
[0074] The recognition data 454 and the input text may be processed to generate an augmented language output 456. The augmented language output 456 may include an enhanced image caption, which may include: "This image is about a white humidifier on a white background. It is about QLX Drip XA-Q2 or QLX Drip Humidifier. Image comes from a page with title 'Qax Cool-Mist Humidifier, 1 Gal. - Clear & White'. Text on the image says: QLX." The enhanced image caption may be generated using a visual language model and may leverage the object recognition output of the object recognition system along with the scene recognition of the visual language model.
[0075] The augmented language output 456 and the input text may then be processed to generate an improved query 458. The improved query 458 may include "What is the use of QLX Drip Humidifier?" The improved query 458 may be generated by leveraging information from the augmented language output 456 to provide a more detailed identifier of what the user has a question about.
[0076] 5 illustrates a diagram of an example generative model response system 500, according to an example embodiment of the present disclosure. In particular, the generative model response system 500 can take an input image 504 and input text 502 to generate a model generated response 506. For example, the input text 502 can include "create a listing to sell this," the input image 504 can show a white chair in a room, and the model generated response 506 can include a model generated listing.
[0077] The generative model response system 500 can leverage an object recognition system to determine the particular product depicted in the input image 504. The generative model can then process the object recognition and the input text 502 to generate a model generated response 506. In some implementations, the generative model can utilize one or more application programming interfaces to obtain additional information associated with the recognized product that can be leveraged to generate a detailed product listing.
[0078] 6A-6B show diagrams of example scene description generation templates according to an example embodiment of the present disclosure. In particular, FIG. 6A can show a template 600 for query generation. The template 600 can be utilized to prompt generative models for question and answer tasks, image captioning, and / or query generation. The template 600 can be utilized to prompt generative models for masked language tasks.
[0079] The templates 600 may include a question and answer prompt template 602, an image captioning template 604, and a query template 606. The question and answer prompt template 602 may be utilized to instruct the generative model to generate questions and answer questions based on the input data. The image captioning template 604 may be utilized to instruct the generative model to generate image captions that fill mask tokens with details determined based on the input image, which may include details associated with the entire image, selected objects in the image, and / or other objects in the image. The query template 606 may be utilized to instruct the generative model to generate queries based on the multimodal input data.
[0080] In some implementations, the question and answer prompt template 602, the image captioning template 604, and / or the query template 606 may be processed to generate detailed prompts. For example, the question and answer prompt template 602 may be utilized as a preamble for prompt composition, the image captioning template 604 may be utilized as a scene description, and the query template 606 may be utilized to distill intent.
[0081] FIG. 6B illustrates different recognition output templates 650. The recognition output templates 650 may include a dominant object template 652, a multi-object template 654, and / or a scene description template 656. The dominant object template 652 may include an itemized table to be completed based on the output of the object recognition system and the visual language model. The multi-object template 654 may include a natural language template including masked tokens to be replaced by text tokens associated with the recognition data generated using the object recognition system and the visual language model. The scene description template 656 may include a natural language template in the form of a scene description that may be provided by an individual. The scene description template 656 may include masked tokens to be replaced by text tokens associated with the recognition data generated using the object recognition system and the visual language model.
[0082] 7 shows a flow chart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. Although FIG. 7 shows steps performed in a particular order for purposes of illustration and explanation, the method of the present disclosure is not limited to the specifically shown order or arrangement. Various steps of method 700 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0083] At 702, the computing system may obtain image data. The image data may include an input image. The input image may describe an object in an environment with one or more additional objects. In some implementations, the computing system may include text data along with the image data. The text data may describe prompts for the generative model (e.g., "generate a short story based on the object in this lake" and / or "what is origin of this structure?"). The text data may include input text that references the object and / or the environment.
[0084] At 704, the computing system may process the input image to generate an object recognition output. In some implementations, the computing system may process the input image using an object recognition model to generate the object recognition output. Alternatively and / or additionally, the computing system may process the input image using a visual search engine to determine one or more visual search results, which may then be processed to determine the object recognition output. The object recognition output may describe details of identity information about an object shown in the input image. The object recognition output may be associated with the object. In some implementations, the object recognition model may include one or more classification models. The object recognition output may include instance-level object recognition associated with the object. The computing system may segment a portion of the input image associated with the object, which will then be processed to generate the object recognition output. In some implementations, the image segmentation may be based on text data. The object recognition output may be generated based on visual search, optical character recognition, feature recognition, object classification, and / or one or more other techniques.
[0085] At 706, the computing system can process the input image with the visual language model to generate a scene recognition output. The scene recognition output can include a linguistic output. The linguistic output can include a predicted set of words predicted to describe the input image. The set of words can include terms describing a predicted identity of an object shown in the input image. The linguistic output can be associated with the object and an environment with one or more additional objects. In some implementations, the visual language model can be trained on a training dataset including a plurality of image caption pairs. The plurality of image caption pairs can include a plurality of training images and a plurality of respective captions associated with the plurality of training images. In some implementations, the visual language model can include one or more text encoders, one or more image encoders, and one or more decoders. The linguistic output can include a scene understanding associated with the input image. In some implementations, the computing system can process the image data and the text data with the visual language model to generate a linguistic output. The linguistic output can be based on a prompt of the text data. The linguistic output can be formatted as a generative model prompt, a query, a dialog message, a question, and / or a response. An input image may be processed in parallel with the object recognition model and the visual language model to perform parallel determination of the object recognition output and the linguistic output. In some implementations, the object recognition output may be determined independently of the determination of the linguistic output, and the linguistic output may be determined independently of the determination of the object recognition output. Both the object recognition model and the visual language model may process the input image separately to perform their respective determinations.
[0086] At 708, the computing system can process the object recognition output and the scene recognition output with the visual language model to generate an augmented language output. The augmented language output can include a set of words with terms replaced with the object recognition output. Alternatively and / or additionally, the augmented language output can include a query and / or prompt (e.g., a prompt of the language output) augmentation based on the object recognition output. In some implementations, the object recognition output, the scene recognition output, and / or the text data can be processed with a generative model to generate a model-generated response. The model-generated response can include text data, image data, audio data, latent coding data, and / or multimodal data.
[0087] 8 shows a flow chart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. Although FIG. 8 shows steps performed in a particular order for purposes of illustration and explanation, the method of the present disclosure is not limited to the specifically shown order or arrangement. Various steps of method 800 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0088] At 802, the computing system may acquire image data. The image data may include an input image. The input image may describe a particular environment and / or one or more particular objects. The input image may be acquired and / or generated using a visual search application. In some implementations, the image data may be acquired and / or generated using a smart wearable (e.g., a smart watch, smart glasses, a smart helmet, etc.). The image data may be acquired along with other input data, which may include text data and / or audio data associated with the question.
[0089] At 804, the computing system may process the input image to determine an object recognition output. The object recognition output may describe identifying details about an object shown in the input image. Processing the input image to determine the object recognition output may include processing the input image with a search engine to determine text data describing the object identifying information. The search engine may perform a feature search, an embedding-based search, an optical character recognition search, and / or other search techniques. In some implementations, the search engine may identify multiple search results responsive to the image query. The multiple search results may be parsed and / or processed to determine one or more object labels. The object labels may be based on processing the multiple search results with a semantic understanding model, a language model, and / or other models. In some implementations, the input image and / or the multiple search results may be embedded to determine text embeddings associated with the generated embeddings. The text embeddings may be decoded to determine the object labels.
[0090] Alternatively and / or additionally, processing the input image to determine the object recognition output may include processing the input image with the embedding model to generate image embeddings and determining one or more object labels based on the image embeddings. The one or more object labels may include identity details for objects shown in the input image. The object labels may be indexed labels associated with one or more data clusters.
[0091] At 806, the computing system can process the input image with the visual language model to generate a linguistic output. The linguistic output can include a predicted set of words predicted to describe the input image. The set of words can include terms describing predicted identities of objects shown in the input image. The visual language model can include a language model tuned to process image data and / or multimodal data. The visual language model can be one trained to encode images and output text. The visual language model can include multiple separate models utilized in series and / or in parallel.
[0092] At 808, the computing system can process the object recognition output and the linguistic output to generate an augmented linguistic output. The augmented linguistic output may include a set of words with terms replaced with the object recognition output. The object recognition output and / or the linguistic output may be processed with a generative model (e.g., a generative language model, a text-to-image generative model, etc.) to generate a model-generated output. The model-generated output may include additional information associated with one or more particular environments and / or one or more particular objects.
[0093] At 810, the computing system may determine one or more search results associated with the augmented language output. The one or more search results may be associated with one or more web resources. In some implementations, determining the one or more search results associated with the augmented language output may include determining that a plurality of search results are responsive to a search query that includes the augmented language output. Additionally and / or alternatively, the computing system may provide the plurality of search results for display.
[0094] 9A illustrates a block diagram of an example computing system 100 for performing instance-level scene recognition, according to an example embodiment of the present disclosure. The system 100 includes a user computing system 102, a server computing system 130, and / or a third-party computing system 150, communicatively coupled via a network 180.
[0095] The user computing system 102 may include any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0096] The user computing system 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing system 102 to perform operations.
[0097] In some implementations, the user computing system 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or otherwise include various machine-learned models, such as neural networks (e.g., deep neural networks) or other types of machine-learned models including nonlinear and / or linear models. The neural networks can include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.
[0098] In some implementations, one or more machine-learned models 120 may be received from the server computing system 130 over the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, the user computing system 102 may implement multiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel machine-learned model processing across multiple instances of input data and / or detected features).
[0099] More specifically, the one or more machine-learned models 120 may include one or more detection models, one or more classification models, one or more segmentation models, one or more augmentation models, one or more generative models, one or more natural language processing models, one or more optical character recognition models, and / or one or more other machine-learned models. The one or more machine-learned models 120 may include one or more transformer models. The one or more machine-learned models 120 may include one or more neural radiance field models, one or more diffusion models, and / or one or more autoregressive language models.
[0100] One or more machine-learned models 120 may be utilized to detect one or more object features. The detected object features may be classified and / or embedded. The classification and / or embedding may then be utilized to perform a search to determine one or more search results. Alternatively and / or additionally, one or more detected features may be utilized to determine that an indicator (e.g., a user interface element indicating the detected feature) will be provided to indicate that the feature has been detected. A user may then select the indicator to cause the feature classification, embedding, and / or search to be performed. In some implementations, the classification, embedding, and / or search may be performed before the indicator is selected.
[0101] In some implementations, the one or more machine-learned models 120 may process image data, text data, audio data, and / or latent coding data to generate output data, which may include image data, text data, audio data, and / or latent coding data. The one or more machine-learned models 120 may perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, context determination, action prediction, image correction, image enhancement, text enhancement, sentiment analysis, object detection, error detection, inpainting, video stabilization, audio correction, audio enhancement, and / or data segmentation (e.g., mask-based segmentation).
[0102] Additionally or alternatively, one or more machine-learned models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing system 102 according to a client-server relationship. For example, the machine-learned models 140 may be implemented by the server computing system 130 as part of a web service (e.g., a viewfinder service, a visual search service, an image processing service, an ambient computing service, and / or an overlay application service). Thus, one or more models 120 may be stored and implemented in the user computing system 102 and / or one or more models 140 may be stored and implemented in the server computing system 130.
[0103] The user computing system 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component can be useful for implementing a virtual keyboard. Other exemplary user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.
[0104] In some implementations, the user computing system may store and / or provide one or more user interfaces 124, which may be associated with one or more applications. The one or more user interfaces 124 may be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, augmented reality experiences, virtual reality experiences, and / or other data for display). The user interfaces 124 may be associated with one or more other computing systems (e.g., the server computing system 130 and / or the third-party computing system 150). The user interfaces 124 may include a viewfinder interface, a search interface, a generative model interface, a social media interface, and / or a media content gallery interface.
[0105] The user computing system 102 may include and / or receive data from one or more sensors 126. The one or more sensors 126 may be housed in a housing component that houses one or more processors 112, memory 114, and / or one or more hardware components that may cause one or more software packets to be stored and / or executed. The one or more sensors 126 may include one or more image sensors (e.g., cameras), one or more lidar sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biological sensors (e.g., heart rate sensors, pulse sensors, retinal sensors, and / or fingerprint sensors), one or more infrared sensors, one or more location sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. The one or more sensors may be utilized to obtain data associated with the user's environment (eg, an image of the user's environment, a record of the environment, and / or the user's location).
[0106] The user computing system 102 may include and / or be a part of a user computing device 104. The user computing device 104 may include a mobile computing device (e.g., a smartphone or tablet), a desktop computer, a laptop computer, a smart wearable, and / or a smart appliance. Additionally and / or alternatively, the user computing system may acquire data from and / or generate data with one or more user computing devices 104. For example, a smartphone camera may be utilized to capture image data describing the environment, and / or an overlay application of the user computing device 104 may be utilized to track and / or process data being provided to the user. Similarly, one or more sensors associated with a smart wearable may be utilized to acquire data about the user and / or about the user's environment (e.g., image data may be acquired with a camera housed in the user's smart glasses). Additionally and / or alternatively, data may be acquired and uploaded from other user devices that may be dedicated to data acquisition or generation.
[0107] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0108] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0109] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or otherwise include a variety of machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer nonlinear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Exemplary models 140 are described with reference to FIG. 9B.
[0110] Additionally and / or alternatively, the server computing system 130 may include and / or be communicatively connected to a search engine 142 that may be utilized to crawl one or more databases (and / or resources). The search engine 142 may process data from the user computing system 102, the server computing system 130, and / or the third-party computing system 150 to determine one or more search results associated with the input data. The search engine 142 may perform term-based searches, label-based searches, Boolean-based searches, image searches, embedding-based searches (e.g., nearest neighbor searches), multi-modal searches, and / or one or more other search techniques.
[0111] The server computing system 130 may store and / or provide one or more user interfaces 144 for obtaining input data and / or providing output data to one or more users. The one or more user interfaces 144 may include one or more user interface elements, which may include input fields, navigation tools, content chips, selectable tiles, widgets, data display carousels, dynamic animations, information popups, image augmentation, text to speech, speech to text, augmented reality, virtual reality, feedback loops, and / or other interface elements.
[0112] User computing system 102 and / or server computing system 130 can train models 120 and / or 140 through interaction with a third party computing system 150 that is communicatively coupled via network 180. Third party computing system 150 can be separate from server computing system 130 or can be part of server computing system 130. Alternatively and / or additionally, third party computing system 150 can be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more contexts.
[0113] The third party computing system 150 may include one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 that are executed by the processor 152 to cause the third party computing system 150 to perform operations. In some implementations, the third party computing system 150 includes or is otherwise implemented by one or more server computing devices.
[0114] Network 180 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications over network 180 may be carried over any type of wired and / or wireless connections using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0115] The machine-learned models described herein may be used in a variety of tasks, applications, and / or use cases.
[0116] In some implementations, an input to the machine-learned model of the present disclosure may be image data. The machine-learned model may process the image data to generate an output. As an example, the machine-learned model may process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine-learned model may process the image data to generate an image segmentation output. As another example, the machine-learned model may process the image data to generate an image classification output. As another example, the machine-learned model may process the image data to generate an image data modification output (e.g., alteration of the image data, etc.). As another example, the machine-learned model may process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine-learned model may process the image data to generate an upscaled image data output. As another example, the machine-learned model may process the image data to generate a prediction output.
[0117] In some implementations, the input to the machine-learned models of the present disclosure may be text or natural language data. The machine-learned models may process the text or natural language data to generate an output. As an example, the machine-learned models may process the natural language data to generate a language encoding output. As another example, the machine-learned models may process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned models may process the text or natural language data to generate a translation output. As another example, the machine-learned models may process the text or natural language data to generate a classification output. As another example, the machine-learned models may process the text or natural language data to generate a text segmentation output. As another example, the machine-learned models may process the text or natural language data to generate a semantic intent output. As another example, the machine-learned models may process the text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is of higher quality than the input text or natural language). As another example, the machine-learned models may process the text or natural language data to generate a prediction output.
[0118] In some implementations, the input to the machine-learned model of the present disclosure may be audio data. The machine-learned model may process the audio data to generate an output. As an example, the machine-learned model may process the audio data to generate a speech recognition output. As another example, the machine-learned model may process the audio data to generate a speech translation output. As another example, the machine-learned model may process the audio data to generate a latent embedding output. As another example, the machine-learned model may process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine-learned model may process the audio data to generate an upscaled audio output (e.g., audio data that is of higher quality than the input audio data, etc.). As another example, the machine-learned model may process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine-learned model may process the audio data to generate a predicted output.
[0119] In some implementations, the input to the machine-learned models of the present disclosure may be sensor data. The machine-learned models may process the sensor data to generate an output. As an example, the machine-learned models may process the sensor data to generate a recognition output. As another example, the machine-learned models may process the sensor data to generate a prediction output. As another example, the machine-learned models may process the sensor data to generate a classification output. As another example, the machine-learned models may process the sensor data to generate a segmentation output. As another example, the machine-learned models may process the sensor data to generate a segmentation output. As another example, the machine-learned models may process the sensor data to generate a visualization output. As another example, the machine-learned models may process the sensor data to generate a diagnostic output. As another example, the machine-learned models may process the sensor data to generate a detection output.
[0120] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing a likelihood that one or more images show an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in one or more images and, for each region, a likelihood that the region shows an object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories may be foreground and background. As another example, the set of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines, for each pixel in one or more images, a respective depth value. As another example, the image processing task may be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel of one of the input images, the scene motion represented in pixels between images in the network input.
[0121] A user computing system may include several applications (e.g., applications 1-N). Each application may include its own respective machine learning library and machine learned model. For example, each application may include a machine learned model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.
[0122] Each application may communicate with several other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application may communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0123] The user computing system 102 may include several applications (e.g., applications 1-N). Each application is in communication with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0124] The central intelligence layer may include several machine-learned models. For example, a respective machine-learned model (e.g., a model) may be provided to each application and managed by the central intelligence layer. In other implementations, two or more applications may share a single machine-learned model. For example, in some implementations, the central intelligence layer may provide a single model (e.g., a single model) to all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing system 100.
[0125] The central intelligence layer can communicate with a central device data layer, which can be a centralized repository of data for the computing system 100. The central device data layer can communicate with several other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0126] FIG. 9B illustrates a block diagram of an exemplary computing system 50 that performs instance-level scene recognition, according to an exemplary embodiment of the present disclosure. In particular, the exemplary computing system 50 can include one or more computing devices 52 that can be utilized to acquire and / or generate one or more datasets, which can be processed by the sensor processing system 60 and / or the output determination system 80 to provide a user with feedback that can provide information about features in the one or more acquired datasets. The one or more datasets can include image data, text data, audio data, multimodal data, latent coding data, and the like. The one or more datasets can be acquired via one or more sensors associated with the one or more computing devices 52 (e.g., one or more sensors in the computing device 52). Additionally and / or alternatively, the one or more datasets can be stored data and / or retrieved data (e.g., data retrieved from a web resource). For example, images, text, and / or other content items can be interacted with by a user. The interacted content items can then be utilized to generate one or more decisions.
[0127] The one or more computing devices 52 may acquire and / or generate one or more datasets based on image capture, sensor tracking, data storage retrieval, content download (e.g., downloading images or other content items from web resources via the Internet), and / or via one or more other techniques. The one or more datasets may be processed using a sensor processing system 60. The sensor processing system 60 may perform one or more processing techniques using one or more machine-learned models, one or more search engines, and / or one or more other processing techniques. The one or more processing techniques may be performed in any combination and / or individually. The one or more processing techniques may be performed serially and / or in parallel. In particular, the one or more datasets may be processed using a context determination block 62, which may determine a context associated with one or more content items. The context determination block 62 may identify and / or process metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data to determine a particular context associated with the user. A context may be associated with an event, a determined trend, a particular action, a particular type of data, a particular environment, and / or another context associated with the user and / or the retrieved or acquired data.
[0128] The sensor processing system 60 may include an image pre-processing block 64. The image pre-processing block 64 may be utilized to adjust one or more values of the acquired and / or received image to prepare the image to be processed by one or more machine-learned models and / or one or more search engines 74. The image pre-processing block 64 may resize the image, adjust saturation values, adjust resolution, remove and / or add metadata, and / or perform one or more other operations.
[0129] In some implementations, the sensor processing system 60 may include one or more machine-learned models, which may include a detection model 66, a segmentation model 68, a classification model 70, an embedding model 72, and / or one or more other machine-learned models. For example, the sensor processing system 60 may include one or more detection models 66 that may be utilized to detect particular features in a processed dataset. In particular, one or more images may be processed with the one or more detection models 66 to generate one or more bounding boxes associated with detected features in the one or more images.
[0130] Additionally and / or alternatively, one or more segmentation models 68 may be utilized to segment one or more portions of a dataset from one or more datasets. For example, one or more segmentation models 68 may utilize one or more segmentation masks (e.g., one or more segmentation masks generated manually and / or based on one or more bounding boxes) to segment a portion of an image, a portion of an audio file, and / or a portion of text. Segmentation may include isolating one or more detected objects and / or removing one or more detected objects from an image.
[0131] The one or more classification models 70 may be utilized to process image data, text data, audio data, latent coding data, multimodal data, and / or other data to generate one or more classifications. The one or more classification models 70 may include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. The one or more classification models 70 may process the data to determine one or more classifications.
[0132] In some implementations, data may be processed with one or more embedding models 72 to generate one or more embeddings. For example, one or more images may be processed with one or more embedding models 72 to generate one or more image embeddings in the embedding space. The one or more image embeddings may be associated with one or more image features of the one or more images. In some implementations, the one or more embedding models 72 may be configured to process multimodal data to generate multimodal embeddings. The one or more embeddings may be utilized for classification, retrieval, and / or learning embedding spatial distributions.
[0133] The sensor processing system 60 may include one or more search engines 74 that may be utilized to perform one or more searches. The one or more search engines 74 may crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more proprietary databases, and / or one or more comprehensive databases) to determine one or more search results. The one or more search engines 74 may perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), metadata-based search, multimodal search, web resource search, image search, text search, and / or application search.
[0134] Additionally and / or alternatively, the sensor processing system 60 may include one or more multimodal processing blocks 76 that may be utilized to aid in processing the multimodal data. The one or more multimodal processing blocks 76 may include generating multimodal queries and / or multimodal embeddings for processing by one or more machine-learned models and / or one or more search engines 74.
[0135] The output of the sensor processing system 60 may then be processed with an output determination system 80 to determine one or more outputs for provision to a user. The output determination system 80 may include heuristic-based decisions, machine-learned model-based decisions, user-selection-based decisions, and / or context-based decisions.
[0136] The output determination system 80 may determine how and / or where to provide one or more search results in the search result interface 82. Additionally and / or alternatively, the output determination system 80 may determine how and / or where to provide one or more machine-learned model outputs in the machine-learned model output interface 84. In some implementations, the one or more search results and / or the one or more machine-learned model outputs may be provided for display via one or more user interface elements. The one or more user interface elements may be overlaid on top of the displayed data. For example, one or more detection indicators may be overlaid on top of a detected object in a viewfinder. The one or more user interface elements may be selectable to perform one or more additional searches and / or one or more additional machine-learned model processes. In some implementations, the user interface elements may be provided as dedicated user interface elements for a particular application and / or may be provided uniformly across different applications. The one or more user interface elements may include pop-up displays, interface overlays, interface tiles and / or chips, carousel interfaces, audio feedback, animations, interactive widgets, and / or other user interface elements.
[0137] Additionally and / or alternatively, data associated with the output of the sensor processing system 60 may be utilized to generate and / or provide an augmented reality experience and / or a virtual reality experience 86. For example, one or more acquired datasets may be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which may then be utilized to provide an augmented reality experience and / or a virtual reality experience 86 to a user. The augmented reality experience may render information associated with an environment into a respective environment. Alternatively and / or additionally, objects related to the processed dataset may be rendered into the user environment and / or the virtual environment. The rendering dataset generation may include training one or more neural radiance field models to learn three-dimensional representations for one or more objects.
[0138] In some implementations, one or more action prompts 88 may be determined based on the output of the sensor processing system 60. For example, a search prompt, a purchase prompt, a create prompt, a reservation prompt, a call prompt, a redirect prompt, and / or one or more other prompts may be determined to be associated with the output of the sensor processing system 60. The one or more action prompts 88 may then be provided to the user via one or more selectable user interface elements. In response to selection of the one or more selectable user interface elements, a respective action of the respective action prompt may be executed (e.g., a search may be performed, a purchasing application programming interface may be utilized, and / or another application may be opened).
[0139] In some implementations, one or more datasets and / or outputs of the sensor processing system 60 may be processed with one or more generative models 90 to generate model-generated content items, which may then be provided to a user. Generation may be prompted based on user selection and / or may be performed automatically (e.g., automatically based on one or more conditions that may be associated with not identifying a threshold amount of search results).
[0140] The one or more generative models 90 may include a language model (e.g., a large-scale language model and / or a visual language model), an image generation model (e.g., a text-to-image generation model and / or an image augmentation model), an audio generation model, a video generation model, a graph generation model, and / or other data generation models (e.g., other content generation models). The one or more generative models 90 may include one or more transformer models, one or more convolutional neural networks, one or more recurrent neural networks, one or more feedforward neural networks, one or more generative adversarial networks, one or more self-attention models, one or more embedding models, one or more encoders, one or more decoders, and / or one or more other models. In some implementations, the one or more generative models 90 may include one or more autoregressive models (e.g., machine-learned models trained to generate predicted values based on prior behavioral data) and / or one or more diffusion models (e.g., machine-learned models trained to generate predicted data based on generating and processing distribution data associated with input data).
[0141] One or more generative models 90 may be trained to process input data and generate model-generated content items, which may include predicted words, pixels, signals, and / or other data. The model-generated content items may include new content items that are not identical to any existing work. The one or more generative models 90 may leverage learned expressions, sentences, and / or probability distributions to generate the content items, which may include phrases, plots, settings, objects, characters, beats, lyrics, and / or other aspects not included in the existing content items.
[0142] The one or more generative models 90 may include a visual language model. The visual language model may be trained, tuned, and / or configured to process image data and / or text data to generate natural language output. The visual language model may leverage a pre-trained large-scale language model (e.g., a large-scale autoregressive language model) along with one or more encoders (e.g., one or more image encoders and / or one or more text encoders) to provide detailed natural language output that emulates natural language constructed by humans.
[0143] Visual language models may be utilized for zero-shot image classification, few-shot image classification, image captioning, multimodal query distillation, multimodal question and answering, and / or may be tuned and / or trained for multiple different tasks. Visual language models may perform visual question answering, image caption generation, feature detection (e.g., content monitoring (e.g., for inappropriate content), object detection, scene recognition, and / or other tasks.
[0144] The visual language model may leverage a pre-trained language model, which may then be tuned for multi-modality. The training and / or tuning of the visual language model may include image-text matching, masked language modeling, multi-modal fusion with cross-attention, contrastive learning, prefix language model training, and / or other training techniques. For example, the visual language model may be trained to process an image to generate predicted text that is similar to ground truth text data (e.g., a ground truth caption for the image). In some implementations, the visual language model may be trained to replace masked tokens of a natural language template with text tokens that describe features shown in the input image. Alternatively and / or additionally, the training, tuning, and / or model inference may include multi-layer concatenation of visual and text embedding features. In some implementations, the visual language model may be trained and / or tuned through jointly learning image embeddings and text embedding generation, which may include training and / or tuning the system to map embeddings to a joint feature embedding space that maps text features and image features to a shared embedding space. The joint training may include image-text pair parallel embedding and / or may include triplet training. In some implementations, images may be utilized and / or processed as a prefix to the language model.
[0145] The output determination system 80 may process one or more data sets and / or the output of the sensor processing system 60 using a data augmentation block 92 to generate augmented data. For example, one or more images may be processed using the data augmentation block 92 to generate one or more augmented images. Data augmentation may include data correction, data cropping, removal of one or more features, addition of one or more features, resolution adjustment, lighting adjustment, saturation adjustment, and / or other enhancements.
[0146] In some implementations, one or more data sets and / or outputs of the sensor processing system 60 may be stored based on the determination in a data storage block 94.
[0147] The output of the output determination system 80 may then be provided to a user via one or more output components of the user computing device 52. For example, one or more user interface elements associated with the one or more outputs may be provided for display via a visual display of the user computing device 52.
[0148] The process may be performed iteratively and / or continuously, with one or more user inputs to provided user interface elements conditioning and / or influencing the successive processing loops.
[0149] The techniques described herein refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes described herein may be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components may operate sequentially or in parallel.
[0150] Although the present subject matter has been described in detail with respect to various specific exemplary embodiments thereof, each example is provided as an explanation, not a limitation of the present disclosure. Once a person skilled in the art arrives at the above understanding, he or she can easily create modifications, variations, and equivalents of such embodiments. Thus, the present disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the present subject matter as would be readily apparent to a person skilled in the art. For example, features shown or described as part of one embodiment may also be used with another embodiment to produce a further embodiment. Thus, it is intended that the present disclosure covers such modifications, variations, and equivalents. [Explanation of symbols]
[0151] 10 Detailed Image Captioning System 12, 212 Image data 14, 214 Object Recognition Block 16,216 Object Recognition Output 18, 218, 436 Visual language models 20, 220 language output 22, 222, 456 extended language output 50 Computing Systems 52 Computing Devices, User Computing Devices 60 Sensor Processing System 62 Context Decision Block 64 Image Pre-processing Blocks 66 Detection Model 68 Segmentation Models 70 Classification Models 72 Embedded Model 74 Search Engines 76 Multimodal Processing Block 80 Output Determination System 82 Search Results Interface 84 Machine learning model output interface 86 Augmented and / or Virtual Reality Experiences 88 Action Prompt 90 Generative Model 92 Data Extension Block 94 Data Storage Blocks 100 Computing system, system 102 User Computing System 104 User Computing Devices 112, 132, 152 processors 114 Memory, user computing device memory 116, 136, 156 Data 118, 138, 158 instructions 120, 140 Machine learning models, models 122 User Input Components 124, 144 User Interface 126 Sensors 130 Server Computing System 134, 154 memory 142 Search Engines 150 Third Party Computing Systems 180 Network 200 Search System Using Generative Models 224 Expansion Model 226 Results 228 Generative Model 230 Generated Response, Model Generated Response 232 Text Data 402, 502 Input text 404, 434, 452, 504 Input images 406, 454 Recognition Data 408 Improved prompts 410, 506 Model generation response 432 Input text, Text input 438 Preliminary Content Items 440 Instance Level Recognition Block 442 Instance-Level Awareness 444 Image-Aware Model-Generated Content Items 458 Improved Queries 500 Generative Model Response System 600 Templates 602 Question and Answer Prompt Template 604 Image Captioning Templates 606 Query Templates 650 recognition output templates 652 Main Object Templates 654 Multi-Object Templates 656 Scene Description Template
Claims
1. 1. A computer-implemented method comprising: acquiring, by a computing system comprising one or more processors, image data, the image data comprising an input image; processing, by the computing system, the input image with an object recognition model to generate fine-grained object recognition output, the fine-grained object recognition output describing identifying details about objects depicted in the input image; processing, by the computing system, the input image with a visual language model to generate linguistic output, the linguistic output comprising a set of predicted words predicted to describe the input image, the set of predicted words comprising coarse-grained terms describing predicted identities of the objects shown in the input image; processing, by the computing system, the fine-grained object recognition output and the linguistic output to generate an extended linguistic output, the extended linguistic output comprising the set of predicted words with the coarse-grained terms replaced by the fine-grained object recognition output; The method includes:
2. processing, by the computing system, the input image with the object recognition model to generate the fine-grained object recognition output, detecting the object in the input image; generating an object embedding; determining an image cluster associated with the object embedding; processing web resources associated with the image clusters to determine identifying details about the objects; 2. The method of claim 1, comprising:
3. 20. The method of claim 19, further comprising: generating a bounding box associated with a position of the object in the input image; generating an image segment based on the bounding box; processing the image segments with an embedding model to generate the object embeddings; 3. The method of claim 2, comprising:
4. processing, by the computing system, the fine-grained object recognition output and the linguistic output to generate the augmented linguistic output; processing, by the computing system, the linguistic output to determine a number of text tokens associated with features in the input image; determining, by the computing system, that a particular token of the plurality of text tokens is associated with the object; replacing, by the computing system, the particular token with the fine-grained object recognition output; 2. The method of claim 1, comprising:
5. determining, by the computing system, that the particular token of the plurality of text tokens is associated with the object, processing, by the computing system, the fine-grained object recognition output with an embedding model to generate an instance-level embedding; processing, by the computing system, the plurality of text tokens with the embedding model to generate a plurality of token embeddings; determining, by the computing system, that the instance-level embedding is associated with a particular embedding associated with the particular token; 5. The method of claim 4, comprising:
6. Processing, by the computing system, the augmented language output with a second language model to generate a natural language response to the augmented language output. The method of claim 1, further comprising:
7. The method of claim 6 , wherein the natural language response comprises additional information associated with the augmented language output.
8. The method of claim 1 , wherein the coarse-grained terminology comprises an object type and the fine-grained object recognition output comprises detailed identification information of the object.
9. Providing, by the computing system, the augmented language output in an augmented reality experience. The method of claim 1, further comprising:
10. The method of claim 9 , wherein the augmented reality experience comprises the augmented language output overlaid on top of a live video feed of an environment.
11. 1. A computing system for image captioning, comprising: one or more processors; and one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: acquiring image data, the image data comprising an input image; processing the input image with an object recognition model to generate an object recognition output, the object recognition output detailing identifying information about objects depicted in the input image; processing the input image with a visual language model to generate a linguistic output, the linguistic output comprising a set of words predicted to describe the input image, the set of words comprising terms describing predicted identities of the objects shown in the input image; processing the object recognition output and the language output with the visual language model to generate an augmented language output, the augmented language output comprising the set of words with the terms replaced by the object recognition output; Including, the system.
12. 2. The system of claim 1, wherein the input image describes the object in an environment with one or more additional objects, the object recognition output is associated with the object, and the language output is associated with the object and the environment with the one or more additional objects.
13. 12. The system of claim 11, wherein the visual language model is trained on a training dataset comprising a plurality of image-caption pairs, the plurality of image-caption pairs comprising a plurality of training images and a plurality of respective captions associated with the plurality of training images.
14. 12. The system of claim 11, wherein the input image is processed in parallel using the object recognition model and the visual language model to perform parallel determination of the object recognition output and the language output, the visual language model comprising one or more text encoders, one or more image encoders, and one or more decoders.
15. The system of claim 11 , wherein the object recognition model comprises one or more classification models.
16. 12. The system of claim 11, wherein the object recognition output comprises an instance-level object recognition associated with the object, and the language output comprises a scene understanding associated with the input image.
17. One or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, including: acquiring image data, the image data comprising an input image; processing the input image to determine an object recognition output, the object recognition output describing identifying details about objects depicted in the input image; processing the input image with a visual language model to generate a linguistic output, the linguistic output comprising a set of words predicted to describe the input image, the set of words comprising terms describing predicted identities of the objects shown in the input image; processing the object recognition output and the linguistic output to generate an augmented linguistic output, the augmented linguistic output comprising the set of words with the terms replaced by the object recognition output; determining one or more search results associated with the augmented language output, the one or more search results being associated with one or more web resources; [0023] In one or more non-transitory computer readable media,
18. 20. The one or more non-transitory computer-readable media of claim 17, wherein processing the input image to determine the object recognition output includes processing the input image with a search engine to determine text data describing object identification information.
19. processing the input image to determine the object recognition output; processing the input image with an embedding model to generate an image embedding; determining one or more object labels based on the image embeddings, the one or more object labels comprising details of the identity of the objects depicted in the input image; 20. The one or more non-transitory computer-readable media of claim 17, comprising:
20. determining the one or more search results associated with the augmented language output includes determining that a plurality of search results are responsive to a search query that comprises the augmented language output; 20. The one or more non-transitory computer-readable media of claim 17, wherein the operations further comprise providing the plurality of search results for display.
Citation Information
Patent Citations
Expanded reality-sense display device
JP2009211251A
Image identification using facial recognition
JP2010511938A
Anchors for location-based navigation and augmented reality applications
JP2015520461A
Image recognition device, image recognition method, and program
JP2017005389A
Face detection device
JP2023012283A
Cited By
Wildlife detection system
JP7924722B1