Instance-Level Scene Recognition Using Vision-Language Models

Through the computing system's object recognition model and visual language model to process images in parallel, and generate enhanced language output is solved, which solves the problem that users find it difficult to determine image objects and their information, and achieves efficient and accurate information acquisition.

CN118587623BActive Publication Date: 2025-05-27GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410631660.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-10-27
Filing Date
2024-05-21
Publication Date
2025-05-27
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

It is difficult for users to determine objects and their detailed information in the image through simple text searches, and the content requested by the user may be difficult to obtain due to the unspecified search location or the absence of content.

Method used

The computing system uses the object recognition model and visual language model to process the input images in parallel, generates fine-grained object recognition output and language output, and then generates enhanced language output for search result determination and model content generation.

Benefits of technology

It realizes detailed identification and description of objects in the image, improves the efficiency of users to obtain required information, and ensures the accuracy and relevance of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118587623B_ABST
    Figure CN118587623B_ABST
Patent Text Reader

Abstract

Systems and methods for image understanding can include one or more object recognition systems and one or more vision-language models to generate enhanced language output, where the enhanced language output can be both scene-aware and object-aware. The systems and methods can process an input image with an object recognition model to generate an object recognition output that describes the identification details of the objects depicted in the input image. The systems and methods can include processing the input image with a vision-language model to generate a language output that describes a predicted scene description. Then, the object recognition output can be utilized to enhance the language output to generate an enhanced language output that includes scene understanding of the language output with the specificity of the object recognition output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to enhancing the output of a vision - language model based on instance - level object recognition. More specifically, the present disclosure relates to using a vision - language model and object recognition processing to generate a detailed output that can be used for search result determination and / or generation model content generation. Background Art

[0002] Understanding the entire world can be difficult. Whether an individual is trying to understand what the object in front of them is, trying to determine where else the object can be found, or trying to determine where an image on the Internet was captured, a text - only search can be difficult. Specifically, users may have difficulty determining which words to use. Additionally, the descriptiveness and / or richness of the words may not be sufficient to generate the desired results.

[0003] Furthermore, the content requested by the user may not be easily accessible to the user due to the user not knowing where to search, due to the storage location of the content, and / or due to the content not existing. The user may be requesting search results based on an imagined concept without a clear way to express the imagined concept. Summary of the Invention

[0004] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned by practice of the embodiments.

[0005] One example aspect of the present disclosure relates to a computer - implemented method. The method may include: obtaining, by a computing system including one or more processors, image data. The image data may include an input image. The method may include: processing, by the computing system, the input image with an object recognition model to generate a fine - grained object recognition output. The fine - grained object recognition output may describe identification details of the object depicted in the input image. The method may include: processing, by the computing system, the input image with a vision - language model to generate a language output. The language output may include a set of predicted words predicted to describe the input image. In some implementations, the set of predicted words may include coarse - grained terms that describe the predicted identification of the object depicted in the input image. The method may include: processing, by the computing system, the fine - grained object recognition output and the language output to generate an enhanced language output. The enhanced language output may include the set of predicted words with the coarse - grained terms replaced by the fine - grained object recognition output.

[0006] In some implementations, the processing of an input image by a computing system with an object recognition model to generate a fine-grained object recognition output may include: detecting an object in the input image; generating an object embedding; determining an image cluster associated with the object embedding; and processing web resources associated with the image cluster to determine identification details of the object. Generating an object embedding may include: generating a bounding box associated with the location of the object within the input image; generating an image patch based on the bounding box; and processing the image patch with an embedding model to generate an object embedding.

[0007] In some implementations, the processing of the fine-grained object recognition output and a language output by a computing system to generate an enhanced language output may include: the computing system processing the language output to determine a plurality of text tokens associated with features in the input image; the computing system determining that a particular token among the plurality of text tokens is associated with the object; and the computing system replacing the particular token with the fine-grained object recognition output. The computing system determining that a particular token among the plurality of text tokens is associated with the object may include: the computing system processing the fine-grained object recognition output with an embedding model to generate an instance-level embedding; the computing system processing the plurality of text tokens with an embedding model to generate a plurality of token embeddings; and the computing system determining that the instance-level embedding is associated with a particular embedding associated with the particular token.

[0008] In some implementations, the method may include: the computing system processing the enhanced language output with a second language model to generate a natural language response to the enhanced language output. The natural language response may include additional information associated with the enhanced language output. The coarse-grained term may include an object type. The fine-grained object recognition output may include a detailed identification of the object. In some implementations, the method may include: the computing system providing the enhanced language output in an augmented reality experience. The augmented reality experience may include the enhanced language output superimposed over a live video feed of the environment.

[0009] Another example aspect of the present disclosure relates to a computing system for image captioning. The system may include: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include: obtaining image data. The image data may include an input image. The operations may include: processing the input image with an object recognition model to generate an object recognition output. The object recognition output may describe identification details of the objects depicted in the input image. The operations may include: processing the input image with a vision-language model to generate a language output. The language output may include a set of words predicted to describe the input image. In some implementations, the set of words may include terms that describe the predicted identification of the objects depicted in the input image. The operations may include: processing the object recognition output and the language output with the vision-language model to generate an enhanced language output. The enhanced language output may include a set of words with the terms replaced by the object recognition output.

[0010] In some implementations, the input image may depict an object in an environment with one or more additional objects. The object recognition output may be associated with the object. The language output may be associated with the object and the environment with one or more additional objects. The vision-language model may have been trained based on a training dataset including a plurality of image-caption pairs. The plurality of image-caption pairs may include a plurality of training images and a plurality of corresponding captions associated with the plurality of training images. The input image may be processed in parallel with the object recognition model and the vision-language model to perform parallel determination of the object recognition output and the language output. The vision-language model may include one or more text encoders, one or more image encoders, and one or more decoders. In some implementations, the object recognition model may include one or more classification models. The object recognition output may include instance-level object recognition associated with the object. The language output may include scene understanding associated with the input image.

[0011] Another example aspect of the present disclosure relates to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations may include: obtaining image data. The image data may include an input image. The operations may include: processing the input image to determine an object recognition output. The object recognition output may describe identification details of an object depicted in the input image. The operations may include: processing the input image with a vision-language model to generate a language output. The language output may include a set of words predicted to describe the input image. In some implementations, the set of words may include terms that describe a predicted identification of an object depicted in the input image. The operations may include: processing the object recognition output and the language output to generate an enhanced language output. The enhanced language output may include a set of words with terms replaced by the object recognition output. The operations may include: determining one or more search results associated with the enhanced language output. The one or more search results may be associated with one or more web resources.

[0012] In some implementations, processing the input image to determine an object recognition output may include: processing the input image with a search engine to determine text data describing the object identification. Processing the input image to determine an object recognition output may include: processing the input image with an embedding model to generate an image embedding; and determining one or more object labels based on the image embedding. The one or more object labels may include the identification details of the object depicted in the input image. In some implementations, determining one or more search results associated with the enhanced language output may include: determining that a plurality of search results are in response to a search query that includes the enhanced language output. The operations may further include: providing the plurality of search results for display.

[0013] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0014] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] With reference to the drawings, a detailed discussion of the embodiments directed to those of ordinary skill in the art is set forth in this specification, in which:

[0016] Figure 1 A block diagram of an example detailed image caption generation system in accordance with an example embodiment of the present disclosure is depicted.

[0017] Figure 2 A block diagram of an example search system utilizing a generative model according to an example embodiment of the present disclosure is depicted.

[0018] Figure 3 A flowchart of an example method for performing enhanced language output generation according to an example embodiment of the present disclosure is depicted.

[0019] Figures 4A to 4C An illustration of an example fine-grained object recognition with scene understanding according to an example embodiment of the present disclosure is depicted.

[0020] Figure 5 An illustration of an example generative model response system according to an example embodiment of the present disclosure is depicted.

[0021] Figures 6A to 6B An illustration of an example scene description generation template according to an example embodiment of the present disclosure is depicted.

[0022] Figure 7 A flowchart of an example method for performing object-aware scene recognition according to an example embodiment of the present disclosure is depicted.

[0023] Figure 8 A flowchart of an example method for performing a search utilizing a generative model according to an example embodiment of the present disclosure is depicted.

[0024] Figure 9A A block diagram of an example computing system for performing instance-level scene recognition according to an example embodiment of the present disclosure is depicted.

[0025] Figure 9B A block diagram of an example computing system for performing instance-level scene recognition according to an example embodiment of the present disclosure is depicted.

[0026] Reference numerals repeated on multiple figures are intended to identify the same features in various implementations. Detailed Description

[0027] Generally, the present disclosure relates to systems and methods for detailed instance-level scene recognition. Specifically, the systems and methods disclosed herein can utilize an object recognition system and a vision-language model to generate detailed captions, queries, and / or prompts associated with an input image. For example, an object recognition system (e.g., a system having one or more object recognition models) can process the input image to generate an object recognition output that describes the recognition of a particular object of a particular object class. The object recognition output can describe the detailed identification of a particular object instance. Additionally, a vision-language model can process the input image to generate a language output that describes the scene recognition of the scene depicted in the input image. The language output can include details describing the environment and one or more objects in the environment. The language output may not include the granularity and / or specificity of the object recognition output. The object recognition model and the vision-language model can process the input image in parallel to reduce latency. Then, the object recognition output and the language output can be processed to generate an enhanced language output that describes the scene recognition of the language output with the specificity and / or particularity of the object recognition output. For example, the language output can include the identification of the particular object class of the object depicted in the input image, while the enhanced language output can include a specific indication of the instance-level identification of the depicted object (e.g., the brand and model name of a product, the name of the depicted person, the name of an artwork, and / or the species and subspecies identification of a plant or animal).

[0028] Then, the enhanced language output can be used as a query and / or prompt for obtaining additional information associated with the scene and / or objects depicted in the input image. In some implementations, input text can be received along with the input image, and the language output and / or the enhanced language output can be generated partially based on the input text. Thus, a user can ask questions about the depicted scene, a detailed scene recognition can be generated, and a detailed query and / or prompt can be generated that includes the semantic intent of the question and the recognition information of the enhanced language output. The enhanced language output can be processed with a search engine and / or a generative model (e.g., a large language model, a vision-language model, an image generation model, etc.) to generate additional information that can respond to the input question.

[0029] Visual language models can utilize learned image and language associations to generate natural language captions for images; however, visual language models may have difficulty with details including object specificity. Lack of specificity may result in the generation of general queries and / or prompts, which may not provide results specific to and / or applicable to the features depicted in the image. For example, a user may provide an image along with the question "how do I take care of this?". The visual language model can process the image to determine that the image depicts a plant and can then utilize that determination to generate a refined query "what do plants need to stay alive and grow?". The refined query can be processed to determine search results that can be associated with general care instructions for plants, which may include watering twice a week, half a day of direct sunlight, and loamy soil. However, the general care instructions may not be suitable for the specific plant depicted in the image (e.g., succulents (e.g., agave plants) require less water and different soil, while Matteuccia struthiopteris can thrive in the shade away from direct sunlight). Thus, utilizing general information for an object class may be disadvantageous for care and contrary to the original intent of the input.

[0030] The systems and methods disclosed herein can process an image in parallel with a visual language model and a fine-grained object recognition model to generate an output that is scene-aware and object-aware while being formatted in a natural language format. The parallel processing can be separate and independent such that the scene-aware output and the object-aware output are determined separately and not affected by the other. Token replacement can be utilized to replace the coarse-grained object recognition of the visual language model (e.g., object class recognition (e.g., plant, human, car, building, etc.)) with the fine-grained recognition of an instance-level object recognition system (e.g., specific object recognition indicating a specific object identity (e.g., tiger lily, George Washington, Model T Roadster with a 5L engine, Monticello, etc.)). For example, the systems and methods can include processing an input image with an object recognition system to generate an object recognition output that describes the identity details of the specific object depicted in the input image. The identity details can include instance-level identities that describe the specific and detailed identity of the object. The systems and methods can also process the input image with a visual language model to generate a language output that describes the scene recognition of the entire scene depicted in the input image. The scene recognition may not be as specific as the object recognition output. Thus, the systems and methods can process the object recognition output and the language output to generate an enhanced language output that utilizes the scene recognition of the language output and the specificity of the object recognition output.

[0031] Pairing instance-level object recognition with visual language model processing can be used to generate detailed text descriptions, queries, and / or prompts. Combining scene understanding with instance understanding can be used for image search, image indexing, automated content generation and / or understanding, and / or other image understanding tasks. For example, enhanced language output can be used as and / or used to generate detailed queries and / or detailed prompts for obtaining and / or generating additional information. Specificity can lead to improved customization of search results and / or generated prompts.

[0032] Different objects within the same object class may have different characteristics with respect to maintenance, use, assembly, and / or repair, which may cause a general search query to generate search results that may not be relevant for the particular object. Thus, utilizing scene understanding together with object understanding may generate an output that may be processed with a search engine and / or a machine learning model to generate object perception information.

[0033] A multimodal large language model (e.g., a large visual language model) can be tuned and / or trained to have a rough understanding of the image. For example, processing an image with a large language model may be able to output "This is a black dogsitting on a beach". However, an object recognition system may be trained and / or configured to recognize objects at instance-level granularity. In the same image, the object recognition system may identify the breed of the dog as an Australian Kelpie and the beach as Bondi Beach. When the two systems are coupled, the system and method can teach and / or adjust the large language model so that the scene includes an Australian Kelpie sitting on Bondi Beach. The large language model can then learn and / or be prompted to describe the scene at instance-level granularity. The systems and methods disclosed herein can be used to identify each product in an aisle as a user walks past the product, which can then help the user find products that meet their dietary restrictions and / or other preferences and criteria.

[0034] In some implementations, object recognition and / or scene recognition techniques (including visual search) can be used to tune and / or train a visual language model for instance-level recognition, which can include training a visual language model for attribute specificity based on output from visual search.

[0035] The systems and methods disclosed herein can be used to process a variety of different data types (e.g., image data, text data, video data, audio data, statistical data, graphical data, latent encoded data, and / or multimodal data) to generate outputs that can be in a variety of different data formats (e.g., image data, text data, video data, audio data, statistical data, graphical data, latent encoded data, and / or multimodal data). For example, the input data can include a video that can be processed to generate an overview of the video, which can include a natural language overview, a timeline, a flowchart, an audio file in the form of a podcast, and / or a comic book. An object recognition system can be used for object-specific details, while a scene understanding model (e.g., a vision language model) can be used for scene recognition and / or understanding of groups of frames. In some implementations, one or more additional models can be used for context understanding. For example, a hierarchical video encoder can be used for frame understanding, understanding of frame sequences, and / or understanding of the entire video. Audio input processing can include using a text-to-speech model, which can be implemented as part of a language model.

[0036] The systems and methods of the present disclosure provide a variety of technical effects and benefits. As an example, the systems and methods can be used to generate instance-level scene recognition outputs. Specifically, the systems and methods disclosed herein can utilize a vision language model in parallel with an object recognition system to generate a natural language output that is both scene-aware and object-specific. The enhanced language output can then be used as a query for searching and / or a prompt for generating model content generation.

[0037] Another example technical effect and benefit relates to improving the computational efficiency of a computing system and enhancing its operation. For example, a technical benefit of the systems and methods of the present disclosure is the ability to reduce the computational resources required for generating detailed queries and / or detailed prompts. Specifically, training and / or tuning a language model for instance-level object recognition can be computationally expensive and may require large training datasets. Additionally, training a language model for such specificity can be computationally expensive for model inference. The methods disclosed herein can reduce the training time and resource costs for generating detailed image captions to generate instance-level scene recognition outputs. In some implementations, an input image can be processed in parallel with an object recognition model and a vision language model to reduce latency. Alternatively and / or additionally, if executed by a computing device with limited processing capabilities, the input image can be processed with the object recognition model and the vision language model at different times.

[0038] Referring now to the drawings, example embodiments of the present disclosure will be discussed in more detail.

[0039] Figure 1FIG. 0 depicts a block diagram of an example detailed image captioning generation system 10 in accordance with an example embodiment of the present disclosure. In some implementations, the detailed image captioning generation system 10 is configured to receive and / or obtain an input data set that includes image data 12 depicting an environment having one or more objects, and as a result of receiving the image data 12, generate, determine, and / or provide an enhanced language output 22 that describes instance-level scene recognition of the objects. Thus, in some implementations, the detailed image captioning generation system 10 may include a vision-language model 18 operable to perform scene recognition and an object recognition block 14 operable to perform object recognition.

[0040] Specifically, the detailed image captioning generation system 10 may obtain input data that may include image data 12 depicting one or more input images. The one or more input images may depict an environment and one or more objects. The environment may include a room, landscape, city, town, sky, and / or other environments. In some implementations, the environment is described by a user environment generated by one or more image sensors of a user computing device. The one or more objects may include products, people, plants, animals, artworks, structures, landmarks, and / or other objects.

[0041] The image data 12 may be processed with a parallel processing pipeline. A first pipeline may process the image data 12 to generate an object recognition output (e.g., instance-level object recognition). A second pipeline may process the image data 12 to generate a language output 20 that describes scene recognition. Then, the detailed image captioning generation system 10 may process the outputs of the pipelines to generate an enhanced language output 22 that describes the detailed image caption.

[0042] For example, the object recognition block 14 (e.g., an object recognition system) may process the image data 12 to generate an object recognition output 16 that describes fine-grained object recognition. The object recognition output 16 may include identification details of one or more objects depicted in the one or more input images. The identification details may describe instance-level recognition associated with a particular object in a particular object class, which may include a product model name, a plant species and / or subspecies, a particular person's name, an artwork's name, a location name, and / or other instance-level identifiers.

[0043] The object recognition block 14 may include an object recognition model, which may include one or more machine learning models. The object recognition model may be trained and / or configured to process an image, detect an object, segment a portion of the image that includes the object, and then process the image segment to generate a recognition output. The object recognition model may include a detection model that processes an input image to generate a bounding box indicating the location of the detected object. Then, the segmentation model of the object recognition model may segment the detected object based on the bounding box to generate an image segment of the detected object. Then, the image segment may be processed with the classification model of the object recognition model to generate an object classification. Then, the object classification may be processed to generate an object recognition output 16.

[0044] Alternatively and / or additionally, the object recognition block 14 may include one or more embedding models. The one or more embedding models may process the image data 12 and / or the image segments to generate one or more image embeddings. The one or more image embeddings may be used to query an embedding space to obtain similar embeddings, neighbor embeddings, embedding clusters, and / or embedding labels (e.g., labels of learned features that describe the learned distribution in the embedding space). Multiple web resources determined to be associated with the objects depicted in the one or more input images may be obtained using the similar embeddings, neighbor embeddings, embedding clusters, and / or embedding labels. The multiple web resources may be processed to determine details associated with the objects, which may include product names, object sources, object lists, object locations, other instances of the objects, other identifiers, and / or other details. Then, the details may be used to generate an object recognition output 16. The multiple web resources may be sources of content items that are embedded to generate other embeddings in the similar embeddings, neighbor embeddings, and / or embedding clusters.

[0045] The object recognition block 14 may generate an object recognition output 16 for each object depicted in the input image. Alternatively and / or additionally, the object recognition block 14 may determine a focus object and / or an object of interest based on object location, object size, image semantics, image focus, appearance in an input image sequence, and / or other context attributes.

[0046] The vision language model 18 may process the image data 12 to generate a language output 20. The language output 20 may be a natural language text string that describes the scene recognition of the scene (e.g., the environment and one or more objects) depicted in the one or more input images. The language output 20 may include a coarse-grained recognition output associated with the location and / or one or more objects, which may include the location and the class identification of the one or more objects.

[0047] The vision-language model 18 can include a language model that is trained, configured, and / or tuned to process multimodal data, which can include tuning for image understanding tasks. For example, the vision-language model 18 can be trained based on a training data set that includes image-caption pairs. The image-caption pairs can include training images and corresponding training captions for the specific training images. The training and / or tuning can include processing the training images with the vision-language model 18 to generate predicted text strings. The predicted text strings and the corresponding training captions can be processed to evaluate a loss function to generate gradient descent. Then, the gradient descent can be backpropagated to adjust one or more parameters of the vision-language model.

[0048] Alternatively and / or additionally, the vision-language model can include a text encoder and an image encoder that may have been jointly trained and / or jointly tuned to encode input data, and the encoded input data can then be processed with a decoder to generate a vision-language model output. In some implementations, an image embedding model can be trained to process images and generate image embeddings, which can then be processed with a large language model. The image embeddings can describe representations associated with image features.

[0049] Then, the object recognition output 16 and the language output 20 can be processed to generate an enhanced language output 22. The enhanced language output 22 can include a scene understanding of the language output 20 with the object recognition granularity of the object recognition output 16. In some implementations, the enhanced language output 22 can describe a detailed image caption for one or more input images.

[0050] Figure 2 A block diagram of a search system 200 that utilizes a generative model according to an example embodiment of the present disclosure is depicted. The search system 200 that utilizes a generative model is similar to Figure 1 the detailed image caption generation system 10, except that the search system 200 that utilizes a generative model further includes a search result 226 determination and a generative model 228 for generating a generative response 230.

[0051] Specifically, a search system 200 utilizing a generative model can obtain input data, which can include image data 212 describing one or more input images (e.g., one or more images of beef wellington with kale on a red tablecloth on a large plate) and text data 232 describing a request for specific information (e.g., a request for a recipe for the depicted beef wellington). One or more input images can describe an environment and one or more objects. The environment can include a room, a landscape, a city, a town, the sky, and / or other environments. In some implementations, the environment is described using one or more image sensors of a user computing device. One or more objects can include products, people, plants, animals, artworks, structures, landmarks, and / or other objects.

[0052] The image data 212 and / or the text data 232 can be processed using one or more image processing pipelines. A first pipeline can process the image data 212 to generate an object recognition output (e.g., instance-level object recognition). A second pipeline can process the image data 212 and / or the text data to generate a language output 220 describing scene recognition. The pipelines can be executed in parallel, serially, and / or in a self-attention loop. Then, the search system 200 utilizing the generative model can process the output of the pipelines and / or the text data 232 to generate an enhanced language output 222 describing the detailed image. Then, the enhanced language output 222 can be processed using a search engine and / or the generative model 228 to obtain and / or generate additional data (e.g., one or more search results 226 and / or one or more model-generated responses 230).

[0053] For example, an object recognition block 214 (i.e., an object recognition system) can process the image data 212 to generate an object recognition output 216 describing fine-grained object recognition. The object recognition output 216 can include identification details of one or more objects depicted in the one or more input images (e.g., "beef wellington", Jane Doe, Mona Lisa, Sixteenth Chapel, Washington Monument, Brand X Model YZ Smartphone, etc.). The identification details can describe instance-level recognition associated with a specific object in a specific object class (e.g., the recognition of the specific object depicted), which can include a product model name, a plant species and / or subspecies, the name of a specific person, the name of an artwork, a location name, and / or other instance-level identifiers.

[0054] The object recognition block 214 may include an object recognition model, which may include one or more machine learning models (e.g., one or more embedding models, one or more detection models, one or more segmentation models, one or more classification models, one or more semantic understanding models, one or more feature extractors, and / or one or more other models). The object recognition model may be trained and / or configured to process an image, detect an object, segment a portion of the image that includes the object (e.g., segment the portion of the image within the object and / or segment the object from the image), and then process the image fragment to generate a recognition output. The object recognition model may include a detection model that processes the input image to generate a bounding box indicating the location of the detected object. Then, the segmentation model of the object recognition model may segment the detected object based on the bounding box to generate an image fragment of the detected object. Then, the image fragment may be processed by the classification model of the object recognition model to generate an object classification. Then, the object classification may be processed to generate an object recognition output 216.

[0055] Alternatively and / or additionally, the object recognition block 214 may include one or more embedding models. The one or more embedding models may process the image data 212 and / or the image fragment to generate one or more image embeddings. The one or more image embeddings may be utilized to query an embedding space to obtain similar embeddings, neighbor embeddings, embedding clusters, and / or embedding labels (e.g., labels that describe learned characteristics of the learned distribution in the embedding space). Multiple web resources may be obtained that are determined to be associated with the object depicted in the one or more input images using the similar embeddings, neighbor embeddings, embedding clusters, and / or embedding labels. The multiple web resources may be processed to determine details associated with the object, which may include product names, object sources, object lists, object locations, other instances of the object, other identifiers, and / or other details. Then, the details may be utilized to generate an object recognition output 16. The multiple web resources may be sources of content items that are embedded to generate other embeddings in the similar embeddings, neighbor embeddings, and / or embedding clusters.

[0056] The object recognition block 214 may generate an object recognition output 16 for each object depicted in the input image. Alternatively and / or additionally, the object recognition block 14 may determine the focus object and / or the object of interest based on object location, object size, image semantics, image focus, appearance in the input image sequence, and / or other contextual attributes. In some implementations, selection of a particular object for processing may be based on text data 232 (e.g., "what are recipes for this item?" causes food to be processed, while "what is that on the right?" causes objects on the right side of the input image to be processed).

[0057] The visual language model 218 can process the image data 212 and / or the text data 232 to generate a language output 220. The language output 220 can be a natural language text string describing a scene recognition of a scene (e.g., an environment and one or more objects) depicted in one or more input images. The language output 220 can include a coarse-grained recognition output associated with a location and / or one or more objects, and the coarse-grained recognition output can include a class identification of the location and one or more objects. In some implementations, the language output 220 can be a scene recognition output based on processing the text data 232, and the scene recognition output includes a structure, format, and / or additional language. For example, the text data 232 can include "What is the origin of this food item?", and the language output can include "The image depicts a formal dinner, in which the food item is a pastry, which is presented on a plate with a vegetable on a red table (the image depicts a formal dinner, where the food is a pastry on a red table with vegetables on a plate)". Object recognition output for the example image may include "beef wellington," "kale," and / or "maroon tablecloth."

[0058] The vision - language model 218 can include a language model that is trained, configured, and / or tuned to process multimodal data, which can include tuning for image understanding tasks. For example, the vision - language model 218 can be trained based on a training data set that includes image - caption pairs. The image - caption pairs can include training images and corresponding training captions for the specific training images. Training and / or tuning can include processing the training images with the vision - language model 218 to generate predicted text strings. The predicted text strings and the corresponding training captions can be processed to evaluate a loss function to generate gradient descent. Then, the gradient descent can be backpropagated to adjust one or more parameters of the vision - language model.

[0059] Alternatively and / or additionally, the vision - language model can include a text encoder and an image encoder that may have been jointly trained and / or jointly tuned to encode input data, and the encoded input data can then be processed with a decoder to generate the vision - language model output. In some implementations, an image embedding model can be trained to process images and generate image embeddings, which can then be processed with a large - language model. The image embeddings can describe representations associated with image features.

[0060] Then, the object recognition output 216, the language output 220, and / or the text data 232 can be processed with an enhancement model 224 to generate an enhanced language output 222. The enhancement model 224 can include a vision - language model, other language models, and / or other generative models. The enhancement model 224 can be trained and / or tuned to identify text tokens associated with the same object and enhance the language output to replace coarse - grained tokens of the language output 220 with fine - grained terms of the object recognition output 216. The enhanced language output 222 can include a scene understanding of the language output 220 with the object recognition granularity of the object recognition output 216. In some implementations, the enhanced language output 222 can describe a detailed image caption for one or more input images. For example, the enhanced language output 222 can include “The image depicts a formal dinner, in which the food item is a beef wellington, which is presented on a plate with kale on a maroon tablecloth”.

[0061] In some implementations, a search engine can process the enhanced language output 222 and / or the text data 232 to determine one or more search results 226. The enhanced language output 222 and / or the text data 232 can include and / or be formatted as a query (e.g., "recipe for a beef wellington"). The one or more search results 224 can be in response to the query posed by the enhanced language output 222 and / or the text data 232 (e.g., web pages, videos, and / or books having recipes for beef wellington).

[0062] Alternatively and / or additionally, a generative model 228 (e.g., a large language model, a text-to-image model, and / or other generative models) can process the enhanced language output 222 and / or the text data 232 to generate one or more model-generated responses 230. The enhanced language output 222 and / or the text data 232 can include and / or be formatted as a prompt (e.g., "generate step by step instructions for a beef wellington"). The one or more model-generated responses 230 can be in response to the prompt of the enhanced language output 222 and / or the text data 232 (e.g., the model-generated response can include step-by-step instructions for preparing beef wellington based on processing one or more recipes for beef wellington and can include an image with steps summarized in text (e.g., an image pulled from a web resource and / or an image generated with a text-to-image generative model)).

[0063] Figure 3 A flowchart depicting an example method for execution in accordance with an example embodiment of the present disclosure is shown. Although Figure 3 the steps are depicted in a particular order for purposes of illustration and discussion, the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 300 can be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of the present disclosure.

[0064] At 302, the computing system can obtain image data. The image data can include an input image. The input image can include one or more objects in an environment. The environment can include a room, a landscape, and / or other environments. The one or more objects can include people, structures, animals, plants, monuments, artworks, products, and / or other objects. The computing system can obtain and / or generate image data with a computing device, which can include a mobile computing device, a smart wearable device, a smart appliance, and / or other computing devices.

[0065] At 304, a computing system can process an input image with an object recognition model to generate a fine-grained object recognition output. The fine-grained object recognition output can describe the identification details of the object depicted in the input image. The object can include a product, and the fine-grained object recognition output can include a specific product label, which can include a model name, a registration number, a model number, a product-specific name, and / or identification details. In some implementations, the object can include a person, and the fine-grained object recognition output can include the person's name. Alternatively and / or additionally, the object can include a work of art (e.g., a sculpture, a painting, a photograph, etc.), and the fine-grained object recognition output can include the name of the work of art. The object recognition model can include one or more embedding models, one or more detection models, one or more feature extractors, one or more classification models, one or more segmentation models, one or more search engines, and / or one or more other models.

[0066] In some implementations, processing an input image with an object recognition model to generate a fine-grained object recognition output can include: detecting an object in the input image; generating an object embedding; determining an image cluster associated with the object embedding; and processing web resources associated with the image cluster to determine the identification details of the object. Generating an object embedding can include: generating a bounding box associated with the position of the object within the input image; generating an image patch based on the bounding box; and processing the image patch with an embedding model to generate an object embedding. Additionally and / or alternatively, the object embedding can be used to search an embedding space to obtain a nearest neighbor embedding and / or a similar embedding. In some implementations, one or more embedding clusters can be determined to be associated with the object embedding. The nearest neighbor embedding, the similar embedding, and / or the embedding cluster can be associated with an image embedding, a text embedding, a document embedding, a multimodal embedding, and / or other embeddings. Content items and / or web resources associated with the nearest neighbor embedding, the similar embedding, and / or the embedding cluster can be obtained and / or processed to determine the identification details. The identification details can include the exact name (and / or classification) of a specific object. In some implementations, the input image can depict multiple objects, and multiple object recognition outputs can be generated based on the multiple objects.

[0067] At 306, the computing system can process the input image with a vision language model to generate a language output. The language output can include a set of predicted words predicted to describe the input image. The set of predicted words can include coarse-grained terms that describe the predicted identities of the objects depicted in the input image. In some implementations, the coarse-grained terms can include object types. The fine-grained object recognition output can include the detailed identity of the object. The set of predicted words can include one or more sentences that include the coarse-grained terms. The coarse-grained terms can include a general description of the fine-grained object recognition output (e.g., the coarse-grained term can include "vacuum", while the fine-grained recognition output can include "an RF-600 cordless Dilred XL Vacuum"). The vision language model can include one or more transformer models, one or more autoregressive language models, one or more image encoder models, one or more text encoder models, one or more decoder models, one or more diffusion models, and / or one or more other models. The vision language model can have been trained based on image-caption pairs, can have been trained by contrastive learning, can include a prefix language model, can include masked language modeling, can include image-text matching, can include learned sequence representations, can have been trained by black-box optimization, and / or can include multimodal fusion with cross-attention. In some implementations, the vision language model can be trained based on multiple templates, multiple scene types, multiple object classes, and / or multiple natural language examples.

[0068] At 308, the computing system can process the fine-grained object recognition output and the language output to generate an enhanced language output. The enhanced language output can include the set of predicted words with the coarse-grained terms replaced by the fine-grained object recognition output. The enhanced language output can be generated by processing the fine-grained object recognition output and the language output with a vision language model, an enhancement model, a natural language processing model, and / or one or more other models. The replacement can be determined by determining that the coarse-grained term is associated with the object described by the fine-grained object recognition output. Alternatively and / or additionally, the enhanced language output can include a different structure, syntax, diction, and / or orientation from the language output based on processing the fine-grained object recognition output.

[0069] In some implementations, processing the fine-grained object recognition output and the language output to generate an enhanced language output may include: processing the language output to determine a plurality of text tokens associated with features in the input image; determining that a particular token among the plurality of text tokens is associated with an object; and replacing the particular token with the fine-grained object recognition output. In some implementations, determining that a particular token among the plurality of text tokens is associated with the object may include: processing the fine-grained object recognition output with an embedding model to generate an instance-level embedding; processing the plurality of text tokens with the embedding model to generate a plurality of token embeddings; and determining that the instance-level embedding is associated with a particular embedding associated with the particular token.

[0070] In some implementations, the computing system may process the enhanced language output with a second language model to generate a natural language response to the enhanced language output. The natural language response may include additional information associated with the enhanced language output. In some implementations, the enhanced language output may be processed with a visual language model to generate a response. The response may include text data, image data, audio data, potential coding data, and / or multimodal data. The language output, object recognition output, enhanced language output, and / or response may be conditioned on and / or based on the input text available from the input image.

[0071] Additionally and / or alternatively, the computing system may provide augmented language output in an augmented reality experience. The augmented reality experience may include augmented language output superimposed on a live video feed of the environment. The augmented reality experience may be provided via a viewfinder of a mobile computing device, a smart wearable device, and / or other computing device. The augmented reality experience may be utilized to query and receive additional information about the user's environment via an augmented reality interface.

[0072] Figures 4A to 4C Depicted is an illustration of an example fine-grained object recognition with scene understanding according to an example embodiment of the present disclosure. Specifically, Figure 4A Depicted is processing input text 402 and input image 404 to generate object-aware cues for generating model-generated responses 410 .

[0073] Input text 402 may include "How much water does this need?" Input image 404 may include a plant in a pot on a windowsill. Input text 402 and input image 404 may be obtained via a mobile computing device, a smart wearable device, and / or other computing devices. Input may be obtained via one or more user interfaces, which may include a viewfinder interface, a search interface, an augmented reality interface, and / or an assistant interface.

[0074] An input image 404 can be processed with an object recognition system and a vision language model to generate recognition data 406 that describes object recognition and scene recognition. For example, the recognition data 406 can include instance-level object recognition generated with the object recognition system (e.g., Fasciated haworthia). Additionally, the recognition data 406 can include predicted image captions generated with the vision language model (e.g., “aloe vera plant in a black pot on a windowsill”, which can be associated with scene recognition). In some implementations, the recognition data 406 can include web resources identified as being associated with the objects in the input image 404. For example, a web page titled “How to care for a zebra succulent aka haworthia an easy step by step guide to everything you need to know to save a dying zebra succulent sundazesaltair” can be determined to be associated with the objects in the input image 404 based on a visual search, which can include reverse image search and / or embedding-based search. The web resource can be processed to determine that the web page is associated with the plant “Fasciated haworthia”. Then, entity recognition can be utilized to determine and / or confirm fine-grained object recognition.

[0075] The recognition data 406 and the input text 402 can be processed to generate a refined prompt 408. The refined prompt can utilize the recognition data 406 and the input text 402 to generate a prompt that includes the intent of the input text 402 and the details determined based on the input image 404. The refined prompt 408 can include “How much water does Fasciated haworthia need?”

[0076] Refine prompt 408 may be processed with one or more generative models to generate model-generated response 410. Model-generated response 410 may be responsive to input text 402 and input image 404. Model-generated response 410 may be generated by obtaining and processing web data associated with search results for refine prompts. Generative models may then utilize information from the web data to generate a summary of responses to refine prompt 408. Refined tip 408 may include "Haworthia plants don't need to be watered often because they store water efficiently. You should only water them when the soil has been completely dry for a few days. This could be every two weeks, but in warmer months or warmer climates, it could be more often. In more humid environments, it may not be as frequent."

[0077] Figure 4B Depicted is processing input text 432 and input image 434 to generate image-aware model-generated content item 444. Input text 432 can describe a prompt for generating model-generated content. For example, input text 432 can include "Write a haiku about this place". Input image 434 can depict a location with a house and a landscape with grass, bushes, and trees.

[0078] The vision-language model 436 can process the input text 432 and the input image 434 to generate a preliminary content item 438. The vision-language model 436 can process the input image 434 to generate a scene recognition, which can be processed together with the input text 432 to generate the preliminary content item 438. The preliminary content item 438 can be in response to a prompt of the text input 432 and include details from the scene recognition. The preliminary content item 438 can include “A quaint village in the countryside, with stone houses and pub. A peaceful place to stay”.

[0079] The instance-level recognition block 440 can process the input image 434 to generate an instance-level recognition 442 of the depicted location. The instance-level recognition block 440 can include a visual search system for granular recognition of objects and / or locations. The instance-level recognition 442 can include “Palant, Denmark”.

[0080] Then, the preliminary content item 438 and the instance-level recognition 442 can be processed to generate an image-aware model-generated content item 444. The image-aware model-generated content item 444 can be generated by processing the preliminary content item 438 and the instance-level recognition 442 with a generation model (e.g., the vision-language model 436). In some implementations, the image-aware model-generated content item 444 can be generated based on text token identification and replacement. The image-aware model-generated content item 444 can include “Palant in the Denmark countryside, with stone houses and pub. A peaceful place to stay”.

[0081] Figure 4C Processing of the input image 452 and the input text to generate a refined query 458 is depicted. The input image 453 can depict a specific humidifier. The input text can include a question about the depicted product. For example, the input text can include “what’s it use?”.

[0082] An input image 452 can be processed with an object recognition system and a vision language model to generate recognition data 454. For example, the recognition data 454 can include instance-level object recognition generated with the object recognition system (e.g., QLX Drip XA-Q2 and QLX Drip Humidifier). Additionally, the recognition data 454 can include predicted image captions generated with the vision language model (e.g., "white humidifier on a white background", which can be associated with scene recognition). In some implementations, the recognition data 454 can include web resources identified as being associated with the objects in the input image 452. For example, a web page titled "Qax Cool-Mist Humidifier, 1 Gal.-Clear&White" can be determined to be associated with the objects in the input image 452 based on a visual search, which can include reverse image search and / or embedding-based search. The web resources can be processed to determine that the web page is associated with the product "QLX Drip Humidifier". Then, entity recognition can be utilized to determine and / or confirm fine-grained object recognition. In some implementations, the recognition data 454 can include optical character recognition generated by performing optical character recognition on the input image 452 to determine that the input image 452 includes the text "QLX".

[0083] The recognition data 454 and the input text may be processed to generate an enhanced language output 456. The enhanced language output 456 may include a refined image text description, which may include "This image is about a white humidifier on a white background. It is about QLX Drip XA-Q2 or QLX DripHumidifier. Image comes from a page with title'Qax Cool-Mist Humidifier, 1Gal.–Clear&White'. Text on the image says: QLX". The refined image text description may be generated using a visual language model, and may utilize the object recognition output of an object recognition system and scene recognition of the visual language model.

[0084] The enhanced language output 456 and the input text may then be processed to generate a refined query 458. The refined query 458 may include "What is the use of QLX Drip Humidifier?" The refined query 458 may be generated by utilizing information from the enhanced language output 456 to provide a detailed identifier of what the user is questioning.

[0085] Figure 5 Depicted is an illustration of an example generative model response system 500 according to an example embodiment of the present disclosure. Specifically, the generative model response system 500 can obtain an input image 504 and input text 502 to generate a model-generated response 506. For example, the input text 502 can include "create a listing to sell this", the input image 504 can depict a white chair in a room, and the model-generated response 506 can include a model-generated listing.

[0086] The generation model response system 500 can utilize an object recognition system to identify a specific product depicted in the input image 504. Then, the generation model can process the object recognition and the input text 502 to generate a model-generated response 506. In some implementations, the generation model can utilize one or more application programming interfaces to obtain additional information associated with the identified product, and the additional information can be used to generate a detailed product list.

[0087] Figures 6A to 6B FIG. depicts an illustration of an example scenario description generation template in accordance with an example embodiment of the present disclosure. Specifically, Figure 6A A template 600 for query generation can be depicted. The template 600 can be utilized to prompt the generation model for question answering tasks, image captioning, and / or query generation. The template 600 can be used to prompt the generation model for masked language tasks.

[0088] The template 600 can include a question answering prompt template 602, an image caption generation template 604, and a query template 606. The question answering prompt template 602 can be used to instruct the generation model to generate and answer questions based on the input data. The image caption generation template 604 can be used to instruct the generation model to generate an image caption that fills in masked tokens with details determined based on the input image, and the details can include details associated with the entire image, selected objects in the image, and / or other objects in the image. The query template 606 can be used to instruct the generation model to generate a query based on multimodal input data.

[0089] In some implementations, the question answering prompt template 602, the image caption generation template 604, and / or the query template 606 can be processed to generate a detailed prompt. For example, the question answering prompt template 602 can be used as a preamble for prompt construction, the image caption generation template 604 can be used as a scenario description, and the query template 606 can be used to refine the intent.

[0090] Figure 6B FIG. depicts different recognition output templates 650. The recognition output templates 650 can include a dominant object template 652, a multi-object template 654, and / or a scenario description template 656. The dominant object template 652 can include a detailed list to be completed based on the outputs of an object recognition system and a vision language model. The multi-object template 654 can include a natural language template that includes masked tokens to be replaced by text tokens associated with the recognition data generated by the object recognition system and the vision language model. The scenario description template 656 can include a natural language template in the form of a scenario description that can be provided by an individual. The scenario description template 656 can include masked tokens to be replaced by text tokens associated with the recognition data generated by the object recognition system and the vision language model.

[0091] Figure 7 depicts a flowchart of an example method for execution in accordance with example embodiments of the present disclosure. Although Figure 7 the steps are depicted for purposes of illustration and discussion as being executed in a particular order, the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 700 may be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of the present disclosure.

[0092] At 702, a computing system may obtain image data. The image data may include an input image. The input image may depict an object in an environment with one or more additional objects. In some implementations, the computing system may include text data along with the image data. The text data may describe a prompt to a generative model (e.g., “generate a short story based on the object in this lake” and / or “what is origin of this structure?”). The text data may include input text that refers to the object and / or the environment.

[0093] At 704, the computing system may process the input image to generate an object recognition output. In some implementations, the computing system may process the input image with an object recognition model to generate the object recognition output. Alternatively and / or additionally, the computing system may process the input image with a visual search engine to determine one or more visual search results, and then may process the one or more visual search results to determine the object recognition output. The object recognition output may describe identification details of the object depicted in the input image. The object recognition output may be associated with the object. In some implementations, the object recognition model may include one or more classification models. The object recognition output may include instance-level object recognition associated with the object. The computing system may segment a portion of the input image associated with the object and then process the portion to generate the object recognition output. In some implementations, the image segmentation may be based on the text data. The object recognition output may be generated based on visual search optical character recognition features, object classification, and / or one or more other techniques.

[0094] At 706, the computing system can process the input image with a vision-language model to generate a scene recognition output. The scene recognition output can include a language output. The language output can include a set of predicted words predicted to describe the input image. The set of words can include terms that describe the predicted identities of the objects depicted in the input image. The language output can be associated with the objects and the environment having one or more additional objects. In some implementations, the vision-language model may have been trained based on a training data set including multiple image-caption pairs. The multiple image-caption pairs can include multiple training images and multiple corresponding captions associated with the multiple training images. In some implementations, the vision-language model can include one or more text encoders, one or more image encoders, and one or more decoders. The language output can include a scene understanding associated with the input image. In some implementations, the computing system can process image data and text data with the vision-language model to generate a language output. The language output can be a prompt based on the text data. The language output can be formatted as a generative model prompt, query, conversation message, question, and / or response. The input image can be processed in parallel with an object recognition model and a vision-language model to perform a parallel determination of an object recognition output and a language output. In some implementations, the object recognition output can be determined independently of the determination of the language output, and the language output can be determined independently of the determination of the object recognition output. Both the object recognition model and the vision-language model can process the input image separately to perform their respective determinations.

[0095] At 708, the computing system can process the object recognition output and the scene recognition output with the vision-language model to generate an enhanced language output. The enhanced language output can include a set of words with terms replaced by the object recognition output. Alternatively and / or additionally, the enhanced language output can include an enhancement of a query and / or prompt (e.g., a prompt of the language output) based on the object recognition output. In some implementations, the object recognition output, the scene recognition output, and / or the text data can be processed with a generative model to generate a model-generated response. The model-generated response can include text data, image data, audio data, latent encoding data, and / or multimodal data.

[0096] Figure 8 A flowchart of an example method for execution in accordance with an example embodiment of the present disclosure is depicted. Although Figure 8 the steps are depicted for purposes of illustration and discussion as being performed in a particular order, the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 800 can be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of the present disclosure.

[0097] At 802, the computing system may obtain image data. The image data may include an input image. The input image may depict a particular environment and / or one or more particular objects. The input image may be obtained and / or generated with a visual search application. In some implementations, the image data may be obtained and / or generated with a smart wearable device (e.g., a smartwatch, smart glasses, a smart helmet, etc.). The image data may be obtained together with other input data, which may include text data and / or audio data associated with the problem.

[0098] At 804, the computing system may process the input image to determine an object recognition output. The object recognition output may describe the identification details of the object depicted in the input image. Processing the input image to determine the object recognition output may include: processing the input image with a search engine to determine text data describing the object identification. The search engine may perform feature search, embedding-based search, optical character recognition search, and / or other search techniques. In some implementations, the search engine may identify multiple search results in response to an image query. The multiple search results may be parsed and / or processed to determine one or more object labels. The object labels may be based on processing the multiple search results with a semantic understanding model, a language model, and / or other models. In some implementations, the input image and / or the multiple search results may be embedded to determine a text embedding associated with the generated embedding. The text embedding may be decoded to determine the object labels.

[0099] Alternatively and / or additionally, processing the input image to determine the object recognition output may include: processing the input image with an embedding model to generate an image embedding; and determining one or more object labels based on the image embedding. The one or more object labels may include the identification details of the object depicted in the input image. The object labels may be index labels associated with one or more data clusters.

[0100] At 806, the computing system may process the input image with a vision-language model to generate a language output. The language output may include a set of predicted words predicted to describe the input image. The set of words may include terms describing the predicted identification of the object depicted in the input image. The vision-language model may include a language model tuned to process image data and / or multimodal data. The vision-language model may have been trained to encode images and output text. The vision-language model may include multiple individual models used serially and / or in parallel.

[0101] At 808, a computing system can process object recognition output and language output to generate enhanced language output. The enhanced language output can include a set of words with terms replaced by the object recognition output. The object recognition output and / or the language output can be processed with a generative model (e.g., a generative language model, a text-to-image generative model, etc.) to generate a model-generated output. The model-generated output can include additional information associated with one or more specific environments and / or one or more specific objects.

[0102] At 810, the computing system can determine one or more search results associated with the enhanced language output. The one or more search results can be associated with one or more web resources. In some implementations, determining one or more search results associated with the enhanced language output can include: determining that multiple search results are in response to a search query that includes the enhanced language output. Additionally and / or alternatively, the computing system can provide the multiple search results for display.

[0103] Figure 9A A block diagram of an example computing system 100 that performs instance-level scene recognition in accordance with example embodiments of the present disclosure is depicted. System 100 includes a user computing system 102, a server computing system 130, and / or a third-party computing system 150 communicatively coupled via a network 180.

[0104] The user computing system 102 can include any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or a desktop computer), a mobile computing device (e.g., a smart phone or a tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0105] The user computing system 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing system 102 to perform operations.

[0106] In some implementations, the user computing system 102 may store or include one or more machine learning models 120. For example, the machine learning model 120 may be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. The neural network may include a feed-forward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks.

[0107] In some implementations, one or more machine learning models 120 may be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, the user computing system 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel machine learning model processing across multiple instances of input data and / or detected features).

[0108] More specifically, one or more machine learning models 120 may include one or more detection models, one or more classification models, one or more segmentation models, one or more enhancement models, one or more generation models, one or more natural language processing models, one or more optical character recognition models, and / or one or more other machine learning models. One or more machine learning models 120 may include one or more transformer models. One or more machine learning models 120 may include one or more neural radiance field models, one or more diffusion models, and / or one or more autoregressive language models.

[0109] One or more machine learning models 120 may be utilized to detect one or more object features. The detected object features may be classified and / or embedded. Then, the classification and / or embedding may be utilized to perform a search to determine one or more search results. Alternatively and / or additionally, one or more detected features may be utilized to determine an indicator to provide (e.g., a user interface element indicating the detected feature) to indicate that a feature has been detected. Then, the user may select the indicator to cause the feature classification, embedding, and / or search to be performed. In some implementations, the classification, embedding, and / or search may be performed before the indicator is selected.

[0110] In some implementations, one or more machine learning models 120 can process image data, text data, audio data, and / or latent encoded data to generate output data, which can include image data, text data, audio data, and / or latent encoded data. One or more machine learning models 120 can perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, scene determination, action prediction, image correction, image enhancement, text enhancement, sentiment analysis, object detection, error detection, image restoration, video stabilization, audio correction, audio enhancement, and / or data segmentation (e.g., mask-based segmentation).

[0111] Additionally or alternatively, one or more machine learning models 140 can be included in or otherwise stored and implemented by a server computing system 130 that communicates with a user computing system 102 according to a client-server relationship. For example, a machine learning model 140 can be implemented by the server computing system 130 as part of a web service (e.g., a viewfinder service, a visual search service, an image processing service, an environmental computing service, and / or an overlay application service). Thus, one or more models 120 can be stored at and implemented at the user computing system 102, and / or one or more models 140 can be stored at and implemented at the server computing system 130.

[0112] The user computing system 102 can also include one or more user input components 122 that receive user input. For example, a user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other components through which a user can provide user input.

[0113] In some implementations, the user computing system can store and / or provide one or more user interfaces 124 that can be associated with one or more applications. One or more user interfaces 124 can be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, an augmented reality experience, a virtual reality experience, and / or other data for display). The user interface 124 can be associated with one or more other computing systems (e.g., the server computing system 130 and / or a third-party computing system 150). The user interface 124 can include a viewfinder interface, a search interface, a generative model interface, a social media interface, and / or a media content gallery interface.

[0114] The user computing system 102 may include one or more sensors 126 and / or receive data from one or more sensors. The one or more sensors 126 may be housed in a housing assembly that houses one or more processors 112, a memory 114, and / or one or more hardware components that may store and / or cause the execution of one or more software packages. The one or more sensors 126 may include one or more image sensors (e.g., cameras), one or more lidar sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biosensors (e.g., heart rate sensors, pulse sensors, retina sensors, and / or fingerprint sensors), one or more infrared sensors, one or more position sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. One or more sensors may be utilized to obtain data associated with the user environment (e.g., an image of the user environment, a recording of the environment, and / or the location of the user).

[0115] The user computing system 102 may include a user computing device 104 and / or may be part of a user computing device. The user computing device 104 may include a mobile computing device (e.g., a smart phone or a tablet computer), a desktop computer, a laptop computer, a smart wearable device, and / or a smart appliance. Additionally and / or alternatively, the user computing system may obtain data from one or more user computing devices 104 and / or generate data with one or more user computing devices. For example, the camera of a smart phone may be utilized to capture image data depicting the environment, and / or an overlay application of the user computing device 104 may be utilized to track and / or process data provided to the user. Similarly, one or more sensors associated with a smart wearable device may be utilized to obtain data about the user and / or about the user's environment (e.g., image data may be obtained with a camera housed in the user's smart glasses). Additionally and / or alternatively, data may be obtained and uploaded from other user devices that may be dedicated to data acquisition or generation.

[0116] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138, which are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0117] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by one or more server computing devices. In the case where the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0118] As described above, the server computing system 130 can store or otherwise include one or more machine learning models 140. For example, the model 140 can be or otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Refer to Figure 9B for a discussion of example model 140.

[0119] Additionally and / or alternatively, the server computing system 130 can include a search engine 142 and / or can be communicatively connected to a search engine, which can be used to crawl one or more databases (and / or resources). The search engine 142 can process data from the user computing system 102, the server computing system 130, and / or the third-party computing system 150 to determine one or more search results associated with the input data. The search engine 142 can perform term-based searches, tag-based searches, boolean-based searches, image searches, embedding-based searches (e.g., nearest neighbor searches), multimodal searches, and / or one or more other search techniques.

[0120] The server computing system 130 may store and / or provide one or more user interfaces 144 for obtaining input data and / or providing output data to one or more users. The one or more user interfaces 144 may include one or more user interface elements, and the one or more user interface elements may include input fields, navigation tools, content chips, selectable tiles, widgets, data display carousels, dynamic animations, information pop-up windows, image enhancement, text-to-speech, speech-to-text, augmented reality, virtual reality, feedback loops, and / or other interface elements.

[0121] The user computing system 102 and / or the server computing system 130 may train the models 120 and / or 140 via interaction with a third-party computing system 150 communicatively coupled via the network 180. The third-party computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130. Alternatively and / or additionally, the third-party computing system 150 may be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more contexts.

[0122] The third-party computing system 150 may include one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operably connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158, which are executed by the processor 152 to cause the third-party computing system 150 to perform operations. In some implementations, the third-party computing system 150 includes one or more server computing devices or is otherwise implemented by one or more server computing devices.

[0123] The network 180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. Generally, communication over the network 180 may use a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL), and may be carried via any type of wired and / or wireless connection.

[0124] The machine learning models described in this specification can be used for various tasks, applications, and / or use cases.

[0125] In some implementations, the input to the machine learning models of the present disclosure can be image data. The machine learning models can process the image data to generate an output. As an example, the machine learning models can process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning models can process the image data to generate an image segmentation output. As another example, the machine learning models can process the image data to generate an image classification output. As another example, the machine learning models can process the image data to generate an image data modification output (e.g., change of the image data, etc.). As another example, the machine learning models can process the image data to generate an encoded image data output (e.g., encoded and / or compressed representation of the image data, etc.). As another example, the machine learning models can process the image data to generate an enlarged image data output. As another example, the machine learning models can process the image data to generate a prediction output.

[0126] In some implementations, the input to the machine learning models of the present disclosure can be text or natural language data. The machine learning models can process the text or natural language data to generate an output. As an example, the machine learning models can process the natural language data to generate a language encoding output. As another example, the machine learning models can process the text or natural language data to generate a latent text embedding output. As another example, the machine learning models can process the text or natural language data to generate a translation output. As another example, the machine learning models can process the text or natural language data to generate a classification output. As another example, the machine learning models can process the text or natural language data to generate a text segmentation output. As another example, the machine learning models can process the text or natural language data to generate a semantic intent output. As another example, the machine learning models can process the text or natural language data to generate an upgraded text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language, etc.). As another example, the machine learning models can process the text or natural language data to generate a prediction output.

[0127] In some implementations, the input to the machine learning model of the present disclosure can be speech data. The machine learning model can process the speech data to generate an output. As an example, the machine learning model can process the speech data to generate a speech recognition output. As another example, the machine learning model can process the speech data to generate a speech translation output. As another example, the machine learning model can process the speech data to generate a latent embedding output. As another example, the machine learning model can process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, the machine learning model can process the speech data to generate an enhanced speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine learning model can process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model can process the speech data to generate a prediction output.

[0128] In some implementations, the input to the machine learning model of the present disclosure can be sensor data. The machine learning model can process the sensor data to generate an output. As an example, the machine learning model can process the sensor data to generate a recognition output. As another example, the machine learning model can process the sensor data to generate a prediction output. As another example, the machine learning model can process the sensor data to generate a classification output. As another example, the machine learning model can process the sensor data to generate a segmentation output. As another example, the machine learning model can process the sensor data to generate a segmentation output. As another example, the machine learning model can process the sensor data to generate a visualization output. As another example, the machine learning model can process the sensor data to generate a diagnostic output. As another example, the machine learning model can process the sensor data to generate a detection output.

[0129] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to the object class. The image processing task can be object detection, where the image processing output identifies one or more regions in one or more images and, for each region, identifies the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in one or more images, the corresponding likelihood of each of a set of predefined classes. For example, a set of classes can be foreground and background. As another example, a set of classes can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines a corresponding depth value for each pixel in one or more images. As another example, the image processing task can be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel in one of the input images, the motion of the scene depicted at the pixel between the images in the network input.

[0130] A user computing system can include multiple applications (e.g., Application 1 through N). Each application can include its own respective machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.

[0131] Each application can communicate with multiple other components of the computing device (such as, for example, one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a common API). In some implementations, the API used by each application is specific to the application.

[0132] User computing system 102 can include multiple applications (e.g., Application 1 through N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).

[0133] The central intelligence layer may include multiple machine learning models. For example, the central intelligence layer management may provide corresponding machine learning models for each application and manage the corresponding machine learning models. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model for all application programs. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of computing device 100.

[0134] The central intelligence layer may communicate with the central device data layer. The central device data layer may be a centralized data repository of computing device 100. The central device data layer may communicate with multiple other components of the computing device (such as, for example, one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, the central device data layer may communicate with each device component using an API (e.g., a private API).

[0135] Figure 9B A block diagram of an example computing system 50 that performs instance-level scene recognition according to an example embodiment of the present disclosure is depicted. Specifically, the example computing system 50 may include one or more computing devices 52 that may be used to obtain and / or generate one or more data sets, which may be processed by a sensor processing system 60 and / or an output determination system 80 and fed back to a user, who may provide information about features in one or more of the obtained data sets. The one or more data sets may include image data, text data, audio data, multimodal data, latent encoded data, and the like. The one or more data sets may be obtained via one or more sensors associated with the one or more computing devices 52 (such as one or more sensors in computing device 52). Additionally and / or alternatively, the one or more data sets may be stored data and / or retrieved data (such as data retrieved from a web resource). For example, an image, text, and / or other content item may interact with the user. Then, the interaction with the content item may be utilized to generate one or more determinations.

[0136] One or more computing devices 52 may obtain and / or generate one or more data sets based on image capture, sensor tracking, data storage retrieval, content download (e.g., downloading images or other content items from web resources via the Internet) and / or via one or more other techniques. The one or more data sets may be processed by a sensor processing system 60. The sensor processing system 60 may perform one or more processing techniques using one or more machine learning models, one or more search engines, and / or one or more other processing techniques. The one or more processing techniques may be performed in any combination and / or individually. The one or more processing techniques may be performed serially and / or in parallel. Specifically, the one or more data sets may be processed by a context determination block 62, which may determine a context associated with one or more content items. The context determination block 62 may identify and / or process metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data to determine a specific context associated with the user. The context may be associated with an event, a determined trend, a specific action, a specific data type, a specific environment, and / or another context associated with the user and / or the retrieved or obtained data.

[0137] The sensor processing system 60 may include an image preprocessing block 64. The image preprocessing block 64 may be used to adjust one or more values of the obtained and / or received images to prepare the images for processing by one or more machine learning models and / or one or more search engines 74. The image preprocessing block 64 may resize the image, adjust the saturation value, adjust the resolution, strip and / or add metadata, and / or perform one or more other operations.

[0138] In some implementations, the sensor processing system 60 may include one or more machine learning models, which may include a detection model 66, a segmentation model 68, a classification model 70, an embedding model 72, and / or one or more other machine learning models. For example, the sensor processing system 60 may include one or more detection models 66, which may be used to detect specific features in the processed data set. Specifically, one or more images may be processed by one or more detection models 66 to generate one or more bounding boxes associated with the detected features in the one or more images.

[0139] Additionally and / or alternatively, one or more segmentation models 68 can be utilized to segment one or more portions of a dataset from one or more datasets. For example, one or more segmentation models 68 can utilize one or more segmentation masks (e.g., one or more segmentation masks generated manually and / or based on one or more bounding boxes) to segment a portion of an image, a portion of an audio file, and / or a portion of text. Segmentation can include isolating one or more detected objects and / or removing one or more detected objects from an image.

[0140] One or more classification models 70 can be used to process image data, text data, audio data, latent encoding data, multimodal data, and / or other data to generate one or more classifications. One or more classification models 70 can include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. One or more classification models 70 can process data to determine one or more classifications.

[0141] In some implementations, data can be processed with one or more embedding models 72 to generate one or more embeddings. For example, one or more images can be processed with one or more embedding models 72 to generate one or more image embeddings in an embedding space. One or more image embeddings can be associated with one or more image features of one or more images. In some implementations, one or more embedding models 72 can be configured to process multimodal data to generate multimodal embeddings. One or more embeddings can be used for classification, search, and / or learning the embedding space distribution.

[0142] The sensor processing system 60 can include one or more search engines 74, which can be used to perform one or more searches. One or more search engines 74 can crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more dedicated databases, and / or one or more general databases) to determine one or more search results. One or more search engines 74 can perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), metadata-based search, multimodal search, web resource search, image search, text search, and / or application search.

[0143] Additionally and / or alternatively, the sensor processing system 60 may include one or more multimodal processing blocks 76, which may be used to assist in processing multimodal data. The one or more multimodal processing blocks 76 may include generating multimodal queries and / or multimodal embeddings for processing by one or more machine learning models and / or one or more search engines 74.

[0144] Then, the output of the sensor processing system 60 can be processed by an output determination system 80 to determine one or more outputs to be provided to the user. The output determination system 80 may include heuristic-based determination, machine learning model-based determination, user selection-based determination, and / or context-based determination.

[0145] The output determination system 80 can determine how and / or where to provide one or more search results in a search result interface 82. Additionally and / or alternatively, the output determination system 80 can determine how and / or where to provide one or more machine learning model outputs in a machine learning model output interface 84. In some implementations, one or more search results and / or one or more machine learning model outputs can be provided for display via one or more user interface elements. The one or more user interface elements may be superimposed on the displayed data. For example, one or more detection indicators may be superimposed on detected objects in a viewfinder. The one or more user interface elements may be selectable to perform one or more additional searches and / or one or more additional machine learning model processes. In some implementations, the user interface elements may be provided as dedicated user interface elements for a specific application and / or may be provided uniformly across different applications. The one or more user interface elements may include pop-up window displays, interface overlays, interface tiles and / or chips, carousel interfaces, audio feedback, animations, interactive widgets, and / or other user interface elements.

[0146] Additionally and / or alternatively, data associated with the output of the sensor processing system 60 can be utilized to generate and / or provide an augmented reality experience and / or a virtual reality experience 86. For example, one or more acquired data sets can be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which can then be utilized to provide the user with an augmented reality experience and / or a virtual reality experience 86. The augmented reality experience can render information associated with the environment into the corresponding environment. Alternatively and / or additionally, objects associated with the processed data set can be rendered into the user environment and / or virtual environment. Rendering data set generation may include training one or more neural radiance field models to learn three-dimensional representations of one or more objects.

[0147] In some implementations, one or more action prompts 88 can be determined based on the output of the sensor processing system 60. For example, a search prompt, a purchase prompt, a generation prompt, a reservation prompt, a call prompt, a redirection prompt, and / or one or more other prompts can be determined to be associated with the output of the sensor processing system 60. Then, the one or more action prompts 88 can be provided to the user via one or more selectable user interface elements. In response to the selection of one or more selectable user interface elements, the corresponding action of the corresponding action prompt can be performed (e.g., a search can be performed; a purchase application programming interface can be utilized; and / or another application can be opened).

[0148] In some implementations, one or more generation models 90 can process one or more data sets and / or the output of the sensor processing system 60 to generate model-generated content items, and then the model-generated content items can be provided to the user. The generation can be prompted based on user selection and / or can be performed automatically (e.g., automatically based on one or more conditions that can be associated with a threshold amount of unidentifiable search results).

[0149] One or more generation models 90 can include language models (e.g., large language models and / or vision language models), image generation models (e.g., text-to-image generation models and / or image enhancement models), audio generation models, video generation models, graphics generation models, and / or other data generation models (e.g., other content generation models). One or more generation models 90 can include one or more transformer models, one or more convolutional neural networks, one or more recurrent neural networks, one or more feedforward neural networks, one or more generative adversarial networks, one or more self-attention models, one or more embedding models, one or more encoders, one or more decoders, and / or one or more other models. In some implementations, one or more generation models 90 can include one or more autoregressive models (e.g., machine learning models trained to generate predicted values based on previous behavioral data) and / or one or more diffusion models (e.g., machine learning models trained to generate predicted data based on generating and processing distribution data associated with input data).

[0150] One or more generative models 90 can be trained to process input data and generate generative content items of the model, and the generative content items of the model can include multiple predicted words, pixels, signals, and / or other data. The content items generated by the model can include novel content items that are not the same as any existing works. One or more generative models 90 can utilize learned representations, sequences, and / or probability distributions to generate content items, which can include phrases, storylines, occasions, objects, characters, beats, lyrics, and / or other aspects not included in existing content items.

[0151] One or more generative models 90 can include vision-language models. The vision-language models can be trained, tuned, and / or configured to process image data and / or text data to generate natural language outputs. The vision-language models can utilize pre-trained large language models (e.g., large autoregressive language models) having one or more encoders (e.g., one or more image encoders and / or one or more text encoders) to provide detailed natural language outputs that mimic natural language written by humans.

[0152] The vision-language models can be used for zero-shot image classification, few-shot image classification, image caption generation, multi-modal query refinement, multi-modal question answering, and / or can be tuned and / or trained for multiple different tasks. The vision-language models can perform visual question answering, image caption generation, feature detection (e.g., content monitoring (e.g., for inappropriate content)), object detection, scene recognition, and / or other tasks.

[0153] The vision-language models can utilize pre-trained language models and then can tune the pre-trained language models for multi-modal. The training and / or tuning of the vision-language models can include image-text matching, masked language modeling, multi-modal fusion with cross-attention, contrastive learning, prefix language model training, and / or other training techniques. For example, the vision-language models can be trained to process images to generate predicted text similar to the ground truth text data (e.g., the ground truth caption of the image). In some implementations, the vision-language models can be trained to replace masked tokens of a natural language template with text tokens that describe the features depicted in the input image. Alternatively and / or additionally, the training, tuning, and / or model inference can include multi-level concatenation of visual embedding features and text embedding features. In some implementations, the vision-language models can be trained and / or tuned via joint learning of image embedding and text embedding generation, which can include training and / or tuning a system that maps embeddings to a joint feature embedding space, and the system maps text features and image features to a shared embedding space. The joint training can include parallel embedding of image-text pairs and / or can include triplet training. In some implementations, an image can be used as and / or processed as a prefix of the language model.

[0154] The output determination system 80 can process one or more data sets and / or the output of the sensor processing system 60 with the data enhancement block 92 to generate enhanced data. For example, one or more images can be processed with the data enhancement block 92 to generate one or more enhanced images. Data enhancement can include data correction, data cropping, removal of one or more features, addition of one or more features, resolution adjustment, lighting adjustment, saturation adjustment, and / or other enhancements.

[0155] In some implementations, it can be determined based on the data storage block 94 to store one or more data sets and / or the output of the sensor processing system 60.

[0156] Then, the output of the output determination system 80 can be provided to the user via one or more output components of the user computing device 52. For example, one or more user interface elements associated with one or more outputs can be provided for display via the visual display of the user computing device 52.

[0157] The process can be performed iteratively and / or continuously. One or more user inputs to the provided user interface elements can adjust and / or affect the continuous processing loop.

[0158] The techniques discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for many possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0159] Although the subject matter has been described in detail with respect to various specific example embodiments of the subject matter, each example is provided by way of explanation and not as a limitation of the disclosure. Those skilled in the art can readily produce changes, alterations, and equivalents of such embodiments upon understanding the foregoing. Accordingly, the disclosure does not exclude including such modifications, alterations, and / or additions to the subject matter that would be readily understood by one of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used with another embodiment to produce yet another embodiment. Accordingly, the disclosure is intended to cover such changes, alterations, and equivalents.

Claims

1. A computer-implemented method, the method comprising: obtaining, by a computing system including one or more processors, image data, wherein the image data includes an input image; processing, by the computing system, the input image with an object recognition model to generate a fine-grained object recognition output, wherein the fine-grained object recognition output describes identifying details of an object depicted in the input image; processing, by the computing system, the input image with a visual language model to generate a language output, wherein the language output comprises a set of predicted words predicted to describe the input image, wherein the set of predicted words comprises coarse-grained terms describing a predicted identity of the object depicted in the input image; as well as The fine-grained object recognition output and the language output are processed by the computing system to generate an enhanced language output, wherein the enhanced language output includes the set of predicted words with the coarse-grained terms replaced with the fine-grained object recognition output.

2. The method of claim 1 , wherein processing, by the computing system, the input image with the object recognition model to generate the fine-grained object recognition output comprises: detecting the object in the input image; Generate object embeddings; determining image clusters associated with the object embedding; as well as Web resources associated with the image cluster are processed to determine identification details of the object.

3. The method of claim 2, wherein generating the object embedding comprises: generating a bounding box associated with a location of the object within the input image; generating an image segment based on the bounding box; as well as The image segment is processed with an embedding model to generate the object embedding.

4. The method of claim 1 , wherein processing, by the computing system, the fine-grained object recognition output and the language output to generate the enhanced language output comprises: processing, by the computing system, the language output to determine a plurality of text words associated with features in the input image; determining, by the computing system, that a particular word-gram of the plurality of text words-grams is associated with the object; as well as The particular word-gram is replaced by the computing system with the fine-grained object recognition output.

5. The method of claim 4, wherein determining, by the computing system, that the particular word-gram of the plurality of text words-grams is associated with the object comprises: Processing, by the computing system, the fine-grained object recognition output with an embedding model to generate an instance-level embedding; Processing the plurality of text tokens using the embedding model by the computing system to generate a plurality of token embeddings; as well as The computing system determines that the instance-level embedding is associated with a particular embedding associated with the particular word-gram.

6. The method of claim 1, further comprising: The enhanced language output is processed by the computing system with a second language model to generate a natural language response to the enhanced language output.

7. The method of claim 6, wherein the natural language response includes additional information associated with the enhanced language output.

8. The method of claim 1, wherein the coarse-grained terms include object types, and wherein the fine-grained object recognition output includes a detailed identification of the object.

9. The method of claim 1, further comprising: The augmented language output is provided by the computing system in an augmented reality experience.

10. The method of claim 9, wherein the augmented reality experience includes the augmented language output superimposed on a live video feed of an environment.

11. A computing system for generating text descriptions of images, the system comprising: one or more processors; as well as one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising: obtaining image data, wherein the image data comprises an input image; processing the input image with an object recognition model to generate an object recognition output, wherein the object recognition output describes identification details of an object depicted in the input image; processing the input image with a visual language model to generate a language output, wherein the language output comprises a set of words predicted to describe the input image, wherein the set of words comprises terms describing the predicted identification of the object depicted in the input image; and The object recognition output and the language output are processed with the visual language model to generate an enhanced language output, wherein the enhanced language output includes the set of words with the terms replaced with the object recognition output.

12. The system of claim 11, wherein the input image depicts the object in an environment with one or more additional objects, wherein the object recognition output is associated with the object, and wherein the language output is associated with the object and the environment with the one or more additional objects.

13. The system of claim 11, wherein the visual language model is trained based on a training dataset comprising a plurality of image-caption pairs, wherein the plurality of image-caption pairs comprises a plurality of training images and a plurality of corresponding captions associated with the plurality of training images.

14. The system of claim 11, wherein the input image is processed in parallel with the object recognition model and the visual language model to perform parallel determination of the object recognition output and the language output, and wherein the visual language model includes one or more text encoders, one or more image encoders, and one or more decoders.

15. The system of claim 11, wherein the object recognition model comprises one or more classification models.

16. The system of claim 11, wherein the object recognition output comprises instance-level object recognition associated with the object, and wherein the language output comprises scene understanding associated with the input image.

17. One or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising: obtaining image data, wherein the image data comprises an input image; processing the input image to determine an object recognition output, wherein the object recognition output describes identification details of an object depicted in the input image; processing the input image with a visual language model to generate a language output, wherein the language output comprises a set of words predicted to describe the input image, wherein the set of words comprises terms describing a predicted identity of the object depicted in the input image; processing the object recognition output and the language output to generate an enhanced language output, wherein the enhanced language output includes the set of words with the lexical items replaced by the object recognition output; as well as One or more search results associated with the enhanced language output are determined, wherein the one or more search results are associated with one or more web resources.

18. The one or more non-transitory computer-readable media of claim 17, wherein processing the input image to determine the object recognition output comprises: The input image is processed with a search engine to determine textual data describing an identification of an object.

19. The one or more non-transitory computer-readable media of claim 17, wherein processing the input image to determine the object recognition output comprises: Processing the input image with an embedding model to generate an image embedding; One or more object tags are determined based on the image embedding, wherein the one or more object tags include the identifying details of the object depicted in the input image.

20. The one or more non-transitory computer-readable media of claim 17, wherein determining the one or more search results associated with the enhanced language output comprises: determining that a plurality of search results are responsive to a search query including the enhanced language output; and The operations further include providing the plurality of search results for display.

Citation Information

Patent Citations

  • Video description method and system based on multistage prediction architecture

    CN110674783A

  • Training method and device of visual language pre-training model, equipment and medium

    CN114022735A