Combining retrieval and generative models into a hybrid system

The hybrid system addresses computational inefficiencies by combining retrieval and generative models to optimize resource use and enhance content quality and variety in digital applications.

WO2025165363A1PCT designated stage Publication Date: 2025-08-07GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/013971
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing systems rely solely on either retrieval or generation of digital content, leading to computational inefficiencies and limitations in variety and quality of output, especially in real-time applications like digital assistants.

Method used

A hybrid system that intelligently combines retrieval and generative models based on the relevance of user queries to determine the most efficient use of resources, reusing existing content where possible and leveraging both models to create a wider variety of outputs.

Benefits of technology

This approach reduces computational overhead, enhances response times, and increases the quality and variety of content by selectively using generative models, optimizing CPU/GPU/TPU resources and minimizing memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024013971_07082025_PF_FP_ABST
    Figure US2024013971_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems for providing intelligent hybrid search are disclosed herein. The method can include receiving a user query as an input, the user query including one or more portions and processing the one or more portions of the user query to determine a relevance of the one or more portions of the user query to a retrieval system and one or more generative models. The method can also include providing at least one of the one or more portions of the user query to the retrieval system and the one or more generative models based on the determined relevance and receiving content from the retrieval system and one or more generative models. The method can further include providing at least a portion of the received content as results to the user query.
Need to check novelty before this filing date? Find Prior Art

Description

COMBINING RETRIEVAL AND GENERATIVE MODELS INTO A HYBRID SYSTEMFIELD[1] The present disclosure relates generally to processing of queries. More particularly, the present disclosure relates to a hybrid query response system that intelligently combines both retrieval and generative models to provide content in response to a query received from a user.BACKGROUND[2] In the field of digital content generation and retrieval, there has been a significant technological advancement with the advent of generative models. These models have the capability to create a vast array of original content such as text, images, video, and music. However, the generation of such content is computationally intensive, necessitating the use of powerful hardware resources like CPU / GPU / TPU and substantial memory allocation.[3] A technical challenge, therefore, is to develop a more efficient and resource-saving method to generate and retrieve digital content, without compromising on the quality and relevance of the output. This problem is particularly pertinent when dealing with applications like digital assistants, which often need to respond to user queries in real time and generate or retrieve the required content promptly.[4] In particular, in certain existing systems, generative models are triggered for every query, regardless of whether the required content already exists or not. This approach can result in a significant computational burden.SUMMARY[5] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.[6] One example embodiment of the present disclosure is directed to a computer- implemented method for performing intelligent hybrid search. The method can include receiving, at one or more processors, a user query as an input, the user query including one or more portions and processing, by the one or more processors, the one or more portions of the user query to determine a relevance of the one or more portions of the user query to a retrieval system and one or more generative models. The method can also include providing, by the one or more processors, at least one of the one or more portions of the user query tothe retrieval system and the one or more generative models based on the determined relevance and receiving, by the one or more processors, content from the retrieval system and one or more generative models. The method can further include providing, by the one or more processors, at least a portion of the received content as results to the user query.[7] Another example embodiment of the present disclosure is directed to a computing system. The computing system can include one or more processors and a non-transitory, computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations can include receiving a user query as an input, the user query including one or more portions and processing the one or more portions of the user query to determine a relevance of the one or more portions of the user query to a retrieval system and one or more generative models. The operations can also include providing at least one of the one or more portions of the user query to the retrieval system and the one or more generative models based on the determined relevance and receiving content from the retrieval system and one or more generative models. The operations can further include providing at least a portion of the received content as results to the user query.[8] A further example embodiment of the present disclosure is directed to a non- transitory, computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations can include receiving a user query as an input, the user query including one or more portions and processing the one or more portions of the user query to determine a relevance of the one or more portions of the user query to a retrieval system and one or more generative models. The operations can also include providing at least one of the one or more portions of the user query to the retrieval system and the one or more generative models based on the determined relevance and receiving content from the retrieval system and one or more generative models. The operations can further include providing at least a portion of the received content as results to the user query.[9] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0010] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Detailed discussion of embodiments directed to one of ordinary skill in the art is set forth in the specification, which makes reference to the appended figures, in which:

[0012] Figure 1 A depicts a block diagram of an example content search system according to example embodiments of the present disclosure.

[0013] Figure IB depicts a block diagram of an example content search system according to example embodiments of the present disclosure.

[0014] Figure 2 depicts a flow chart diagram of an example method to perform content retrieval and generation according to example embodiments of the present disclosure.

[0015] Figure 3 A depicts a block diagram of an example computing system that performs content retrieval and generation according to example embodiments of the present disclosure.

[0016] Figure 3B depicts a block diagram of an example computing device that performs content retrieval and generation according to example embodiments of the present disclosure.

[0017] Figure 3C depicts a block diagram of an example computing device that performs content retrieval and generation according to example embodiments of the present disclosure.

[0018] Reference numerals that are repeated across plural figures are intended to identify the same features in various implementations.DETAILED DESCRIPTIONOverview

[0019] The present disclosure describes a method for combining retrieval of existing content with generative models to create a virtually infinite corpus of media, including text, images, audio, and video. This technology can be integrated into various applications such as digital assistants, editor applications, or video editors. It can operate in different modes, such as combining retrieved media with generated content, retrieval instead of generation, and merging generated and retrieved results.

[0020] One problem with prior techniques is that they relied solely on either retrieved or generated content. This could lead to limitations in the variety and quality of output, and could also result in unnecessary computational overhead when generating content that already exists in a retrievable form. The present disclosure addresses these issues by intelligently combining retrieval and generation based on the specific user query and available resources.

[0021] The advantages of this new technique include saving computational resources by reusing existing content where possible, providing a wider variety’ of output by combining real and generated content, and offering more flexibility- in responding to user queries. The system can process queries in sequence or in parallel, depending on the query and the available content. It also includes a ranking system to determine which results should be returned to the user, and can merge different media channels, such as adding a generated audio track to an existing video. The system can also attribute content used as a generative seed for revenue sharing purposes.

[0022] Thus, the present disclosure provides a method for combining retrieval and generative models into a single, hybrid system, offering significant improvements in terms of computational efficiency, output variety, and user experience. In particular, the proposed hybrid system can provide the highest possible qualify content from both existing content and generated content while enabling the user of the system to be as expressive as possible in terms of input prompt.

[0023] More particularly, an example system can process various inputs and prompts to generate a textual prompt for one or more retrieval models and / or one or more generative models. For example, the system can receive text input from a user, voice input from a user that is processed by natural language processing, and the like. In some embodiments, the user can provide the input into a software application, such as a digital assistant, a document editing application, a video editing application, and the like.

[0024] In some embodiments, the input can also include images, video, audio, or other content. For example, a user may yvish to find or generate music to play over a video clip. The user can therefore provide as an input “please generate relaxing music for this video” and a link to a video or a video file. The system can process both the text input and the video to identify and / or generate content as described below.

[0025] After receiving the input, the system can process the input to determine if the input is relevant to one or more available generative models. In particular, the proposed system can analyze or process a user's query (or certain portions thereof) to ascertain its relevance to retrieval systems versus generative systems. In some implementations, an initial task is to parse and interpret the user's query to extract meaningful information and determine the context of the query.

[0026] For example, the proposed system can analyze the query by applying natural language processing techniques to break down the query into understandable components. This includes identifying critical terms, phrases, or instructions that indicate the user's intent.For example, the use of phrases like "generate a...", "show me...", or "play..." might imply a preference for generative models, while other queries might be more open-ended and relevant to both retrieval and generative systems.

[0027] The proposed system can proceed to determine the relevance of these components to the existing retrieval and generative models. This stage can include matching the query components to the capabilities of the retrieval system and the generative models. For example, each component of the query can be assessed based on predefined criteria or algorithms to evaluate its suitability for retrieval or generation. For instance, a request for a specific song would likely be more suitable for a retrieval system that can search for existing versions of the song, while an open-ended request for a "calm Italian song" might be more suited to a generative model capable of creating such content.

[0028] In some implementations, the system's ability to discern the relevance of the uery to retrieval versus generative systems is not a binary process but rather a spectrum of relevance. For example, the system can assign scores to each component of the query based on its relevance to both systems. These scores can be used to make intelligent decisions about the best approach to satisfy the user's query, whether it involves retrieving existing content, generating new content, or a combination of both. This adaptive, context-aware mechanism ensures the most efficient use of resources while delivering content that accurately meets the user's request.

[0029] To provide an example, using natural language processing, the input can be processed to determine if any portion of the input is relevant to a model, such as determining that the input contains the phrase “play relaxing jazz” or “show me a photo of a robot making pizza,” which can be relevant to audio generative models or image generative models. In some embodiments, the input can explicitly ask for the use of a generative model by using phrases such as “generate a” or “show me” or “make a for me” or other similar phrases. If the input explicitly asks for generation, the entire input or a portion of the input can be provided directly to the proper generative model and the result can be returned directly to the user without the need for further processing. Thus, in some implementations, the hybrid system can leverage only the generative models.

[0030] In other embodiments, the system can determine that the query relates only to retrieval of existing content. For example, in scenarios where the retrieved content fully satisfies the user query, the system can opt to skip the generation process, thus saving computational resources.

[0031] In yet further embodiments, the input can be applicable to both retrievable or generative content. In these embodiments, the input can be provided to both retrieval models and generative models.

[0032] In some embodiments, this use of both retrieval models and generative models can be performed in sequence. First, the input can be provided to retrieval models, which can generate an embedding or token associated with one or more portions of the input and, based on the embeddings or tokens, identify one or more pieces of retrievable content in one or more data stores that are similar to the embedding or token. If there are no items that are considered similar enough to the input embedding or token, a request can then be issued for portions of the input without a similar enough match in the one or more data stores to the generative models to generate content based on the portions of the input.

[0033] In some embodiments, the input can be processed in parallel. For example, the input can be provided to both the retrieval models and the generative models, and both the retrieval models and the generative models can process the input to identify or generate relevant content, respectively. In some embodiments, if highly relevant content is found in one or more data stores by the retrieval models, the retrieval models can send a notification to the generative models that highly relevant content has been identified, and that there is no need for generated content. This can “short-circuit” or otherw ise stop generative model processing of the input, which can save processing cycles, network bandwidth, and memory.

[0034] Based on the retrieved and / or generated content, the system can then rank results to determine which result(s) should be returned as a response to the user. Results can be ranked based on a variety of factors: similarity to the input, similarity to one or more user preferences, and the like. In some embodiments, retrieved content can be given priority over generated content, as it may more accurately reflect the desires of the user. Various quality metrics can be used in addition to or in replacement of similarity metrics in order to rank the results.

[0035] In some embodiments, different media channels and / or items of retrieved content can be merged. As one example, the system could combine the output of a retrieval system that identifies a suitable video clip with the output of a generative system that produces an audio track. For instance, a user may request a video of a sunset with a calming piano soundtrack. The retrieval system might locate a video of a sunset from an existing database, but the original audio may not meet the user's request. The generative system might then be engaged to create a calming piano track. The system could then combine the retrieved video with the generated audio to produce a final output that meets the user's requirements.

[0036] In another example, the system could combine the outputs of both retrieval and generative systems in response to a user's request for a unique image composition. For example, a user may request an image of a cat sitting on a moon. The retrieval system could locate an image of a cat from a database, while the generative system could be tasked with creating an image of a moon. The system could then merge the outputs of the retrieval and generative systems, positioning the retrieved image of the cat onto the generated image of the moon to create the desired composition.

[0037] In yet another example, the system could generate a new song based on a user's musical preferences and an existing song from a database. The retrieval system could identify the existing song that matches the user's preference, and the generative system could then generate a new melody or lyrics based on the identified song. The system could then combine the generated melody or lyrics with the retrieved song, creating a new song that meets the user's musical preference.

[0038] The final content can then be returned to the user. In some embodiments, if multiple content items have a high ranking, the multiple content items can be returned to the user, which enables the user to pick a desired content item from the multiple content items.

[0039] In some embodiments, the type of content can be indicated to the user. For example, if the content is retrieved, the system can indicate to the user “Here is content I found during a search. I found the content at (location)'’ or another suitable indication that the content was retrieved. If the content is generated, the system can indicate “I could not find anything suitable on the web, so here’s something I generated” or “Here is what you requested be generated” or another suitable indication that the content was generated.

[0040] Depending on the application and the user interface, the user can issue subsequent queries to refine outputs, especially generated outputs. For example, a user may ask “Make it faster” about a generated song. The request can be applied to the generated content item by providing the generated content item and the new prompt as conditioned inputs into the generative models.

[0041] Aspects of the present invention provide a number of technical advantages over existing systems. First, by using and reusing existing content, the number of computing cycles (processing cycles, CPU / GPU / TPU cycles) can be reduced, and the amount of required memory' needed for storing content can be reduced. In particular, by intelligently identify ing when a user's query' can be satisfied by retrieval of existing content, the system avoids the necessity of using generative models for every query. This selective use ofgenerative models reduces the computational burden on the system, optimizing the use of CPU / GPU / TPU resources, and minimizing memory allocation.

[0042] As another example, the present invention enables faster response times, especially in time-critical applications like digital assistants. For example, the parallel processing of user queries by the retrieval and generative systems can reduce the overall processing time, improving the responsiveness of the system. As another example, in certain cases, the system can "short-circuit" or halt the generative model processing once a highly relevant match is found in the content database, further contributing to efficiency and speed.

[0043] As another example, existing content can also be used as conditioned inputs into generative models, which can reduce the number of generative cycles needed by the models to generate content because the models are not starting from pure noise and randomness.

[0044] Additionally, because existing and generated content can be combined at will and on-the-fly when users request content, there is no need to have unnecessary saved content stored on a server or in a memory of a user device, thus further reducing the amount of memory needed to store content for retrieval and presentation to users.

[0045] Finally, the proposed approach brings about an objective increase in the nature of returned content. The system's ability7to process user queries in both retrieval and generative systems and then rank the results based on quality and relevance ensures that the most suitable content is delivered to the user from multiple different channels. Additionally, by merging different media channels, the system can create hybrid content, such as adding a generated audio track to an existing video, thus offering more tailored and high-quality7content to the user. The proposed system therefore represents an objective increase in the types of content that can be returned to a user.

[0046] The systems and methods described herein can be implemented into a number of practical applications. As one example, the present disclosure can be particularly useful in the realm of digital assistants. A digital assistant is an application that can understand natural language voice commands and complete tasks for the user. Current digital assistants are limited in their capabilities and often struggle to provide relevant responses to user queries. The present disclosure can enhance these digital assistants by integrating the hybrid retrieval and generative system.

[0047] In these embodiments, when a user asks a question or makes a request, the digital assistant can utilize the system to either retrieve relevant information from existing databases or generate new content to respond to the user's query . For example, a user may askthe digital assistant to "find a recipe for a vegetarian lasagna" or "create a new recipe for a healthy breakfast smoothie." The digital assistant can use the retrieval system to locate a vegetarian lasagna recipe from an existing database or use the generative system to create a new recipe for a breakfast smoothie.

[0048] Moreover, the present disclosure can be beneficial in the field of education. Educational platforms often provide digital resources to assist students in their learning process. The hybrid system can be incorporated into these platforms to provide a more personalized and efficient learning experience.

[0049] In these embodiments, when a student asks a question or requests information, the educational platform can utilize the system to either retrieve relevant information from existing databases or generate new content to respond to the student's query. For example, a student may ask the platform to "find information about the French Revolution" or "generate a summary of the chapter on cellular respiration." The educational platform can use the retrieval system to locate information about the French Revolution from an existing database or use the generative system to create a summary of the chapter on cellular respiration.

[0050] With reference now to the Figures, example embodiments of the present disclosure will be discussed in further detail.Example System Arrangements

[0051] Figure 1 A depicts a block diagram of an example content search system 100 according to example embodiments of the present disclosure. The content search system 100 can include one or more retrieval models 105, one or more generative models 110, and an output ranking system 115. The content search system 100 can be contained on a computing device, such as a mobile device, a desktop computer, a laptop computer, a server system, and the like.

[0052] The content search system 100 receives a user query 120. The user query 120 can be a text input, an audio input, a video input, and the like, and / or combinations thereof. In some embodiments, the content search system 100 can utilize natural language processing (‘"NLP”) to process text inputs to identify words, phrases, sentences, and other portions of the user query 120.

[0053] In some embodiments, if the user query 120 includes audio or video, the content search system 100 can also include audio and / or video processing sub-systems for extracting spoken utterances from the audio and / or video, which can then be transcribed into text and processed using NLP to identify one or more portions of the user query 120,

[0054] In some embodiments, the content search system 100 can include an analysis system 125 that analyzes the one or more portions of the user query’ 120 to determine if each portion of the one or more portions is relevant to at least one of the one or more generative models 110. For example, in some embodiments, one or more portions of the user query 120 can explicitly ask for use of a generative model by including phrases such as “please generate,'’ “make me,'’ “show me,” and the like. In another example, one or more portions of the user query 120 do not explicitly ask for use of a generative model, but still include language associated with a generative model, such as '‘play relaxing jazz,” “show me a photo of a robot making pizza,” “Sing Happy Birthday in Italian,” and the like.

[0055] The analysis system 125 can include several components, such as a query processing unit, a relevance assessment unit, and a decision-making unit. The query processing unit can receive a user's query and break it down into understandable components using natural language processing techniques. These components can include key terms, phrases, or instructions that indicate the user's intent.

[0056] The relevance assessment unit can then determine the relevance of these components to the retrieval and generative models. This stage can involve comparing the query components to the capabilities of the retrieval and generative systems. For example, the relevance assessment unit can use predefined criteria or algorithms to evaluate the suitability7of each component for retrieval or generation.

[0057] In cases where the query components are relevant to both retrieval and generative systems, the relevance assessment unit can assign scores to each component based on its relevance to both systems. This scoring process can employ various methods, including machine learning algorithms, statistical analysis, or rule-based systems.

[0058] For example, the relevance assessment unit can assign scores to each component of the query based on its relevance to both systems. These scores can be calculated using machine learning algorithms, statistical analysis, or rule-based systems. For instance, a machine learning model can be trained on a large dataset of past user queries and their corresponding responses, learning to predict the relevance of a new query7to both the retrieval and generative sy stems.

[0059] To illustrate, consider a user's query for a "calm Italian song". The parsing algorithm can identify "calm", "Italian", and "song" as key components. The matching algorithm can then assess the relevance of these components to the retrieval and generative systems. If the retrieval system has a large database of Italian songs, the "Italian" and "song" components might be highly relevant to the retrieval system. However, the "calm" component might beless relevant to the retrieval system, as it may not have a clear or consistent definition in the context of Italian songs. On the other hand, the "calm" component might be highly relevant to a generative model that is capable of creating music with different moods, such as calm, energetic, or melancholic. Thus, the matching algorithm can assign different scores to each component of the query, reflecting its relevance to both systems.

[0060] The decision-making unit of the analysis system 125 can then use these scores to decide the best approach to satisfy the user's query, whether it involves retrieving existing content, generating new content, or a combination of both. This decision can be based on various factors, including the scores assigned by the relevance assessment unit, the system's available resources, and the user's preferences or past behavior.

[0061] For example, if a user's query is highly relevant to the retrieval system and less relevant to the generative models, the decision-making unit can decide to prioritize the retrieval system. On the other hand, if a user's query is equally relevant to both the retrieval system and the generative models, the decision-making unit can decide to use both systems in parallel or in sequence, depending on the system's configuration and the user's preferences.

[0062] In some embodiments, the analysis system 125 can also consider the system's current load and available resources when making this decision. For example, if the system's resources are strained due to high demand, the decision-making unit can prioritize the retrieval system, as retrieval operations typically require less computational power compared to generative operations.

[0063] Moreover, the analysis system 125 can adapt to the user's behavior over time. For example, if a user consistently prefers generated content over retrieved content, the decisionmaking unit can adjust its decision-making process to prioritize the generative models for this user. This feature can enhance the personalization of the system and improve the user's experience.

[0064] Thus, in some cases, the analysis system 125 can determine that one or more portions of the user query 120 are associated with a request to use the one or more generative models 110 that are associated with the one or more portions. The one or more portions of the user query 120 can then be provided to the retrieval models 105 and / or the generative models 110. The one or more generative models 110 can include models for generating audio, video, or images, and / or other generative outputs.

[0065] However, in some embodiments, such as the embodiment shown in Figure 1A, the one or more portions of the user query 120 can be provided first to the one or more retrieval models 105. In these embodiments, the one or more portions of the user query 120 are firstprovided to the one or more retrieval models 105 because the one or more retrieval models 105 have a better chance of locating content that matches most similarly to the request from the user query 120.

[0066] The one or more retrieval models 105 can take, as input, the one or more portions of the user query 120. In some embodiments, the one or more retrieval models 105 can generate an embedding or a token representation of each portion of the one or more portions.

[0067] The one or more retrieval models 105 can access one or more data stores 130 to compare the tokens or embeddings associated with each portion of the one or more portions of the user query7120. The one or more data stores 130 can include publicly accessible databases, such as online web databases, and can also include databases associated with the user of the computing device, such as private photo collections, private audio or video files, and the like that only the user may have access to. The one or more data stores 130 can store content items and, optionally, token or embedding representations of the content items for comparison to the tokens or embeddings generated by the one or more retrieval models 105.

[0068] Based on the comparison between the token or embedding representation of the content items in the one or more data stores 130 and the token or embedding associated with each portion of the one or more portions of the user query 120, the one or more retrieval models 105 can determine one or more content items to return. For example, the one or more retrieval models 105 can calculate a similarity score between the token or embedding representation of the content item in the data store 130 and the token or embedding associated with a particular portion of the one or more portions of the user query 120. The similarity score can be determined using Euclidean distance, cosine similarity, dot product similarity7, and other suitable similarity measurements.

[0069] In some embodiments, content items can be selected as search results if the similarityscore between the content item and the portion of the one or more portions of the user query 120 is above a threshold, such as having a 95% similarity score or above. In these embodiments, multiple content items can be returned as search results. In other embodiments, the content item with the highest similarity score can be returned as a lone search result.

[0070] After identity ing one or more content items as search results, the one or more retrieval models 105 returns the identified one or more content items as retrieved search results to the content search system 100. The content search system 100 can then determine if any of the retrieved search results satisfy the user query 120 such that there is no need to provide any portions of the one or more portions of the user query 120 to the one or more generativemodels 110. In these cases, the content search system 100 can simply return the retrieved search results as the search results to the user query 120. For example, if the user query 120 is “Play Happy Birthday in Italian" and the retrieved search results include a video of a popular opera singer singing Happy Birthday in Italian, there is no need to generate any additional content using the one or more generative models 110.

[0071] In some embodiments, where a content item is retrieved, the retrieved search results can also have proper author credits linking back to the original creator of the content item, and can also output a notification to a revenue sharing service that the content was accessed so that appropriate revenue sharing can occur.

[0072] Portions of the one or more portions of the user query 120 that did not have any suitable matches in the retrieved search results and that have an associated model in the one or more generative models 1 10 are then provided to the model of the one or more generative models 110 that the portion is associated with. For example, a portion of the user query 120 that recites “make me an image of a landscape” can be provided to an image generation model as a conditional prompt for the image generation model. In another example, a portion of the user query 120 that recites “Play Happy Birthday in Italian” can be provided as a conditional input into an audio generation model.

[0073] Using the input one or more portions, the one or more generative models 110 can generate new content items based on the input one or more portions. As described above, for example, an audio generation model can generate a voice singing Happy Birthday in Italian. In another example, an image generation model can generate “an image of a robot making a pizza.” The generated content items can then be returned to the content search system 100 as generated search results.

[0074] In some embodiments, an existing content item can be used as a generative seed for the one or more generative models 110. For example, a video file, an image, an audio file, or text input can be accessed from retrieved content items and used with modifiers as a generative seed, such as using an existing image as a generative seed and requesting “make this image have a cartoonish look.” In instances where a content item is used as a generative seed, the generated search results can also have proper author credits linking back to the original creator of the content item used as the generative seed, and can also output a notification to a revenue sharing sendee that the content item was used as a generative seed so that appropriate revenue sharing can occur.

[0075] In some embodiments, instead of providing the one or more portions of the user query 120 to the one or more retrieval models 105 first and then to the one or more generativemodels 110, the one or more portions of the user query 120 can be provided to the one or more retrieval models 105 and the one or more generative models 110 in parallel, or at the same time. This embodiment is illustrated in Figure IB. Each portion of the user query 120 is provided to both the one or more retrieval models 105 and the one or more generative models 110, which in turn process the one or more portions to either identify relevant content items or generate new content associated with the portion, respectively.

[0076] In some embodiments, if the one or more retrieval models 105 identify a strong match (e.g., a similarity score of 95% or above) for a portion of the user uery 120 to a content item, the one or more retrieval models 105 can return this content item as a retrieved search result.The one or more retrieval models 105 can also send a notification to the one or more generative models 110 that a strong match has been found in retrievable content, and that generated content is not required for that portion. The one or more generative models 110 can receive the notification and not perform any further generation for that portion of the user query 120, thus saving intensive processing cycles needed for generating new content.

[0077] Returning to Figure 1 A, the retrieved search results and the generated search results are returned to the content search system 100 as search results to the user query 120. In some embodiments, the search results are provided to the output ranking system 115.

[0078] The output ranking system 115 can decide final output(s) 135 to present to the user. For example, the output ranking system 115 can determine one or more quality metrics for the search results and decide the final output(s) 135 based on these quality metrics. The quality metrics can include similarity scores between the user query 120 and the search result, click-through rate of retrieved search results when the user query 120 is similar to other user queries presented to the content search system 100, a number and quality of keywords associated with the search result that are relevant to the user query 120. and the like.

[0079] In some embodiments, content items can be selected based on consistency metrics. For example, if a particular content item is routinely selected as a search result for similar user queries, the content item can have a better consistency metric and thus can be ranked higher in importance or relevance by the output ranking system 115.

[0080] In some embodiments, content items can be selected based on an affinity score. The affinity score can measure both quality metrics as described above and also can take into account other factors when determining the qualify of the search result, such as modify ing a quality metric by an amount based on a user preference associated with the user who made the user query 120. For example, the user preference can take into account a user’s desire tosee retrieved content instead of generated content, a user’s desire to see only search results with higher quality metrics, and the like.

[0081] The output ranking system 115 can determine, based on the quality metrics and / or the affinity scores associated with different search results, which search results are to be provided to the user by, for example, a user interface.

[0082] In some embodiments, certain search results can be combined to output multimedia search results. For example, if the user query 120 requested “a video of slow clouds with light jazz,” the search results can contain any combination of retrieved video, generated video, retrieved music, and generated music. The audio can be combined with the video to create a singular search result. In other embodiments, original audio from the video can be replaced with generated audio.

[0083] The output ranking system 1 15 can then output the final output(s) 135 to the user. In some embodiments, multiple options (e.g., multiple search results) can be returned as a list, optionally with quality metrics and / or affinity scores displayed such that the user can identify what most closely matches the user query 120. In some embodiments, the type of content (e.g.. retrieved or generated) is also indicated to the user. For example, if no content could be retrieved, the content search system 100 can generate content and indicate to the user “I couldn’t find anything suitable on the web, so here is something I generated” or any other appropriate indication. Additionally, if content items were retrieved or used as generative seeds, the displayed search results can also indicate the source of the content items.

[0084] In some embodiments, the user can request a modification to one or more search results. For example, if a video with music is returned as a search result, the user may provide user input indicating that the user desires for the music in the video to be faster. This user input can be a text input or any other suitable input for processing. The content search system 100 can analyze the user input to determine the desired modifications (e.g., “make it faster” or “start the music sooner”) and then input the desired modifications and the original search result into a generative model as conditioning inputs for the generative model. The generative model can then generate a new content item that includes the desired modifications to the original search result.

[0085] In some embodiments, generated search results can be saved into a memory for later retrieval. For example, if the user requests a video of a landscape with light jazz playing in the background, the one or more generative models 110 can generate the video and audio and create a new content item. This content item can then be stored in a private database associated with the user or in one of the one or more data stores 130 for later retrieval tominimize the need for using the one or more generative models 110 if the user needs to generate something similar to the generated content item. The generated content item can also be used as a generative seed should the user need a similar content item generated, such as needing a landscape at night with light jazz playing in the background.

[0086] The search system 100 can be implemented in a number of different computational arrangements. In an on-device arrangement, the entire functionality of the hybrid query’ response system as described in the present disclosure is integrated into a single device, such as a personal computer, a smartphone, or a tablet. The retrieval models, the generative models, and the associated processing capabilities are all contained within the device.

[0087] In this configuration, when a user issues a query to the system, the system processes the query directly on the device. This involves parsing the query, determining its relevance to the retrieval or generative models, executing the necessary models, and returning the results to the user, all within the confines of the single device.

[0088] In a client-server arrangement, the functionality of the hybrid query response system is split between a client device and a server. The client device, such as a personal computer, smartphone, or tablet, serves as the interface for the user to interact with the system, while the server handles the bulk of the processing tasks.

[0089] In this configuration, when a user issues a query' to the system, the client device transmits the query to the server. The server then processes the query, which involves parsing the query, determining its relevance to the retrieval or generative models, executing the necessary' models, and generating the results. The server then transmits the results back to the client device, which presents the results to the user.

[0090] In a hybrid arrangement, the functionalities of the hybrid query response system are distributed across multiple devices and / or servers. This configuration combines elements of both the on-device and client-server arrangements to achieve a balance of efficiency, flexibility, and privacy. In such a hybrid arrangement, specific components of the hybrid query response system can be located at different computing devices based on capabilities of the computing devices. For example, because retrieval services require less processing power than generative services, retrieval services can be located on-device while generative services can be located on a server accessible by the computing device. Additionally, services can be distributed in this hybrid arrangement based on access to data. For example, portions of retrieval services that utilize personal data associated with a user of the computing device can be located on-device, thus ensuring data privacy for personal data, while retrieval services for public content can be located at a server or other computing device.

[0091] In some implementations of this configuration, when a user issues a query' to the system, the initial processing tasks, such as parsing the query and determining its relevance to the retrieval or generative models, can be performed on the user's device. Depending on the results of this initial processing, the system may either execute the necessary models on the device (in cases where the device has sufficient capabilities and the query' is suitable for on- device processing), or transmit the query’ to a server for further processing.Example Methods

[0092] Figure 2 depicts a flow chart diagram of an example method 200 to perform according to example embodiments of the present disclosure. Although Figure 2 depicts steps performed in a particular order for purposes of illustration and discussion, the methods of the present disclosure are not limited to the particularly illustrated order or arrangement. The various steps of the method 200 can be omitted, rearranged, combined, and / or adapted in various ways without deviating from the scope of the present disclosure.

[0093] At block 202, a computing system can receive a user query as an input from a computing device, the user query including one or more portions. In some embodiments, the user query can include textual input. In other embodiments, the user uery can include at least one of audio input and video input. The computing system can process the audio input or the video input to identify the one or more portions of the user query by, for example, determining a transcript of the audio input or the video input and then analyzing the transcript to identify the one or more portions.

[0094] In some embodiments, the user query can include text and at least one of audio input and video input. The computing system can process the audio input or the video input along with the text to obtain the one or more portions of the user query.

[0095] At block 204, the computing system can process the one or more portions of the user query' to determine whether each portion of the one or more portions is relevant to one or more available generative models. In some embodiments, processing the one or more portions of the user query can include determining that at least one portion of the one or more portions includes a word or phrase associated with at least one of the one or more generative models. For example, if the user query includes words such as ‘‘show me a video,” it can be determined that the user query’ is relevant to a video generation model.

[0096] In some embodiments, processing the one or more portions of the user query can include determining that at least one portion of the one or more portions includes an explicit request for generation. For example, if the user query' includes phrases such as “create me animage’' or “generate audio,” it can be determined that the user query is explicitly a request for the use of generative models.

[0097] At block 206, the computing system can provide the user query to one or more available retrieval models. In some embodiments, providing the user query to the one or more available retrieval models can include generating a token or embedding for each portion of the one or more portions of the user query’ and providing the tokens or embeddings for each portion of the one or more portions to the one or more retrieval models for comparison to available content items in one or more data stores.

[0098] At block 208, the computing system can provide at least one portion of the one or more portions determined to be relevant to the one or more generative models to the one or more generative models. In some embodiments, providing the at least one portion of the one or more portions can include generating a token or embedding for each portion of the one or more portions determined to be relevant to the one or more generative models providing the tokens or embeddings for each portion of the one or more portions determined to be relevant to the one or more generative models to a generative model of the one or more generative models that is associated with the portion.

[0099] In some embodiments, the user query is provided to the one or more retrieval models before the user query’ is provided to the one or more generative models.

[0100] In some embodiments, the user query is provided to the one or more retrieval models and the one or more generative models in parallel.

[0101] At block 210, the computing system can receive retrieved content from the one or more retrieval models. In some embodiments, the retrieved content is retrieved from one or more data stores, the one or more data stores including at least one of a public database and a database associated with a user of the computing device.

[0102] In some embodiments, a portion of the one or more portions is not provided to the one or more generative models if the portion has content retrieved by the one or more retrieval models.

[0103] At block 212, the computing system can receive generated content from the one or more generative models.

[0104] At block 214, the computing system can return at least a portion of the retrieved content or the generated content as search results to the user query’.

[0105] In some embodiments, the search results can be ranked based on at least one of a quality metric and a consistency metric.

[0106] In some embodiments, the search results can be ranked based on an affinity score determined based on at least one user preference.

[0107] In some embodiments, at least a portion of the retrieved results and at least a portion of the generative results are combined into a single returned result, such as combining retrieved video and generated audio into a singular audiovisual content item.

[0108] In some embodiments, the method 200 can also include receiving a requested update to the search results as user input and providing the requested update to the search results and the search results to the one or more generative models as a conditioning input. The method 200 can also include receiving updated generated contents from the one or more generative models based on the requested update to the search results providing the updated generated contents to the user.Example Devices and Systems

[0109] Figure 3 A depicts a block diagram of an example computing system 300 that performs content retrieval and generation according to example embodiments of the present disclosure. The system 300 includes a user computing device 302, a server computing system 330, and a training computing system 350 that are communicatively coupled over a network 380.

[0110] The user computing device 302 can be any type of computing device, such as, for example, a personal computing device (e.g.. laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0111] The user computing device 302 includes one or more processors 312 and a memory 314. The one or more processors 312 can be any suitable processing device (e.g.. a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 314 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 314 can store data 316 and instructions 318 which are executed by the processor 312 to cause the user computing device 302 to perform operations.

[0112] In some implementations, the user computing device 302 can store or include one or more content retrieval and generation models 320. For example, the content retrieval and generation models 320 can be or can otherwise include various machine-learned models such as neural networks (e.g., deep neural networks) or other types of machine-learned models,including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory’ recurrent neural networks), convolutional neural networks or other forms of neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example content retrieval and generation models 320 are discussed with reference to Figures 1 and 2.

[0113] In some implementations, one or more content retrieval and generation models 320 can be received from the server computing system 330 over network 380, stored in the user computing device memory 314, and then used or otherwise implemented by the one or more processors 312. In some implementations, the user computing device 302 can implement multiple parallel instances of a single content retrieval and generation model 320 (e.g., to perform parallel content retrieval and generation across multiple instances of user queries).

[0114] More particularly, the one or more content retrieval and generation models 320 are configured to retrieve relevant content and / or generate new content based on user queries. The user query can be tokenized or an embedding can be generated to encapsulate the meaning of the user query, and in turn the token or embedding can be used to identify similar existing and retrievable content in databases or as a generative seed or conditioning input into a generative model.

[0115] The one or more content retrieval and generation models 320 can access a retrieval system 324. The retrieval system 324 can take as input search terms, including text, audio, video, images, and the like, and access one or more data stores to retrieve content items from the one or more data stores. Content items can then be provided to the one or more content retrieval and generation models 320 as results to the search terms.

[0116] Additionally or alternatively, one or more content retrieval and generation models 340 can be included in or otherwise stored and implemented by the server computing system 330 that communicates with the user computing device 302 according to a clientserver relationship. For example, the one or more content retrieval and generation models 340 can be implemented by the server computing system 340 as a portion of a web serv ice (e.g., a content generation and retrieval service). Thus, one or more models 320 can be stored and implemented at the user computing device 302 and / or one or more models 340 can be stored and implemented at the server computing system 330 and retrieval system 342 can be used by the one or more models 320.

[0117] The user computing device 302 can also include one or more user input components 322 that receives user input. For example, the user input component 322 can be a touch- sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can sen e to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0118] The server computing system 130 includes one or more processors 332 and a memory 334. The one or more processors 332 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 334 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 334 can store data 336 and instructions 338 which are executed by the processor 332 to cause the server computing system 330 to perform operations.

[0119] In some implementations, the server computing system 330 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 330 includes plural server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0120] As described above, the server computing system 330 can store or otherwise include one or more content retrieval and generation models 340. For example, the models 140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networks include feed forward neural networks, deep neural netw orks, recurrent neural networks, and convolutional neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine- learned models can include multi-headed self-attention models (e.g., transformer models). Example models 140 are discussed with reference to Figures 1 and 2.

[0121] The user computing device 302 and / or the server computing system 330 can train the models 320 and / or 340 via interaction with the training computing system 350 that is communicatively coupled over the network 380. The training computing system 350 can be separate from the server computing system 330 or can be a portion of the server computing system 330.

[0122] The training computing system 350 includes one or more processors 352 and a memory 354. The one or more processors 352 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 354 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 354 can store data 356 and instructions 358 which are executed by the processor 352 to cause the training computing system 350 to perform operations. In some implementations, the training computing system 350 includes or is otherwise implemented by one or more server computing devices.

[0123] The training computing system 350 can include a model trainer 360 that trains the machine-learned models 320 and / or 340 stored at the user computing device 302 and / or the server computing system 330 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.

[0124] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 360 can perform a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.

[0125] In particular, the model trainer 360 can train the one or more content retrieval and generation models 320 and / or 340 based on a set of training data 362. The training data 362 can include, for example, one or more content items and user queries associated with the one or more content items. Retrieval models can be trained to generate a token or embedding for the user queries based on known embeddings of the user queries and then trained to identify similar embeddings or tokens associated with the one or more content items. Generative models can be trained to use Gaussian noise and conditioning inputs or generative seeds to generate the desired content item.

[0126] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 302. Thus, in such implementations, the model 320 provided to the user computing device 302 can be trained by the training computing system350 on user-specific data received from the user computing device 302. In some instances, this process can be referred to as personalizing the model.

[0127] The model trainer 360 includes computer logic utilized to provide desired functionality. The model trainer 360 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 360 includes program files stored on a storage device, loaded into a memory and executed by one or more processors. In other implementations, the model trainer 360 includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.

[0128] The network 380 can be any type of communications network, such as a local area network (e.g.. intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 380 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP. HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0129] Machine-learned models described in this specification may be used in a variety of tasks, applications, and / or use cases.

[0130] In some implementations, the input to the machine-learned model(s) of the present disclosure can be image data. The machine-learned model(s) can process the image data to generate an output. As an example, the machine-learned model(s) can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an image segmentation output. As another example, the machine-learned model(s) can process the image data to generate an image classification output. As another example, the machine-learned model(s) can process the image data to generate an image data modification output (e.g., an alteration of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data. etc.). As another example, the machine-learned model(s) can process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) can process the image data to generate a prediction output.

[0131] In some implementations, the input to the machine-learned model(s) of the present disclosure can be text or natural language data. The machine-learned model(s) can processthe text or natural language data to generate an output. As an example, the machine-learned model(s) can process the natural language data to generate a language encoding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a translation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a classification output. As another example, the machine-learned model(s) can process the text or natural language data to generate a textual segmentation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a semantic intent output. As another example, the machine-learned model(s) can process the text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is higher quality than the input text or natural language, etc.). As another example, the machine-learned model(s) can process the text or natural language data to generate a prediction output.

[0132] In some implementations, the input to the machine-learned model(s) of the present disclosure can be speech data. The machine-learned model(s) can process the speech data to generate an output. As an example, the machine-learned model(s) can process the speech data to generate a speech recognition output. As another example, the machine-learned model(s) can process the speech data to generate a speech translation output. As another example, the machine-learned model(s) can process the speech data to generate a latent embedding output. As another example, the machine-learned model(s) can process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate an upscaled speech output (e.g., speech data that is higher quality than the input speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate a textual representation output (e.g., a textual representation of the input speech data, etc.). As another example, the machine- learned model(s) can process the speech data to generate a prediction output.

[0133] In some implementations, the input to the machine-learned model(s) of the present disclosure can be latent encoding data (e.g., a latent space representation of an input, etc.). The machine-learned model(s) can process the latent encoding data to generate an output. As an example, the machine-learned model(s) can process the latent encoding data to generate a recognition output. As another example, the machine-learned model(s) can process the latent encoding data to generate a reconstruction output. As another example, the machine-learnedmodel(s) can process the latent encoding data to generate a search output. As another example, the machine-learned model(s) can process the latent encoding data to generate a reclustering output. As another example, the machine-learned model(s) can process the latent encoding data to generate a prediction output.

[0134] In some implementations, the input to the machine-learned model(s) of the present disclosure can be statistical data. Statistical data can be. represent, or otherwise include data computed and / or calculated from some other data source. The machine-learned model(s) can process the statistical data to generate an output. As an example, the machine-learned model(s) can process the statistical data to generate a recognition output. As another example, the machine-learned model(s) can process the statistical data to generate a prediction output. As another example, the machine-learned model(s) can process the statistical data to generate a classification output. As another example, the machine-learned model(s) can process the statistical data to generate a segmentation output. As another example, the machine-learned model(s) can process the statistical data to generate a visualization output. As another example, the machine-learned model(s) can process the statistical data to generate a diagnostic output.

[0135] In some implementations, the input to the machine-learned model(s) of the present disclosure can be sensor data. The machine-learned model(s) can process the sensor data to generate an output. As an example, the machine-learned model(s) can process the sensor data to generate a recognition output. As another example, the machine-learned model(s) can process the sensor data to generate a prediction output. As another example, the machine- learned model(s) can process the sensor data to generate a classification output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a visualization output. As another example, the machine-learned model(s) can process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) can process the sensor data to generate a detection output.

[0136] In some cases, the machine-learned model(s) can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data and the output may comprise compressed audio data. In another example, the input includes visual data (e.g. one or more images or videos), the output comprises compressed visual data, and the task is a visual data compression task. In anotherexample, the task may comprise generating an embedding for input data (e.g. input audio or visual data).

[0137] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that the one or more images depict an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images and, for each region, a likelihood that region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in the one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in the one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel of one of the input images, a motion of the scene depicted at the pixel between the images in the network input.

[0138] In some cases, the input includes audio data representing a spoken utterance and the task is a speech recognition task. The output may comprise a text output which is mapped to the spoken utterance. In some cases, the task comprises encry pting or decrypting input data. In some cases, the task comprises a microprocessor performance task, such as branch prediction or memory address translation.

[0139] Figure 3 A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 302 can include the model trainer 360 and the training dataset 362. In such implementations, the models 320 can be both trained and used locally at the user computing device 302. In some of such implementations, the user computing device 302 can implement the model trainer 360 to personalize the models 320 based on user-specific data.

[0140] Figure 3B depicts a block diagram of an example computing device 400 that performs content retrieval and generation according to example embodiments of the present disclosure. The computing device 400 can be a user computing device or a server computing device.

[0141] The computing device 400 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a brow ser application, etc.

[0142] As illustrated in Figure 3B, each application can communicate with a number of other components of the computing device, such as. for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0143] Figure 4C depicts a block diagram of an example computing device 500 that performs content retrieval and generation according to example embodiments of the present disclosure. The computing device 500 can be a user computing device or a server computing device.

[0144] The computing device 500 includes a number of applications (e.g., applications 1 through N). Each application is in communication with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a brow ser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).

[0145] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in Figure 3C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 500.

[0146] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 500. As illustrated in Figure 3C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).Additional Disclosure

[0147] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0148] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations and / or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method for performing intelligent hybrid search, the method comprising: receiving, at one or more processors, a user query as an input, the user query including one or more portions; processing, by the one or more processors, the one or more portions of the user query to determine a relevance of the one or more portions of the user query to a retrieval system and one or more generative models; providing, by the one or more processors, at least one of the one or more portions of the user query to the retrieval system and the one or more generative models based on the determined relevance; receiving, by the one or more processors, content from the retrieval system and one or more generative models; and providing, by the one or more processors, at least a portion of the received content as results to the user query.

2. The computer-implemented method of claim 1, wherein at least one portion of the one or more portions of the user query is a text portion.

3. The computer-implemented method of claim 1, wherein at least one portion of the one or more portions of the user query is an audio portion or an image portion, and wherein the computer-implemented method further comprises processing audio input or image input to identify at least one of the audio portion or the image portion.

4. The computer-implemented method of claim 1, wherein at least a first portion of the one or more portions of the user query is a text portion and a second portion of the one or more portions of the user query7is an audio or image portion.

5. The computer-implemented method of claim 1, wherein processing the one or more portions of the user query comprises determining that at least one portion of the one or more portions includes a word or phrase associated with at least one of the one or more generative models.

6. The computer-implemented method of claim 1 , wherein processing the one or more portions of the user query comprises determining that at least one portion of the one or more portions includes an explicit request for use of one or more generative models.

7. The computer-implemented method of claim 1, wherein providing the one or more portions of the user query to the one or more retrieval models comprises: generating, by the one or more processors, a token or embedding for each portion of the one or more portions; and providing, by the one or more processors, the tokens or embeddings for each portion of the one or more portions to the one or more retrieval models for comparison to available content items.

8. The computer-implemented method of claim 1, wherein providing the one or more portions determined to be relevant to the one or more generative models to the one or more generative models comprises: generating, by the one or more processors, a token or embedding for each portion of the one or more portions determined to be relevant to the one or more generative models; and providing, by the one or more processors, the tokens or embeddings for each portion of the one or more portions determined to be relevant to the one or more generative models to a generative model of the one or more generative models that is associated with the portion.

9. The computer-implemented method of claim 1, wherein the retrieved content is retrieved from one or more data stores, the one or more data stores including at least one of a public database and a database associated with a user of the computing device.

10. The computer-implemented method of claim 1 , wherein the user query is provided to the one or more retrieval models before the user query is provided to the one or more generative models.

11. The computer-implemented method of claim 10, wherein the user query is provided to the one or more retrieval models and the one or more generative models in parallel.

12. The computer-implemented method of claim 11, wherein a portion of the one or more portions is not provided to the one or more generative models if the portion has content retrieved by the one or more retrieval models.

13. The computer-implemented method of claim 1, wherein the search results are ranked based on at least one of a quality metric and a consistency metric.

14. The computer-implemented method of claim 1, wherein the search results are ranked based on an affinity' score determined based on at least one user preference.

15. The computer-implemented method of claim 1, wherein at least a portion of the retrieved results and at least a portion of the generative results are combined into a single returned result.

16. The computer-implemented method of claim 1, the computer-implemented method further comprising: receiving, by the one or more processors, a requested update to the search results as user input; providing, by the one or more processors, the requested update to the search results and the search results to the one or more generative models as a conditioning input; receiving, by the one or more processors, updated generated contents from the one or more generative models based on the requested update to the search results; and providing, by the one or more processors, the updated generated contents to the user.

17. A computing system, the computing system compnsing: one or more processors; and a non-transitory, computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising:receiving a user query' as an input, the user query' including one or more portions; processing the one or more portions of the user query to determine a relevance of the one or more portions of the user query to a retrieval system and one or more generative models; providing at least one of the one or more portions of the user query to the retrieval system and the one or more generative models based on the determined relevance; receiving content from the retrieval system and one or more generative models; and providing at least a portion of the received content as results to the user query.

18. The computing system of claim 17, wherein processing the one or more portions of the user query comprises determining that at least one portion of the one or more portions includes a word or phrase associated with at least one of the one or more generative models.

19. The computing system of claim 17, the operations further comprising: receiving a requested update to the search results as user input; providing the requested update to the search results and the search results to the one or more generative models as a conditioning input; receiving updated generated contents from the one or more generative models based on the requested update to the search results; and providing the updated generated contents to the user.

20. A non-transitory, computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations comprising: receiving a user query' as an input, the user query' including one or more portions;processing the one or more portions of the user query to determine a relevance of the one or more portions of the user query to a retrieval system and one or more generative models; providing at least one of the one or more portions of the user query to the retrieval system and the one or more generative models based on the determined relevance; receiving content from the retrieval system and one or more generative models; and providing at least a portion of the received content as results to the user query.

Citation Information

Patent Citations

  • Multi-mode database query method and device, electronic equipment and storage medium

    CN117290411A

  • Planning-based automated fusing of data from multiple heterogeneous sources

    WO2012018475A2