Querying proprietary data using generative models

The framework integrates proprietary data into generative models, addressing the cost and access limitations of existing LLMs by enabling efficient querying and monetization of proprietary data sources.

WO2026102401A1PCT designated stage Publication Date: 2026-05-15X DEVELOPMENT LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
X DEVELOPMENT LLC
Filing Date
2025-11-10
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Generative models like large language models (LLM) are costly and time-consuming to fine-tune due to their large size and often lack access to proprietary data sources, leading to delays in incorporating new information and limited query capabilities.

Method used

A framework that enables proprietary data from specialized sources to be integrated into generative model input prompts, allowing users to query these models without prior training or fine-tuning, using techniques such as information gain scoring and preprocessing to condition model outputs on proprietary data.

Benefits of technology

Enables efficient and effective querying of proprietary data sources, providing enhanced model outputs and allowing publishers to control access and monetize their data through subscription models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025054828_15052026_PF_FP_ABST
    Figure US2025054828_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Implementations are described herein for querying proprietary data using generative model(s). In various implementations, a first input prompt may be assembled to include a query for generative model(s). The first input prompt may be processed using generative model(s) to generate reference model output. Proprietary data source(s) may be identified for use in prompting generative model(s), and a second input prompt may be assembled to include the query and proprietary data from one or more of the proprietary data sources. The second input prompt may be processed using generative model(s) to generate at least partially exogenous model output based at least in part on the proprietary data. The reference and at least partially exogenous model outputs may then be compared to determine information gain score(s) associated with the proprietary data source(s). These information gain scores may be used for various purposes.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. XDEV-0025-WO-01QUERYING PROPRIETARY DATA USING GENERATIVE MODELSBackground

[0001] Generative models such as large language models (LLM) and multi-modal LLMs are increasingly being used to power search engines and other similar systems. Generative models can have hundreds of billions of parameters or more, and may be trained and / or fine-tuned using massive, web-scale datasets of publicly available information. Given their size, it can be costly and time consuming to fine-tune generative models. As a consequence, there is usually a delay between events giving rise to new information and the training of that new information into the generative model. Moreover, generative models are often trained using publicly available information, and may not have access to proprietary data sources hosted by private providers or publishers.SUMMARY

[0002] Implementations described herein relate to a framework for making proprietary and / or specialized data from proprietary data sources available for use with generative models. More particularly, but not exclusively, techniques described herein relate to assembling proprietary and / or specialized (hereinafter “proprietary”) data from proprietary data sources into generative model input prompts, e.g., along with user queries (e.g., in natural language and / or other modalities). When these generative model input prompts are processed using generative models, the resulting generative model output may be conditioned on the proprietary data. Consequently, a user can query proprietary data sources using generative models that have not been trained or fine-tuned using those proprietary data sources.

[0003] It should be appreciated that all combinations of the foregoing concepts and additional concepts described in greater detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein.Attorney Docket No. XDEV-0025-WO-01BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Fig. 1 depicts a block diagram of an example environment that demonstrates various aspects of the present disclosure, and in which some implementations disclosed herein can be implemented.

[0005] Fig. 2 schematically depicts an example of how various components may exchange data to facilitate multi-modal assistant engagement, in accordance with various implementations.

[0006] Fig. 3 schematically depicts an example of how selected aspects of the present disclosure may be implemented.

[0007] Fig. 4 schematically depicts an example graphical user interface on which selected aspects of the present disclosure have been implemented.

[0008] Fig. 5 depicts a flowchart illustrating an example method of practicing selected aspects of the present disclosure, in accordance with various implementations.

[0009] Fig. 6 depicts an example architecture of a computing device, in accordance with various implementations.DETAILED DESCRIPTION

[0010] Implementations described herein relate to a framework for making proprietary and / or specialized data from proprietary data sources available for use with generative models. More particularly, but not exclusively, techniques described herein relate to assembling proprietary and / or specialized (hereinafter “proprietary”) data from proprietary data sources into generative model input prompts, e.g., along with user queries (e.g., in natural language). When these generative model input prompts are processed using generative models, the resulting generative model output may be conditioned on the proprietary data. Consequently, a user is able to query proprietary data sources using generative models that have not been trained or fine-tuned using those proprietary data sources.

[0011] When additional data beyond a basic query (e.g., a question or command), such as proprietary data from proprietary data source(s), is included in a generative model’s context window, the resulting model output may be referred to herein as “exogenous.” Put another way, “exogenous” model output is conditioned not only on the user’s query and intrinsic parameters of the generative model, but is also conditioned on other data included in the generative model’sAttorney Docket No. XDEV-0025-WO-01 context window, such as proprietary data from proprietary data source(s). By contrast, when only a query is included in the generative model’s context window, and no additional information is included to condition the generative model’s output, the generative model will generate what will be referred to herein as “endogenous” model output. Put another way, “endogenous” model output is conditioned solely based on the user’s query and intrinsic parameters of the generative model.

[0012] In various implementations, publishers of proprietary data may make their proprietary data sources available, e.g., as part of an online marketplace, to any subscriber who wishes to query these proprietary data sources using generative models. Users who wish to query particular proprietary data sources using the generative models may subscribe to those proprietary data sources, e.g., for a fee. To determine whether to make their proprietary data available in such a marketplace, publishers may evaluate how their proprietary data will likely be received, e.g., by determining the value their proprietary data is likely to add relative to what non-subscribing users may obtain querying the generative model endogenously (or exogenously using a different publisher’s proprietary data). Potential subscribers may desire the same information. This may allow the potential subscribers to determine whether the added benefit obtained from subscribing to the proprietary data source and querying the generative model exogenously — relative to what the potential subscriber would obtain by querying the generative model endogenously — is worth the subscription’s cost.

[0013] Accordingly, in various implementations, proprietary data from proprietary data sources may be evaluated to determine information gain score(s) associated with the proprietary data sources. These information gain score(s) may represent added value (e.g., information gain) associated with exogenous model output, compared to endogenous model output, or compared to other exogenous model output generated using different proprietary data. These information gain scores may serve a variety of purposes. In some implementations, the information gain scores may be used to inform users of how much information they are likely to gain by subscribing to proprietary data source(s) for use with generative model (s), e.g., compared to not subscribing to those proprietary source(s). Alternatively, the information gain score(s) may inform information providers (also referred to herein as “publishers”) how much their proprietary data is likely to increase an amount and / or quality of information generated for subscribers using generative model(s). The publishers can use this information to alter which ofAttorney Docket No. XDEV-0025-WO-01 their proprietary data sources are made available to subscribers, to determine appropriate pricing for their proprietary data sources, etc.

[0014] In some implementations, to determine information gain score(s) between endogenous and exogenous model outputs, a first input prompt may be assembled to include a query and nothing else, to ensure the resulting model output is predicted based only on the query and intrinsic parameters of the generative model. This first input prompt may be processed using the generative model to generate endogenous model output. A second input prompt may be assembled to include the query and proprietary data from proprietary data sources(s), and may be processed using the generative model to generate at least partially exogenous model output. The latter model output is referred to as “at least partially exogenous” because it also will be dependent on intrinsic parameters of the generative model.

[0015] The endogenous model output and the at least partially exogenous model output may then be compared to determine information gain score(s) associated with the proprietary data source(s). This comparison may be performed in various ways. In some implementations, both the model outputs may be processed using the same generative model or a different machine learning model (which may or may not be a generative model). When using a generative model, the generative model’s context window may include a command to determine what information, if any, is provided in the at least partially exogenous model output that is lacking in the endogenous model output.

[0016] Additionally or alternatively, the model outputs may be compared using other techniques that rely on heuristics and / or machine learning, such as edit distances, differences in bags of words, phrase indexing, term frequency -inverse document frequency (TF-IDF), etc. In some implementations, techniques such as word2vec or doc2vec may be used to generate semantic vectors representing the words and / or whole model outputs. These semantic vectors may then be compared, e.g., using techniques such as cosine similarity or Euclidean distance, to determine how similar they are. The more similar the documents, the less new information one document likely provides over the other. In other implementations, the two model outputs may be rendered on a screen, and the publisher or potential subscriber may study the two model outputs to identify differences and estimate value associated with those differences.

[0017] Techniques described herein are not limited to comparing what a generative model might predict natively (i.e. endogenously) versus exogenously. In some cases, a publisher mayAttorney Docket No. XDEV-0025-WO-01 wish to determine how much value their proprietary data may add compared to another publisher’s proprietary data. In some such implementations, a first input prompt may be assembled to include proprietary data from the first publisher’s proprietary data source(s) and a query. A second input prompt may be assembled to include proprietary data from the second publisher’s proprietary data sources and the query. The first and second input prompts may be processed using a generative model to generate, respectively, first at least partially exogenous model output and second at least partially exogenous model output. These at least partially exogenous model outputs may be compared as described above to determine information gain scores associated with the first publisher’s proprietary data source(s), relative to the second publisher’s proprietary data source(s). In this manner, the first publisher can determine how much to charge subscribers, whether to share additional proprietary data from their proprietary data sources, etc.

[0018] In some implementations, a publisher’s proprietary data source(s) may be vetted automatically to determine information gain scores and / or other metrics, such as credibility metrics. For example, historical queries may be retrieved, and / or synthetic queries may be generated automatically, e.g., using generative model(s). These queries may be systematically included in a generative model’s context window, both alone and in combination with proprietary data of a publisher, to generate respective endogenous and exogenous model outputs. These model outputs may be compared as described previously to generate metrics such as information gain scores. These information gain scores can be used by publishers for various purposes, such as ongoing tracking of the added value of their proprietary content.

[0019] Proprietary data may take various forms, some which may be more structured than others. For instance, some proprietary data may be stored in database tables and / or spreadsheets. Other proprietary data may be stored in structured text files such as JSON or XML files. One particular entity might deploy Internet of Things (loT) sensors around a facility, and may store sensor data generated by those loT sensors in database(s). Yet other proprietary data may be less structured. For example, a publisher may have a corpus of digital images and / or videos that it wishes to make accessible to a generative model-powered search interface. As another example, a publisher may have large numbers of heterogeneous digital files (e.g., word processing documents, spreadsheets, images, efc.) stored in an online drive that it wants to make available to a generative model.Attorney Docket No. XDEV-0025-WO-01

[0020] Accessing heterogeneous data from heterogeneous proprietary data sources may be a challenge. Accordingly, in various implementations, proprietary data from proprietary data sources may be preprocessed (e.g., indexed) so that it can be more easily retrieved and used to condition generative models as described herein. For example, in some implementations, what will be referred to herein as an “indexing” input prompt may be assembled to include raw proprietary data retrieved from proprietary data source(s), and a request to index, summarize and / or organize the proprietary data retrieved from one or more of the proprietary data sources so that it can be more effectively used to condition generative models. Additionally or alternatively, the indexing prompt may include a request to generate steps for accessing the proprietary data source. This indexing prompt may then be processed using generative model(s) to generate indexing output. The indexing output may include, for example, an indexed view, summary, or other representation of the proprietary data in a format that can be more readily incorporated into a downstream generative model input prompt. Alternatively, or additionally, the indexing output may include steps that can then be taken to access the proprietary data.

[0021] This preprocessing may be particularly beneficial in situations where the proprietary data is natively stored in a format that cannot be readily queried. For example, a publisher may make a corpus of digital images accessible to subscribers so that those subscribers can include data derived from those images in the context window of a generative model. To this end, in some implementations, an indexing input prompt may be assembled to include data indicative of the publisher’s digital images and a request to caption and / or summarize the digital images. The indexing input prompt may be processed using, for instance, a vision-language model (VLM) to generate indexing output. The indexing output in this scenario might include caption(s) of the digital images, lists of objects (or more generally, visual features) detected in the digital images, a textual summary of one or more of the digital images, etc.

[0022] In various implementations, the indexing output may then be made available for inclusion into a context window of a generative model. When a subscriber to the publisher’s images later queries a generative model for information about these images, the caption(s), list of objects, and / or the summary can be included in the generative model’s context window. Consequently, the resulting generative model output may be conditioned on the caption(s), list of objects, and / or the summary.Attorney Docket No. XDEV-0025-WO-01

[0023] In some implementations, a publisher’s proprietary data may be accessible via an application programming interface (API). In some such implementations, documentation about interfacing with, and / or examples of how to interface with, the API may be made available by the publisher, e.g., to be incorporated into the generative model’s context window, e.g., at the same time as a user’s query or as a preprocessing step. This documentation may include, for instance, descriptions of functions and / or other API interfaces that are accessible via the API, examples of source code written in one or more programming languages that can be used to connect to the API, etc. In some implementations, this documentation may be processed using a generative model to generate model output that facilitates interfacing with the publisher’s API. For example, the model output may include source code that is executable to connect to the API, one or more steps (e.g., expressed in natural language) that may be performed to connect to the API, a graph such as a control flow graph (CFG) that demonstrates how to connect to the API, etc.

[0024] Metrics other than information gain scores may be calculated for proprietary data and / or its source(s). In some implementations, credibility metrics may be calculated for proprietary data and / or its sources. Credibility metrics may be calculated in various ways. In some implementations, a credibility input prompt may be assembled to include, for instance, a description of how data from one or more proprietary data sources was obtained and a request to determine a credibility metric for this proprietary data based at least in part on the description. The credibility input prompt may be processed using the same generative model as is used to process other requests described herein, or a different machine learning model (e.g., that is specifically trained and / or fine-tuned to generate credibility feedback), to generate credibility model output. The credibility model output may include credibility metric(s) associated with one or more of the proprietary data sources.

[0025] In some implementations, exogenous model output generated using proprietary data may include an attribution to the proprietary data’s publisher. For example, a portion of generative model output that is generated endogenously from intrinsic parameters of the generative model may be rendered in one way (e.g., a first font, first color, efc.). Another portion of the generative model output that is attributable to the proprietary data’s publisher — e.g., the portion of the model output generated at least in part exogenously based on proprietary content added to the generative model’s context window — may be rendered in another way (e.g., aAttorney Docket No. XDEV-0025-WO-01 second font or color, highlighted or otherwise made conspicuous). In some implementations where the publisher’s proprietary data includes multimedia content such as videos, those videos may be embedded into the model output, e.g., so that when the model output is rendered, the video appears as an embedded applet that can be separately controlled within the same interface.

[0026] In some implementations, the portion of the generative model output that is attributable to the publisher, rather than to intrinsic parameters of the generative model, may be rendered as a selectable element such as a hyperlink. A user may be able to actuate this selectable element to be provided additional content related to the publisher and / or the publisher’s proprietary data. To ensure the portion of the model output attributable to the publisher is identifiable, in some implementations, the input prompt may be assembled with a command to annotate or otherwise flag exogenous portion(s) of the model output that are attributable to the proprietary data added to the context window — as opposed to other endogenous portion(s) of the model output that are attributable to the intrinsic parameters of the generative model.

[0027] In some implementations, a user who has not subscribed to a publisher’s proprietary data may nonetheless be presented with generative model output that is conditioned on the publisher’s proprietary data. For example, the publisher may attempt to persuade the user to become a subscriber by allowing at least some proprietary data to be incorporated into the generative model’s context window as it is used by the user. In some such implementations, portion(s) of the model output that are attributable to the publisher may be at least partially obfuscated, e.g., to inform the user of the potential proprietary data that would be available should the user become a subscriber. For example, the generative model may be conditioned to generate a partial summary of proprietary data of the publisher, e.g., with the summary ending in the middle of a sentence or passage (e.g., using ellipsis). In some implementations, this partial summary may be rendered as a selectable element that is operable by the user to be taken to an interface (e.g., an online marketplace of publishers’ proprietary data sources) that allows the user to subscribe to the publisher’s proprietary data source(s).

[0028] In some implementations, publishers may be able to control how their proprietary data is accessed by subscribers. This may allow publishers to, among other things, prevent coordinated exfiltration of proprietary data. For example, a subscriber may be limited to how many times during a given time period (e.g., an hour, a day, a week, efc.) they are permitted toAttorney Docket No. XDEV-0025-WO-01 query the publisher’s proprietary data using generative model(s). Such limits may be set by the publisher on an individual subscriber basis and / or organization-wide basis (e.g., employees of company X may be limited to querying the proprietary data one hundred times total each day). The publisher may limit access to their proprietary data in other ways as well. For example, the publisher may perform rate limiting based on the type of proprietary data being requested, an amount or quantity of proprietary data being requested, trust level(s) assigned to the proprietary data, whether the proprietary data contains personally identifiable information (PII), etc.

[0029] Turning now to Fig. 1, a block diagram of an example environment that demonstrates various aspects of the present disclosure, and in which implementations disclosed herein can be implemented, is depicted. The example environment includes a client device 110, a knowledge system 120, and some number of publisher systems 170. Fig. 1 depicts a first publisher system 170A and a second publisher system 170B, but this is for illustrative purposes only; there may be any number of publisher systems. Publisher systems 170 may be controlled by publishers. Publishers may be, for instance, individuals, companies, governmental organizations, universities, social clubs, sports teams, etc.

[0030] In some implementations, all or aspects of the knowledge system 120 can be implemented locally at the client device 110. In additional or alternative implementations, all or aspects of the knowledge system 120 can be implemented remotely from the client device 110 as depicted in Fig. 1 (e.g., at remote server(s)). In those implementations, the client device 110 and the knowledge system 120, as well as any publisher systems 170, can be communicatively coupled with each other via one or more networks 199, such as one or more wired or wireless local area networks (“LANs,” including Wi-Fi, mesh networks, Bluetooth, near-field communication, etc. or wide area networks (“WANs”, including the Internet).

[0031] The client device 110 can be, for example, one or more of: a desktop computer, a laptop computer, a tablet, a mobile phone, a computing device of a vehicle (e.g., an in-vehicle communications system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (optionally having a display), a smart appliance such as a smart television, and / or a wearable apparatus of the user that includes a computing device (e.g., a watch of the user having a computing device, glasses of the user having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.Attorney Docket No. XDEV-0025-WO-01

[0032] The client device 110 can execute one or more software applications, via application engine 115, through which input can be submitted and / or output that is responsive to the input can be rendered (e.g., audibly and / or visually). The application engine 115 can execute one or more software applications that are separate from an operating system of the client device 110 (e.g., one installed “on top” of the operating system) - or can alternatively be implemented directly by the operating system of the client device 110. For example, the application engine 115 can execute a web browser or automated assistant installed on top of the operating system of the client device 110. As another example, the application engine 115 can execute a web browser software application or automated assistant software application that is integrated as part of the operating system of the client device 110. The application engine 115 (and the one or more software applications executed by the application engine 115) can interact with the knowledge system 120.

[0033] In various implementations, the client device 110 can include a user input engine 111 that is configured to detect user input provided by a user of the client device 110 using one or more user interface input devices. For example, the client device 110 can be equipped with one or more microphones that capture audio data, such as audio data corresponding to spoken utterances of the user or other sounds in an environment of the client device 110. Additionally, or alternatively, the client device 110 can be equipped with one or more vision components that are configured to capture vision data corresponding to digital images and / or movements (e.g., gestures) detected in a field of view of one or more of the vision components. Additionally, or alternatively, the client device 110 can be equipped with one or more touch sensitive components (e.g., a keyboard and mouse, a stylus, a touch screen, a touch panel, one or more hardware buttons, etc.) that are configured to capture signal(s) corresponding to touch input directed to the client device 110.

[0034] Some instances of an input described herein can be a natural language (NL) input that is formulated based on user input provided by a user of the client device 110 and detected via user input engine 111. For example, the query can be a typed query that is typed via a physical or virtual keyboard, a suggested query that is selected via a touch screen or a mouse of the client device 110, a spoken voice query that is detected via microphone(s) of the client device 110 (and optionally directed to an automated assistant executing at least in part at the client device 110), an image or video query that is based on vision data captured by vision component(s) of theAttorney Docket No. XDEV-0025-WO-01 client device 110 (or based on NL input generated based on processing the image using, for example, object detection model(s), captioning model(s), efc.), or any combination thereof.

[0035] Other instances of a NL based input described herein can be a prompt for proprietary content that is formulated based on user input provided by a user of the client device 110 and detected via the user input engine 111. For example, the prompt can be a typed prompt that is typed via a physical or virtual keyboard, a suggested prompt that is selected via a touch screen or a mouse of the client device 110, a spoken prompt that is detected via microphone(s) of the client device 110, or an image prompt that is based on an image captured by a vision component of the client device 110. In some implementations, a spoken NL input may be transcribed into text using speech-to-text (STT) processing, and the text may be used by downstream components to perform responsive actions. Alternatively, in some implementations, a spoken NL input may be processed using a machine learning model that is trained to directly map audio (e.g., waveform, phonemes, etc.) to responsive action(s), without performing STT to generate text.

[0036] In various implementations, the client device 110 can include a rendering engine 112 that is configured to render responsive content (e.g., NL based output, an indication of proprietary source(s)) associated with generative model output, and / or other content) for audible and / or visual presentation to a user of the client device 110 using one or more user interface output devices. For example, the client device 110 can be equipped with one or more speakers that enable the responsive content to be provided for audible presentation to the user via the client device 110. Additionally, or alternatively, the client device 110 can be equipped with a display or projector that enables the responsive content to be provided for visual presentation to the user via the client device 110.

[0037] In various implementations, the client device 110 can include a context engine 113 that is configured to determine a context (e.g., current or recent context) of the client device 110 and / or of a user of the client device 110 (e.g., an active user of the client device 110 when the client device 110 is associated with multiple users). In some of those implementations, context engine 113 can determine a context based on data stored in client device database 110A. The data stored in the client device database 110A can include, for example, user interaction data that characterizes current or recent interact! on(s) of the client device 110 and / or a user of the client device 110, location data that characterizes a current or recent location(s) of the clientAttorney Docket No. XDEV-0025-WO-01 device 110 and / or a user of the client device 110, user attribute data that characterizes one or more attributes of a user of the client device 110, user preference data that characterizes one or more preferences of a user of the client device 110, user profile data that characterizes a profile of a user of the client device 110, and / or any other data accessible to the context engine 113 via the client device database 110A or otherwise.

[0038] In some implementations, client device database 110A may include information such as metadata about proprietary data sources or “datasets” that are available to client device 110 and / or to a user of client device 110. A proprietary dataset may be any structured or unstructured source of proprietary data that can be accessed and queried using generative models as described herein. As will be explained in more detail below, publishers 170 may make proprietary data sources available, e.g., for subscription by users and / or organizations. In various implementations, a dataset may take the form of, for instance, an API that is accessible to retrieve or otherwise interact with proprietary data, a structured database (e.g., accessible using SQL), vector or raster image files, digital images, videos, satellite images, or aerial images, spreadsheets, free-form documents such as articles, reports, papers, a shared drive that contains homogenous and / or heterogenous digital content, and so forth. A user may have access to a proprietary dataset by virtue of that dataset being stored locally in client device database 110A, or by virtue of that dataset being accessible at another remote computing device / system (e.g., the cloud) using data stored in client device database 110A. In some implementations, a user may be “subscribed” or otherwise authorized to access specific proprietary datasets, e.g., via an API or other interface. Such a subscription may be provided to the user in various scenarios, such as part of the user’s employment or enrollment at a university or other research institute, as an on demand (e.g., paid) service, via a government agency, and so forth.

[0039] In various implementations, the context engine 113 can determine a current context based on a current state of a dialog session (e.g., considering one or more recent inputs provided by a user or outputs provided to the user during the human-to-computer dialog session), profile data, and / or a current location of the client device 110. For instance, the context engine 113 can determine a current context of “arborist seeking a count of trees in Cherokee Park in Louisville, Kentucky” based on a recently issued query, profile data, and an anticipated future location of the client device 110 (e.g., based on recently booked hotel accommodations in Louisville). As another example, the context engine 113 can determine a current context based on whichAttorney Docket No. XDEV-0025-WO-01 software application is active in the foreground of the client device 110, a current or recent state of the active software application, and / or content currently or recently rendered by the active software application. As yet another example, the context engine 113 can determine a current context based on which dataset(s) are available in or via client device database 110A.

[0040] A context determined by the context engine 113 can be utilized, for example, in supplementing or rewriting NL based input (or any type of input) that is formulated based on user input, in generating an intermediate input (e.g., an implicit query or prompt formulated independent of any explicit NL based input provided by a user of the client device 110), and / or in determining to submit an intermediate input and / or to render result(s) for an implicit NL based input. In some implementations, a context determined by the context engine 113 can be utilized in selecting one or more relevant or applicable datasets from the datasets that are available in or via client device database 110A. As will be described below, these datasets may be evaluated to determine which contain proprietary data that is potentially responsive or otherwise relevant to a user’s query.

[0041] Further, the client device 110 and / or the knowledge system 120 can include one or more memories for storage of data and / or software applications, one or more processors for accessing data and executing the software applications, and / or other components that facilitate communication over one or more of the networks 199. In some implementations, one or more of the software applications can be installed locally at the client device 110, whereas in other implementations one or more of the software applications can be hosted remotely (e.g., by one or more servers) and can be accessible by the client device 110 over one or more of the networks 199.

[0042] Although aspects of Fig. 1 are illustrated or described with respect to a single client device having a single user, it should be understood that is for the sake of example and is not meant to be limiting. For example, one or more additional client devices of a user and / or of additional user(s) can also implement the techniques described herein. For instance, the client device 110, the one or more additional client devices, and / or any other computing devices of a user can form an ecosystem of devices that can employ techniques described herein. These additional client devices and / or computing devices may be in communication with the client device 110 (e.g., over the network(s) 199). As another example, a given client device can beAttorney Docket No. XDEV-0025-WO-01 utilized by multiple users in a shared setting (e.g., a group of users, a household, a workplace, a hotel, etc.).

[0043] The knowledge system 120 is illustrated in Fig. 1 as including a prompt engine 139, an input processing engine 140, a dataset processing engine 152, and proprietary data metric engine 156. Some of these engines can be combined and / or omitted in various implementations. Further, these engines can include various sub-engines. For instance, the input processing engine 140 is illustrated in Fig. 1 as including a GM engine 142, a dialog context engine 146, and an output engine 150. Similarly, some of these sub-engines can be combined and / or omitted in various implementations. Accordingly, it should be understood that the various engines and sub-engines of the knowledge system 120 illustrated in Fig. 1 are depicted for the sake of describing certain functionalities and are not meant to be limiting.

[0044] The knowledge system 120 is illustrated in Fig. 1 as interfacing with various databases, such as GM(s) database 144, dialog context(s) database 148, and dataset GM(s) database 154. Although particular engines and / or sub-engines are depicted as having access to particular databases, it should be understood that is for the sake of example and is not meant to be limiting. For instance, in some implementations, each of the various engines and / or subengines of the knowledge system 120 may have access to each of the various databases.Further, some of these databases can be combined and / or omitted in various implementations. Accordingly, it should be understood that the various databases interfacing with the knowledge system 120 illustrated in Fig. 1 are depicted for the sake of describing certain data that is accessible to the knowledge system 120 and is not meant to be limiting.

[0045] In various implementations, knowledge system 120 can cause prompt engine 139 to generate an input prompt. GM engine 142 may then process this input prompt using a generative model such as a VLM or LLM stored in the GM(s) database 144 to generate model output. The model output may be provided by output engine 150 (which may be configured to cause the generative model output to be rendered at client device 110) and / or may be used to prompt the same GM or different GM(s) to generate additional information.

[0046] A generative model contained in any of databases 144 and 154 can include model(s) such as PaLM, BERT, LaMDA, Meena, and / or any other generative model, such as any other generative model that is encoder-only based, decoder-only based, sequence-to-sequence based and that optionally includes an attention mechanism or other memory, diffusion model(s), etc.Attorney Docket No. XDEV-0025-WO-01Generative models may have hundreds of millions, or even hundreds of billions of parameters. In some implementations, generative models may include multi-modal models such as a VLM and / or a visual question answering (VQA) model, which can have any of the aforementioned architectures, and which can be used to process multiple modalities of data, particularly images and text, and / or images and audio for example, to generate one or more modalities of output. Non-limiting examples of VLMs that may be applied as described herein include Gemini and / or Flamingo, to name a few.

[0047] The generative model output can include, for example, a probability distribution over a sequence of tokens, such as words, phrases, or other semantic units, which are predicted to be responsive to input processed using the GM. Notably, any generative model described herein can include billions of weights and / or parameters that are learned through training the generative model on enormous amounts of diverse data. This enables the generative model to generate output as the probability distribution over the sequence of tokens. In various implementations, knowledge system 120 may cause dialog context engine 146 to manage dialog contexts based on data stored in dialog context database 148, including identifying new dialog contexts, shifting between existing dialog contexts, managing the cascade of information across multiple dialog turns, etc.

[0048] In some implementations, GM engine 142 may be configured to process a natural language request (e.g., a query) received by user input engine 111 of client device 110. The natural language request may convey or otherwise identify a question to be answered and / or task to be performed. As will be explained in more detail herein, in some cases, the natural language request may be assembled, e.g., by a prompt engine 139, into both an endogenous prompt and an exogenous prompt. These prompts may then be processed by GM engine 142 using one or more generative models in database 144 to generate, respectively, endogenous and exogenous model outputs.

[0049] In some implementations, one or more generative models in GM(s) database 144 may be a “workflow” generative model that is trained to generate, based on natural language requests, output tokens that convey, correspond to, or otherwise identify, directly or indirectly, high-level actions for completing various tasks. In the context of the present disclosure, such a task often may include steps for accessing proprietary data of one or more publishers 170, so that the proprietary data can be queried using a generative model. In some implementations, theAttorney Docket No. XDEV-0025-WO-01 steps, which may or may not include a sequence of natural language commands (e.g., “go to xyz.com / proprietary.xml, download the documents that are linked to on that page”), may be incorporated into an input prompt, as part of in-context learning. In other implementations, the steps may be performed prior to application of the generative model, e.g., by prompt engine 139 while assembling a generative model input prompt.

[0050] In some implementations, for knowledge system 120 to generate allow users to query proprietary data (i.e., exogenous data that was not used to train or fine tune the applicable generative model(s)), knowledge system 120 may seek out proprietary data that is relevant and / or responsive to the user’s query. To this end, knowledge system 120 may be configured to identify, e.g., using context engine 113, proprietary dataset(s) that are (a) available to it and (b) contain potentially responsive proprietary data.

[0051] To accomplish this, in some implementations, knowledge system 120 may cause the dataset processing engine 152 to process, using a dataset GM stored in the dataset GM(s) database 154, data indicative of one or more proprietary datasets that are available to client device 110 and / or to a user of client device 110. For example, in some implementations, dataset processing engine 152 may be prompted with metadata indicative of one or more available proprietary datasets.

[0052] In some cases, proprietary datasets to which a user is not yet subscribed may nonetheless be at least partially available. This allows GM engine 142 to apply the generative model to the at least partially available proprietary data (or to complete proprietary data) to generate partially exogenous model output. These partially exogenous model outputs may be operable to navigate to an interface that allows a user to subscribe to the proprietary data source(s), such as a publisher’s webpage, an online marketplace of proprietary data sources, etc.

[0053] In some implementations, once workflow(s) such as those set forth above are identified from workflow output tokens, data indicative of the workflow(s), such as the workflow output tokens directly, embedding(s) generated therefrom, etc., may be processed by dataset processing engine 152 using a dataset GM stored in dataset GM(s) database 154. A dataset GM may be trained to identify suitable dataset(s) for accomplishing various tasks. In some such implementations, data indicative of the workflow(s) and data indicative of candidate datasets (e.g., dataset metadata) may be used to assemble a prompt for the dataset GM. TheAttorney Docket No. XDEV-0025-WO-01 dataset GM may then generate dataset output tokens that identify one or more responsive datasets that likely contain proprietary data that is responsive to the user’s initial query.

[0054] A proprietary data metric engine 156 may be configured to calculate a variety of metrics, scores, or other quantifications associated with proprietary data sources. For example, proprietary data metric engine 156 may compare different model outputs generated using the same query to determine one or more information gain scores associated with one or the other of the model outputs, and / or with one or more proprietary data sources used to generate one or both outputs. For example, proprietary data metric engine 156 may compare an endogenous model output generated without adding proprietary data to the generative model’s context window with an exogenous model output generated by adding proprietary data to the generative model’s context window to calculate an information gain score associated with the exogenous model output, relative to the endogenous model output.

[0055] Such an information gain score may be used for a variety of purposes. In some implementations, an information gain score may inform a publisher 170 of how much added value their proprietary data adds, compared to an endogenous model output generated using the same query. A publisher 170 may use this information in a variety of ways.

[0056] For instance, a publisher 170 could obtain or synthesize a large number of queries related to subject matter and / or topic areas in which the publisher seeks to publicize their proprietary data. Those queries may be assembled with proprietary data into multiple endogenous input prompts (without including additional proprietary data in the context window) and corresponding exogenous input prompts (with proprietary data included in the context window). Those pairs of endogenous and exogenous input prompts may be processed, e.g., by GM engine 142 using generative model(s) 144, to generate pairs of endogenous and exogenous model outputs. The endogenous and exogenous model outputs of each pair may be compared to each other to determine information gain score(s). Queries that yielded model output pairs with relatively low information gain scores may be identified by the publisher 170 as subject matter and / or topic areas in which their proprietary data may be lacking, relative to what the generative model is able to generate endogenously.

[0057] An information gain score may alternatively be used by users, who may be potential subscribers to proprietary data of publishers 170. In some implementations, when a user queries a generative model, in addition to endogenous model output being generated, one or moreAttorney Docket No. XDEV-0025-WO-01 exogenous model outputs may be generated using one or more publishers’ proprietary data. Information gain scores determined for these exogenous model outputs, relative to the endogenous model output, may be presented to the user, so that the user can decide whether to subscribe to one or more of the publishers’ proprietary data. In some cases, partial exogenous model output (e.g., truncated or abridged so that some content is not made available without subscribing) may be presented to the user, in addition to or instead of the corresponding information gain score. This may further help the user decide whether to subscribe to the publisher’s proprietary data.

[0058] In some implementations, the information gain scores may be used, with or without being explicitly presented to the user, to determine whether to present partial exogenous model output to the user. For example, if an information gain score for a particular exogenous model output is relatively high — e.g., satisfying some threshold or deviating sufficiently from other information gain scores associated with other exogenous model outputs — then partial exogenous model output may be presented to the user. In some implementations, this partial exogenous model output may include a selectable element that is operable by the user to see the rest of the exogenous model output and / or to subscribe to the publisher’s proprietary data. On the other hand, exogenous model outputs having lower information gain scores may not be partially presented to the user. Intuitively, if there is not much information to be gained by a particular exogenous output, then the user is not likely to subscribe, or is more likely to be unhappy if they do subscribe.

[0059] In various implementations, each publisher 170 may host or otherwise control one or more proprietary datasets 172. Thus, in Fig. 1, first publisher 170A has one or more proprietary datasets 172A, second publisher 170B has one or more proprietary datasets 172B, and so on. Each publisher 170A, B may also preprocess their proprietary datasets 172A, B to generate indexed proprietary data 174 A, B. Indexed proprietary data 174 A, B may take various forms, such as structured data that is more easily queried than proprietary datasets 172A, B that may be less structured or unstructured. For example, if a publisher’s proprietary datasets 172 include a plurality of digital images, then the publisher 170 and / or GM engine 142 may preprocess these images, e.g, using a VLM, to generate captions, textual summaries, lists of detected objects, list of visual features, etc. These captions, textual summaries, lists of detected objects, list of visual features, etc. may be preemptively stored as indexed proprietary data 174 A, B.Attorney Docket No. XDEV-0025-WO-01

[0060] Additionally or alternatively, each publisher 170A, B may or may not have additional metadata 176A, B available that describes proprietary datasets 172A, B and / or indexed proprietary data 174A, B. For instance, if proprietary dataset 172A includes data that is accessible via an API, then metadata 176 A may include a description of the API and / or documentation of how to interface with the API.

[0061] In some implementations, one or more of proprietary dataset(s) 172, indexed proprietary data 174, or metadata 176 may be used to determine whether proprietary data may be potentially relevant to a user’s query. For example, one or more of these sources may be used to generate a semantic embedding that can then be compared to a semantic embedding generated from the user’s query, e.g., using techniques such as cosine or Jaccard similarity. In some implementations, those proprietary data sources that are found to be sufficiently similar to the user’s query may be leveraged to generate at least partial exogenous outputs. These at least partial exogenous outputs may be designed to entice potential subscribers to subscribe to the proprietary data sources. Fig. 4 depicts two such examples.

[0062] Fig. 2 schematically depicts an example of how various components of Fig. 1 may cooperate to implement selected aspects of the present disclosure. Starting at top, client device 110 may provide data indicative of a natural language request 260 typed or spoken by a user (not depicted) to knowledge system 120. In various implementations, the natural language request 260 may convey a query or request to perform a task. In other implementations, other modalities of data may be provided as input(s), such as visual data, audio data, etc.

[0063] Knowledge system 120 may cause GM engine 142 to process the data indicative of the natural language request 260, e.g., as input tokens for a workflow GM stored in GM(s) database 144. The workflow GM may generate workflow output tokens 262 that identify high- level actions for responding to the natural language request. In some implementations, these high-level actions may be expressed as natural language statement(s).

[0064] The workflow output tokens 262 may then be used by dataset processing engine 152 to assemble an input prompt for a dataset GM stored in dataset GM(s) database 154. In some implementations, the workflow output tokens 262 may be used directly as input tokens for the dataset GM. In other implementations, the high-level actions identified from the workflow output tokens 262 may be used. In some implementations, the input prompt for the dataset GMAttorney Docket No. XDEV-0025-WO-01 may also include metadata (e.g., 176) associated with candidate datasets that are available to the user and / or client device 110.

[0065] Dataset processing engine 152 may generate, using the dataset GM, dataset output tokens 264 that identify one or more responsive datasets that likely contain data responsive to the natural language request 260. For example, the dataset output tokens 264 may identify which of the candidate datasets are likely to contain suitable data for responding to the natural language request 260. In other implementations, the dataset GM may be used as a classifier that separately processes metadata for each candidate dataset, and outputs a pass or fail.

[0066] The dataset output tokens 264 may be used by a data manipulation (DM) processing engine 265 (not depicted in Fig. 1) to assemble an input prompt for a data manipulation GM from data manipulation GM(s) database 267. As shown by the arrow at right, in some implementations, the workflow output tokens (or data indicative thereof) may also be provided to DM processing engine 265 and used to assemble the input prompt for the data manipulation GM from DM GM(s) database 267. Using the data manipulation GM, DM processing engine 265 may generate data manipulation output tokens 266 that identify data manipulation instructions for assembling data from the one or more responsive datasets into a response that fulfills the natural language request 260.

[0067] In various implementations, the data manipulation output tokens 266 may then be processed, e.g., by DM processing engine 265 or another component of knowledge system 120, to generate the response 268 to the natural language request 260. Data indicative of the response 268 may then be provided to client device 110, e.g., so that rendering engine 112 can render appropriate audible or visual output.

[0068] In some implementations, DM processing engine 265 may evaluate the data manipulation output tokens 266 to determine whether they satisfy various criteria, e.g., before assembling the response 268 and sending it back to client device 110. For instance, if the data manipulation output tokens 266 include or are indicative of source code in a high level programming language, DM processing engine 265 may attempt to compile the source code. If compiler errors are thrown, DM processing engine 265 may take various remedial actions. In some implementations, the source code and / or compiler errors may be provided as inputs to an error correcting LLM (not depicted), which in turn will generate, as output tokens, new source code. The new source code may once again be evaluated by DM processing engine 265, and theAttorney Docket No. XDEV-0025-WO-01 process may repeat until, for instance, no more compiler errors are thrown, or until a maximum number of loops has been reached. If the criteria are satisfied, the response 268 may then be generated and provided to client device 110.

[0069] In some implementations, knowledge system 120 may provide a multi -turn human-to- computer dialog that enables the user to engage with an automated assistant iteratively until the user can achieve one or more goals. For instance, the process depicted in Fig. 2 may be performed iteratively to eventually provide a response 268 to the user’s natural language request 260. If the user is dissatisfied with response 268, the user may provide, e.g., as spoken or typed natural language, feedback about the response.

[0070] Fig. 3 schematically depicts an example of how techniques described herein may be used to generate information gain scores that can then be used for various purposes. Starting at top, client device 110 once again receives a natural language request 360, similar to what was depicted in Fig. 2. Knowledge system 120 may cause prompt engine 139 to assemble an endogenous input prompt 380 that includes data indicative of the natural language request 360, without any extra proprietary data.

[0071] Knowledge system 120 may also cause prompt engine 139 to assemble one or more exogenous prompts so that the model exogenous output generated from those exogenous prompts can be compared to the endogenous model output generated using the endogenous input prompt 380. In Fig. 3, prompt engine 139 uses two different sets of proprietary data 382A, 382B, to assemble two different exogenous prompts 384A, 384B, respectively. First proprietary data 382A may be published by a first publisher, e.g., first publisher 170A in Fig. 1. Similarly, second proprietary data 382B may be published by a second publisher, e.g., second publisher 170B in Fig. 1. In other implementations, other numbers of exogenous input prompts may be generated, so that the respective resulting exogenous model outputs can be compared to each other and / or to the endogenous model output generated using endogenous prompt 380. In yet other implementations, endogenous prompt 380 may be omitted, and one or more exogenous prompts 384 may be assembled, e.g., for comparison of their respective exogenous model outputs.

[0072] GM engine 142 may use one or more generative models from GM database 144 to process endogenous input prompt 380, first exogenous prompt 384 A, and second exogenousAttorney Docket No. XDEV-0025-WO-01 prompt 384B. This processing may result in the generation of an endogenous model output 386, as well as first exogenous model output 388A and second exogenous model output 388B.

[0073] These model outputs 386, 388A-B may be compared by proprietary data metric engine 156 to generate information gain scores 390 that demonstrate how much more information one or more of model outputs 386, 388A-B contains versus the others. Information gain scores 390 may be calculated in various ways, such as edit distances, distances between latent space embeddings (e.g., using cosine similarity, Euclidean distance, Jaccard distance, etc.).

[0074] In some implementations, proprietary data metric engine 156 may calculate an information gain score 390 between two model outputs, say 386 and 388A in Fig. 3, as follows. First, proprietary data metric engine 156 may perform tokenization by breaking down both model outputs 386 and 388 A into individual words or tokens. Proprietary data metric engine 156 may then perform vocabulary creation in which it creates a combined vocabulary of all unique words from both documents. Proprietary data metric engine 156 may then calculate term frequency (TF), e.g., by calculating the frequency of each word in each model output 386 and 388A.

[0075] Proprietary data metric engine 156 may next calculate inverse document frequency(IDF) for each word, which measures how rare the word is across the entire corpus. With the TD and IDF calculated, proprietary data metric engine 156 may then calculate the TF-IDF score for each word in each document. The TF-IDF score combines term frequency and inverse document frequency to weigh words based on their importance. In some implementations, proprietary data metric engine 156 may then calculate the cosine similarity between the TF-IDF vectors of the two documents. Cosine similarity measures the angle between the vectors, with a value closer to 1 indicating higher similarity. Finally, proprietary data metric engine 156 may calculate the information gain as the difference between an entropy of the combined vocabulary of the model outputs 386 and 388 A, and the entropy of the vocabulary based on one of the model outputs, e.g., 386. In some implementations, proprietary data metric engine 156 may calculate entropy using Shannon entropy, e.g., with the following formula: logp(x)where p(x) is the probability of occurrence of the xth character in the document,Attorney Docket No. XDEV-0025-WO-01

[0076] These information gain scores 390 may be used for a variety of purposes. In some implementations, the information gain scores 390 may be rendered at one or more output devices, e.g., to inform a publisher 170 of the information gained in an exogenous model output from their proprietary data compared to an endogenous model output, or compared to another exogenous model output generated using different proprietary data. Alternatively, a potential subscriber of proprietary data may be presented with the information gain score to help the user assess whether to subscribe.

[0077] In some implementations, information gain scores may not be sufficient by themselves to determine whether proprietary data will be useful to the user who issued the query being answered. For example, it may be the case that proprietary data that is only tangentially relevant to an initial query nonetheless provides a significant amount of information beyond endogenous output. This gained information may not necessarily be interesting to the user. Accordingly, in some implementations, in addition to or instead of an information gain score, a similarity score may be calculated between the extra information returned in an exogenous model output and the query that was used to generate the exogenous model output. This similarity score may be calculated using techniques such as cosine similarity, Jaccard distance, Euclidean distance, etc. In some implementations, the similarity score may be used to boost or reduce an information gain score. In other implementations, the similarity score and the information gain score may be used in tandem to determine whether to present a particular exogenous model output.

[0078] Fig. 4 depicts a non-limiting example of how information gain scores may be used to influence output generated for a user of a generative model-powered automated assistant. In Fig. 4, a client device 410 takes the form of a tablet computer or smartphone that includes a touchscreen 492. Touchscreen 492 is being used to render a graphical user interface (GUI) that includes a natural language input field 493 and a model output portion 494.

[0079] In the example of Fig. 4, the user has provided (e.g., typed, spoken and transcribed) the natural language input, “What is the tree coverage in Cherokee Park?” In the model output portion 494, an endogenous model output 486 provides the following information based solely on knowledge trained into the generative model:Cherokee park currently has a canopy coverage of 80%. This means that about 80% of the park's ground area is shaded by trees, providingAttorney Docket No. XDEV-0025-WO-01 a significant amount of green space and environmental benefits to the city.This output indicates that at some point, the generative model(s) used to generate endogenous model output 486 were trained and / or fine tuned based on documents that included this information.

[0080] The model output portion also includes additional exogenous information that is presented to give the user an opportunity to subscribe to proprietary data that may not have been used to train the generative model(s), and therefore does not appear in endogenous output 486. For example, a first partial exogenous output 488A includes more granular detail, beginning with “This coverage includes 35% maple trees. . .” This suggests that the proprietary data underlying first exogenous output 488A includes information about individual tree species populations in the part. First exogenous output 488A also includes a selectable element in the form of a hyperlink (“SUBSCRIBE TO LEARN MORE”) that if clicked by the user, may navigate client device 310 to website of the corresponding publisher and / or to an online marketplace of proprietary data sources to which users can subscribe.

[0081] A second partial exogenous output 488B has been generated based on proprietary data about temperature readings measured by sensors (e.g., loT sensors) deployed in the park. Second partial exogenous output 488B includes the following text:Temperature sensors provided by ACME CORP demonstrate that this coverage influences the measured ground temperatures in Cherokee Park significantly. For example, in the open area around Hogan ’s fountain, the average daytime temperature in July is 90°, whereas in the surrounding tree-covered areas, the average daytime temperature in July is...Again, the user is presented with a selectable element (“SUBSCRIBE TO LEARN MORE”) that, if selected, navigates client device 310 to an online marketplace of proprietary data sources to which users can subscribe.

[0082] In some implementations, the partial exogenous model outputs presented during any given session between user(s) and a generative model-powered automated assistant may be presented conditionally / selectively, e.g., based on various signals, such as information gain scores, quality metrics, credibility metrics, similarity scores between the exogenous model outputs and the original query, etc. In Fig. 4, there may have been one or more additionalAttorney Docket No. XDEV-0025-WO-01 proprietary data sources with data that is at least tangentially relevant to the user’s original query. For example, there may be a historical society that publishes proprietary data in the form of Cherokee Park’s history. However, it may have been determined that this historical data that is not directly related to tree coverage is less similar to the natural language query than the other examples of proprietary data demonstrated in Fig. 4. Consequently, the partial exogenous model outputs 488A and 488B are presented instead.

[0083] Fig. 5 depicts a flowchart illustrating an example method of practicing selected aspects of the present disclosure, in accordance with various implementations. For convenience, the operations of method 500 are described with reference to a system that performs the operations. This system may include one or more processors, memory, and / or other component(s) of computing device(s). Moreover, while operations of the method 500 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.

[0084] At block 502, the system, e.g., by way of prompt engine 139, may assemble, as a first input prompt, at least one query (e.g., natural language query 360 or 460) for a generative model (e.g., 144). While natural language queries have been described herein, this is not meant to be limiting. In various implementations, queries having additional or alternative modalities of data, such as images, video, audio, etc., may be assembled by prompt engine 139 into input prompts.

[0085] At block 504, the system, e.g., by way of GM engine 142, may process the first input prompt using the generative model to generate “reference” model output. This “reference” model output is so named because it may be used for comparison with other, exogenous model outputs. To this end, the reference model output may be an endogenous model output (e.g., 486) or an exogenous model output (e.g., 488 A or B). In the former case, there may be an exogenous- to-endogenous model output comparison. In the latter case, there may be a downstream exogenous-to-exogenous model output comparison.

[0086] At block 506, the system, e.g., by way of dataset processing engine 152, may identify one or more proprietary data sources for use in prompting one or more generative models. For example, and as demonstrated in Fig. 2, dataset processing engine 152 may process workflow output tokens 262 to generate dataset output tokens 264, which can then be used by downstream processes (e.g., data manipulation engine 265) to access the proprietary data.Attorney Docket No. XDEV-0025-WO-01

[0087] At block 508, the system, e.g., by way of prompt engine 139, may assemble, as a second, exogenous input prompt, the at least one query and proprietary data retrieved, obtained, and / or derived from one or more of the proprietary data sources. At block 510, the system, e.g., by way of GM engine 142, may process the second input prompt using the generative model to generate at least partially exogenous model output based at least in part on the proprietary data.

[0088] At block 512, the system, e.g., by way of proprietary data metric engine 156, may compare the reference model output generated at block 504 with the at least partially exogenous model output generated at block 510, to determine one or more information gain scores associated with the one or more proprietary data sources. At block 514, the system, e.g., by way of proprietary data metric engine 156, may determine whether the information gain scores satisfy one or more criteria. For example, proprietary data metric engine 156 may determine whether an information gain score of an exogenous model output satisfies some predetermined threshold. While not shown in Fig. 5, in some implementations, other metrics and / or signals may be considered, in addition to (or instead of) the information gain scores, such as credibility metrics, similarity of the exogenous information to the user’s original query, etc.

[0089] If the answer at block 514 is yes, then method 500 may proceed to block 516, At block 516, the system, e.g., by way of output engine 150, may cause one or more of the information gain scores and / or the at least partial exogenous model output to be presented at one or more output devices. For example, in Fig. 4, two partial exogenous model outputs 488A and 488B were rendered on touchscreen 494 because their respective information gain scores (and other signals if applicable) satisfied the criteria of block 514. While not shown in Fig. 5, in some cases, the endogenous (e.g., reference) model output may be rendered, too. If the answer at block 514 is no, on the other hand, then the reference model output may be rendered, e.g., without any additional exogenous output.

[0090] Turning now to Fig. 6, a block diagram of an example computing device 610 that may optionally be utilized to perform one or more aspects of techniques described herein is depicted. In some implementations, one or more of a client device, cloud-based automated assistant component(s) or other cloud-based software application component(s), and / or other component(s) may comprise one or more components of the example computing device 610.

[0091] Computing device 610 typically includes at least one processor 614 which communicates with a number of peripheral devices via bus subsystem 612. These peripheralAttorney Docket No. XDEV-0025-WO-01 devices may include a storage subsystem 624, including, for example, a memory subsystem 625 and a file storage subsystem 626, user interface output devices 620, user interface input devices 622, and a network interface subsystem 616. The input and output devices allow user interaction with computing device 610. Network interface subsystem 616 provides an interface to outside networks and is coupled to corresponding interface devices in other computing devices.

[0092] User interface input devices 622 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into computing device 610 or onto a communication network.

[0093] User interface output devices 620 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual display such as via audio output devices. In general, use of the term "output device" is intended to include all possible types of devices and ways to output information from computing device 610 to the user or to another machine or computing device.

[0094] Storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 may include the logic to perform selected aspects of the methods disclosed herein, as well as to implement various components depicted in Figs. 1 or 2.

[0095] These software modules are generally executed by processor 614 alone or in combination with other processors. Memory 625 used in the storage subsystem 624 can include a number of memories including a main random-access memory (RAM) 630 for storage of instructions and data during program execution and a read only memory (ROM) 632 in which fixed instructions are stored. A file storage subsystem 626 can provide persistent storage for program and data files, and may include a hard disk drive, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical drive, or removable media cartridges. The modules implementing the functionality of certain implementations may be stored by fileAttorney Docket No. XDEV-0025-WO-01 storage subsystem 626 in the storage subsystem 624, or in other machines accessible by the processor(s) 614.

[0096] Bus subsystem 612 provides a mechanism for letting the various components and subsystems of computing device 610 communicate with each other as intended. Although bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem 612 may use multiple buses.

[0097] Computing device 610 can be of varying types including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 610 depicted in Fig. 6 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computing device 610 are possible having more or fewer components than the computing device depicted in Fig. 6.

[0098] In various implementations, a method may be implemented using one or more processors and may include: assembling, as a first input prompt, at least one query for a generative model; processing the first input prompt using the generative model to generate reference model output; identifying one or more proprietary data sources for use in prompting one or more generative models; assembling, as a second input prompt, the at least one query and proprietary data from one or more of the proprietary data sources; processing the second input prompt using the generative model to generate at least partially exogenous model output based at least in part on the proprietary data; comparing the reference model output with the at least partially exogenous model output to determine one or more information gain scores associated with the one or more proprietary data sources; and causing one or more of the information gain scores to be presented at one or more output devices.

[0099] In various implementations, the reference model output may include endogenous model output. In various implementations, the reference model output may be at least partially exogenous. In various implementations, the first input prompt may be further assembled to include additional proprietary data from one or more other proprietary data sources.

[0100] In various implementations, method may further include: assembling, as an indexing input prompt, data retrieved from one or more of the proprietary data sources, and a request to summarize the data retrieved from one or more of the proprietary data sources; and processing the indexing input prompt using one or more of the generative models to generate indexingAttorney Docket No. XDEV-0025-WO-01 output, wherein the indexing output includes the proprietary data that is assembled into the second input prompt.

[0101] In various implementations, one or more of the proprietary data sources may include one or more digital images. In various implementations, the method may include: assembling, as an indexing input prompt, data indicative of the one or more digital images and a request to caption and / or summarize the one or more digital images; and processing the indexing input prompt using a vision-language model (VLM) to generate indexing output, wherein the indexing output includes: one or more captions of the one or more digital images, and / or a summary of the one or more digital images; wherein the proprietary data that is assembled into the second input prompt comprises the one or more captions and / or the summary.

[0102] In various implementations, one or more of the proprietary data sources may include digital audio content. In various implementations, the method may include: assembling, as an indexing input prompt, data indicative of the digital audio content; and processing the indexing input prompt using a speech-to-text (STT) model to generate indexing output, wherein the indexing output includes a transcript of at least part of the digital audio content; wherein the proprietary data that is assembled into the second input prompt comprises the transcript.

[0103] In various implementations, one or more of the proprietary data sources may include one or more Internet of Things (loT) sensors. In various implementations, the method may include: assembling, as an indexing input prompt, sensor data obtained from the one or more loT sensors; and processing the indexing input prompt using the same generative model or a different generative model to generate indexing output, wherein the indexing output includes a table view or summary of, or statistics about, the sensor data obtained from the one or more loT sensors; wherein the proprietary data that is assembled into the second input prompt comprises the table view or summary of, or statistics about, the sensor data obtained from the one or more loT sensors.

[0104] In various implementations, the method may include: assembling, as a credibility input prompt, a description of how data from one or more of the proprietary data sources was obtained; processing the credibility input prompt using the same generative model or a different machine learning model to generate credibility model output, wherein the credibility model output includes one or more credibility metrics associated with one or more of the proprietary data sources; and causing one or more of the credibility metrics to be presented at one or moreAttorney Docket No. XDEV-0025-WO-01 of the output devices. In various implementations, the causing may be performed conditionally based on a comparison of one or more of the information gain scores to one or more thresholds.

[0105] In another aspect, a method may be implemented using one or more processors and may include: assembling, as a first input prompt, data that is descriptive of one or more proprietary data sources; processing the first input prompt using one or more generative model(s) to generate first model output that is usable to access one or more of the proprietary data sources; based on the first model output, retrieving proprietary data from one or more of the proprietary data sources; assembling, as a second input prompt, at least one query and the proprietary data; processing the second input prompt using one or more of the generative models to generate at least partially exogenous model output; and causing one or more output devices to render output indicative of the at least partially exogenous model output.

[0106] In various implementations, the first model output may include one or more natural language snippets describing one or more steps to take to retrieve the proprietary data from one or more of the proprietary data sources. In various implementations, the first model output may include synthetic source code that is executable to retrieve the proprietary data from one or more of the proprietary data sources.

[0107] In various implementations, the data that is descriptive of one or more proprietary data sources may include documentation about an application programming interface (API) that is usable to retrieve proprietary data from one or more of the proprietary data sources. In various implementations, retrieving proprietary data from one or more of the proprietary data sources may include: retrieving one or more digital images from one or more of the proprietary data sources; and processing the one or more digital images using a vision language model (VLM) to generate textual content that describes one or more visual features of the one or more digital images, wherein the textual content forms at least part of the proprietary data that is assembled into the second input prompt.

[0108] In addition, some implementations include one or more processors (e.g., central processing unit(s) (CPU(s)), graphics processing unit(s) (GPU(s)), neural processing unit(s) (NPU(s)), and / or tensor processing unit(s) (TPU(s)) of one or more computing devices, where the one or more processors are operable to execute instructions stored in associated memory, and where the instructions are configured to cause performance of any of the aforementioned methods. Some implementations also include one or more transitory or non-transitory computerAttorney Docket No. XDEV-0025-WO-01 readable storage media storing computer instructions executable by one or more processors to perform any of the aforementioned methods.

[0109] While several implementations have been described and illustrated herein, a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein may be utilized, and each of such variations and / or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.

Claims

Attorney Docket No. XDEV-0025-WO-01CLAIMSWhat is claimed is:

1. A method implemented using one or more processors and comprising: assembling, as a first input prompt, at least one query for one or more generative models; processing the first input prompt using one or more of the generative models to generate reference model output; identifying one or more proprietary data sources for use in prompting one or more of the generative models; assembling, as a second input prompt, the at least one query and proprietary data from one or more of the proprietary data sources; processing the second input prompt using one or more of the generative models to generate at least partially exogenous model output based at least in part on the proprietary data; comparing the reference model output with the at least partially exogenous model output to determine one or more information gain scores associated with the one or more proprietary data sources; and causing one or more of the information gain scores to be presented at one or more output devices.

2. The method of claim 1, wherein the reference model output comprises endogenous model output.

3. The method of claim 1 or 2, wherein the reference model output is at least partially exogenous.

4. The method of claim 3, wherein the first input prompt is further assembled to include additional proprietary data from one or more other proprietary data sources.

5. The method of any of the preceding claims, further comprising: assembling, as an indexing input prompt, data retrieved from one or more of the proprietary data sources, and a request to summarize the data retrieved from one or more of the proprietary data sources; and processing the indexing input prompt using one or more of the generative models to generate indexing output, wherein the indexing output includes the proprietary data that is assembled into the second input prompt.Attorney Docket No. XDEV-0025-WO-016. The method of any of the preceding claims, wherein one or more of the proprietary data sources comprise one or more digital images.

7. The method of claim 6, further comprising: assembling, as an indexing input prompt, data indicative of the one or more digital images and a request to caption and / or summarize the one or more digital images; and processing the indexing input prompt using a vision-language model (VLM) to generate indexing output, wherein the indexing output includes: one or more captions of the one or more digital images, and / or a summary of the one or more digital images; wherein the proprietary data that is assembled into the second input prompt comprises the one or more captions and / or the summary.

8. The method of any of the preceding claims, wherein one or more of the proprietary data sources comprises digital audio content.

9. The method of claim 8, further comprising: assembling, as an indexing input prompt, data indicative of the digital audio content; and processing the indexing input prompt using a speech-to-text (STT) model to generate indexing output, wherein the indexing output includes a transcript of at least part of the digital audio content; wherein the proprietary data that is assembled into the second input prompt comprises the transcript.

10. The method of any of the preceding claims, wherein one or more of the proprietary data sources comprises one or more Internet of Things (loT) sensors.

11. The method of claim 10, further comprising: assembling, as an indexing input prompt, sensor data obtained from the one or more loT sensors; and processing the indexing input prompt using the same one or more of the generative models to generate indexing output, wherein the indexing output includes a table view or summary of, or statistics about, the sensor data obtained from the one or more loT sensors; wherein the proprietary data that is assembled into the second input prompt comprises the table view or summary of, or statistics about, the sensor data obtained from the one or more loT sensors.Attorney Docket No. XDEV-0025-WO-0112. The method of any of the preceding claims, further comprising: assembling, as a credibility input prompt, a description of how data from one or more of the proprietary data sources was obtained; processing the credibility input prompt using one or more of the generative models to generate credibility model output, wherein the credibility model output includes one or more credibility metrics associated with one or more of the proprietary data sources; and causing one or more of the credibility metrics to be presented at one or more of the output devices.

13. The method of any of the preceding claims, wherein the causing is performed conditionally based on a comparison of one or more of the information gain scores to one or more thresholds.

14. A method implemented using one or more processors and comprising: assembling, as a first input prompt, data that is descriptive of one or more proprietary data sources; processing the first input prompt using one or more generative models to generate first model output that is usable to access one or more of the proprietary data sources; based on the first model output, retrieving proprietary data from one or more of the proprietary data sources; assembling, as a second input prompt, at least one query and the proprietary data; processing the second input prompt using one or more of the generative models to generate at least partially exogenous model output; and causing one or more output devices to render output indicative of the at least partially exogenous model output.

15. The method of claim 14, wherein the first model output comprises one or more natural language snippets describing one or more steps to take to retrieve the proprietary data from one or more of the proprietary data sources.

16. The method of claim 14, wherein the first model output comprises synthetic source code that is executable to retrieve the proprietary data from one or more of the proprietary data sources.

17. The method of any of claims 14-16, wherein the data that is descriptive of one or more proprietary data sources includes documentation about an application programmingAttorney Docket No. XDEV-0025-WO-01 interface (API) that is usable to retrieve proprietary data from one or more of the proprietary data sources.

18. The method of any of claims 14-17, wherein retrieving proprietary data from one or more of the proprietary data sources comprises: retrieving one or more digital images from one or more of the proprietary data sources; and processing the one or more digital images using a vision language model (VLM) to generate textual content that describes one or more visual features of the one or more digital images, wherein the textual content forms at least part of the proprietary data that is assembled into the second input prompt.

19. A method implemented by one or more processors, the method comprising: constructing a first input prompt comprising at least one query directed to one or more generative models; providing the first input prompt to at least one of the one or more generative models to elicit a reference model output; identifying one or more proprietary data sources containing data determined to be contextually relevant to the at least one query; constructing a second input prompt comprising the at least one query and proprietary data retrieved from the one or more proprietary data sources; providing the second input prompt to the at least one of the one or more generative models to elicit an at least partially exogenous model output conditioned on the proprietary data; calculating one or more information gain metrics associated with the one or more proprietary data sources by performing a comparative analysis between the reference model output and the at least partially exogenous model output; and transmitting data indicative of the one or more information gain metrics for presentation at one or more output devices.

20. A method implemented by one or more processors, the method comprising: receiving, from a user device associated with a user account, at least one query for one or more generative models; identifying a proprietary data source containing proprietary data that is contextually relevant to the at least one query;Attorney Docket No. XDEV-0025-WO-01 determining that the user account lacks a subscription to the identified proprietary data source; assembling an input prompt comprising the at least one query, at least a portion of the proprietary data from the identified proprietary data source, and an instruction to generate a partial or obfuscated response; processing the input prompt using at least one of the one or more generative models to generate a partial exogenous model output; and causing the partial exogenous model output to be presented at the user device, wherein the partial exogenous model output includes a selectable element operable to initiate a subscription to the identified proprietary data source.

21. A method implemented using one or more processors, the method comprising: receiving data indicative of a request; processing, using a first generative model, the data indicative of the request to generate a first model output, the first model output identifying one or more high-level actions for responding to the request; processing, using a second generative model, the first model output to generate a second model output, the second model output identifying one or more proprietary data sources determined to be relevant to the one or more high-level actions; processing, using a third generative model, the second model output to generate a third model output, the third model output identifying one or more data manipulation instructions for generating a response to the request using proprietary data from the one or more proprietary data sources; and causing a response to be generated based at least in part on the one or more data manipulation instructions.

22. A method implemented by one or more processors, the method comprising: receiving a request, associated with a subscriber account, to access proprietary data from a proprietary data source; determining one or more access control policies associated with the proprietary data source, wherein at least one of the one or more access control policies defines a rate limit for accessing the proprietary data;Attorney Docket No. XDEV-0025-WO-01 evaluating the request against the one or more access control policies to determine whether the request complies with the rate limit, wherein the rate limit is based on at least one of: a quantity of queries permitted over a defined time period, a type of the proprietary data being requested, an amount of the proprietary data being requested, a trust level assigned to the proprietary data, or a determination that the proprietary data contains personally identifiable information (PII); and in response to determining that the request complies with the rate limit, causing the proprietary data to be provided for use in conditioning an output of one or more generative models.

23. A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform any of the methods of claims 1-22.

24. At least one transitory or non-transitory computer-readable medium comprising instructions that, in response to execution by one or more processors, cause the one or more processors to perform any of the methods of claims 1-22.