Ambient multi-device framework for agent companions

The multi-device framework optimizes computing device interactions by using environment-based prompt generation and generative models to enhance output delivery across devices, addressing interconnectivity issues and improving computational efficiency.

JP2025121898AActive Publication Date: 2025-08-20GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025061256
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2025-04-02
Publication Date
2025-08-20
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

Existing computing devices in a user's environment lack interconnectivity, leading to inefficient utilization of their capabilities, as they often fail to determine the optimal device for providing context-specific output based on user queries.

Method used

A multi-device framework that leverages environment-based prompt generation and generative models to determine the most suitable computing device for output, utilizing input and environmental data to generate prompts that are processed to provide responses tailored to the device's capabilities.

Benefits of technology

Enhances the utilization of diverse computing device capabilities by providing context-specific output across multiple devices, improving computational efficiency and functionality by centralizing data processing and reducing repetitive tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025121898000001_ABST
    Figure 2025121898000001_ABST
Patent Text Reader

Abstract

To provide an ambient multi-device framework for agent companions.SOLUTION: Systems and methods for generating and providing outputs in a multi-device system can include leveraging environment-based prompt generation and generative model response generation to provide dynamic response generation and display. The systems and methods can obtain input data associated with one or more computing devices within an environment, can obtain environment data descriptive of the plurality of computing devices within the environment, and can generate a prompt based on the input data and environment data. The prompt can be processed with a generative model to generate a model-generated output. The model-generated output can then be transmitted to a particular computing device of the plurality of computing devices.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority claims This application claims priority to U.S. Non-provisional Application No. 18 / 390,768, filed December 20, 2023, which is incorporated herein by reference.

[0002] The present disclosure relates generally to generating and providing output in a multi-device system, and more particularly to defining output generation based on devices in an environment and determining which computing devices provide the output. [Background technology]

[0003] Computing devices may be found throughout a user's living room, bedroom, study, office, and / or other environment. A composite device environment may provide a user with multiple computing devices to interact with, which may include smart TVs, smart speakers, smart appliances, virtual assistant devices, tablets, smart wearables, smartphones, and / or other computing devices, throughout the user's environment. However, the capabilities of these devices may not be utilized based on a lack of interconnectivity between the devices. For example, a user may be performing a search on a smartphone, which may cause video search results to play on the smartphone even though the smart TV is several feet away.

[0004] Understanding the entire situation can be difficult. Whether an individual is trying to understand what an object in front of them is, whether they are trying to determine if the object can be found elsewhere, and / or whether an internet image was captured from where, text search alone can be challenging. In particular, users may struggle to decide what words to use. Furthermore, those words may not be descriptive and / or rich enough to produce the desired results. Summary of the Invention

[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned by practice of the embodiments.

[0006] One exemplary aspect of the present disclosure is directed to a computing system for determining an output device for providing a query response. The system may include one or more processors and one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include obtaining input data. The input data may include a query associated with a particular user. The operations may include obtaining environmental data. The environmental data may describe multiple computing devices in the user's environment. In some implementations, the multiple computing devices may be associated with multiple different output components. The operations may include generating a prompt based on the input data and the environmental data. The prompt may include data describing the query and device information associated with at least a subset of the multiple computing devices. The operations may include processing the prompt with a generative model to generate a model-generated output. The model-generated output may include a response to the query. In some implementations, the model-generated output may be generated to be provided to a particular computing device of the multiple computing devices. The operations may include sending the model-generated output to the particular computing device.

[0007] In some implementations, the model-generated output may be generated to be provided to a particular output component of a plurality of different output components. The generative model may generate output device instructions. The output device instructions may describe a particular computing device of a plurality of computing devices that provides the model-generated output. The particular computing device may be associated with a particular output component. The model-generated output may be sent to the particular computing device based on the output device instructions.

[0008] In some implementations, processing the prompt with the generative model to generate a model-generated output may include generating a plurality of model outputs. The plurality of model outputs may include a plurality of candidate responses. Sending the model-generated output to a particular computing device may include sending a first model output of the plurality of model outputs to a first computing device of the plurality of computing devices and sending a second model output of the plurality of model outputs to a second computing device of the plurality of computing devices. The first model output may include visual data for display via a visual display. The second model output may include audio data for playback via a speaker component. In some implementations, the first computing device may include a smart television and the second computing device includes a smart speaker.

[0009] In some implementations, generating a prompt based on the input data and the environmental data may include determining an environment-specific device configuration based on the environmental data, retrieving a prompt template from a prompt library based on the environment-specific device configuration, and expanding the prompt template based on the input data to generate the prompt. The environment-specific device configuration may describe a respective output type and a respective output quality of a plurality of computing devices. The prompt library may include a plurality of different prompt templates associated with a plurality of different device configurations.

[0010] In some embodiments, the plurality of computing devices can be connected via a cloud computing system. Each of the plurality of computing devices can be registered with a platform of the cloud computing system. Environmental data can be acquired through the cloud computing system. Model-generated output can be transmitted via the cloud computing system. In some embodiments, the plurality of computing devices can be located proximate to each other computing device within the plurality of computing devices. The plurality of computing devices can be communicatively connected across a local network. A particular computing device of the plurality of computing devices can facilitate acquisition of input data and transmission of model-generated output.

[0011] In some embodiments, the Generative Model may be communicatively coupled to a search engine via an application programming interface. Processing the prompt with the Generative Model to generate the model-generated output may include generating application programming interface calls based on the prompt, determining a plurality of search results with the search engine based on the application programming interface calls, and processing the plurality of search results with the Generative Model to generate the model-generated output.

[0012] Another exemplary aspect of the present disclosure is directed to a computer-implemented method. The method may include acquiring, by a computing system including one or more processors, input data. The input data may include a query associated with a particular user. The method may include acquiring, by the computing system, environmental data. The environmental data may describe a plurality of computing devices in the user's environment. The plurality of computing devices may be associated with a plurality of different output components. The method may include generating, by the computing system, a prompt based on the input data and the environmental data. The prompt may include data describing the query and device information associated with at least a subset of the plurality of computing devices. The method may include processing, by the computing system, the prompt with the generative model to generate a model-generated output and output device instructions. The model-generated output may include a response to the query. In some implementations, the model-generated output may be generated to be provided to a particular output component of a plurality of different output components. The output device instructions may describe a particular computing device of the plurality of computing devices that provides the model-generated output. The particular computing device may be associated with a particular output component. The method may include sending, by the computing system, the model-generated output to the particular computing device based on the output device instructions.

[0013] In some implementations, a plurality of different output components may be associated with a plurality of respective output capabilities associated with a plurality of computing devices. Each of the plurality of respective output capabilities may describe an output type and output quality available via the respective computing device. The output device instructions may include application programming interface calls for sending the model-generated output to a particular computing device. In some implementations, the plurality of different output components may include a speaker associated with a first device and a visual display associated with a second device. The method may include determining, by the computing system, that a particular output component is associated with the intent of the query. A prompt may be generated based on the particular output component associated with the intent of the query.

[0014] In some implementations, the method may include determining, by the computing system, an output hierarchy based on specification information for a plurality of different output components based on the environmental data. A prompt may be generated based on the output hierarchy and the query.

[0015] Another exemplary aspect of the present disclosure is directed to one or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations can include obtaining environmental data. The environmental data can describe a plurality of computing devices in an environment associated with a particular user. The operations can include processing the environmental data to determine a plurality of respective input functions and a plurality of respective output functions associated with the plurality of computing devices. The plurality of respective input functions can be associated with candidate input types associated with the plurality of computing devices. The plurality of respective input functions can be associated with candidate output types associated with the plurality of computing devices. The operations can include generating a plurality of respective interfaces for the plurality of computing devices based on the plurality of respective input functions and the plurality of respective output functions. The plurality of respective interfaces can be specialized for the plurality of computing devices based on the plurality of respective input functions and the plurality of respective output functions. The operations can include providing the plurality of respective interfaces to the plurality of computing devices.

[0016] In some implementations, the plurality of respective interfaces may include a plurality of device indicators that indicate a plurality of computing devices in an environment associated with a particular user. The plurality of computing devices may be configured as a user-specific device ecosystem that is communicatively connected to receive input and provide output. Each of the plurality of respective interfaces may be configured to receive a particular input type and provide a particular output type based on the respective input capabilities and respective output capabilities of a particular computing device of the plurality of computing devices.

[0017] In some implementations, the operations may include obtaining user input via a first interface of a first computing device of the plurality of computing devices, processing the user input with a search engine to determine a plurality of search results, processing the plurality of search results with a generative model to generate model output, and providing the model output for display via a second interface of a second computing device of the plurality of computing devices.

[0018] Other aspects of the present disclosure are directed to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.

[0019] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the detailed description, serve to explain associated principles.

[0020] Detailed descriptions of embodiments directed to those skilled in the art are set forth herein with reference to the accompanying drawings. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 illustrates a block diagram of an exemplary generative model system, according to an exemplary embodiment of the present disclosure. [Figure 2] 1 illustrates a block diagram of an exemplary response generation system, according to an exemplary embodiment of the present disclosure. [Figure 3] FIG. 1 illustrates a flowchart diagram of an exemplary method for response generation, according to an exemplary embodiment of the present disclosure. [Figure 4] 1 illustrates a block diagram of an exemplary multi-device management system, in accordance with an exemplary embodiment of the present disclosure. [Figure 5]1 illustrates a block diagram of an exemplary environment personalization system, according to an exemplary embodiment of the present disclosure. [Figure 6A] 1 illustrates a diagram of an exemplary interface according to an exemplary embodiment of the present disclosure. [Figure 6B] 1 illustrates a diagram of an exemplary image capture entry point, according to an exemplary embodiment of the present disclosure. [Figure 6C] 1 shows a diagram of an exemplary smart TV entry point according to an exemplary embodiment of the present disclosure. [Figure 6D] 1 illustrates a diagram of an exemplary calendar interface, according to an exemplary embodiment of the present disclosure. [Figure 6E] 1 illustrates a diagram of an exemplary video conferencing interface, according to an exemplary embodiment of the present disclosure. [Figure 6F] 1 illustrates a diagram of an exemplary email interface, according to an exemplary embodiment of the present disclosure. [Figure 6G] 1 shows a diagram of an exemplary video player interface, according to an exemplary embodiment of the present disclosure. [Figure 7] FIG. 1 illustrates a flowchart diagram of an exemplary method for output generation and routing, according to an exemplary embodiment of the present disclosure. [Figure 8] FIG. 1 illustrates a flowchart diagram of an exemplary method for performing interface generation, according to an exemplary embodiment of the present disclosure. [Figure 9A] 1 illustrates a block diagram of an exemplary computing system for performing multi-device output management, according to an exemplary embodiment of the present disclosure. [Figure 9B] 1 illustrates a block diagram of an exemplary computing system for performing multi-device output management, according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0022] Reference numbers repeated among the drawings are intended to identify like features in the various embodiments.

[0023] In general, the present disclosure is directed to a multi-device framework for managing input acquisition and output generation for an environment including multiple computing devices. Specifically, the systems and methods disclosed herein can leverage environment-based prompt generation and generative model processing to generate and / or provide output that can be generated based on computing devices available in the environment. For example, the multi-device framework can include a computing system that facilitates acquiring input data and environmental data associated with multiple computing devices and generating one or more model-generated outputs configured to be provided to one or more specific computing devices of the multiple computing devices. The computing system can be utilized to leverage the diverse input and / or output capabilities of different computing devices in the environment, which can include automatically acquiring different types of input from and / or providing different types of output via different devices (e.g., voice commands can be received via a smartphone, while output can be provided via a television's visual display and / or sound system speakers).

[0024] Environmental data associated with multiple computing devices in an environment can be utilized for prompt generation, environmental understanding, and / or interface generation. The environmental data can be processed to generate prompts that can specify output generation of a generative model that is compatible with the user's environment and / or optimized for display and / or playback in the user's environment. Prompt generation can include retrieving a particular prompt template from a prompt library based on the computing devices associated with the environment and / or by processing input data and / or environmental data with a machine-learned model to generate the prompt. Additionally and / or alternatively, the environmental data can be processed to determine which computing devices associated with the environment to utilize for a particular type of input acquisition and / or a particular type of output playback. In some implementations, the environmental data can be processed to understand the computing devices in the environment and generate respective interfaces for the multiple computing devices based on the determined input and / or output capabilities of the computing devices associated with the environment.

[0025] The ambient multi-device framework can be used to provide an immersive virtual assistant that can receive input from and provide output on multiple different computing devices. The ambient multi-device framework can receive queries and / or prompts from a user and provide responses with relevant content types using computing devices in the environment that provide the relevant content types, which may be of higher quality than other devices in the environment. The determination of the computing devices can be based on a determined hierarchy, which can be determined based on device specification information (and / or device capability information).

[0026] Computing devices may be found throughout a user's environment, whether in the office, home, and / or other locations. Smart TVs, smart speakers, smart appliances, virtual assistant devices, tablets, smart wearables, smartphones, and / or other computing devices may be present throughout a user's environment. However, the capabilities of these devices may not be utilized based on a lack of interconnectivity and / or lack of cooperation between the devices. In particular, computing devices may be unable to determine when and / or how to interact to obtain input and / or provide context-specific output.

[0027] The ambient multi-device framework can include obtaining and / or generating information describing the computing devices of the environment, including information associated with input and / or output capabilities. The information can then be utilized in generating prompts for a generative model (e.g., a large language model) that can specify what data to obtain and / or generate for output and / or which computing devices to utilize for output associated with a user's query. In some implementations, the generative model can be fine-tuned for the multi-device framework (e.g., efficient fine-tuning of parameters and / or soft prompt adjustment), which can include adjusting the generative model to process prompts with environmental context and generate output based on both the query and environmental context.

[0028] In particular, different computing devices may have different components for capturing different forms of input data (e.g., text input, voice commands, gesture input, etc.) and / or for providing different types of output data to a user (e.g., visual display, audio playback, etc.) The ambient multi-device framework can leverage information of capabilities associated with computing devices in an environment to provide an immersive, multifaceted computing system that can obtain input and provide output across multiple different computing systems in the environment.

[0029] Smartphones, smart watches, smart speakers, smart TVs, smart assistant devices, and / or other computing devices are always around users. However, interconnectivity may be limited, and the highest quality input and / or output data may not be efficiently utilized. For example, a user may input a query via their smartphone, which may provide video and / or audio in response on the smartphone, but a high-quality smart TV and / or smart speaker may be easily accessible and nearby to the user. Whether the content is entertainment, educational, and / or for other purposes in nature, a multi-device framework can be leveraged to obtain additional forms of input and / or provide additional forms of output that may be of higher quality than a single-device system.

[0030] The systems and methods disclosed herein can be used to manage the acquisition of input from and / or the provision of output via multiple devices. In some implementations, input can be received from a smartphone and a smartwatch, and output responsive to the input can be provided via a smart television and a smart speaker. Additionally and / or alternatively, the systems and methods can determine that a particular computing device in the environment has the highest processing power, and can then utilize that particular computing device to perform model estimation. In some implementations, processing tasks can be divided among multiple computing devices in the environment and / or performed by a server computing system.

[0031] The systems and methods disclosed herein provide several technical effects and advantages. As an example, the systems and methods can be utilized to provide an interconnected multi-device ecosystem. In particular, the systems and methods disclosed herein can acquire input data from one or more devices within the ecosystem. The systems and methods can acquire environmental data describing devices in an environment and then generate prompts based on the input data and environmental data, which can be processed using a generative model to generate responses to the generated input data based on the devices in the environment. The responses can then be sent to one or more specific devices in the environment. For example, text input can be acquired via a tablet and output can be provided via a smart speaker.

[0032] Other exemplary technical effects and benefits relate to improved computational efficiency and improved functionality of computing systems. For example, a technical advantage of the systems and methods of the present disclosure is the ability to reduce the computational resources required to interact with multiple computing devices in an environment. In particular, a multi-device framework can centralize data processing and flow, thereby reducing instances of repetitive processing across devices within an ecosystem.

[0033] The systems and methods of the present disclosure provide several technical effects and advantages. As an example, the systems and methods may provide an interface generation system that can be utilized to generate environment- and device-specific interfaces that can be custom-generated based on input and output capabilities available within a multi-device ecosystem.

[0034] Another technical advantage of the systems and methods of the present disclosure is that interface generation can be utilized to provide an immersive multi-device environment. In particular, interfaces can be generated and provided to provide a user with the ability to obtain multiple different input types from multiple different devices in an interconnected system, and can provide multiple different output types via multiple different devices in the interconnected system.

[0035] Referring now to the drawings, exemplary embodiments of the present disclosure will be discussed in further detail.

[0036] 1 illustrates a block diagram of an exemplary generative model system 10 according to an exemplary embodiment of the present disclosure. In some implementations, the generative model system 10 is configured to receive and / or obtain a set of input data 12 describing a query and / or context, and to generate, determine, and / or provide a model generation output 20 describing a response to the query as a result of receiving the input data 12. Thus, in some implementations, the generative model system 10 may include a generative model 18 operable to process prompts 16 to generate a response to a query configured to be provided by a particular computing device in an environment.

[0037] In particular, the generative model system 10 can acquire input data 12 and environmental data 14. The input data 12 can describe one or more inputs. The input data 12 can describe queries (e.g., "Who's playing today?", "How do you make focaccia?", "What song played at the end of movies this year?"). The input data 12 can include directly entered input (e.g., a user typing on a graphical keyboard, capturing an image, selecting a graphical user interface element, providing a voice command, etc.) and / or contextual input (e.g., a user's search history, the time of day, a user's habit data, a user's browsing history, application activity, currently open applications, temperature, a user's location, and / or data acquired from other computing devices in the environment).

[0038] The environment data 14 can describe multiple computing devices in an environment. The environment can be a room, a set of rooms, a space containing multiple computing devices in proximity to a user, and / or in proximity to each other. The environment data 14 can include specification information for each of the multiple computing devices in the environment. Alternatively and / or additionally, the environment data 14 can be descriptive identification data of the multiple computing devices (e.g., device names, device serial numbers, device labels, registration data, etc.), functional data of the multiple computing devices (e.g., input capabilities for the multiple computing devices, processing capabilities for the multiple computing devices, and / or output capabilities of the multiple computing devices), and / or other environment data.

[0039] The generative model system 10 can process the input data 12 and the environmental data 14 to generate a prompt 16. The prompt 16 can include information associated with a query and multiple computing devices. The prompt 16 can include a query and an indicator of one or more candidate computing devices that will provide the output. The prompt 16 can be generated by querying a prompt library to determine a prompt template based on the environmental data 14. The prompt template can then be filled in based on the input data 12. Alternatively and / or additionally, the prompt 16 can be generated by processing the input data 12 and the environmental data 14 with a prompt generation model. The prompt generation model can include a machine-learned model, which can include a generative model (e.g., an autoregressive language model, a diffusion model, a visual language model, and / or other generative model).

[0040] A generative model 18 can then process the prompt 16 to generate a model-generated output 20. The generative model 18 can include a generative language model (e.g., a large language model, a visual language model, and / or other language model), an image generation model (e.g., a text-to-image generative model), and / or other generative models. The generative model 18 can include a transformer model, a convolutional model, a feedforward model, a recurrent model, a self-attention model, and / or other models.

[0041] The model-generated output 20 may include a predicted response to a query of the input data 12. The model-generated output 20 may be a novel generated output including a plurality of predicted features (e.g., predicted text characters, predicted audio signals, predicted pixels, etc.). The model-generated output 20 may be generated to be rendered by a particular computing device among a plurality of computing devices. For example, the model-generated output 20 may be generated to be of a particular size, quality, and / or content type based on a preferred computing device (having particular capabilities) determined based on prompts and / or processing of the generative model 18.

[0042] The model-generated output 20 may then be provided to a particular computing device within the environment. Providing the model-generated output 20 to a particular computing device may include transmitting data, which may include generating application programming interface calls using the generative model and executing API calls using one or more application programming interfaces. The particular computing device may then provide the model-generated output 20 to a user (e.g., via a visual display, speaker, and / or other output component of the particular computing device).

[0043] 2 illustrates a block diagram of an exemplary response generation system 200 according to an exemplary embodiment of the present disclosure. Response generation system 200 is similar to generative model system 10 of FIG. 1, except that response generation system 200 further includes a prompt library 224 and a search engine 226.

[0044] Specifically, the response generation system 200 can acquire input data 212 and environmental data 214. The input data 212 can include text data, image data, audio data, embedded data, signal data, search history data, browsing history data, application interaction data, latent coding data, multimodal data, global data, and / or other data. The input data 212 can be a description of one or more inputs, which can include acquiring input from multiple computing devices in the environment. The input data 212 can describe queries (e.g., "Who's playing today?", "How do you make focaccia?", "What song played at the end of this year's movie?"). The input data 212 can include directly entered input (e.g., a user typing on a graphical keyboard, capturing an image, selecting a graphical user interface element, providing a voice command, etc.) and / or contextual input (e.g., a user's search history, time of day, user habit data, user browsing history, application activity, currently open applications, temperature, the user's location, and / or data acquired from other computing devices in the environment). The input data 12 may include contextual data determined to be relevant to the query. Alternatively and / or additionally, the input data 12 may include predictive queries that may be generated based on predictions of what a user may be interested in based on one or more determined user contexts (e.g., determining an interest in a particular football team, utilizing a series of searches to determine a likelihood of an interest in horror movies, and / or determining that a particular product for purchase may be of interest based on similar products being purchased). In some implementations, the response generation system 200 may suggest content to a user by generating input data based on the personalized and / or contextualized signals.

[0045] The environment data 214 can describe the plurality of computing devices 222 in an environment. The environment can be a room, a set of rooms, a space containing a plurality of computing devices in proximity to a user, and / or in proximity to each other. The environment data 214 can include specification information for each of the plurality of computing devices 222 in the environment. Alternatively and / or additionally, the environment data 214 can be descriptive identification data of the plurality of computing devices 222 (e.g., device names, device serial numbers, device labels, registration data, etc.), functional data of the plurality of computing devices 222 (e.g., input capabilities for the plurality of computing devices 222, processing capabilities for the plurality of computing devices 222, and / or output capabilities of the plurality of computing devices 222), and / or other environment data. The plurality of computing devices 222 may include smartphones, smart wearables (e.g., smart watches and / or smart glasses), smart speakers, smart TVs, laptops, desktops, smart appliances (e.g., smart refrigerators, smart washers, smart dryers, and / or smart dispensers), virtual assistant devices (e.g., smart home panels, room-based assistants, and / or other assistant devices), tablets, and / or other computing devices. The plurality of computing devices 222 may include mobile computing devices and / or stationary computing devices. In some implementations, the plurality of computing devices 222 may include devices in proximity to a user, devices registered with a user's mobile device, devices registered in a user's profile, devices connected to a particular internet hub, devices registered with a particular assistant device, and / or other correlated devices.

[0046] The response generation system 200 can process the input data 212 and / or the environmental data 214 to generate a prompt 216. The prompt 216 can include information associated with a query and multiple computing devices. The prompt 216 can include an indicator of the query and one or more candidate computing devices that will provide the output. The prompt 216 can be generated by querying a prompt library 224 to determine a prompt template based on the environmental data 214. Querying the prompt library 224 can include determining a prompt template associated with an environmental configuration associated with the user's environment. Alternatively and / or additionally, querying the prompt library 224 can include generating an embedding based on the environmental data 214 and then determining a prompt template associated with the embedding based on a nearest neighbor search. The prompt template can then be filled in based on the input data 212. Alternatively and / or additionally, the prompt 216 can be generated by processing the input data 212 and the environmental data 214 with a prompt generation model. The prompt generation model may include a machine-learned model, which may include a generative model (e.g., an autoregressive language model, a diffusion model, a visual language model, and / or other generative model). The prompts 216 may include text data, image data, audio data, latent coding data, multimodal data, embedded data, and / or other data. The prompts 216 may include hard prompts (e.g., text strings and / or other data inputs) and / or soft prompts (e.g., learned sets of parameters). In some implementations, the prompts 216 may be zero-shot prompts and / or few-shot prompts. The prompts 216 may separate a user request into multiple tasks, which may be provided as multiple prompts.

[0047] The generative models 218 can then process the prompts 216 to generate model generation outputs 220. The generative models 218 can include generative language models (e.g., large language models, visual language models, and / or other language models), image generation models (e.g., text-to-image generative models), and / or other generative models. The generative models 218 can include transformer models, convolutional models, feedforward models, recurrent models, self-attention models, and / or other models.

[0048] In some implementations, the generative model 218 may be communicatively coupled to one or more processing engines (e.g., a search engine, a rendering engine, and / or other engines) and / or one or more machine-learned models (e.g., classification models, augmentation models, segmentation models, detection models, embedding models, and / or other models). For example, the generative model 218 may communicate with a search engine 226 to obtain search results based on the input data 212, the environmental data 214, and / or inferences. The search engine 226 may output search results, which may then be processed by the generative model 218 to generate a summary of the search results and / or determine a response to a query based on information from the search results, which may then be utilized to generate the model generation output 220. The model generation output 220 may include text data, image data, audio data, latent coding data, multimodal data, and / or other data.

[0049] The model generated output 220 may include a predicted response to a query of the input data 212. The model generated output 220 may be a novel generated output including a plurality of predicted features (e.g., predicted text characters, predicted audio signals, predicted pixels, etc.). The model generated output 220 may be generated to be provided by a particular computing device of a plurality of computing devices. For example, the model generated output 220 may be generated to be of a particular size, quality, and / or content type based on a preferred computing device (having particular capabilities) determined based on prompts and / or processing of the generative model 218.

[0050] The generative model 218 may process the prompt 216 to generate output device instructions 230. The output device instructions 230 may include instructions for sending the model-generated output 220 to a particular computing device 228 of the plurality of computing devices 222. The output device instructions 230 may include application programming interface calls, registration files, and / or instructions for the particular computing device 228.

[0051] The model-generated output 220 may then be provided to a particular computing device in the environment based on and / or in conjunction with the output device instructions 230. Providing the model-generated output 220 to a particular computing device may include transmitting data, which may include generating application programming interface calls using the generative model and executing API calls using one or more application programming interfaces. The particular computing device may then provide the model-generated output 220 to a user (e.g., via a visual display, speaker, and / or other output component of the particular computing device).

[0052] 3 shows a flowchart diagram of an exemplary method for functioning in accordance with an exemplary embodiment of the present disclosure. While FIG. 3 shows steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly shown order or arrangement. Various steps of method 300 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0053] At 302, the computing system can obtain input data. The input data can include a query associated with a particular user. The input data can include text data, audio data, image data, gesture data, latent coding data, multimodal data, and / or other data. The input data can be obtained from and / or generated at one or more computing devices in the environment. The input data can include prompts for a generative model. In some implementations, the input data can be written questions associated with one or more topics. The questions can be associated with content viewed on one or more computing devices and / or features in the environment.

[0054] At 304, the computing system may acquire environmental data. The environmental data may describe multiple computing devices in the user's environment. The multiple computing devices may be associated with multiple different output components. In some implementations, the multiple computing devices may be located proximate to each other computing device within the multiple computing devices. The multiple computing devices may be communicatively connected across a network. In some implementations, the environmental data may describe specifications for the multiple computing devices. For example, the environmental data may indicate components of the computing devices, which may include input components and / or output components. Additionally and / or alternatively, the environmental data may include information describing input and / or output capabilities of multiple different computing devices.

[0055] At 306, the computing system can generate a prompt based on the input data and the environmental data. The prompt can include data describing the query and device information associated with at least a subset of the plurality of computing devices. The prompt can include text data, image data, audio data, embedded data, statistical data, graphical representation data, latent coding data, semantic data, multimodal data, and / or other data. The prompt can be generated by a machine-learned model based on deterministic functions, heuristics, and / or hybrid techniques. The prompt can be generated by querying a prompt template library based on the environmental data and populating a prompt template selected based on the input data.

[0056] In some implementations, generating a prompt based on the input data and the environmental data may include determining an environment-specific device configuration based on the environmental data, retrieving a prompt template from a prompt library based on the environment-specific device configuration, and expanding the prompt template based on the input data to generate the prompt. The environment-specific device configuration may describe a respective output type and a respective output quality of a plurality of computing devices. The prompt library may include a plurality of different prompt templates associated with a plurality of different device configurations.

[0057] At 308, the computing system can process the prompt with the generative model to generate a model-generated output. The model-generated output can include a response to a query. In some implementations, the model-generated output may be generated to be provided to a particular computing device of a plurality of computing devices. The model-generated output can be generated to be provided to a particular output component of a plurality of different output components. The generative model can generate output device instructions. The output device instructions can describe a particular computing device of the plurality of computing devices that provides the model-generated output. The particular computing device can be associated with a particular output component. For example, the model-generated output can be generated to be provided to a user via a respective output component of a particular computing device (e.g., speakers of a smart surround system and / or a display screen of a smart television).

[0058] In some implementations, the Generative Model may be communicatively coupled to a search engine via an application programming interface. Processing the prompt with the Generative Model to generate the model-generated output may include generating application programming interface calls based on the prompt, determining a plurality of search results at the search engine based on the application programming interface calls, and processing the plurality of search results with the Generative Model to generate the model-generated output.

[0059] At 310, the computing system can send the model-generated output to a specific computing device. The model-generated output can be sent to a specific computing device based on output device instructions. Sending the model-generated output to a specific computing device can include executing an application programming interface call generated by the generative model. The model-generated output can be sent to a virtual assistant device in the user environment, which can then control the specific computing device to provide the model-generated output to a specific user.

[0060] In some implementations, the plurality of computing devices can be connected via a cloud computing system. Each of the plurality of computing devices can be registered with a platform of the cloud computing system. Environmental data can be acquired through the cloud computing system. Model-generated output can be transmitted via the cloud computing system. Alternatively and / or additionally, the plurality of computing devices can be communicatively connected across a network. A particular computing device of the plurality of computing devices can facilitate acquisition of input data and transmission of model-generated output.

[0061] In some embodiments, processing the prompt with the generative model to generate a model-generated output may include generating a plurality of model outputs. The plurality of model outputs may include a plurality of candidate responses. Additionally and / or alternatively, transmitting the model-generated output to a particular computing device may include transmitting a first model output of the plurality of model outputs to a first computing device of the plurality of computing devices and transmitting a second model output of the plurality of model outputs to a second computing device of the plurality of computing devices. The first model output may include visual data for display via a visual display. The second model output may include audio data for playback via a speaker component. In some embodiments, the first computing device may include a smart television. Additionally and / or alternatively, the second computing device may include a smart speaker.

[0062] 4 depicts a block diagram of an exemplary multi-device management system 400, in accordance with an exemplary embodiment of the present disclosure. Specifically, the multi-device management system 400 may include a computing system including a plurality of user computing devices and a server computing system 420. The plurality of user computing devices and the server computing system 420 may be communicatively connected via a network 410.

[0063] The plurality of user computing devices may include a first computing device 402, a second computing device 404, a third computing device 406, and / or an nth computing device 408. The plurality of user computing devices may include multiple different computing devices that may have different input, processing, and / or output capabilities. For example, the first computing device 402 may include a smartphone equipped with an image sensor, an audio sensor, a touch sensor, a motion sensor, a speaker, a visual display, a haptic component, and / or a light. The second computing device 404 may include a smart wearable (e.g., a smart watch) that may include a biometric sensor, a motion sensor, a touch sensor, a visual display, and / or a haptic component. The third computing device 406 may include a smart speaker that may include a high-quality speaker and / or a Bluetooth transmitter. The nth computing device 408 may include a smart television, which may include an infrared sensor receiver, a transmitter-receiver, a speaker, a visual display, and / or multiple input ports. Multiple user computing devices can connect to the network 410 via Ethernet, WiFi, and / or Bluetooth connections to companion devices.

[0064] Multiple user computing devices may be associated with the environment based on device registration, location, and / or proximity. Environmental data may be generated based on the multiple user computing devices, which may include acquiring signals from the multiple user computing devices.

[0065] The server computing system 420 can obtain input data and / or environmental data from multiple user computing devices via the network 410. The server computing system 420 can include multiple processing services for processing the input data and / or environmental data. For example, the server computing system 420 can include one or more generative models 422, one or more search engines 424, one or more prompt generation models 426, one or more interface models 428, and / or one or more other models. The one or more generative models 422 can be configured, trained, and / or tuned to process prompts associated with the input data and / or environmental data to generate model-generated output responsive to the input data and can be configured to be of a specific content type based on the environmental data. The one or more search engines 424 can be communicatively connected to obtain search results that can be utilized to understand the input data and / or environmental data and / or respond to the input data and / or environmental data. The one or more prompt generation models 426 can be configured, trained, and / or tuned to process the input data and / or environmental data to generate prompts for the one or more generative models 422. One or more interface models 428 may be configured, trained, and / or tuned to process environmental data associated with a plurality of user computing devices and generate a plurality of respective interfaces for the plurality of user computing devices, each of which may be generated based on the input and / or output capabilities of the devices in the environment.

[0066] 5 illustrates a block diagram of an exemplary environment personalization system 500, in accordance with an exemplary embodiment of the present disclosure. Specifically, the environment personalization system 500 can leverage information from multiple applications and / or platforms to provide a personalized response and / or a personalized experience. For example, data from a search assistant 502, a document assistant 504, an operating system assistant 506, a video player assistant 508, a browser assistant 510, and / or a chat interface assistant 512 can determine a user context and / or generate queries and / or suggestions.

[0067] The search assistant 502 can be associated with a search application (and / or platform) and can retrieve current queries, session states, search history, trend data, and / or other search data.

[0068] The document assistant 504 can be associated with one or more document applications and can retrieve data associated with the current document being viewed and / or edited, driving data (e.g., stored document information), sharing permission data, and / or other document data.

[0069] The operating system assistant 506 may be associated with the operating systems of one or more computing devices. The operating system assistant 506 may obtain data associated with content currently being provided for display, application data (e.g., app deep links, application programming interfaces (APIs), and / or usage data), and / or other operational data.

[0070] The video player assistant 508 may be associated with a video player application (and / or platform) and may be utilized to obtain video data for the currently displayed video, video storage, viewing history, followed media providers, subscriptions, comment history, and / or other video player data.

[0071] A browser assistant 510 may be associated with a browser application and may retrieve data associated with the current page, bookmark data, tab data, browsing history data, and / or other browser data.

[0072] The chat interface assistant 512 can be associated with one or more chat robots. The chat interface assistant can obtain session history data, response history, input history, topics, links, and / or other chat robot data.

[0073] The multi-device management system 514 can obtain data from multiple applications and / or platforms and provide the data to one or more other systems, which may include a core model 516, a ground services model 518, and / or an individualized model 520.

[0074] For example, the core model 516 can be utilized for summarization, planning and reasoning, and / or function calls. The ground services model 518 can be utilized for utilizing tool libraries, utilizing external connector application programming interfaces (APIs), accessing search results, and / or accessing context and / or memory. The core model 516 and / or ground services model 518 can interact with one or more other models and / or services via one or more cloud APIs. The output of the core model 516 and / or ground services model 518 can be provided back to the multi-device management system 514 and then to the personalization model 520.

[0075] The personalization model 520 may process the data to generate personalized output, which may include predictive prompts (and / or suggestive prompts), which may be enhanced prompt responses based on user preferences, interactions, and / or user data associated with the device.

[0076] FIG. 6A illustrates a diagram of an exemplary interface, according to an exemplary embodiment of the present disclosure. Specifically, the systems and methods disclosed herein can be utilized to generate and process environmental data associated with computing devices to generate multiple respective interfaces for multiple computing devices in an environment. The multiple respective interfaces can take into account multiple other computing devices and may obscure and / or blend with existing interfaces. FIG. 6A illustrates three exemplary interfaces that can be associated with different devices in an environment and / or associated with different environments.

[0077] For example, a first interface 602 may be associated with a first computing device (e.g., a mobile device) in the environment, a second interface 604 may be associated with a second computing device (e.g., a smart watch) in the environment, and a third interface 606 may be associated with a third computing device (e.g., a smart refrigerator) in the environment. Alternatively and / or additionally, the same computing device may have different interfaces based on being in different environments and / or based on different user contexts.

[0078] 6B shows a diagram of an exemplary image capture entry point, according to an exemplary embodiment of the present disclosure. In particular, one entry point for a search and / or assistance interface that leverages multi-device input and / or output may be provided via a contextually provided user interface element.

[0079] For example, a user may capture an image 610 (e.g., capturing an image of a refrigerator). The image may be processed to determine (and / or identify) the object in the image (e.g., classify the object as a brand X model Y refrigerator). Selectable user interface elements may then be provided in the viewfinder 612. The image and / or identification may then be processed to generate a response, which may include search results 614 associated with objects similar to the identified object. Other options may also be provided for interacting with the response. Other options may include an augmented reality experience, which may include rendering the object into the image of the user's home 616.

[0080] 6C shows a diagram of an exemplary smart TV entry point, in accordance with an exemplary embodiment of the present disclosure. In particular, one entry point can include providing suggested search entry point indicators to one device, which can interact with other devices.

[0081] For example, content provided for display on a first computing device 620 (e.g., a smart TV) can be determined to include features that a user may find interesting in a search. Accordingly, suggested entry point user interface elements may be rendered over the content. The user can then select a selectable user interface element 622 on a second computing device (e.g., a mobile computing device) to view a model-generated output 624. The model-generated output can be provided via the first computing device 620, the second computing device, and / or a third computing device. The model-generated output 624 can include search results, a generated model output, one or more renderings, one or more suggestions, and / or other options.

[0082] FIG. 6D illustrates a diagram of an exemplary calendar interface, according to an exemplary embodiment of the present disclosure. In particular, FIG. 6D illustrates an exemplary calendar interface including a search window that can be defined based on calendar data. For example, the search window can include a text entry field and two suggested actions. The suggested actions can include finding a room for a meeting (630) and / or suggesting time to block off for focus time (632). The suggestions can be generated based on information from multiple applications and / or multiple computing devices. The search window can be provided on the device on which the calendar application is open and / or can be provided on other devices.

[0083] FIG. 6E illustrates a diagram of an exemplary videoconferencing interface 634, according to an exemplary embodiment of the present disclosure. Specifically, FIG. 6E illustrates an exemplary videoconferencing interface 634 including a search window 636 that can be defined based on conference data. For example, the search window 636 can include a text entry field and one or more suggested actions. The suggested actions can include taking notes about the conference (e.g., writing in the conference and / or opening a note-taking application), setting a reminder, rescheduling the conference, and / or retrieving notes from a similar conference. The suggestions can be generated based on information from multiple applications and / or multiple computing devices. The search window 636 can be provided on the device on which the videoconferencing application is open and / or can be provided on other devices.

[0084] FIG. 6F illustrates a diagram of an exemplary email interface, according to an exemplary embodiment of the present disclosure. Specifically, FIG. 6F illustrates an exemplary email interface 638 including a search window 640 that can be defined based on email data and / or conference data. For example, the search window 640 can include a text entry field and one or more suggested actions. The suggested actions can include summarizing meeting notes and / or minutes, scheduling another meeting, obtaining information about a subject from the meeting, drafting a follow-up email based on context associated with the meeting notes, and / or drafting an email based on email data and / or conference data. The suggestions can be generated based on information from multiple applications and / or multiple computing devices. The search window 640 can be provided on the device on which the email application is open and / or on other devices.

[0085] FIG. 6G shows a diagram of an exemplary video player interface according to an exemplary embodiment of the present disclosure. In particular, FIG. 6G shows an exemplary video player interface 650 including a search window 654 that can be defined based on video data and / or viewing history data. For example, the search window 654 can include a text entry field and one or more suggested actions. The suggested actions can include obtaining a product listing for similar products depicted in the displayed video 652, summarizing the video, obtaining an entity label for the displayed video, finding similar videos, and / or obtaining additional information associated with the displayed video 652. For example, a user may request more information about a dress depicted in the displayed video 652. One or more frames can be segmented and retrieved from the video, which may include frame cropping. Alternatively and / or additionally, entity labels associated with the depicted frames can be obtained and retrieved. Search results can be provided for display in the search window 654 overlaid over the displayed video. The search results can include product listings, web links, and / or other data. Suggestions may be generated based on information from multiple applications and / or multiple computing devices. Search window 654 may be provided on the device on which the video player application is open and / or may be provided on other devices.

[0086] 7 shows a flowchart diagram of an exemplary method for functioning in accordance with an exemplary embodiment of the present disclosure. While FIG. 7 shows steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly shown order or arrangement. Various steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0087] At 702, the computing system can acquire input data. The input data can include a query associated with a particular user. The query can include a question associated with one or more topics, and the query can request a response to the question. The input data can include a voice command acquired via a microphone, a text string entered via a graphical keyboard, and / or a gesture acquired via a camera, an inertial measurement unit, and / or a touch sensor. In some implementations, the input data can include data describing input acquired from multiple computing devices in the environment (e.g., a voice command acquired via a microphone on a smartphone, a gesture via a touch sensor on a smartwatch, image data from a smart refrigerator, and / or a viewing history acquired from a smart TV).

[0088] At 704, the computing system can acquire environmental data. The environmental data can describe a plurality of computing devices in the user's environment. The plurality of computing devices can be associated with a plurality of different output components. The plurality of different output components can be associated with a plurality of respective output capabilities associated with the plurality of computing devices. In some implementations, each of the plurality of respective output capabilities can describe an output type and output quality available via the respective computing device. The plurality of output components can include a speaker associated with the first device and a visual display associated with the second device. The environmental data can include registration data associated with computing devices, WiFi routers, virtual assistant devices, user computing devices, and / or user profiles registered in the environment. In some implementations, the environmental data can include an output hierarchy for a plurality of candidate output types, and the output hierarchy can include a hierarchical representation of the performance capabilities of the plurality of computing devices for a plurality of different output types (e.g., visual display, audio output, haptic feedback, etc.).

[0089] At 706, the computing system can generate a prompt based on the input data and the environmental data. The prompt can include data describing the query and device information associated with at least a subset of the plurality of computing devices. The prompt can be generated by processing the input data and the environmental data with a prompt generation model. The prompt generation model can include a language model (e.g., a generative language model (e.g., a large language model)). The prompt generation model can be trained and / or tuned to generate the prompt based on understanding the intent of the query and determining, based on the environmental data, an output type associated with the intent and / or available output types. The prompt generation model can generate a prompt embedding that specifies the output generation of the generative model.

[0090] In some implementations, the computing system can determine that a particular output component is associated with the intent of the query, and a prompt can be generated based on the particular output component associated with the intent of the query (e.g., a request for a song can be associated with a speaker output component, and a request to play a video can be associated with a visual display on a television).

[0091] Additionally and / or alternatively, the computing system may determine an output hierarchy based on specification information for a plurality of different output components based on the environmental data. A prompt may be generated based on the output hierarchy and the query. For example, the prompt may include text and / or embeddings that specify output generation based on the output capabilities of computing devices in the environment.

[0092] At 708, the computing system can process the prompt with the generative model to generate a model-generated output and output device instructions. The model-generated output can include a response to the query. In some implementations, the model-generated output can be generated to be provided to a particular output component of a plurality of different output components. The output device instructions can describe a particular computing device of a plurality of computing devices that provides the model-generated output. The particular computing device can be associated with a particular output component. In some implementations, the output device instructions can include an application programming interface call to send the model-generated output to the particular computing device.

[0093] At 710, the computing system can transmit the model-generated output to a particular computing device based on the output device instructions. In some implementations, the transmission can be performed via signal transmission over a network. Alternatively and / or additionally, a notification can be provided to the input computing device indicating that an output configured for another device is available, and a user can then interact with the notification before the model-generated output is transmitted to a particular computing device. In some implementations, a generative model can generate multiple model-generated outputs, and the computing system can transmit the multiple model-generated outputs to multiple different computing devices in the environment (e.g., a slideshow can be transmitted to a smart TV for playback, a text document can be transmitted to an e-reader or personal computing device (e.g., a smartphone or tablet), an audio file can be transmitted to a smart speaker, and / or a scheduled set of color and brightness instructions can be transmitted to an RGB smart light setup).

[0094] 8 shows a flowchart of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. While FIG. 8 shows steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly shown order or arrangement. Various steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0095] At 802, the computing system can obtain environmental data. The environmental data can describe multiple computing devices in an environment associated with a particular user. The environmental data can describe specification information for the multiple computing devices. In some implementations, the multiple computing devices can be determined based on registration of the device with a particular network, registration of the device with a particular user device, registration of the device with a particular user profile, proximity to a user device, and / or signal exchange between the device and the user device. The user device may be a smartphone, tablet, smartwatch, smartglasses, and / or other computing device.

[0096] At 804, the computing system may process the environmental data to determine a plurality of respective input capabilities and a plurality of respective output capabilities associated with the plurality of computing devices. The plurality of respective input capabilities may be associated with candidate input types associated with the plurality of computing devices. In some implementations, the plurality of respective output capabilities may be associated with candidate input types associated with the plurality of computing devices. The plurality of input capabilities and / or the plurality of output capabilities may be determined based on specification information (and / or component information) associated with the plurality of computing devices. In some implementations, the plurality of input capabilities and / or the plurality of output capabilities may be determined using machine-learned models, based on one or more searches, based on heuristics, and / or based on one or more other decision techniques. The plurality of input capabilities may describe types of inputs available along with a range of qualities associated with obtaining and / or generating the input types (e.g., microphone decibel range and / or camera resolution). The plurality of output capabilities may describe types of outputs available along with a range of qualities associated with providing output of that output type (e.g., audio quality (e.g., volume range, frequency range, etc.)).

[0097] At 806, the computing system may generate a plurality of respective interfaces for the plurality of computing devices based on the plurality of respective input capabilities and the plurality of respective output capabilities. The plurality of respective interfaces may be specialized for the plurality of computing devices based on the plurality of respective input capabilities and the plurality of respective output capabilities. In some implementations, the plurality of respective interfaces may include a plurality of device indicators that indicate a plurality of computing devices in an environment associated with a particular user. The plurality of computing devices may be configured as a user-specific device ecosystem communicatively connected to receive input and provide output. Each of the plurality of respective interfaces may be configured to receive a particular input type and provide a particular output type based on the respective input capabilities and respective output capabilities for a particular computing device of the plurality of computing devices.

[0098] At 808, the computing system can provide the plurality of respective interfaces to the plurality of computing devices. The plurality of respective interfaces can be transmitted to the plurality of computing devices in response to generation of the interfaces and / or in response to a user interacting with a particular computing device. The plurality of respective interfaces can be stored by a server computing system and provided to the computing devices upon use. Alternatively and / or additionally, the respective interfaces for each computing device can be downloaded locally and provided during offline and online states.

[0099] In some implementations, the computing system may obtain user input via a first interface of a first computing device of the plurality of computing devices, process the user input with a search engine to determine a plurality of search results, process the plurality of search results with a generative model to generate model output, and provide the model output for display via a second interface of a second computing device of the plurality of computing devices.

[0100] Increasingly, users will have composite devices that may or may not share operating systems, apps, and / or platforms. Foundational models (e.g., large foundational models that may include generative models) can be used as the primary technology by which users interact with their devices. Systems can offer a wide range of new experiences enabled by these foundational models that can specify both input and output across all devices in a unified way.

[0101] A multi-device framework utilizing the basic model can be utilized for a number of different companion tasks, which may include defining output, generating interfaces, facilitating input acquisition, and / or other tasks.

[0102] For example, a multi-device framework can be utilized for an ambient agent companion with fluid interaction mechanisms that are device and surface dependent. A user may own several devices, potentially using a disjointed ecosystem within which an ambient LLM companion exists. An agent (e.g., a computing system including a multi-device framework) may include foundational models (e.g., generative and / or prompt-generative foundational models) that adapt input / output interaction models based on device specifics. For example, a speaker may have a voice-only interface as the primary interaction model, a device with a non-touch screen may have a hybrid text / voice interface, and / or phones and / or smart wearables (e.g., smartwatches, smart jackets, and / or smart glasses) may have user interfaces that are ad hoc generated based on their form factor.

[0103] The interface may be generated to blend with existing surfaces (e.g., the user interface may be minimized to one or more app elements (e.g., an agent interface may be transformed into text and / or animation within a particular app interface)). The generated interface may have multi-device / surface awareness elements that can highlight which devices the ambient companion is actively listening to and / or monitoring based on the user's proximity to the devices. When interactions with the companion on one device require context from other devices and / or actions to be performed on other devices, the user interface may include user interface elements that highlight that these other devices are being utilized and / or considered.

[0104] Additionally and / or alternatively, systems and methods may specify the output of the ambient companion model with respect to the device's form factor and characteristics. For example, a user of an environment with a smart TV and a smart home device (e.g., a virtual assistant device) may provide input (e.g., via voice command) that poses a query (e.g., a query of the type "What movies were nominated for an Oscar for Best Original Score"). The LLM-enabled search engine / assistant / companion can obtain additional input to the query, which may describe the device in detail (e.g., an input specifying the environment, including the screen and speakers, along with their specific parameters). A higher-quality smart speaker may be provided as part of the LLM prompt. Based on the prompt, a generative model (e.g., an LLM) may generate a response targeted to the smart speaker (e.g., "List movies and play samples of their award-winning soundtracks on the speakers"). Meanwhile, other users issuing the same query may be in an environment with the same smart TV but with lower quality speakers, and thus may receive a response in the form of a visual display via the smart TV (e.g., the response may include, "We'll play part of the movie trailer on your TV, and you'll hear part of its award-winning soundtrack later.").

[0105] In some implementations, systems and methods can be utilized to obtain and / or determine the availability and context of nearby devices and then modify actions performed on behalf of the user. For example, a user may be in an environment with multiple Internet of Things (IoT) devices, mobile computing devices (e.g., smartphones, tablets, smartwatches, etc.), automobiles, and / or laptops equipped with ambient companions. The user can then ask the companion (via input sensors on the computing devices) about driving to a nearby recreation area. The ambient companion can then obtain and / or determine input in a manner that can take into account all devices in the environment. Determining input and / or providing a response can include responding to queries across multiple devices. For example, the companion can identify typical routes and / or destinations and generate specific visualizations that can be provided within a map application on one or more specific computing devices. A companion variant of a car agent can perform local inference and determine that the car needs to be recharged to reach one or more of the destinations. The maps app used by the main agent companion handling may process the query and output relevant charging stations that are highlighted on the depicted route.

[0106] In some implementations, the systems and methods disclosed herein may be utilized to generate and / or provide interfaces across different computing devices that may be interconnected and / or have similar style, layout, and / or semantics, regardless of the manufacturer and / or operating system of the computing device. For example, the systems and methods may generate native interfaces for different computing devices from different manufacturers and / or interfaces that obscure differences between operating systems.

[0107] The systems and methods can include fluid cross-device representations. Cross-device representations can be provided based on the output of foundational models that can specify input / output behavior on composite devices and surfaces (e.g., apps). Cross-device representations can be implemented via (1) a centralized foundational model running in the cloud to which device nodes report directly; (2) a hybrid model in which the centralized foundational model functions with a distributed, local, larger model that can perform inference using device-specific context and coordinate the resolution of higher-level tasks in the main model; and / or (3) a distributed architecture in which individual devices have companion versions, all of which are grouped into the same physical space or the same logical unit by the user who owns them. A certain level of inference may be required (e.g., >10B params) so that the models can interact.

[0108] The hybrid approach may be implemented through several different configurations, which may include adaptations to centralized and / or distributed architectures.

[0109] For example, a user may have several registered, interconnected devices. The characteristics of these devices may be known and / or determined. The degree of interoperability may vary and may include an API that allows full control of the device when plugged in. An API may be used by a device's speaker to output audio. An API may be used to render elements directly in the operating system (and / or for complex functionality within an app, etc.).

[0110] The APIs may be exposed in various configurations based on whether a local foundational model exists. If no foundational model exists, the raw functionality may be described in some accessible document library, and / or alternatively, if a local model exists and is provided, a natural language interface may be available. The device may also provide the ability to interface with other external systems (e.g., sensors that can be used to read the ambient temperature, operate window blinds, and / or autonomously navigate through the home to perform specific tasks). Examples and information about the APIs may be provided in the foundational model that powers the ambient companion agent in ways that can be used to prescribe response generation and task resolution on behalf of the user.

[0111] Prompt generation can include obtaining and / or generating zero or a few shots of prompts for the underlying model to process to understand how to use the device's API. If that is not enough, devices can have a small data set associated with them (e.g., about 1000 examples that can be used to adjust prompts and / or weight the underlying model to understand how to operate the device). These examples can include task->decomposition using the API.

[0112] Once surfaced, devices can be located together through a network, where each device is a node. The edges of this network may be persistent (e.g., always on) and / or may have some weight associated with them that quantifies the relatedness of two devices at a given time. For example, two devices may be very close to each other, and this information can be quantified by the weight between the two devices having a smaller / larger numerical value. Also, two devices may share certain context (e.g., if they run and / or display the same app together (e.g., state can also be dynamically encoded on the graph edges via message passing)).

[0113] Available devices can be queried at inference time. The queries can be provided as a list with metadata and / or represented via a graph network, which can also be provided to the ambient companion base model at inference time. Prior to model inference, the base model may be fine-tuned to work with the device network topology and / or characteristics. The graph network can be the modality in which the base model can operate. In some implementations, the network can be directly serialized and passed to a generative model (such as an LLM) as part of a prompt. If insufficient, the large base model powering the ambient companion can be fine-tuned using examples of the multi-device graph network (e.g., even 1000 or more examples, each with several task->step-by-step decompositions on how to use composite devices in solving a task better).

[0114] In some implementations, the systems and methods can include prompting the LLM (or other generative model) with an example (e.g., "[Device Context][Response][Metadata: This response is appropriate for a smart speaker at breakfast time] and / or "[Device Context][Response][Metadata: The response is visualized by rendering a UI with three checkboxes for each response on the phone]). In a device decomposition example, the prompt might include "[Device Context][Response][Metadata: This response should be passed to the smart speaker, and a small notification with a summary should be displayed on the phone because the user may not be near the speaker." The user can then face one of their devices and decide to interface with an ambient companion.

[0115] Input may depend particularly on form factor details. For example, when the user pulls out the phone, the companion may activate in a voice-only manner. When the phone is locked or when the user unlocks the phone, the companion may render itself with a UI that allows keyboard input. Modalities that depend on the state of the device may be defined relative to the user's context in relation to nearby devices. For example, if the user's watch is available, voice input may be activated there instead.

[0116] Issued queries can be resolved with the cooperation of all devices. For example, if a user says something along the lines of "I need to go to bed now. I wonder what time I have to get up," the system can trigger the underlying model to act on the query and / or context using all available context from the device. Based on the determined action, the response can first process surfaces that have access to work and / or calendar context and respond to the speaker with some actual suggested time.

[0117] Processing may continue, for example, in a smart home, a lighting system may determine the user's intent to go to bed and activate its local inference companion to slowly adjust the lights to the user's defined bedtime routine. Information may be conveyed visually and rendered as UI elements on the user's phone and / or watch to make them aware of the decision.

[0118] There may be a proactive component to an ambient LLM companion. For example, some part of the device may be interacted with to determine if user input is needed. For example, a companion-enabled device that has sensors and can operate the device autonomously may determine, based on the user's context, that input should be received and / or that a prompt should be generated.

[0119] In some implementations, a car companion may be built internally that may be prompted, configured, and / or trained to schedule the user to check the temperature, rainfall, etc., 15 minutes before the user's estimated departure time, and in response to the prompt, configuration, and / or training, the car companion may remind the user to perform an action (e.g., pick up X) and / or remind them to bring their raincoat.

[0120] The determined context may be received by an ambient companion, and a surface through which the context is communicated to the user may be determined there. For example, the companion may decide to use a smart speaker (e.g., a voice notification, "Even if it's raining, you can have your car pick you up instead of walking to the parking lot"). Alternatively and / or additionally, the system may render the same information and / or actions on the user's watch based on the device usage context.

[0121] Across all devices, there may be uniform behavior for input and / or output that may enable branding elements. Uniform behavior may be possible through device-specific prompts and several example shots available. A device may come with 10 seconds of audio for five voice samples. A device may also come with five UI examples of how the assistant may be rendered on that particular device. An ambient foundation model can be specified to generate responses for those examples and produce a uniform output of an interface that the user may be familiar with from that manufacturer.

[0122] A hybrid framework approach may include devices that have some degree of autonomy but may generally interact with each other through a centralized companion. A variation may include cases where devices do not have the ability to execute local underlying models and decisions may be made entirely by a centralized companion. A variation may include continuous streaming of information. Another variation may include cases where devices have full autonomy and there is no central companion. In an approach without a central companion, the network backbone may become more important and may rely on simpler signals to guide which devices are interconnected at a particular time to resolve a given query.

[0123] The systems and methods disclosed herein can unlock new value for users by using generative artificial intelligence models to connect fragmented systems and provide a single service level companion.

[0124] The functionality of a conversational interface (chat robot) can extend beyond Q&A associated with a single device and / or application. For example, a user might be viewing a shopping blog and say, "Show me reviews of the products recommended in this article / video." The system may need to understand what the products are to retrieve the reviews in a shopping graph and summarize them for easier comparison. In another example, a user might be viewing a recipe and say, "Add the ingredients to my shopping app cart." The system may then need to extract the ingredients and call a shopping app API. In another example, a user might be viewing a travel vlog and say, "Show me on a map the places mentioned in this article / video." The system may need to understand the places mentioned and call an API to extract addresses and create pins in a custom maps app. In another example, while viewing a documents app, a user might be writing an essay about Abraham Lincoln by asking, "Write a 1,000-word essay about Lincoln that focuses on the Civil War and explains your difficulties in a way that a fifth-grader can understand." Thus, the system may need to understand key facts about Lincoln, a list of Lincoln's works, the contents of those works, historical treatises, and summarize all of that information. In another example, while viewing a maps app, a user may look at a map at a zoom level and say, "Show me videos about things to do around here." The system may need to understand the important things to do at that location, the places mentioned in the video, and retrieve the appropriate video(s). In another example, while viewing content on a phone, a user may be viewing a screen from a trail hiking app and say, "Show me restaurants near trailheads." Thus, the system may need to understand the location of the trailhead being viewed and retrieve local results for restaurants near that location.

[0125] The underlying dependencies are that chat robots in these products, regardless of how they are integrated, may rely on common elements such as inference, reasoning, retrieval, function calls, user state, and personal preferences, many of which originate from search and / or knowledge graph services.

[0126] The systems and methods disclosed herein may utilize a chat robot large language model (LLM), a search understanding large language model and search engine, a cloud service large language model, and / or one or more other model-enabled systems. While the architectures for the chat robot large language model (LLM), the search understanding large language model, and the search engine may appear identical, they may be implemented independently with differences in each block, such as different RLHF (reinforcement learning from human feedback) training, different planning algorithms, different sets of third-party plug-ins, different interfaces for searching the backend, and / or other differences. The different pipelines and / or systems may exchange information for inference, understanding, and / or context determination.

[0127] Separate architectures may allow products to be developed independently and iteratively. However, while adherence to optimal processes may ensure the success of individual products, failure to adopt a holistic, service-oriented approach may ultimately lead to disconnects between products.

[0128] FIG. 5 shows an example architecture for building a layer cake of chat robot functionality that can be implemented and a set of instantiations of a chat robot service that imparts appropriate context and personality that can be associated with the product in which the service is instantiated.

[0129] For example, there may be a generic chat robot service built with a capable LLM, fine-tuned for common tasks such as command following, with access to a search backend, access to 1p / 3p APIs, function calling capabilities, etc. Some examples of services that can be leveraged to obtain input and / or determine context include a search platform (e.g., accessing user search history to obtain concise factual data), a document application (e.g., accessing document files, emails, current email / file being displayed, favorites, etc. to obtain redundant and original data), and / or a browser application (e.g., accessing current tab, other tabs, bookmarks, history, etc. to obtain concise and factual data).

[0130] In some implementations, the systems and methods can include using use-case specific examples and content to fine-tune the underlying model differently for different products.

[0131] The systems and methods disclosed herein can include companion models (e.g., a prompt generation model and / or a foundation model that can include a generative model).

[0132] By integrating AI and LLM functionality as standalone entities within the user interface (UI), separate from all other user activity, the systems and methods can ensure that they are permanently available for immediate access whenever the user needs them. This integration allows the AI and LLM to remain open and seamlessly connected with the user's actions as they transition between various applications and tasks.

[0133] This configuration can be useful for users who want to quickly access information or perform a task without leaving the app they are currently using. Additionally, a persistent system can be useful for users who want to use the companion with other products. For example, a user may use the companion to search for information on a topic while also using a maps application to navigate.

[0134] The system and method may provide two additional benefits: First, the system eliminates the need to build on existing solutions with potentially incompatible architectures, thereby avoiding potential integration difficulties; Second, by not changing existing solutions that users are already familiar with, the disruption caused by introducing significant, potentially suboptimal changes may be minimized.

[0135] If a user is looking at a product in a store, the user can use the LLM to get information about the product (e.g., price, reviews, and specifications). The user can also use the camera to take a photo of the product and then use the LLM to search for similar products online. A multifaceted system may allow the user to get all the information they need to know about a product without having to leave the store.

[0136] Another example would be if a user were to view a painting in a museum, they could use the LLM to obtain information about the artist, the painting, and the painting's history. The user could also use a camera to take a picture of the painting and then use the LLM to search for other paintings by the same artist or paintings with a similar theme. This system allows users to learn more about the artwork they are viewing without relying on information provided by the museum.

[0137] In some implementations, a user may be looking for a product but may not be able to find what they are looking for, they can then open their mobile phone and describe the product they are looking for, and the mobile phone will search and find the product.

[0138] The systems and methods can leverage the ambient ecosystem by allowing users to connect their other devices as companions. This connection can enable users to have a more seamless and integrated experience across devices. For example, a user can start a task on a mobile phone and then continue the task on a laptop without having to re-enter any information. The ambient ecosystem can create a more personalized and convenient experience for users.

[0139] 9A depicts a block diagram of an exemplary computing system 100 for performing multi-device output management in accordance with an exemplary embodiment of the present disclosure. System 100 includes a user computing system 102, a server computing system 130, and / or a third computing system 150 communicatively coupled via a network 180.

[0140] The user computing system 102 may include any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0141] The user computing system 102 includes one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively connected processors. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing system 102 to perform operations.

[0142] In some implementations, the user computing system 102 may store or include one or more machine-learned models 120. For example, the machine-learned models 120 may be or otherwise include various machine-learned models, such as neural networks (e.g., deep neural networks), or other types of machine-learned models, including nonlinear and / or linear models. The neural networks may include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other types of neural networks.

[0143] In some implementations, one or more machine-learned models 120 may be received from server computing system 130 over network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing system 102 may implement multiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel machine-learned model processing across multiple instances of input data and / or detected features).

[0144] More specifically, the one or more machine-learned models 120 may include one or more detection models, one or more classification models, one or more segmentation models, one or more augmentation models, one or more generative models, one or more natural language processing models, one or more optical property recognition models, and / or one or more other machine-learned models. The one or more machine-learned models 120 may include one or more Transformer models. The one or more machine-learned models 120 may include one or more neural radiance field models, one or more diffusion models, and / or one or more autoregressive language models.

[0145] One or more machine-learned models 120 may be used to detect one or more object features. The detected object features may be classified and / or embedded. A search may then be performed using the classification and / or embedding to determine one or more search results. Alternatively and / or additionally, one or more detected features may be used to determine whether an indicator (e.g., a user interface element indicating the detected feature) should be provided to indicate that the feature was detected. A user may then select an indicator to cause the feature to be classified, embedded, and / or searched. In some implementations, the classification, embedding, and / or search may be performed before the indicator is selected.

[0146] In some implementations, one or more machine-learned models 120 may process image data, text data, audio data, and / or latent coded data to generate output data, which may include image data, text data, audio data, and / or latent coded data. The one or more machine-learned models 120 may perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, context determination, action prediction, image correction, image enhancement, text enhancement, sentiment analysis, object detection, error detection, inpainting, video stabilization, audio correction, audio enhancement, and / or data segmentation (e.g., mask-based segmentation).

[0147] Additionally or alternatively, one or more machine-learned models 140 may be included in or otherwise stored on and implemented by a server computing system 130 that communicates with the user computing system 102 according to a client-server relationship. For example, the machine-learned models 140 may be implemented by the server computing system 130 as part of a web service (e.g., a viewfinder service, a visual search service, an image processing service, an ambient computing service, and / or an overlay application service). Thus, one or more models 120 may be stored and implemented at the user computing system 102 and / or one or more models 140 may be stored and implemented at the server computing system 130.

[0148] The user computing system 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may function to implement a virtual keyboard. Other exemplary user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0149] In some implementations, the user computing system may store and / or provide one or more user interfaces 124, which may be associated with one or more applications. The one or more user interfaces 124 may be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, augmented reality experiences, virtual reality experiences, and / or other data for display). The user interfaces 124 may be associated with one or more other computing systems (e.g., server computing system 130 and / or third-party computing system 150). The user interfaces 124 may include a viewfinder interface, a search interface, a generative model interface, a social media interface, and / or a media content gallery interface.

[0150] The user computing system 102 may include and / or receive data from one or more sensors 126. The one or more sensors 126 may be housed in a housing component that houses one or more processors 112, memory 114, and / or one or more hardware components that may store and / or execute one or more software packages. The one or more sensors 126 may include one or more image sensors (e.g., cameras), one or more lidar sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biological sensors (e.g., heart rate sensors, pulse sensors, retinal sensors, and / or fingerprint sensors), one or more infrared sensors, one or more position sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. The one or more sensors may be utilized to obtain data associated with the user's environment (e.g., an image of the user's environment, a record of the environment, and / or the user's location).

[0151] The user computing system 102 may include and / or be a part of a user computing device 104. The user computing device 104 may include a mobile computing device (e.g., a smartphone or tablet), a desktop computer, a laptop computer, a smart wearable, and / or a smart appliance. Additionally and / or alternatively, the user computing system may acquire data from and / or generate data using one or more user computing devices 104. For example, a smartphone camera may be utilized to capture image data describing the environment, and / or an overlay application on the user computing device 104 may be utilized to track and / or process data being provided to the user. Similarly, one or more sensors associated with a smart wearable may be utilized to acquire data about the user and / or about the user's environment (e.g., image data may be acquired by a camera housed in the user's smart glasses). Additionally and / or alternatively, data may be acquired from and uploaded to other user devices, which may be specialized in acquiring or generating data.

[0152] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively connected processors. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0153] In some implementations, server computing system 130 includes or is otherwise implemented by one or more server computing devices. When server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0154] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or otherwise include various machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Exemplary models 140 are described with reference to FIG. 9B .

[0155] Additionally and / or alternatively, server computing system 130 may include and / or be communicatively connected to a search engine 142 that may be utilized to crawl one or more databases (and / or resources). Search engine 142 may process data from user computing system 102, server computing system 130, and / or third-party computing system 150 to determine one or more search results associated with the input data. Search engine 142 may perform term-based searches, label-based searches, Boolean-based searches, image searches, embedding-based searches (e.g., nearest neighbor searches), multimodal searches, and / or one or more other search techniques.

[0156] The server computing system 130 may store and / or provide one or more user interfaces 144 for obtaining input data and / or providing output data to one or more users. The one or more user interfaces 144 may include one or more user interface elements, which may include input fields, navigation tools, content tips, selectable tiles, widgets, data display carousels, dynamic animations, information popups, image augmentations, text-to-speech, speech-to-text, augmented reality, virtual reality, feedback loops, and / or other interface elements.

[0157] User computing system 102 and / or server computing system 130 can train models 120 and / or 140 through interaction with a third-party computing system 150 that is communicatively coupled via network 180. Third-party computing system 150 may be separate from server computing system 130 or may be part of server computing system 130. Alternatively and / or additionally, third-party computing system 150 may be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more contexts.

[0158] The third-party computing system 150 may include one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operably connected processors. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 executed by the processor 152 to cause the third-party computing system 150 to perform operations. In some implementations, the third-party computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0159] Network 180 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. Generally, communications over network 180 may occur over any type of wired and / or wireless connection, using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0160] The machine-learned models described herein may be used in a variety of tasks, applications, and / or use cases.

[0161] In some implementations, the input to the machine-learned model(s) of the present disclosure may be image data. The machine-learned model(s) may process the image data to generate an output. As an example, the machine-learned model(s) may process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine-learned model(s) may process the image data to generate an image segmentation output. As another example, the machine-learned model(s) may process the image data to generate an image classification output. As another example, the machine-learned model(s) may process the image data to generate an image data modification output (e.g., a modification of the image data, etc.). As another example, the machine-learned model(s) may process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine-learned model(s) may process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) may process the image data to generate a prediction output.

[0162] In some implementations, input to the machine-learned model(s) of the present disclosure may be text or natural language data. The machine-learned model(s) may process the text or natural language data to generate an output. As another example, the machine-learned model(s) may process the natural language data to generate a language-encoded output. As another example, the machine-learned model(s) may process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) may process the text or natural language data to generate a translation output. As another example, the machine-learned model(s) may process the text or natural language data to generate a classification output. As another example, the machine-learned model(s) may process the text or natural language data to generate a text segmentation output. As another example, the machine-learned model(s) may process the text or natural language data to generate a semantic intent output. As another example, the machine-learned model(s) can process text or natural language data to generate upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine-learned model(s) can process text or natural language data to generate predicted outputs.

[0163] In some implementations, the input to the machine-learned model(s) of the present disclosure can be audio data. The machine-learned model(s) can process the audio data to generate an output. As an example, the machine-learned model(s) can process the audio data to generate a speech recognition output. As another example, the machine-learned model(s) can process the audio data to generate a speech translation output. As another example, the machine-learned model(s) can process the audio data to generate a latent embedding output. As another example, the machine-learned model(s) can process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate an upscaled audio output (e.g., audio data of higher quality than the input audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate a predicted output.

[0164] In some implementations, the input to the machine-learned model(s) of the present disclosure may be sensor data. The machine-learned model(s) may process the sensor data to generate an output. As an example, the machine-learned model(s) may process the sensor data to generate a recognition output. As another example, the machine-learned model(s) may process the sensor data to generate a prediction output. As another example, the machine-learned model(s) may process the sensor data to generate a classification output. As another example, the machine-learned model(s) may process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) may process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) may process the sensor data to generate a visualization output. As another example, the machine-learned model(s) may process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) may process the sensor data to generate a detection output.

[0165] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, and the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to that object class. The image processing task can be object detection, and the image processing output identifies one or more regions of one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, and the image processing output determines, for each pixel of one or more images, a respective likelihood of each category of a set of predetermined categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, and the image processing output determines, for each pixel of one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel in one of the input images, the motion of the scene as depicted in pixels between the images in the network input.

[0166] A user computing system may include several applications (e.g., Applications 1-N). Each application may include its own respective machine learning library and machine-learned model(s). For example, each application may include a machine-learned model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0167] Each application can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0168] The user computing system 102 may include several applications (e.g., applications 1-N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API across all applications).

[0169] The central intelligence layer may include several machine-learned models. For example, each machine-learned model (e.g., model) may be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications may share a single machine-learned model. For example, in some embodiments, the central intelligence layer may provide a single model (e.g., a single model) for all of the applications. In some embodiments, the central intelligence layer is included within or otherwise implemented by the operating system of computing system 100.

[0170] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for computing system 100. The central device data layer can communicate with several other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0171] 9B illustrates a block diagram of an exemplary computing system 50 that performs multi-device output management, according to an exemplary embodiment of the present disclosure. Specifically, the exemplary computing system 50 may include one or more computing devices 52 that can be utilized to acquire and / or generate one or more data sets that can be processed by the sensor processing system 60 and / or the output determination system 80 to provide feedback to a user that may provide information regarding characteristics of the one or more acquired data sets. The one or more data sets may include image data, text data, audio data, multimodal data, latent coding data, etc. The one or more data sets may be acquired via one or more sensors associated with the one or more computing devices 52 (e.g., one or more sensors of the computing devices 52). Additionally and / or alternatively, the one or more data sets may be stored data and / or acquired data (e.g., data acquired from a web resource). For example, images, text, and / or other content items may be interacted with by a user. The interacted content of the content items may then be utilized to generate one or more decisions.

[0172] The one or more computing devices 52 may acquire and / or generate one or more datasets based on image capture, sensor tracking, data storage searches, content downloads (e.g., downloading images or other content items over the Internet from web resources), and / or via one or more other techniques. The one or more datasets may be processed by the sensor processing system 60. The sensor processing system 60 may perform one or more processing techniques using one or more machine-learned models, one or more search engines, and / or one or more other processing techniques. The one or more processing techniques may be performed in any combination and / or individually. The one or more processing techniques may be performed serially and / or in parallel. In particular, the one or more datasets may be processed by the context determination block 62, which may determine a context associated with one or more content items. The context determination block 62 may identify and / or process metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data to determine a particular context associated with the user. A context may be associated with an event, a determined trend, a particular action, a particular type of data, a particular environment, and / or other context associated with the user and / or the searched or retrieved data.

[0173] The sensor processing system 60 may include an image pre-processing block 64. The image pre-processing block 64 may be utilized to adjust one or more values of the captured and / or received images to prepare the images for processing by one or more machine-learned models and / or one or more search engines 74. The image pre-processing block 64 may resize the images, adjust saturation values, adjust resolution, remove and / or add metadata, and / or perform one or more other operations.

[0174] In some implementations, sensor processing system 60 may include one or more machine-learned models, which may include a detection model 66, a segmentation model 68, a classification model 70, an embedding model 72, and / or one or more other machine-learned models. For example, sensor processing system 60 may include one or more detection models 66 that can be utilized to detect particular features in a processed dataset. In particular, one or more images may be processed with one or more detection models 66 to generate one or more bounding boxes associated with features detected in the one or more images.

[0175] Additionally and / or alternatively, one or more segmentation models 68 may be utilized to segment one or more portions of a dataset from one or more datasets. For example, one or more segmentation models 68 may utilize one or more segmentation masks (e.g., one or more segmentation masks generated manually and / or based on one or more bounding boxes) to segment portions of an image, portions of an audio file, and / or portions of text. Segmentation may include isolating one or more detected objects and / or removing one or more detected objects from an image.

[0176] One or more classification models 70 may be utilized to process image data, text data, audio data, latent coding data, multimodal data, and / or other data to generate one or more classifications. The one or more classification models 70 may include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. The one or more classification models 70 may process the data to determine one or more classifications.

[0177] In some implementations, data may be processed with one or more embedding models 72 to generate one or more embeddings. For example, one or more images may be processed with one or more embedding models 72 to generate embeddings of the one or more images in an embedding space. The embeddings of the one or more images may be associated with one or more image features of the one or more images. In some implementations, the one or more embedding models 72 may be configured to process multimodal data to generate multimodal embeddings. The one or more embeddings may be utilized for classification, retrieval, and / or learning embedding spatial distributions.

[0178] The sensor processing system 60 may include one or more search engines 74 available for performing one or more searches. The one or more search engines 74 may crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more proprietary databases, and / or one or more general databases) to determine one or more search results. The one or more search engines 74 may perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), meta-database search, multimodal search, web resource search, image search, text search, and / or application search.

[0179] Additionally and / or alternatively, the sensor processing system 60 may include one or more multimodal processing blocks 76 that can be utilized to assist in processing multimodal data. The one or more multimodal processing blocks 76 may include generating multimodal queries and / or multimodal embeddings that are processed by one or more machine-learned models and / or one or more search engines 74.

[0180] The output(s) of the sensor processing system 60 may then be processed by an output determination system 80 to determine one or more outputs to provide to a user. The output determination system 80 may include heuristic-based determinations, machine-learned model-based determinations, user-selection-based determinations, and / or context-based determinations.

[0181] The output determination system 80 may determine how and / or where to provide one or more search results in the search result interface 82. Additionally and / or alternatively, the output determination system 80 may determine how and / or where to provide one or more machine-learned model outputs in the machine-learned model output interface 84. In some implementations, the one or more search results and / or one or more machine-learned model outputs may be provided for display via one or more user interface elements. The one or more user interface elements may be overlaid on the displayed data. For example, one or more detection indicators may be overlaid over detected objects in a viewfinder. The one or more user interface elements may be selectable to perform one or more additional searches and / or one or more additional machine-learned model processes. In some implementations, the user interface elements may be provided as user interface elements dedicated to a particular application and / or may be provided uniformly across different applications. The one or more user interface elements may include a pop-up display, an interface overlay, an interface tile and / or chip, a carousel interface, audio feedback, animation, an interactive widget, and / or other user interface elements.

[0182] Additionally and / or alternatively, data associated with the output(s) of sensor processing system 60 may be utilized to generate and / or provide an augmented reality and / or virtual reality experience 86. For example, one or more acquired datasets may be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which may then be utilized to provide an augmented reality and / or virtual reality experience 86 to a user. The augmented reality experience may render information associated with an environment into the respective environment. Alternatively and / or additionally, objects associated with the processed dataset(s) may be rendered within the user's environment and / or virtual environment. Generating the rendering datasets may include training one or more neural radiance field models to learn three-dimensional representations of one or more objects.

[0183] In some implementations, one or more action prompts 88 may be determined based on the output(s) of sensor processing system 60. For example, a search prompt, a purchase prompt, a create prompt, a reservation prompt, a call prompt, a redirect prompt, and / or one or more other prompts may be determined to be associated with the output(s) of sensor processing system 60. The one or more action prompts 88 may then be provided to the user via one or more selectable user interface elements. In response to selecting the one or more selectable user interface elements, the respective action of each action prompt may be executed (e.g., a search may be performed, a purchasing application programming interface may be utilized, and / or other applications may be opened).

[0184] In some implementations, one or more datasets and / or output(s) of sensor processing system 60 can be processed by one or more generative models 90 to generate model-generated content items, which can then be provided to a user. Generation can be prompted based on user selection and / or can be performed automatically (e.g., based on one or more conditions, which can be associated with search results of an unidentified threshold quantity).

[0185] The one or more generative models 90 may include a language model (e.g., a large language model and / or a visual language model), an image generation model (e.g., a text-to-image generation model and / or an image augmentation model), an audio generation model, a video generation model, a graph generation model, and / or other data generation models (e.g., other content generation models). The one or more generative models 90 may include one or more Transformer models, one or more convolutional neural networks, one or more recurrent neural networks, one or more feedforward neural networks, one or more generative adversarial networks, one or more self-attention models, one or more embedding models, one or more encoders, one or more decoders, and / or one or more other models. In some implementations, the one or more generative models 90 may include one or more autoregressive models (e.g., machine-learned models trained to generate predictions based on previous behavioral data) and / or one or more diffusion models (e.g., machine-learned models trained to generate predictions based on generating and processing distribution data associated with input data).

[0186] One or more generative models 90 may be trained to process input data and generate model-generated content items, which may include predicted words, pixels, signals, and / or other data. The model-generated content items may include new content items that are not identical to any existing work. The one or more generative models 90 may utilize learned representations, sequences, and / or probability distributions to generate content items, which may include phrases, storylines, settings, objects, characters, beats, lyrics, and / or other aspects not included in existing content items.

[0187] The one or more generative models 90 may include a visual language model. The visual language model may be trained, tuned, and / or configured to process image data and / or text data to generate natural language output. The visual language model may utilize a pre-trained large language model (e.g., a large autoregressive language model) with one or more encoders (e.g., one or more image encoders and / or one or more text encoders) to provide detailed natural language output that emulates natural language produced by humans.

[0188] The visual language model may be utilized for zero-shot image classification, few-shot image classification, image captioning, multimodal query distillation, multimodal question answering, and / or may be tuned and / or trained for multiple different tasks. The visual language model may perform visual question answering, image caption generation, feature detection (e.g., content monitoring (e.g., inappropriate content)), object detection, scene recognition, and / or other tasks.

[0189] The visual language model can leverage a pre-trained language model, which can then be tuned for multimodality. Training and / or tuning the visual language model may include image-text matching, masked language modeling, multimodal fusion with cross-attention, contrastive learning, prefix language model training, and / or other training techniques. For example, the visual language model may be trained to process images and generate predictive text similar to ground truth text data (e.g., ground truth captions for images). In some implementations, the visual language model may be trained to replace masked tokens in natural language templates with text tokens that describe features depicted in the input image. Alternatively and / or additionally, the training, tuning, and / or model estimation may include multi-layer concatenation of visual and text embedding features. In some implementations, the visual language model may be trained and / or tuned by jointly learning image embeddings and generating text embeddings, which may include training and / or tuning a system that maps embeddings to a joint feature embedding space that maps text features and image features to a shared embedding space. The joint training may include parallel embedding of image-text pairs and / or may include triplet training. In some implementations, images may be used and / or processed as prefixes for a language model.

[0190] The output determination system 80 may process one or more data sets and / or output(s) of the sensor processing system 60 using a data augmentation block 92 to generate augmented data. For example, one or more images may be processed in the data augmentation block 92 to generate one or more augmented images. Data augmentation may include data correction, data cropping, removing one or more features, adding one or more features, adjusting resolution, adjusting lighting, adjusting color saturation, and / or other enhancements.

[0191] In some implementations, one or more data sets and / or output(s) of the sensor processing system 60 may be stored based on the determination in the data storage block 94 .

[0192] The output(s) of output determination system 80 may then be provided to a user via one or more output components of user computing device 52. For example, one or more user interface elements associated with the one or more outputs may be provided for display via a visual display of user computing device 52.

[0193] The process may be performed iteratively and / or continuously, and one or more user inputs to provided user interface elements may define and / or affect a continuous processing loop.

[0194] The technology described herein refers to servers, databases, software applications, and other computer-based systems, as well as actions performed on and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionality among components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0195] While the subject matter of the present disclosure has been described in detail with reference to various specific exemplary embodiments thereof, each example is provided for purposes of explanation and not limitation of the present disclosure. Those skilled in the art, upon understanding the foregoing, will readily be able to make modifications, variations, and equivalents to such embodiments. Accordingly, the present disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter as would be readily apparent to one skilled in the art. For example, features illustrated or described as part of one embodiment can be used with other embodiments to yield still further embodiments. Accordingly, the present disclosure is intended to cover such modifications, variations, and equivalents.

Claims

1. 1. A computing system for determining an output device for providing a query response, comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations; and wherein the operation comprises: obtaining input data, the input data including a query associated with a particular user; acquiring environmental data describing a plurality of computing devices in the user's environment, the plurality of computing devices being associated with a plurality of different output components; generating a prompt based on the input data and the environmental data, the prompt including data describing the query and device information associated with at least a subset of the plurality of computing devices; processing the prompt with a generative model to generate a model-generated output, the model-generated output including a response to the query, the model-generated output generated to be provided to a particular computing device of the plurality of computing devices; transmitting the model generation output to the particular computing device; Including, the system.

2. the model-generated output is generated to be provided to a particular output component of the plurality of different output components; the generative model generates output device instructions, the output device instructions describing a particular computing device of the plurality of computing devices that provides the model-generated output, the particular computing device being associated with the particular output component; The system of claim 1 , wherein the model generation output is sent to the particular computing device based on the output device instructions.

3. processing the prompts with the generative model to generate the model-generated outputs, generating a plurality of model outputs, the plurality of model outputs including a plurality of candidate responses; transmitting the model generation output to the particular computing device; transmitting a first model output of the plurality of model outputs to a first computing device of the plurality of computing devices; transmitting a second model output of the plurality of model outputs to a second computing device of the plurality of computing devices; The system of claim 1 , comprising:

4. 4. The system of claim 3, wherein the first model output comprises visual data for display via a visual display and the second model output comprises audio data for playback via a speaker component.

5. The system of claim 4 , wherein the first computing device comprises a smart television and the second computing device comprises a smart speaker.

6. generating the prompt based on the input data and the environmental data, determining an environment-specific device configuration based on the environment data; retrieving a prompt template from a prompt library based on the environment-specific device configuration; expanding the prompt template based on the input data to generate the prompt; The system of claim 1 , comprising:

7. 7. The system of claim 6, wherein the environment-specific device configuration describes an output type and a respective output quality of each of the plurality of computing devices, and the prompt library includes a plurality of different prompt templates associated with a plurality of different device configurations.

8. 2. The system of claim 1, wherein the plurality of computing devices are connected via a cloud computing system, each of the plurality of computing devices is registered with a platform of the cloud computing system, the environmental data is acquired using the cloud computing system, and the model generation output is transmitted via the cloud computing system.

9. 10. The system of claim 1, wherein the plurality of computing devices are located proximate to each other computing device in the plurality of computing devices, the plurality of computing devices are communicatively connected via a local network, and particular computing devices of the plurality of computing devices facilitate obtaining input data and transmitting model generation output.

10. the Generative Model is communicatively coupled to a search engine via an application programming interface; processing the prompts with the generative model to generate the model-generated outputs, generating an application programming interface call based on the prompt; determining a plurality of search results using the search engine based on the application programming interface calls; processing the plurality of search results with the generative model to generate the model-generated output; The system of claim 1 , comprising:

11. 1. A computer-implemented method comprising: obtaining, by a computing system including one or more processors, input data, the input data including a query associated with a particular user; acquiring, by the computing system, environmental data describing a plurality of computing devices in the user's environment, the plurality of computing devices being associated with a plurality of different output components; generating, by the computing system, a prompt based on the input data and the environmental data, the prompt including data describing the query and device information associated with at least a subset of the plurality of computing devices; processing, by the computing system, the prompt with a generative model to generate a model-generated output and output device instructions, the model-generated output including a response to the query, the model-generated output being generated to provide a particular output component of the plurality of different output components, the output device instructions describing a particular computing device of the plurality of computing devices that provides the model-generated output, the particular computing device being associated with the particular output component; transmitting, by the computing system, the model-generated output to the particular computing device based on the output device instructions; A method comprising:

12. the plurality of different output components are associated with a plurality of respective output capabilities associated with the plurality of computing devices; The method of claim 11 , wherein each of the plurality of respective output capabilities describes an output type and output quality available via a respective computing device.

13. The method of claim 11 , wherein the output device instructions include an application programming interface call for sending the model-generated output to the particular computing device.

14. The method of claim 11 , wherein the plurality of different output components comprises a speaker associated with a first device and a visual display associated with a second device.

15. determining, by the computing system, that the particular output component is associated with the intent of the query; The method of claim 11 , wherein the prompt is generated based on the particular output component associated with the intent of the query.

16. determining, by the computing system, an output hierarchy based on specification information of the plurality of different output components based on the environmental data; The method of claim 11 , wherein the prompt is generated based on the output hierarchy and the query.

17. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: acquiring environment data, the environment data describing a plurality of computing devices in an environment associated with a particular user; processing the environmental data to determine a plurality of respective input functions and a plurality of respective output functions associated with the plurality of computing devices, the plurality of respective input functions being associated with candidate input types associated with the plurality of computing devices and the plurality of respective output functions being associated with candidate output types associated with the plurality of computing devices; generating a plurality of respective interfaces for the plurality of computing devices based on the plurality of respective input capabilities and the plurality of respective output capabilities, the plurality of respective interfaces being specialized for the plurality of computing devices based on the plurality of respective input capabilities and the plurality of respective output capabilities; providing the plurality of computing devices with the plurality of respective interfaces; 1. One or more non-transitory computer-readable media, including:

18. 20. The one or more non-transitory computer-readable media of claim 17, wherein the plurality of respective interfaces comprise a plurality of device indicators that represent the plurality of computing devices in the environment associated with the particular user, the plurality of computing devices being configured as a user-specific device ecosystem communicatively connected to receive input and provide output.

19. 20. The one or more non-transitory computer-readable media of claim 17, wherein each of the plurality of respective interfaces is configured to receive a particular input type and provide a particular output type based on a respective input capability and a respective output capability for the particular computing device of the plurality of computing devices.

20. The operation is obtaining user input via a first interface of a first computing device of the plurality of computing devices; processing the user input with a search engine to determine a plurality of search results; processing the plurality of search results with a generative model to generate a model output; providing the model output for display via a second interface of a second computing device of the plurality of computing devices; 20. The one or more non-transitory computer-readable media of claim 17, further comprising:

Citation Information

Patent Citations

  • Interface for in-vehicle equipment

    JP2005178473A

  • Disambiguating heteronyms in speech synthesis

    JP2016122183A

  • Intelligent automatic assistant for tv user interaction

    JP2017530567A

  • Persona chatbot control method and system

    JP2022180282A

  • Multi-device Mediation for Assistant Systems

    US20220358917A1