Ambient multi-device framework for agent companion

The multi-device framework addresses interconnectivity issues by using environmental data and a generative model to optimize input and output across devices, enhancing computational efficiency and delivering high-quality responses through suitable devices.

JP2025098949AActive Publication Date: 2025-07-02GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024203206
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-11-21
Publication Date
2025-07-02
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

Existing computing devices in a user's environment often lack interconnectivity, leading to inefficiencies in utilizing their functions for input and output, particularly when users struggle to determine which device should provide the desired output for their queries due to insufficient descriptive inputs.

Method used

A multi-device framework that utilizes environmental data and a generative model to determine which computing device should provide the output based on input data and device capabilities, enabling coordinated input acquisition and output provision across multiple devices.

Benefits of technology

Enhances computational efficiency by reducing repetitive processing and providing an immersive, interconnected multi-device ecosystem that optimizes input and output across various devices, ensuring high-quality responses are delivered through the most suitable device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098949000001_ABST
    Figure 2025098949000001_ABST
Patent Text Reader

Abstract

To provide a system for conditioning output generation based on the devices in an environment and determining which computing device to provide the output, and a non-transitory computer readable medium.SOLUTION: Methods for generating and providing outputs in a multi-device system can include: leveraging environment-based prompt generation and generative model response generation to provide dynamic response generation and display, to obtain input data associated with one or more computing devices within an environment; obtaining environment data descriptive of the plurality of computing devices within the environment; generating a prompt based on the input data and environment data; processing a prompt with a generative model to generate a model-generated output; and transmitting the model-generated output to a particular computing device of the plurality of computing devices.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Claims of Priority This application claims priority based on U.S. Non-Provisional Application No. 18 / 390,768, filed on December 20, 2023, and is incorporated herein by reference.

[0002] The present disclosure generally relates to generating and providing output in a multi-device system. More specifically, the present disclosure relates to defining output generation based on devices in an environment and determining which computing device provides the output.

Background Art

[0003] Computing devices can be found throughout a user's living room, bedroom, study, office, and / or other environments. A complex device environment may provide a user with multiple interacting computing devices, which may include smart TVs, smart speakers, smart appliances, virtual assistant devices, tablets, smart wearables, smartphones, and / or other computing devices, and can be provided throughout the user's environment. However, the functions of these devices may not be utilized based on the lack of interconnectivity between the devices. For example, a user may be performing a search on a smartphone, but it may not be possible to play video search results on the smartphone even though the smart TV is only a few feet away.

[0004] It may be difficult to understand the overall situation. Even when an individual is trying to understand what an object in front of them is, or trying to determine whether the object can be found elsewhere, and / or trying to determine where an image on the Internet was captured, it may be difficult with only text search. In particular, the user may struggle to determine which words to use. Furthermore, those words may not be descriptive enough and / or rich enough to produce the desired result.

Summary of the Invention

[0005] Aspects and advantages of embodiments of the present disclosure are shown in part in the following description, or can be learned from the description, or can be learned through the practice of the embodiments.

[0006] One exemplary aspect of the present disclosure is directed to a computing system for determining an output device for providing a query response. The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining input data. The input data can include a query associated with a particular user. The operations can include obtaining environmental data. The environmental data can describe a plurality of computing devices in the user's environment. In some embodiments, the plurality of computing devices can be associated with a plurality of different output components. The operations can include generating a prompt based on the input data and the environmental data. The prompt can include data describing the query and device information associated with at least a subset of the plurality of computing devices. The operations can include processing the prompt using a generation model to generate a model-generated output. The model-generated output can include a response to the query. In some embodiments, the model-generated output can be generated to be provided to a particular computing device among the plurality of computing devices. The operations can include sending the model-generated output to the particular computing device.

[0007] In some embodiments, the model-generated output can be generated to be provided to a particular output component among the plurality of different output components. The generation model can generate output device instructions. The output device instructions can describe a particular computing device among the plurality of computing devices that provides the model-generated output. The particular computing device can be associated with the particular output component. The model-generated output can be sent to the particular computing device based on the output device instructions.

[0008] In some embodiments, processing a prompt using a generative model to generate a model-generated output may include generating a plurality of model outputs. The plurality of model outputs can include a plurality of candidate responses. Sending the model-generated output to a particular computing device may include sending a first model output of the plurality of model outputs to a first computing device of the plurality of computing devices and sending a second model output of the plurality of model outputs to a second computing device of the plurality of computing devices. The first model output can include visual data for display via a visual display. The second model output can include audio data for playback via a speaker component. In some embodiments, the first computing device can include a smart TV, and the second computing device can include a smart speaker.

[0009] In some embodiments, generating a prompt based on input data and environmental data may include determining an environment-specific device configuration based on the environmental data, obtaining a prompt template from a prompt library based on the environment-specific device configuration, and expanding the prompt template based on the input data to generate the prompt. The environment-specific device configuration can describe the respective output type and the respective output quality of a plurality of computing devices. The prompt library can include a plurality of different prompt templates associated with a plurality of different device configurations.

[0010] In some embodiments, a plurality of computing devices can be connected via a cloud computing system. Each of the plurality of computing devices can be registered with the platform of the cloud computing system. Environmental data can be obtained in the cloud computing system. Model generation output can be transmitted via the cloud computing system. In some embodiments, the plurality of computing devices can be arranged in proximity to each of the other computing devices within the plurality of computing devices. The plurality of computing devices can be communicatively connected across a local network. A particular computing device among the plurality of computing devices can facilitate the acquisition of input data and the transmission of model generation output.

[0011] In some embodiments, a generative model can be communicatively connected to a search engine via an application programming interface. Processing a prompt using the generative model to generate a model generation output can include generating an application programming interface call based on the prompt, determining a plurality of search results using the search engine based on the application programming interface call, and processing the plurality of search results using the generative model to generate the model generation output.

[0012] Other exemplary aspects of the present disclosure are directed to computer-implemented methods. The method can include obtaining input data by a computing system that includes one or more processors. The input data can include a query associated with a particular user. The method can include obtaining environmental data by the computing system. The environmental data can describe a plurality of computing devices in the user's environment. The plurality of computing devices can be associated with a plurality of different output components. The method can include generating a prompt by the computing system based on the input data and the environmental data. The prompt can include data describing the query and device information associated with at least a subset of the plurality of computing devices. The method can include processing the prompt using a generation model by the computing system to generate a model generation output and an output device instruction. The model generation output can include a response to the query. In some embodiments, the model generation output can be generated to be provided to a particular output component of the plurality of different output components. The output device instruction can describe a particular computing device of the plurality of computing devices that provides the model generation output. The particular computing device can be associated with the particular output component. The method can include transmitting the model generation output to the particular computing device by the computing system based on the output device instruction.

[0013] In some embodiments, a plurality of different output components can be associated with respective output functions associated with a plurality of computing devices. Each of the respective output functions of the plurality can describe the output types and output qualities available via the respective computing devices. Output device instructions can include an application programming interface call for sending model generation output to a particular computing device. In some embodiments, the plurality of different output components can include a speaker associated with a first device and a visual display associated with a second device. The method can include determining, by a computing system, that a particular output component is associated with the intent of a query. A prompt can be generated based on the particular output component associated with the intent of the query.

[0014] In some embodiments, the method can include determining, by a computing system, an output hierarchy based on environmental data and based on specification information of a plurality of different output components. A prompt can be generated based on the output hierarchy and the query.

[0015] Other exemplary aspects of the present disclosure are directed to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations can include obtaining environmental data. The environmental data can describe a plurality of computing devices within an environment associated with a particular user. The operations can include processing the environmental data to determine a plurality of respective input functions and a plurality of respective output functions associated with the plurality of computing devices. The plurality of respective input functions can be associated with candidate input types associated with the plurality of computing devices. The plurality of respective input functions can be associated with candidate output types associated with the plurality of computing devices. The operations can include generating a plurality of respective interfaces for the plurality of computing devices based on the plurality of respective input functions and the plurality of respective output functions. The plurality of respective interfaces can be specialized for the plurality of computing devices based on the plurality of respective input functions and the plurality of respective output functions. The operations can include providing the plurality of respective interfaces to the plurality of computing devices.

[0016] In some embodiments, the plurality of respective interfaces can include a plurality of device indicators indicative of the plurality of computing devices within an environment associated with a particular user. The plurality of computing devices can be configured as a user-specific device ecosystem that is communicatively connected to receive inputs and provide outputs. Each of the plurality of respective interfaces can be configured to receive a particular input type and provide a particular output type based on the respective input function and the respective output function for a particular one of the plurality of computing devices.

[0017] In some embodiments, the operation can include obtaining user input via a first interface of a first computing device of a plurality of computing devices, processing the user input using a search engine to determine a plurality of search results, processing the plurality of search results using a generation model to generate a model output, and providing the model output for display via a second interface of a second computing device of the plurality of computing devices.

[0018] Other aspects of the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0019] These and other features, aspects, and advantages of the various embodiments of the present disclosure will become better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.

[0020] A detailed description of embodiments directed to those skilled in the art is set forth in this specification with reference to the accompanying drawings.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 6E

Figure 6F

Figure 6G

Figure 7

Figure 8

Figure 9A

Figure 9B

DETAILED DESCRIPTION OF THE INVENTION

[0022] Reference numerals repeated throughout multiple drawings are intended to identify the same features in various embodiments.

[0023] Generally, the present disclosure is directed to a multi-device framework for managing the acquisition of inputs and the generation of outputs for an environment that includes a plurality of computing devices. Specifically, the systems and methods disclosed herein can utilize environment-based prompt generation and generative model processing to generate and / or provide outputs that can be generated based on the computing devices available in the environment. For example, the multi-device framework can include a computing system that facilitates obtaining input data and environmental data associated with a plurality of computing devices, and generating one or more model-generated outputs configured to be provided to one or more specific computing devices among the plurality of computing devices. The computing system can be utilized to leverage the diverse input and / or output capabilities of different computing devices within the environment, which can include automatically obtaining different types of inputs from different devices and / or providing different types of outputs via different devices (e.g., voice commands can be received via a smartphone, while outputs can be provided via a television's visual display and / or a sound system's speakers).

[0024] Environmental data associated with multiple computing devices in an environment can be utilized for prompt generation, environment understanding, and / or interface generation. The environmental data may be processed to generate a prompt that defines the output generation of a generative model that is compatible with the user's environment and / or optimized for display and / or playback in the user's environment. Prompt generation may include obtaining a specific prompt template from a prompt library based on the computing device associated with the environment, and / or may include prompt generation by processing input data and / or environmental data using a machine-learned model. Additionally and / or alternatively, the environmental data can be processed to determine which computing devices associated with the environment to utilize for a particular type of input acquisition and / or a particular type of output playback. In some embodiments, the environmental data can be processed to understand the computing devices in the environment and generate respective interfaces for the multiple computing devices based on the determined input and / or output capabilities of the computing devices associated with the environment.

[0025] An immersive virtual assistant can be provided that utilizes an ambient multi-device framework to receive inputs from multiple different computing devices and provide outputs thereby. The ambient multi-device framework can obtain queries and / or prompts from a user and provide responses in relevant content types using a computing device in the environment that can provide relevant content types that may be of higher quality than other devices in the environment. The determination of the computing device can be based on a determined hierarchy that can be determined based on information about the device's specifications (and / or information about the device's capabilities).

[0026] Computing devices can be found throughout a user's environment, regardless of whether it is in an office, at home, and / or other locations. Smart TVs, smart speakers, smart appliances, virtual assistant devices, tablets, smart wearables, smartphones, and / or other computing devices can be provided throughout the user's environment. However, the functionality of these devices may not be utilized due to a lack of interconnectivity and / or lack of coordination between the devices. In particular, computing devices may not be able to determine when and / or how to interact in order to obtain input and / or provide context-specific output.

[0027] An ambient multi-device framework can include obtaining and / or generating information that describes computing devices in an environment that includes information related to input and / or output functions. The information can then be used when generating a prompt for a generative model (e.g., a large language model) that can define what data is obtained and / or generated for output and / or what computing devices are used for output associated with a user query. In some embodiments, the generative model can be fine-tuned for the multi-device framework (e.g., efficient fine-tuning of parameters and / or adjustment of soft prompts), which can include adjusting the generative model to process the prompt using the environmental context and generate an output based on both the query and the environmental context.

[0028] Specifically, different computing devices can have different components for capturing different forms of input data (e.g., text input, voice commands, gesture input, etc.) and / or for providing different types of output data to the user (e.g., visual display, audio playback, etc.). The ambient multi-device framework can utilize information about the functions associated with the computing devices in the environment to provide an immersive and versatile computing system that can obtain inputs and provide outputs in multiple different computing systems within the environment.

[0029] Smartphones, smartwatches, smart speakers, smart TVs, smart assistant devices, and / or other computing devices are always around the user. However, the interconnectivity may be limited, and the highest quality input and / or output data may not be utilized efficiently. For example, a user can input a query via their smartphone, and in response, video and / or audio can be provided on the smartphone. However, high-quality smart TVs and / or smart speakers may be easily accessible and proximate to the user. Regardless of whether the content is essentially for entertainment, education, and / or other purposes, the multi-device framework can be utilized to obtain additional forms of input and / or provide additional output formats that can be of higher quality than a single device system.

[0030] The systems and methods disclosed herein can be utilized to manage the acquisition of inputs from composite devices and / or the provision of outputs via composite devices. In some embodiments, the inputs can be received from smartphones and smartwatches, and the outputs in response to the inputs can be provided via smart TVs and smart speakers. Additionally and / or alternatively, the systems and methods can determine that a particular computing device in the environment has the highest processing capacity, and the systems and methods can then utilize that particular computing device to perform model estimation. In some embodiments, the processing tasks may be divided among the composite computing devices in the environment and / or may be performed by a server computing system.

[0031] The systems and methods of the present disclosure provide several technical effects and advantages. As an example, the systems and methods can be utilized to provide an interconnected multi-device ecosystem. In particular, the systems and methods disclosed herein can acquire input data from one or more devices within the ecosystem. The systems and methods can acquire environmental data that describes the devices in the environment, and can then generate a prompt based on the input data and the environmental data that can be processed using a generation model to generate a response to the input data generated based on the devices in the environment. The response can then be transmitted to one or more particular devices in the environment. For example, text input can be acquired via a tablet, and the output can be provided via a smart speaker.

[0032] Other exemplary technical effects and benefits relate to improvements in computational efficiency and the functionality of computing systems. For example, the technical advantages of the systems and methods of the present disclosure are that they can reduce the computational resources required to interact with multiple computing devices within an environment. In particular, the multi-device framework can unify data processing and flow, thereby reducing instances of repetitive processing between devices within the ecosystem.

[0033] The systems and methods of the present disclosure provide several technical effects and benefits. As an example, the systems and methods can provide an interface generation system. The interface generation system can be utilized to generate environment-specific and device-specific interfaces that can be specifically generated based on the input and output functions available within a multi-device ecosystem.

[0034] Another technical advantage of the systems and methods of the present disclosure is that they can utilize interface generation to provide an immersive multi-device environment. In particular, the interface can be generated and provided to provide the user with the ability to obtain multiple different input types from multiple different devices of an interconnected system, and can provide multiple different output types via multiple different devices of the interconnected system.

[0035] Reference is now made to the drawings, where exemplary embodiments of the present disclosure are discussed in more detail.

[0036] FIG. 1 shows a block diagram of an exemplary generation model system 10 according to an exemplary embodiment of the present disclosure. In some embodiments, the generation model system 10 is configured to receive and / or obtain a set of input data 12 that describes a query and / or context, and as a result of receiving the input data 12, generate, determine, and / or provide a model generation output 20 that describes a response to the query. Thus, in some embodiments, the generation model system 10 may include a generation model 18 operable to process a prompt 16 to generate a response to a query configured to be provided by a particular computing device in an environment.

[0037] In particular, the generation model system 10 can obtain the input data 12 and the environmental data 14. The input data 12 can describe one or more inputs. The input data 12 can describe a query (e.g., "Who is playing today?", "How do you make focaccia?", "What was the song that played at the end of the movie this year?"). The input data 12 can include directly entered inputs (e.g., a user types on a graphical keyboard, captures an image, selects a graphical user interface element, provides a voice command, etc.) and / or context inputs (e.g., a user's search history, time, user habit data, user browsing history, application activity, currently open application, temperature, user location, and / or data obtained from other computing devices in the environment).

[0038] Environmental data 14 can describe multiple computing devices within an environment. The environment may be a room, a set of rooms, those in proximity to a user, and / or a space including multiple computing devices in proximity to each other. Environmental data 14 can include information on the specifications for each of the multiple computing devices within the environment. Alternatively and / or additionally, environmental data 14 can be descriptive identification data of the multiple computing devices (e.g., device name, device serial number, device label, registration data, etc.), functional data of the multiple computing devices (e.g., input functions for the multiple computing devices, processing functions for the multiple computing devices, and / or output functions of the multiple computing devices), and / or other environmental data.

[0039] The generation model system 10 can process the input data 12 and the environmental data 14 to generate a prompt 16. The prompt 16 can include a query and information associated with multiple computing devices. The prompt 16 can include a query and indicators of one or more candidate computing devices that provide an output. The prompt 16 can be generated by querying a prompt library to determine a prompt template based on the environmental data 14. Next, the prompt template can be filled in based on the input data 12. Alternatively and / or additionally, the prompt 16 can be generated by processing the input data 12 and the environmental data 14 by a prompt generation model. The prompt generation model can include a machine-learned model, which can include a generation model (e.g., an autoregressive language model, a diffusion model, a vision-language model, and / or other generation models).

[0040] Next, the generation model 18 can process the prompt 16 to generate a model generation output 20. The generation model 18 can include a generative language model (e.g., a large language model, a vision-language model, and / or other language models), an image generation model (e.g., a text-to-image generation model), and / or other generation models. The generation model 18 can include a transformer model, a convolutional model, a feed-forward model, a regression model, a self-attention model, and / or other models.

[0041] The model generation output 20 can include a predicted response to the query of the input data 12. The model generation output 20 can be a new generation output that includes a plurality of predicted features (e.g., predicted text characters, predicted audio signals, predicted pixels, etc.). The model generation output 20 can be generated to be provided by a specific computing device among a plurality of computing devices. For example, the model generation output 20 can be generated to have a specific size, quality, and / or content type based on a preferred computing device (having a specific function) determined based on the prompt and / or the processing of the generation model 18.

[0042] The model generation output 20 can then be provided to a specific computing device within the environment. Providing the model generation output 20 to a specific computing device can include transmitting data, which can include generating an application programming interface call using the generation model and executing the API call using one or more application programming interfaces. Next, the specific computing device can provide the model generation output 20 to the user (e.g., via a visual display, a speaker, and / or other output components of the specific computing device).

[0043] FIG. 2 shows a block diagram of an exemplary response generation system 200 according to an exemplary embodiment of the present disclosure. The response generation system 200 is similar to the generation model system 10 of FIG. 1, except that the response generation system 200 further includes a prompt library 224 and a search engine 226.

[0044] Specifically, the response generation system 200 can obtain input data 212 and environmental data 214. The input data 212 can include text data, image data, audio data, embedded data, signal data, search history data, browsing history data, application interaction data, potential encoding data, multimodal data, global data, and / or other data. The input data 212 can be a description of one or more inputs that may include obtaining inputs from complex computing devices within the environment. The input data 212 can describe a query (e.g., "Who is playing today?", "How do you make focaccia?", "What was the song that played at the end of the movie this year?"). The input data 212 can include directly inputted inputs (e.g., a user types on a graphical keyboard, captures an image, selects a graphical user interface element, provides a voice command, etc.) and / or context inputs (e.g., the user's search history, time, user habit data, user browsing history, application activity, currently open applications, temperature, user location, and / or data obtained from other computing devices within the environment). The input data 12 can include context data determined to be related to the query. Alternatively and / or additionally, the input data 12 can include a predicted query that can be generated based on a prediction of what the user may be interested in based on one or more determined user contexts (e.g., determining an interest in a specific soccer team, using a series of searches to determine the likelihood of an interest in horror movies, and / or determining that a specific product for purchase may be of interest based on the purchase of similar products). In some embodiments, the response generation system 200 can propose content to the user by generating input data based on individualized signals and / or contextualized signals.

[0045] Environmental data 214 can describe a plurality of computing devices 222 within an environment. The environment can be a room, a set of rooms, things proximate to a user, and / or a space including a plurality of computing devices proximate to each other. The environmental data 214 can include information about the specifications for each of the plurality of computing devices 222 within the environment. Alternatively and / or additionally, the environmental data 214 can be descriptive identification data (e.g., device name, device serial number, device label, registration data, etc.) for the plurality of computing devices 222, functional data for the plurality of computing devices 222 (e.g., input functions for the plurality of computing devices 222, processing functions for the plurality of computing devices 222, and / or output functions for the plurality of computing devices 222), and / or other environmental data. The plurality of computing devices 222 can include smartphones, smart wearables (e.g., smartwatches and / or smart glasses), smart speakers, smart TVs, laptops, desktops, smart appliances (e.g., smart refrigerators, smart washers, smart dryers, and / or smart dispensers), virtual assistant devices (e.g., smart home panels, in-room base assistants, and / or other assistant devices), tablets, and / or other computing devices. The plurality of computing devices 222 can include mobile computing devices and / or fixed computing devices. In some embodiments, the plurality of computing devices 222 can include devices proximate to a user, devices registered with a user's mobile device, devices registered with a user's profile, devices connected to a particular internet hub, devices registered with a particular assistant device, and / or other devices having a correlation relationship.

[0046] The response generation system 200 can process input data 212 and / or environmental data 214 to generate a prompt 216. The prompt 216 can include a query and information associated with a plurality of computing devices. The prompt 216 can include a query and indicators of one or more candidate computing devices that provide an output. The prompt 216 can be generated by querying a prompt library 224 to determine a prompt template based on the environmental data 214. Querying the prompt library 224 can include determining a prompt template associated with an environmental configuration associated with the user's environment. Alternatively and / or additionally, querying the prompt library 224 can include generating an embedding based on the environmental data 214 and then determining a prompt template associated with the embedding based on a nearest neighbor search. Next, the prompt template can be filled in based on the input data 212. Alternatively and / or additionally, the prompt 216 can be generated by processing the input data 212 and the environmental data 214 by a prompt generation model. The prompt generation model can include a machine-learned model, which can include a generation model (e.g., an autoregressive language model, a diffusion model, a vision-language model, and / or other generation models). The prompt 216 can include text data, image data, audio data, latent encoded data, multimodal data, embedding data, and / or other data. The prompt 216 can include a hard prompt (e.g., a text string and / or other data input) and / or a soft prompt (e.g., a set of learned parameters). In some embodiments, the prompt 216 can be a zero-shot prompt and / or a few-shot prompt. The prompt 216 can separate user requests into multiple tasks, and the tasks can be provided as multiple prompts.

[0047] Next, the generation model 218 can process the prompt 216 to generate a model generation output 220. The generation model 218 can include a generative language model (e.g., a large language model, a vision language model, and / or other language models), an image generation model (e.g., a text-to-image generation model), and / or other generation models. The generation model 218 can include a transformer model, a convolutional model, a feedforward model, a regression model, a self-attention model, and / or other models.

[0048] In some embodiments, the generation model 218 can be communicatively connected to one or more processing engines (e.g., a search engine, a rendering engine, and / or other engines) and / or one or more machine-learned models (e.g., a classification model, an expansion model, a segmentation model, a detection model, an embedding model, and / or other models). For example, the generation model 218 can communicate with a search engine 226 to obtain search results based on input data 212, environmental data 214, and / or estimates. The search engine 226 may output search results, which can then be processed by the generation model 218 to generate a summary of the search results and / or determine a response to a query based on information from the search results, which can then be used to generate the model generation output 220. The model generation output 220 can include text data, image data, audio data, latent encoded data, multimodal data, and / or other data.

[0049] The model generation output 220 can include a predicted response to a query of the input data 212. The model generation output 220 can be a new generation output that includes a plurality of predicted features (e.g., predicted text characters, predicted audio signals, predicted pixels, etc.). The model generation output 220 can be generated to be provided by a specific computing device among a plurality of computing devices. For example, the model generation output 220 can be generated to have a specific size, quality, and / or content type based on a preferred computing device (having specific functions) determined based on the prompt and / or the processing of the generation model 218.

[0050] The generation model 218 can process the prompt 216 to generate output device instructions 230. The output device instructions 230 can include instructions for transmitting the model generation output 220 to a specific computing device 228 of the plurality of computing devices 222. The output device instructions 230 can include application programming interface calls, registration files, and / or instructions for the specific computing device 228.

[0051] Next, the model generation output 220 can be provided to a specific computing device in the environment based on and / or together with the output device instructions 230. Providing the model generation output 220 to a specific computing device can include transmitting data, which can include generating an application programming interface call using the generation model and executing the API call using one or more application programming interfaces. Next, the specific computing device can provide the model generation output 220 to the user (e.g., via a visual display, speaker, and / or other output components of the specific computing device).

[0052] FIG. 3 shows a flowchart diagram of an exemplary method for functioning in accordance with an exemplary embodiment of the present disclosure. Although FIG. 3 shows steps executed in a particular order for purposes of illustration and explanation, the method of the present disclosure is not limited to the specifically shown order or arrangement. The various steps of method 300 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0053] At 302, the computing system can obtain input data. The input data can include a query associated with a particular user. The input data can include text data, audio data, image data, gesture data, latent encoded data, multimodal data, and / or other data. The input data can be obtained from one or more computing devices of the environment and / or may be generated by one or more computing devices. The input data can include a prompt for a generative model. In some embodiments, the input data can be a description of a question related to one or more topics. The question can be related to the content viewed by one or more computing devices and / or features within the environment.

[0054] At 304, the computing system can obtain environmental data. The environmental data can describe a plurality of computing devices in the user's environment. The plurality of computing devices can be associated with a plurality of different output components. In some embodiments, the plurality of computing devices can be arranged in proximity to each of the other computing devices within the plurality of computing devices. The plurality of computing devices can be communicatively connected across a network. In some embodiments, the environmental data can describe the specifications of the plurality of computing devices. For example, the environmental data can indicate components of a computing device that can include input components and / or output components. Additionally and / or alternatively, the environmental data can include information describing the input and / or output functions of the plurality of different computing devices.

[0055] At 306, the computing system can generate a prompt based on the input data and the environmental data. The prompt can include data that describes a query and device information associated with at least a subset of the plurality of computing devices. The prompt can include text data, image data, audio data, embedded data, statistical data, graphic representation data, latent encoding data, semantic data, multimodal data, and / or other data. The prompt can be generated by a machine-learned model based on a deterministic function, heuristics, and / or a hybrid approach. The prompt can be generated by executing a query against a prompt template library based on the environmental data and inputting it into a prompt template selected based on the input data.

[0056] In some embodiments, generating a prompt based on input data and environmental data may include determining a device configuration specific to the environment based on the environmental data, obtaining a prompt template from a prompt library based on the device configuration specific to the environment, and expanding the prompt template based on the input data to generate the prompt. The device configuration specific to the environment can describe the output type and output quality of each of the plurality of computing devices. The prompt library can include a plurality of different prompt templates associated with a plurality of different device configurations.

[0057] At 308, the computing system can process the prompt using a generation model to generate a model generation output. The model generation output can include a response to a query. In some embodiments, the model generation output may be generated to be provided to a particular computing device among the plurality of computing devices. The model generation output can be generated to be provided to a particular output component among the plurality of different output components. The generation model can generate output device instructions. The output device instructions can describe a particular computing device among the plurality of computing devices that provides the model generation output. The particular computing device can be associated with a particular output component. For example, the model generation output can be generated to be provided to a user via each output component of a particular computing device (e.g., a speaker of a smart surround system and / or a display screen of a smart TV).

[0058] In some embodiments, the generation model may be communicatively connected to a search engine via an application programming interface. Processing a prompt using the generation model to generate a model generation output may include generating an application programming interface call based on the prompt, determining a plurality of search results with a search engine based on the application programming interface call, and processing the plurality of search results using the generation model to generate a model generation output.

[0059] At 310, the computing system can send the model generation output to a particular computing device. The model generation output can be sent to a particular computing device based on an output device instruction. Sending the model generation output to a particular computing device may include executing an application programming interface call generated by the generation model. The model generation output may be sent to a virtual assistant device of the user environment, and the virtual assistant device may then control a particular computing device to provide the model generation output to a particular user.

[0060] In some embodiments, a plurality of computing devices can be connected via a cloud computing system. Each of the plurality of computing devices can register with the platform of the cloud computing system. Environmental data can be obtained in the cloud computing system. The model generation output can be sent via the cloud computing system. Alternatively and / or additionally, the plurality of computing devices can be communicatively connected across a network. A particular computing device among the plurality of computing devices can facilitate obtaining input data and sending the model generation output.

[0061] In some embodiments, processing a prompt using a generation model to generate a model generation output may include generating a plurality of model outputs. The plurality of model outputs can include a plurality of candidate responses. Additionally and / or alternatively, sending the model generation output to a particular computing device may include sending a first model output of the plurality of model outputs to a first computing device of the plurality of computing devices, and sending a second model output of the plurality of model outputs to a second computing device of the plurality of computing devices. The first model output can include visual data for display via a visual display. The second model output can include audio data for playback via a speaker component. In some embodiments, the first computing device can include a smart TV. Additionally and / or alternatively, the second computing device can include a smart speaker.

[0062] FIG. 4 represents a block diagram of an exemplary multi-device management system 400 according to an exemplary embodiment of the present disclosure. Specifically, the multi-device management system 400 can include a computing system including a plurality of user computing devices and a server computing system 420. The plurality of user computing devices and the server computing system 420 can be communicatively connected via a network 410.

[0063] The plurality of user computing devices can include a first computing device 402, a second computing device 404, a third computing device 406, and / or an nth computing device 408. The plurality of user computing devices can include a plurality of different computing devices that may have different input, processing, and / or output functions. For example, the first computing device 402 can include a smartphone equipped with an image sensor, an audio sensor, a touch sensor, a motion sensor, a speaker, a visual display, a tactile component, and / or a light. The second computing device 404 can include a smart wearable (e.g., a smartwatch) that may include a biometric sensor, a motion sensor, a touch sensor, a visual display, and / or a tactile component. The third computing device 406 can include a smart speaker that may include a high-quality speaker and / or a Bluetooth transmitter. The nth computing device 408 can include a smart TV, and the smart TV can include an infrared sensor receiver, a transmitter-receiver, a speaker, a visual display, and / or a plurality of input ports. The plurality of user computing devices can be connected to a network 410 via Ethernet, WiFi, and / or a Bluetooth connection to a companion device.

[0064] The plurality of user computing devices can be associated with an environment based on device registration, location, and / or proximity. Environmental data can be generated based on the plurality of user computing devices, which can include obtaining signals from the plurality of user computing devices.

[0065] The server computing system 420 can obtain input data and / or environmental data from a plurality of user computing devices via the network 410. The server computing system 420 can include a plurality of processing services for processing the input data and / or environmental data. For example, the server computing system 420 can include one or more generation models 422, one or more search engines 424, one or more prompt generation models 426, one or more interface models 428, and / or one or more other models. The one or more generation models 422 can be configured, trained, and / or adjusted to process a prompt associated with the input data and / or environmental data to generate a model-generated output in response to the input data, and can be configured to be a specific content type based on the environmental data. The one or more search engines 424 can be communicatively connected to obtain search results that can be used to understand the input data and / or environmental data and / or to respond to the input data and / or environmental data. The one or more prompt generation models 426 can be configured, trained, and / or adjusted to process the input data and / or environmental data to generate a prompt for the one or more generation models 422. The one or more interface models 428 can be configured, trained, and / or adjusted to process environmental data associated with a plurality of user computing devices and generate a plurality of respective interfaces for the plurality of user computing devices. Each of the plurality of interfaces can be generated based on the input and / or output functions of the devices within the environment.

[0066] Figure 5 represents a block diagram of an exemplary environment personalization system 500 according to an exemplary embodiment of the present disclosure. Specifically, the environment personalization system 500 can utilize information from multiple applications and / or platforms to provide personalized responses and / or personalized experiences. For example, data from a search assistant 502, a document assistant 504, an operating system assistant 506, a video player assistant 508, a browser assistant 510, and / or a chat interface assistant 512 can determine user context and / or generate queries and / or suggestions.

[0067] The search assistant 502 can be associated with a search application (and / or platform). The search assistant 502 can obtain the current query, session state, search history, trend data, and / or other search data.

[0068] The document assistant 504 can be associated with one or more document applications. The document assistant 504 can obtain data associated with the current document being viewed and / or edited, drive data (e.g., stored document information), sharing permission data, and / or other document data.

[0069] The operating system assistant 506 can be associated with the operating system of one or more computing devices. The operating system assistant 506 can obtain data associated with the content currently being provided for display, application data (e.g., app deep links, application programming interfaces (APIs), and / or usage data), and / or other operation data.

[0070] The video player assistant 508 can be associated with a video player application (and / or platform). The video player assistant 508 can be utilized to obtain video data of the currently displayed video, video storage, viewing history, followed media providers, subscriptions, comment history, and / or other video player data.

[0071] The browser assistant 510 can be associated with a browser application. The browser assistant 510 can obtain data associated with the current page, bookmark data, tab data, browsing history data, and / or other browser data.

[0072] The chat interface assistant 512 can be associated with one or more chatbots. The chat interface assistant can obtain session history data, response history, input history, topics, links, and / or other chatbot data.

[0073] The multi-device management system 514 can obtain data from multiple applications and / or platforms and provide the data to one or more other systems that may include a core model 516, a ground service model 518, and / or a personalization model 520.

[0074] For example, the core model 516 can be utilized for summarization, planning, and inference, and / or function calls. The ground service model 518 can be utilized for the use of tool libraries, the use of application programming interfaces (APIs) of external connectors, access to search results, and / or access to context and / or memory. The core model 516 and / or the ground service model 518 can interact with one or more other models and / or services via one or more cloud APIs. The output of the core model 516 and / or the ground service model 518 can be provided back to the multi-device management system 514 and then provided to the personalization model 520.

[0075] The personalization model 520 can process data to generate a personalized output that may include prediction prompts (and / or proposal prompts). The personalized output may be an extended prompt response based on user preferences, interactions, and / or user data associated with the device.

[0076] FIG. 6A shows a diagram of an exemplary interface according to an exemplary embodiment of the present disclosure. Specifically, the systems and methods disclosed herein can be utilized to generate and process environmental data associated with a computing device to generate a respective interface for each of a plurality of computing devices within the environment. Each of the plurality of respective interfaces can consider a plurality of other computing devices and may obfuscate and / or merge with existing interfaces. FIG. 6A shows three exemplary interfaces that may be associated with different devices in the environment and / or may be associated with different environments.

[0077] For example, the first interface 602 may be associated with a first computing device (e.g., a mobile device) in the environment, the second interface 604 may be associated with a second computing device (e.g., a smartwatch) in the environment, and the third interface 606 may be associated with a third computing device (e.g., a smart refrigerator) in the environment. Alternatively and / or additionally, the same computing device may have different interfaces based on being in different environments and / or different user contexts.

[0078] FIG. 6B shows a diagram of an exemplary image capture entry point according to an exemplary embodiment of the present disclosure. Specifically, one entry point for a search and / or assistant interface that utilizes multi-device input and / or output may be provided via user interface elements provided based on context.

[0079] For example, a user may be capturing an image 610 (e.g., capturing an image of a refrigerator). The image may be processed to determine (and / or identify) objects in the image (e.g., classify the object as a refrigerator of brand X model Y). Next, selectable user interface elements can be provided to the viewfinder 612. Next, the image and / or the identified item can be processed to generate a response that may include search results 614 associated with objects similar to the identified object. Other options may also be provided for interacting with the response. The other options may include an augmented reality experience, which may include rendering the object into an image of the user's residence 616.

[0080] FIG. 6C shows a diagram of an exemplary smart TV entry point according to an exemplary embodiment of the present disclosure. Specifically, one entry point can include providing a proposed search entry point indicator to one device, which can interact with other devices.

[0081] For example, the content provided for display on a first computing device 620 (e.g., a smart TV) can be determined to include features that a user may be interested in during a search. Accordingly, the proposed entry point user interface elements may be rendered across the content. Next, the user can select a selectable user interface element 622 on a second computing device (e.g., a mobile computing device) to view a model generation output 624. The model generation output can be provided via the first computing device 620, the second computing device, and / or a third computing device. The model generation output 624 can include search results, generated model outputs, one or more renderings, one or more proposals, and / or other options.

[0082] FIG. 6D shows a diagram of an exemplary calendar interface according to an exemplary embodiment of the present disclosure. In particular, FIG. 6D shows an exemplary calendar interface that includes a search window that can be defined based on calendar data. For example, the search window can include a text input field and two proposed actions. The proposed actions can include finding a room for a meeting (630) and / or proposing a time to block for focused time (632). The proposals can be generated based on information from multiple applications and / or multiple computing devices. The search window can be provided to the device on which the calendar application is open and / or to other devices.

[0083] Figure 6E shows a diagram of an exemplary video conferencing interface 634 according to an exemplary embodiment of the present disclosure. Specifically, Figure 6E shows an exemplary video conferencing interface 634 that includes a search window 636 that can be defined based on conference data. For example, the search window 636 can include a text input field and one or more proposed actions. The proposed actions can include taking notes about the conference (e.g., writing during the conference and / or opening a note application), setting a reminder, rescheduling the conference, and / or obtaining notes from similar conferences. The proposals can be generated based on information from multiple applications and / or multiple computing devices. The search window 636 can be provided on the device on which the video conferencing application is open and / or on other devices.

[0084] Figure 6F shows a diagram of an exemplary email interface according to an exemplary embodiment of the present disclosure. Specifically, Figure 6F shows an exemplary email interface 638 that includes a search window 640 that can be defined based on email data and / or conference data. For example, the search window 640 can include a text input field and one or more proposed actions. The proposed actions can include summarizing the notes and / or minutes of the conference, scheduling other conferences, obtaining information about the topics from the conference, drafting a follow-up email based on the context associated with the notes of the conference, and / or drafting an email based on email data and / or conference data. The proposals can be generated based on information from multiple applications and / or multiple computing devices. The search window 640 can be provided on the device on which the email application is open and / or on other devices.

[0085] FIG. 6G shows a diagram of an exemplary video player interface according to an exemplary embodiment of the present disclosure. Specifically, FIG. 6G shows an exemplary video player interface 650 that includes a search window 654 that can be defined based on video data and / or viewing history data. For example, the search window 654 can include a text input field and one or more proposed actions. The proposed actions can include obtaining a product list of similar products depicted in the displayed video 652, summarizing the video, obtaining entity labels of the displayed video, discovering similar videos, and / or obtaining additional information associated with the displayed video 652. For example, the user may request further information about the dress depicted in the displayed video 652. One or more frames can be segmented from the video and searched, which may include frame cropping. Alternatively and / or additionally, entity labels associated with the depicted frames can be obtained and searched. The search results can be provided for display in the search window 654 that is overlaid across the displayed video. The search results can include a product list, web links, and / or other data. The proposals can be generated based on information from multiple applications and / or multiple computing devices. The search window 654 can be provided to the device on which the video player application is open and / or to other devices.

[0086] FIG. 7 shows a flowchart diagram of an exemplary method for operating in accordance with an exemplary embodiment of the present disclosure. FIG. 7 shows steps that are executed in a particular order for purposes of illustration and explanation, but the methods of the present disclosure are not limited to the specifically shown order or arrangement. The various steps of method 700 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0087] At 702, the computing system can obtain input data. The input data can include a query associated with a particular user. The query can include a question associated with one or more topics, and the query can request an answer to the question. The input data can include a voice command obtained via a microphone, a text string input via a graphical keyboard, and / or a gesture obtained via a camera, an inertial measurement unit, and / or a touch sensor. In some embodiments, the input data can include data describing input obtained from a composite computing device in the environment (e.g., a voice command obtained via a smartphone's microphone, a gesture via a smartwatch's touch sensor, image data from a smart refrigerator, and / or a viewing history obtained from a smart TV).

[0088] At 704, the computing system can obtain environmental data. The environmental data can describe a plurality of computing devices in the user's environment. The plurality of computing devices can be associated with a plurality of different output components. The plurality of different output components can be associated with the respective output functions associated with the plurality of computing devices. In some embodiments, each of the respective output functions of the plurality can describe the output type and output quality available via the respective computing device. The plurality of output components can include a speaker associated with a first device and a visual display associated with a second device. The environmental data can include registered data associated with computing devices registered in the environment, WiFi routers, virtual assistant devices, user computing devices, and / or user profiles. In some embodiments, the environmental data can include an output hierarchy for a plurality of candidate output types, which can include a hierarchical representation of the performance capabilities of the plurality of computing devices for the plurality of different output types (e.g., visual display, audio output, tactile feedback, etc.).

[0089] At 706, the computing system can generate a prompt based on the input data and the environmental data. The prompt can include data that describes a query and device information associated with at least a subset of the plurality of computing devices. The prompt can be generated by processing the input data and the environmental data using a prompt generation model. The prompt generation model can include a language model (e.g., a generative language model (e.g., a large language model)). The prompt generation model can be trained and / or adjusted to generate a prompt based on understanding the intent of the query and determining the output type and / or available output types associated with the intent based on the environmental data. The prompt generation model can generate a prompt embedding that defines the output generation of the generation model.

[0090] In some embodiments, the computing system can determine that a particular output component is associated with the intent of the query. The prompt can be generated based on the particular output component associated with the intent of the query (e.g., a request for a song may be associated with a speaker output component, and a request to play a video may be associated with the visual display of a television).

[0091] Additionally and / or alternatively, the computing system can determine an output hierarchy based on environmental data and based on specification information of a plurality of different output components. The prompt can be generated based on the output hierarchy and the query. For example, the prompt can include text and / or embeds that define output generation based on the output capabilities of computing devices in the environment.

[0092] At 708, the computing system can process the prompt with a generation model to generate a model generation output and output device instructions. The model generation output can include a response to the query. In some embodiments, the model generation output can be generated to be provided to a particular output component among a plurality of different output components. The output device instructions can describe a particular computing device among a plurality of computing devices that provide the model generation output. The particular computing device can be associated with the particular output component. In some embodiments, the output device instructions can include an application programming interface call to send the model generation output to the particular computing device.

[0093] At 710, the computing system can send model generation output to a particular computing device based on the output device instruction. In some embodiments, the sending can be performed via signal transmission across a network. Alternatively and / or additionally, a notification indicating that an output configured for another device is available may be provided to the input computing device, and then the user may interact with the notification before the model generation output is sent to the particular computing device. In some embodiments, the generated model can generate a plurality of model generation outputs, and the computing system can send the plurality of model generation outputs to a plurality of different computing devices in the environment (e.g., a slide show can be sent to a smart TV for playback, a text document can be sent to an e-reader or a personal computing device (e.g., a smartphone or a tablet), an audio file can be sent to a smart speaker, and / or a scheduled set of color and brightness instructions can be sent to an RGB smart light setup).

[0094] FIG. 8 shows a flowchart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. Although FIG. 8 shows steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 800 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0095] At 802, the computing system can obtain environmental data. The environmental data can describe a plurality of computing devices in an environment associated with a particular user. The environmental data can describe specification information of the plurality of computing devices. In some embodiments, the plurality of computing devices can be determined based on registration of the device with a particular network, registration of the device with a particular user device, registration of the device in a particular user profile, proximity to the user device, and / or signal exchange between the device and the user device. The user device can be a smartphone, a tablet, a smartwatch, smart glasses, and / or other computing devices.

[0096] At 804, the computing system can process the environmental data to determine a plurality of respective input functions and a plurality of respective output functions associated with the plurality of computing devices. The plurality of respective input functions can be associated with candidate input types associated with the plurality of computing devices. In some embodiments, the plurality of respective output functions can be associated with candidate input types associated with the plurality of computing devices. The plurality of input functions and / or the plurality of output functions can be determined based on specification information (and / or component information) associated with the plurality of computing devices. In some embodiments, the plurality of input functions and / or the plurality of output functions can be determined using a machine-learned model, based on one or more searches, based on heuristics, and / or based on one or more other determination techniques. The plurality of input functions can describe the types of inputs available along with a range of quality associated with obtaining and / or generating the input type (e.g., the decibel range of a microphone and / or the resolution of a camera). The plurality of output functions can describe the types of outputs available along with a range of quality associated with providing the output of that output type (e.g., the quality of the audio (e.g., the volume range, the frequency range, etc.)).

[0097] At 806, a computing system can generate a respective interface for each of a plurality of computing devices based on a plurality of respective input functions and a plurality of respective output functions. Each of the plurality of respective interfaces can be specialized for a plurality of computing devices based on the plurality of respective input functions and the plurality of respective output functions. In some embodiments, each of the plurality of respective interfaces can include a plurality of device indicators indicating a plurality of computing devices within an environment associated with a particular user. The plurality of computing devices can be configured as a user-specific device ecosystem that is communicatively connected to receive inputs and provide outputs. Each of the plurality of respective interfaces can be configured to receive a particular input type and provide a particular output type based on the respective input function and the respective output function for a particular one of the plurality of computing devices.

[0098]

[0099] ​In some embodiments, a computing system can obtain user input via a first interface of a first computing device of a plurality of computing devices. The computing system can process the user input with a search engine to determine a plurality of search results, process the plurality of search results with a generation model to generate a model output, and provide the model output for display via a second interface of a second computing device of the plurality of computing devices.

[0100] Users are increasingly able to own complex devices that may or may not share an operating system, application, and / or platform. A base model (e.g., a large-scale base model that may include a generation model) can be used as a primary technology for users to interact with their devices. The system can provide a wide range of new experiences enabled by these basic models that can define input and output in an integrated way across all devices.

[0101] A multi-device framework that leverages base models can be used for a plurality of different companion tasks, which may include output definition, interface generation, input acquisition facilitation, and / or other tasks.

[0102] For example, a multi-device framework can be utilized by an ambient agent companion with a fluid interaction mechanism that depends on devices and surfaces. A user may own several devices that, in the presence of an ambient LLM companion, may use a fragmented ecosystem in some cases. An agent (e.g., a computing system including a multi-device framework) may include a base model (such as a base model for a generative model and / or a prompt generation model) that adapts an input / output interaction model based on device details. For example, a speaker may have an audio-only interface as the primary interaction model, a device with a touchless screen may have a hybrid text / audio interface, and / or a phone and / or smart wearable (e.g., a smartwatch, smart jacket, and / or smart glasses) may have a user interface that is generated ad hoc based on their form factor.

[0103] The interface can be generated to blend with the existing surface (e.g., the user interface can be minimized to one or more app elements (e.g., the agent interface can be changed to text and / or animation within a specific app interface). The generated interface can have multi-device / surface recognition elements, which can highlight, based on the user's proximity to the device, which of the devices the ambient companion is actively listening and / or monitoring. When an interaction with the companion on one device requires context from other devices and / or needs to execute an action on other devices, the user interface may include user interface elements that highlight that these other devices are being utilized and / or considered.

[0104] Additionally and / or alternatively, the system and method may define the output of the ambient companion model with respect to the form factor and characteristics of the device. For example, a user in an environment with a smart TV and smart home devices (e.g., virtual assistant devices) may provide an input (e.g., via a voice command) that issues a query (e.g., a query such as "What are the movies nominated for the Academy Award for Best Original Score?"). The LLM-enabled search engine / assistant / companion may be able to obtain additional input to the query, where the additional input may describe the details of the device (e.g., an input specifying the environment may include the screen and speakers along with their specific parameters). A higher quality smart speaker may be provided as part of the LLM prompt. Based on the prompt, the generative model (e.g., the LLM) may generate a response targeted at the smart speaker (e.g., "List the movies and play a sample of the soundtrack that won the award on the speaker"). On the other hand, another user issuing the same query may be in an environment with the same smart TV but a lower quality speaker, and thus may receive the response in the form of a visual display via the smart TV (e.g., the response may include something like "Stream a part of the trailer of the movie. You can listen to a part of the soundtrack that won the award in it later.").

[0105] In some embodiments, the system and method can be utilized to modify actions performed for a user after obtaining and / or determining the availability and context of nearby devices. For example, a user may be in an environment with a plurality of Internet of Things (IoT) devices with ambient companions, mobile computing devices (e.g., smartphones, tablets, smartwatches, etc.), automobiles, and / or laptops. Next, the user can ask the companion (via the input sensors of the computing device) about driving to a nearby recreation area. Next, the ambient companion can obtain and / or determine the input in a way that takes into account all the devices in the environment. Determining the input and / or providing a response can include responding to queries across multiple devices. For example, the companion can identify typical routes and / or destinations and generate a specific visualization that can be provided within the map application of one or more specific computing devices. A companion variant of the vehicle agent can perform local estimation and may determine that it is necessary to recharge the automobile to head to one or more of the destinations. The map application used by the main agent companion handling can process the query and output relevant charging stations highlighted in the illustrated route.

[0106] In some embodiments, the systems and methods disclosed herein can be utilized to generate and / or provide an interface across different computing devices that can be interconnected and / or have a similar style, layout, and / or semantics regardless of the manufacturer and / or operating system of the computing device. For example, the system and method may generate a native interface for different computing devices of different manufacturers and / or an interface that obscures differences between operating systems.

[0107] The system and method can include a cross - device representation of fluidity. The cross - device representation can be provided based on the output of a base model that can define input / output behavior on composite devices and surfaces (e.g., apps). The cross - device representation can be (1) a centralized base model executed in the cloud that is directly reported by device nodes, (2) a hybrid model in which the centralized base model uses device - specific context to perform inferences and functions with a distributed local large model that can adjust the resolution of higher - level tasks in the main model, and / or (3) a distributed architecture that can be implemented via a distributed architecture in which individual devices have companion versions and are grouped by the user who owns them all into the same physical space or the same logical unit. A certain level of inference may be required for the models to interact (e.g., > 10B params).

[0108] The hybrid approach can be implemented through multiple different configurations that can include compatibility with centralized and / or distributed architectures.

[0109] For example, a user may have several registered and interconnected devices. The characteristics of these devices may be known and / or determined. The degree of interoperability can vary and may include an API that can fully control the device when plugged in. The API can be utilized for the device's speaker to output audio. The API can be utilized to directly render elements in the operating system (and / or for complex functions within an app, etc.).

[0110] The API can be exposed in various configurations based on whether a local base model exists. If no base model exists, the raw functionality may be described in some accessible document libraries and / or alternatively, if a local model exists and is provided, a natural language interface may be available. The device may also provide functionality to interface with other external systems (e.g., sensors available to read the temperature of the environment, operate window blinds, and / or autonomously navigate through a house to perform certain tasks). Examples and information of the API can be fed into a basic model that powers an ambient companion agent in a way that can be used to define the generation of responses and the resolution of tasks on behalf of the user.

[0111] Prompt generation can include obtaining and / or generating zero-shot or few-shot prompts for the base model that are processed to understand how to use the device's API. If insufficient, the device can have a small dataset associated with them (e.g., about 1000 examples that can be used for prompt tuning and / or weighting of the base model to understand how to operate the device). These examples can include task -> decomposition using the API.

[0112] Once surfaced, devices can be placed together via a network, where each device is a node in the network. The edges of this network may be persistent (e.g., always on), and / or some weight quantifying the relatedness of two devices at a given point in time may be associated. For example, two devices may be in very close proximity to each other, and the weight between the two devices can quantify this information by having a smaller / larger numerical value. Also, two devices may share a particular context (e.g., they operate and / or display the same app to each other (e.g., the state can also be encoded dynamically at the graph edge via message passing)).

[0113] Available devices can be queried at inference time. Query execution can be supplied as a list with metadata and / or represented via a graph network, which can also be supplied to an ambient companion base model at inference time. Prior to model estimation, the base model may be fine-tuned to work with the device network topology and / or features. The graph network may be the modality in which the base model operates. In some embodiments, the network can be serialized directly and passed to a generative model (such as an LLM) as part of a prompt. If insufficient, a large-scale base model powering the ambient companion can be fine-tuned using examples of multi-device graph networks (e.g., even with over 1000 examples, each example having some task->step-by-step decomposition regarding how to use composite devices to better solve a task).

[0114] In some embodiments, the system and method can include prompting an LLM (or other generative model), including examples (e.g., "[Device context][Response][Metadata: This response is suitable for a smart speaker at breakfast]" and / or "[Device context][Response][Metadata: The response is visualized by rendering a UI with three checkboxes for each response on the phone]"). In an example of device breakdown, the prompt can include "[Device context][Response][Metadata: This response should be passed to the smart speaker and a small notification with a summary should be displayed on the mobile phone as the user may not be near the speaker]". Next, the user can face one of their devices and decide to interface with the ambient companion.

[0115] The input may particularly depend on the form factor details. For example, when the user pulls out the phone, the companion can be activated in a voice-only manner. When the phone is locked or when the user unlocks the phone, the companion can render itself in a UI that enables keyboard input. The modality depending on the device state can be defined with respect to the user's context in relation to nearby devices. For example, when the user's watch is available, voice input can be activated there instead.

[0116] The issued query can be resolved with the cooperation of all devices. For example, when the user types something like "I have to go to sleep. What time do I have to wake up?", the system can trigger the base model to act based on the query and / or context using all available contexts from the devices. Based on the determined action, the response can first process the surface that can access the work and / or calendar context and respond to the speaker with some actual proposed times.

[0117] The process can continue. For example, in a smart home, the lighting system can determine the intention of the user to go to bed and activate its local inference companion to set the lighting to slowly adjust to the bedtime routine set by the user. The information can be visually communicated and rendered as UI elements on the user's phone and / or watch to let the user know about the decision.

[0118] There may be proactive components in the ambient LLM companion. For example, some parts of the device can interact to determine whether user input is required. For example, a companion-enabled device with sensors that can operate autonomously can determine, based on the user's context, that input should be received and / or a prompt should be generated.

[0119] In some embodiments, a car companion can be built internally. The car companion can be prompted, configured, and / or trained to schedule the user to check the temperature, rainfall, etc. 15 minutes before the estimated departure time of the user, and in response to the prompt, configuration, and / or training, the car companion can notify the user to perform an action (e.g., put on X) and / or notify the user not to forget the raincoat.

[0120] The determined context can be received by the ambient companion, and the surface where the context is communicated to the user can be determined there. For example, the companion may decide to use a smart speaker (e.g., voice notification, "Even if it's raining, instead of walking to the parking lot, you can ask someone to pick you up by car"). Alternatively and / or additionally, the system can render the same information and / or actions on the user's watch based on the usage context of the device.

[0121] There may be unified behavior across all devices with respect to inputs and / or outputs that enable brand elements. The unified behavior may be possible through device-specific prompts and examples of some of the available shots. A device may come with 10 seconds of audio for five voice samples. The device may also be provided with five examples of UIs regarding how the assistant can be rendered on that particular device. The ambient base model can be defined to generate responses regarding those examples and a uniform output of an interface that the user may be familiar with from its manufacturer.

[0122] The hybrid framework approach may include devices that have some degree of autonomy but generally operate with each other through a centralized companion. In a variant, it may include the case where the device does not have the function of running a local base model and the entire decision can be made by the centralized companion. The variant approach may include continuous streaming of information. Other variants may include the case where the device has complete autonomy and there is no central companion. In an approach without a central companion, the network backbone may become more important and, based on simpler signals, it can guide which devices are interconnected at a given time to solve a given query.

[0123] The systems and methods disclosed herein can reveal new value to users by using generative artificial intelligence models to connect fragment systems and providing a single service level companion.

[0124] The functionality of the conversation interface (chatbot) can be extended beyond Q&A associated with a single device and / or application. For example, a user may be viewing a shopping blog and say "Show me reviews of the products recommended in this article / video". The system may need to understand what the product is in order to retrieve reviews in a shopping graph and summarize the reviews to make them easier to compare. In another example, a user may be viewing a recipe and say "Add the ingredients to the cart in the shopping app". Next, the system may need to extract the ingredients and call the shopping app API. In another example, a user may be viewing a travel vlog and say "Show me on a map the places mentioned in this article / video". The system may need to understand the places mentioned, extract the addresses, and may need to call an API to create pins in a custom map app. In another example, while viewing a document app, a user may request an essay about Abraham Lincoln by saying "Write a 1000-word essay about Lincoln that focuses on the Civil War and is understandable by a 5th grader to explain his hardships". Thus, the system may need to understand the important facts about Lincoln, a list of Lincoln's works, the content of those works, understand historical essays, and summarize all that information. In another example, while viewing a map app, a user may look at the map at a zoom level and say "Show me videos about things to do in this area". The system may need to understand the important things to do at that location, the places mentioned in the video, and may need to retrieve the appropriate video(s). In another example, while viewing content on a phone, a user may have the screen of a trail hiking app displayed and say "Show me restaurants near the trailhead". Thus, the system may need to understand the location of the trailhead on the screen being viewed and may retrieve local results for restaurants near that location.

[0125] Underlying dependencies are that the chatbots of these products may depend on common elements such as inference, grounding, acquisition, function calls, user state, personal preferences, etc., regardless of how they are integrated, and many of these are derived from search and / or knowledge graph services.

[0126] The systems and methods disclosed herein can utilize a chatbot large language model (LLM), a retrieval understanding large language model and a search engine, a cloud service large language model, and / or one or more other model - compliant systems. Although the architectures of the chatbot large language model (LLM), the retrieval understanding large language model, and the search engine may appear to be the same, they can be implemented independently due to differences in each block, such as different RLHF (reinforcement learning from human feedback) trainings, different planning algorithms, different sets of third - party plugins, different interfaces for searching the backend, and / or other differences. Different pipelines and / or systems can exchange information for estimation, understanding, and / or context determination.

[0127] Separate architectures may enable products to be developed independently and iteratively. However, by adhering to the optimal process, the success of individual products can be ensured, but due to the inability to adopt an overall service - oriented approach, there may ultimately be a disconnection between products.

[0128] FIG. 5 shows an exemplary architecture for building a layer cake of implementable chatbot functions and a set of instantiations of chatbot services that impart appropriate context and personality related to the products in which the services are instantiated.

[0129] For example, there may exist a general chatbot service built with a compatible LLM, fine-tuned for common tasks such as instruction following, having access to search the backend, accessing 1p / 3p APIs, and having functions such as calling functions. Examples of some services that can be utilized to obtain input and / or determine context include a search platform (e.g., accessing the user's search history to obtain concise factual data), a document application (e.g., accessing document files, emails, the currently displayed email / file, favorites, etc. to obtain redundant and original data), and / or a browser application (e.g., accessing the current tab, other tabs, bookmarks, history, etc. to obtain concise fact-based data).

[0130] In some embodiments, the system and method can include fine-tuning the underlying model differently for different products using use-case specific examples and content.

[0131] The systems and methods disclosed herein can include a companion model (e.g., a base model that can include a prompt generation model and / or a generation model).

[0132] Separate from any other user activity, by integrating AI and LLM capabilities as stand-alone entities within the user interface (UI), the system and method can ensure that they are permanently available for immediate access whenever the user needs them. This integration enables AI and LLM to remain open and be seamlessly connected to the user's actions as the user's actions transition between various applications and tasks.

[0133] This configuration can be useful for users who wish to quickly access information or perform tasks without leaving the currently used application. Additionally, a persistent system can be useful for users who want to use a companion with other products. For example, a user may search for information about a topic using a companion while also using a map application for navigation.

[0134] The system and method can provide two additional advantages. First, the system can eliminate the need to build existing solutions with potentially incompatible architectures, thereby avoiding potential integration difficulties. Second, by not changing existing solutions that users are already familiar with, it is possible to minimize the disruption caused by introducing significant changes that may not be optimal.

[0135] When a user is looking at a product in a store, the user can use an LLM to obtain information about the product (e.g., price, reviews, and specifications). The user can also use the camera to take a photo of the product and then use the LLM to search for similar products online. With a multifaceted system, the user may be able to obtain all the information the user needs to know about the product without having to leave the store.

[0136] Another example is that when a user is viewing a painting in a museum, the user may be able to use an LLM to obtain information about the artist, the painting, and the history of the painting. The user can also use the camera to take a photo of the painting and then use the LLM to search for other paintings by the same artist or paintings on a similar theme. With this system, the user can learn more about the artworks being viewed without relying on the information provided by the museum.

[0137] In some embodiments, a user may be looking for a product but may not be able to find what they are looking for. Next, the user can open their mobile phone and describe the product they are looking for, and the mobile phone will search for and find the product.

[0138] The system and method can utilize the ambient ecosystem by enabling the user to connect their other devices as companions. This connection can potentially allow the user to have a more seamless and integrated experience across devices. For example, a user can start a task on their mobile phone and then continue the task on their laptop without having to re-enter any information. The ambient ecosystem can create a more personalized and convenient experience for the user.

[0139] FIG. 9A represents a block diagram of an exemplary computing system 100 that performs multi-device output management according to an exemplary embodiment of the present disclosure. System 100 includes a user computing system 102, a server computing system 130, and / or a third computing system 150 communicatively coupled via a network 180.

[0140] The user computing system 102 can include any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0141] The user computing system 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing system 102 to perform operations.

[0142] In some embodiments, the user computing system 102 can store or include one or more machine-learned models 120. For example, the machine-learned model 120 can be various machine-learned models such as a neural network (e.g., a deep neural network), or other types of machine-learned models including non-linear models and / or linear models, or can include them. The neural network can include a feed-forward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks.

[0143] In some embodiments, one or more machine-learned models 120 may be received from a server computing system 130 via a network 180, stored in a user computing device memory 114, and then used by, or otherwise implemented by, one or more processors 112. In some embodiments, a user computing system 102 may implement a composite parallel instance of a single machine-learned model 120 (e.g., to perform parallel machine-learned model processing across composite instances of input data and / or detected features).

[0144] More specifically, one or more machine-learned models 120 may include one or more detection models, one or more classification models, one or more segmentation models, one or more augmentation models, one or more generation models, one or more natural language processing models, one or more optical property recognition models, and / or one or more other machine-learned models. One or more machine-learned models 120 can include one or more transformer models. One or more machine-learned models 120 may include one or more neural radiance field models, one or more diffusion models, and / or one or more autoregressive language models.

[0145] One or more pre-trained machine learning models 120 can be utilized to detect one or more object features. The detected features of the object may be classified and / or embedded. Next, a search can be performed using the classification and / or embedding to determine one or more search results. Alternatively and / or additionally, one or more detected features can be utilized to determine whether an indicator (e.g., a user interface element indicating the detected feature) should be provided to indicate that the feature has been detected. Next, the user may select the indicator to cause classification, embedding, and / or searching of the feature to be performed. In some embodiments, classification, embedding, and / or searching may be performed before the indicator is selected.

[0146] In some embodiments, one or more pre-trained machine learning models 120 can process image data, text data, audio data, and / or latent encoded data to generate output data that may include image data, text data, audio data, and / or latent encoded data. One or more pre-trained machine learning models 120 can perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, context determination, action prediction, image correction, image enhancement, text enhancement, sentiment analysis, object detection, error detection, inpainting, video stabilization, audio correction, audio enhancement, and / or data segmentation (e.g., mask-based segmentation).

[0147] Additionally or alternatively, one or more machine-learned models 140 may be included in, or otherwise stored and implemented by, a server computing system 130 that communicates with user computing system 102 according to a client-server relationship. For example, the machine-learned model 140 may be implemented by the server computing system 130 as part of a web service (such as a viewfinder service, a visual search service, an image processing service, an ambient computing service, and / or an overlay application service). Thus, one or more models 120 may be stored and implemented in the user computing system 102, and / or one or more models 140 may be stored and implemented in the server computing system 130.

[0148] The user computing system 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 can be a touch-sensitive component (such as a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (such as a finger or a stylus). The touch sensor-based component can function to implement a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.

[0149] In some embodiments, the user computing system can store and / or provide one or more user interfaces 124 that may be associated with one or more applications. The one or more user interfaces 124 can be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, augmented reality experiences, virtual reality experiences, and / or other data for display). The user interface 124 may be associated with one or more other computing systems (e.g., server computing system 130 and / or third-party computing system 150). The user interface 124 can include a viewfinder interface, a search interface, a generative model interface, a social media interface, and / or a media content gallery interface.

[0150] The user computing system 102 may include one or more sensors 126 and / or may receive data therefrom. The one or more sensors 126 may be housed in a housing component that houses one or more processors 112, memory 114, and / or one or more hardware components, and the hardware components may store and / or execute one or more software packets. The one or more sensors 126 can include one or more image sensors (e.g., cameras), one or more lidar sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biological sensors (e.g., heartbeat sensors, pulse sensors, retinal sensors, and / or fingerprint sensors), one or more infrared sensors, one or more position sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. Data associated with the user's environment (e.g., an image of the user's environment, a record of the environment, and / or the user's position) can be obtained using the one or more sensors.

[0151] The user computing system 102 may include, and / or be part of, a user computing device 104. The user computing device 104 may include a mobile computing device (e.g., a smartphone or tablet), a desktop computer, a laptop computer, smart wearable, and / or a smart appliance. Additionally and / or alternatively, the user computing system may obtain data from, and / or generate data using, one or more user computing devices 104. For example, the camera of a smartphone may be utilized to capture image data describing the environment, and / or the overlay application of the user computing device 104 may be utilized to track and / or process data provided to the user. Similarly, one or more sensors associated with a smart wearable may be utilized to obtain data regarding the user and / or the user's environment (e.g., image data may be obtained by a camera housed in the user's smart glasses). Additionally and / or alternatively, the data may be obtained and uploaded from other user devices specialized for data acquisition or generation.

[0152] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor, or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. The memory 134 can store data 136 and instructions 138 to be executed by the processor 132 to cause the server computing system 130 to perform operations.

[0153] In some embodiments, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. If the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0154] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the model 140 can be or otherwise include various machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer non-linear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. The exemplary model 140 is described with reference to FIG. 9B.

[0155] Additionally and / or alternatively, the server computing system 130 can include and / or be communicatively coupled to a search engine 142 that can be utilized to crawl one or more databases (and / or resources). The search engine 142 can process data from the user computing system 102, the server computing system 130, and / or the third-party computing system 150 to determine one or more search results associated with the input data. The search engine 142 can perform term-based search, label-based search, boolean-based search, image search, embedding-based search (e.g., nearest neighbor search), multimodal search, and / or one or more other search techniques.

[0156] Server computing system 130 may store and / or provide one or more user interfaces 144 for obtaining input data and / or providing output data to one or more users. The one or more user interfaces 144 can include one or more user interface elements, which can include input fields, navigation tools, content tips, selectable tiles, widgets, data display carousels, dynamic animations, information pop-ups, image magnification, text-to-speech, speech-to-text, augmented reality, virtual reality, feedback loops, and / or other interface elements.

[0157] User computing system 102 and / or server computing system 130 can train model 120 and / or 140 via interaction with a third-party computing system 150 communicatively coupled via network 180. The third-party computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130. Alternatively and / or additionally, the third-party computing system 150 may be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more contexts.

[0158] The third-party computing system 150 may include one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be a single processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the third-party computing system 150 to perform operations. In some embodiments, the third-party computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0159] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 180 can be performed via any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL).

[0160] The machine-learned models described herein can be used in a variety of tasks, applications, and / or use cases.

[0161] In some embodiments, the input to the machine-learned model(s) of the present disclosure may be image data. The machine-learned model(s) can process the image data to generate an output. By way of example, the machine-learned model(s) can process the image data to generate an image recognition output (e.g., recognition of the image data, potential embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an image segmentation output. As another example, the machine-learned model(s) can process the image data to generate an image classification output. As another example, the machine-learned model(s) can process the image data to generate an image data modification output (e.g., modification of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an encoded image data output (e.g., encoded and / or compressed representation of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) can process the image data to generate a prediction output.

[0162] In some embodiments, the input to the machine-learned model(s) of the present disclosure may be text or natural language data. The machine-learned model(s) can process the text or natural language data to generate an output. As another example, the machine-learned model(s) can process natural language data to generate a language encoding output. As another example, the machine-learned model(s) can process text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) can process text or natural language data to generate a translation output. As another example, the machine-learned model(s) can process text or natural language data to generate a classification output. As another example, the machine-learned model(s) can process text or natural language data to generate a text segmentation output. As another example, the machine-learned model(s) can process text or natural language data to generate a semantic intent output. As another example, the machine-learned model(s) can process text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine-learned model(s) can process text or natural language data to generate a prediction output.

[0163] In some embodiments, the input to the machine-learned model(s) of the present disclosure can be audio data. The machine-learned model(s) can process the audio data to generate an output. As an example, the machine-learned model(s) can process the audio data to generate an audio recognition output. As another example, the machine-learned model(s) can process the audio data to generate an audio translation output. As another example, the machine-learned model(s) can process the audio data to generate a latent embedding output. As another example, the machine-learned model(s) can process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate an upscaled audio output (e.g., audio data of higher quality than the input audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine-learned model(s) can process the audio data to generate a prediction output.

[0164] In some embodiments, the input to the machine-learned model(s) of the present disclosure may be sensor data. The machine-learned model(s) can process the sensor data to generate an output. As an example, the machine-learned model(s) can process the sensor data to generate a recognition output. As another example, the machine-learned model(s) can process the sensor data to generate a prediction output. As another example, the machine-learned model(s) can process the sensor data to generate a classification output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a visualization output. As another example, the machine-learned model(s) can process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) can process the sensor data to generate a detection output.

[0165] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, and the output is a set of scores, where each score corresponds to a different object class and represents the likelihood that one or more images depict an object belonging to that object class. The image processing task can be object detection, and the image processing output identifies one or more regions of one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, and the image processing output determines, for each pixel of one or more images, the likelihood of each category in a given set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, and the image processing output determines, for each pixel of one or more images, the respective depth value. As another example, the image processing task can be motion estimation, the network input includes a plurality of images, and the image processing output determines, for each pixel of one of the input images, the motion of the scene depicted in the pixels between the images in the network input.

[0166] The user computing system can include several applications (e.g., applications 1 - N). Each application can include its own respective machine learning library and one or more machine - learned models. For example, each application can include a machine - learned model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.

[0167] Each application can communicate with some other components of a computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, each application can communicate with each device component using an API (e.g., a public API). In some embodiments, the API used by each application is specific to that application.

[0168] The user computing system 102 can include several applications (e.g., applications 1 - N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some embodiments, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API across all applications).

[0169] The central intelligence layer can include several machine - learned models. For example, each machine - learned model (e.g., a model) can be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine - learned model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., a single model) for all of the applications. In some embodiments, the central intelligence layer is included within the operating system of the computing system 100 or, otherwise, implemented by the operating system of the computing system 100.

[0170] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository of data for the computing system 100. The central device data layer can communicate with some other components of the computing device, such as, for example, one or more sensors, a context manager, a device status component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0171] FIG. 9B represents a block diagram of an exemplary computing system 50 that performs multi-device output management according to an exemplary embodiment of the present disclosure. Specifically, the exemplary computing system 50 can obtain and / or generate one or more data sets that can be processed by a sensor processing system 60 and / or an output determination system 80, and can include one or more computing devices 52 that can be utilized to provide feedback to a user regarding information about the characteristics of the one or more obtained data sets. The one or more data sets can include image data, text data, audio data, multimodal data, potentially encoded data, and the like. The one or more data sets can be obtained via one or more sensors associated with one or more computing devices 52 (e.g., one or more sensors of the computing device 52). Additionally and / or alternatively, the one or more data sets can be stored data and / or acquired data (e.g., data obtained from a web resource). For example, an image, text, and / or other content item can be interacted with by a user. Next, using the content interacted with the content item, one or more decisions can be generated.

[0172] One or more computing devices 52 can obtain and / or generate one or more data sets based on image capture, sensor tracking, searching of data storage, downloading of content (e.g., downloading an image or other content item from a web resource over the Internet), and / or via one or more other techniques. The one or more data sets can be processed by a sensor processing system 60. The sensor processing system 60 can execute one or more processing techniques using one or more machine-learned models, one or more search engines, and / or one or more other processing techniques. The one or more processing techniques can be executed in any combination and / or individually. The one or more processing techniques can be executed sequentially and / or in parallel. In particular, the one or more data sets can be processed by a context determination block 62, which can determine the context associated with one or more content items. The context determination block 62 can identify and / or process metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data to determine a particular context associated with a user. The context can be associated with an event, a determined trend, a particular action, a particular type of data, a particular environment, and / or other context associated with the user and / or the searched for or obtained data.

[0173] The sensor processing system 60 may include an image preprocessing block 64. The image preprocessing block 64 can be used to adjust one or more values of the acquired image and / or the received image to create an image to be processed by one or more trained machine learning models and / or one or more search engines 74. The image preprocessing block 64 can resize the image, adjust the saturation value, adjust the resolution, remove and / or add metadata, and / or perform one or more other operations.

[0174] In some embodiments, the sensor processing system 60 can include one or more trained machine learning models that can include a detection model 66, a segmentation model 68, a classification model 70, an embedding model 72, and / or one or more other trained machine learning models. For example, the sensor processing system 60 can include one or more detection models 66 that can be used to detect specific features of a processed data set. Specifically, one or more images can be processed with one or more detection models 66 to generate one or more bounding boxes associated with the features detected in the one or more images.

[0175] Additionally and / or alternatively, one or more segmentation models 68 can be used to segment one or more parts of a data set from one or more data sets. For example, one or more segmentation models 68 can use one or more segmentation masks (e.g., one or more segmentation masks generated manually and / or based on one or more bounding boxes) to segment a part of an image, a part of an audio file, and / or a part of text. Segmentation can include separating one or more detected objects and / or removing one or more detected objects from an image.

[0176] One or more classification models 70 can be utilized to process image data, text data, audio data, latent encoded data, multimodal data, and / or other data to generate one or more classifications. The one or more classification models 70 can include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. The one or more classification models 70 can process the data to determine one or more classifications.

[0177] In some embodiments, the data can be processed by one or more embedding models 72 to generate one or more embeddings. For example, one or more images can be processed by one or more embedding models 72 to generate embeddings of the one or more images in an embedding space. The embeddings of the one or more images can be associated with one or more image features of the one or more images. In some embodiments, the one or more embedding models 72 may be configured to process multimodal data to generate multimodal embeddings. The one or more embeddings can be utilized for classification, search, and / or learning of embedding space distribution.

[0178] The sensor processing system 60 may include one or more search engines 74 that can be utilized to perform one or more searches. The one or more search engines 74 may crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more dedicated databases, and / or one or more general databases) to determine one or more search results. The one or more search engines 74 may perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), metadata database search, multimodal search, web resource search, image search, text search, and / or application search.

[0179] Additionally and / or alternatively, the sensor processing system 60 may include one or more multimodal processing blocks 76 that can be utilized to assist in the processing of multimodal data. The one or more multimodal processing blocks 76 may include generating multimodal queries and / or multimodal embeddings that are processed by one or more machine-learned models and / or one or more search engines 74.

[0180] The output(s) of the sensor processing system 60 can then be processed by an output determination system 80 to determine one or more outputs to provide to the user. The output determination system 80 may include heuristic-based determination, machine-learned model-based determination, user-selection-based determination, and / or context-based determination.

[0181] The output determination system 80 may determine a method and / or location for providing one or more search results in the search result interface 82. Additionally and / or alternatively, the output determination system 80 may determine a method and / or location for providing one or more machine-learned model outputs in the machine-learned model output interface 84. In some embodiments, one or more search results and / or one or more machine-learned model outputs may be provided for display via one or more user interface elements. The one or more user interface elements may be overlaid on the displayed data. For example, one or more detection indicators may be overlaid across the detected objects in the viewfinder. The one or more user interface elements may be selectable to perform one or more additional searches and / or one or more additional machine-learned model processes. In some embodiments, the user interface elements may be provided as user interface elements dedicated to a particular application and / or may be provided uniformly across different applications. The one or more user interface elements can include pop-up displays, interface overlays, interface styles and / or chips, carousel interfaces, audio feedback, animations, interactive widgets, and / or other user interface elements.

[0182] Additionally and / or alternatively, an augmented reality experience and / or a virtual reality experience 86 may be generated and / or provided using data associated with the output(s) of the sensor processing system 60. For example, one or more acquired data sets may be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which can then be used to provide an augmented reality experience and / or a virtual reality experience 86 to the user. The augmented reality experience may render information associated with the environment in each environment. Additionally and / or alternatively, objects associated with the processed data set(s) may be rendered within the user's environment and / or virtual environment. Generation of the rendering data set may include training one or more neural radiance field models to learn a three-dimensional representation of one or more objects.

[0183] In some embodiments, one or more action prompts 88 may be determined based on the output(s) of the sensor processing system 60. For example, a search prompt, a purchase prompt, a generation prompt, a reservation prompt, a call prompt, a redirect prompt, and / or one or more other prompts may be determined to be associated with the output(s) of the sensor processing system 60. The one or more action prompts 88 may then be provided to the user via one or more selectable user interface elements. In response to selection of the one or more selectable user interface elements, the respective action of each action prompt may be executed (e.g., a search may be executed, a purchase application programming interface may be utilized, and / or other applications may be launched).

[0184] In some embodiments, one or more datasets and / or outputs of the sensor processing system 60 can be processed by one or more generative models 90 to generate model-generated content items, which can then be provided to the user. Generation can be prompted based on user selection and / or can be executed automatically (e.g., automatically executed based on one or more conditions, where the conditions can be associated with a quantity of search results below an identified threshold).

[0185] One or more generative models 90 can include a language model (e.g., a large language model and / or a vision-language model), an image generation model (e.g., a text-to-image generation model and / or an image enhancement model), an audio generation model, a video generation model, a graph generation model, and / or other data generation models (e.g., other content generation models). One or more generative models 90 can include one or more transformer models, one or more convolutional neural networks, one or more regression neural networks, one or more feed-forward neural networks, one or more generative adversarial networks, one or more self-attention models, one or more embedding models, one or more encoders, one or more decoders, and / or one or more other models. In some embodiments, one or more generative models 90 can include one or more autoregressive models (e.g., machine-learned models trained to generate predicted values based on previous behavioral data) and / or one or more diffusion models (e.g., machine-learned models trained to generate predicted data based on generating and processing distribution data associated with input data).

[0186] One or more generative models 90 may be trained to process input data and generate model-generated content items, which may include a plurality of predicted words, pixels, signals, and / or other data. The model-generated content items may include novel content items that are not the same as any existing works. One or more generative models 90 can utilize learned representations, sequences, and / or probability distributions to generate content items, which can include phrases, storylines, settings, objects, characters, beats, lyrics, and / or other aspects not included in existing content items.

[0187] One or more generative models 90 may include a vision-language model. The vision-language model can be trained, tuned, and / or configured to process image data and / or text data to generate natural language output. The vision-language model can utilize a pre-trained large language model (e.g., a large autoregressive language model) with one or more encoders (e.g., one or more image encoders and / or one or more text encoders) to provide detailed natural language output that emulates human-made natural language.

[0188] The vision-language model may be used for zero-shot image classification, few-shot image classification, image captioning, multimodal query distillation, multimodal question answering, and / or may be tuned and / or trained for a plurality of different tasks. The vision-language model can perform visual question answering, image caption generation, feature detection (e.g., content monitoring (such as inappropriate content)), object detection, scene recognition, and / or other tasks.

[0189] A vision-language model can utilize a pre-trained language model and then adjust the language model for multi-modality. The training and / or adjustment of the vision-language model can include image-text matching, masked language modeling, multi-modal fusion with cross-attention, contrastive learning, prefix language model training, and / or other training techniques. For example, the vision-language model can be trained to process an image to generate predictive text similar to ground truth text data (e.g., the ground truth caption of the image). In some embodiments, the vision-language model can be trained to replace masked tokens of a natural language template with text tokens that describe features depicted in the input image. Alternatively and / or additionally, training, adjustment, and / or model estimation can include multi-layer concatenation of visual and text embedding features. In some embodiments, the vision-language model can be trained and / or adjusted by jointly learning the generation of image embeddings and text embeddings, which can include training and / or adjusting a system that maps the embeddings to a shared feature embedding space that maps text features and image features to the shared embedding space. Joint training can include parallel embeddings of image-text pairs and / or can include triplet training. In some embodiments, an image can be utilized and / or processed as a prefix of the language model.

[0190] The output determination system 80 can process the output(s) of one or more data sets and / or the sensor processing system 60 using the data augmentation block 92 to generate augmented data. For example, one or more images can be processed by the data augmentation block 92 to generate one or more augmented images. Data augmentation can include data modification, data cropping, removal of one or more features, addition of one or more features, resolution adjustment, illumination adjustment, saturation adjustment, and / or other augmentations.

[0191] In some embodiments, one or more datasets and / or outputs(s) of the sensor processing system 60 may be stored based on a determination of the data storage block 94.

[0192] Next, the output(s) of the output determination system 80 may be provided to the user via one or more output components of the user computing device 52. For example, one or more user interface elements associated with the one or more outputs may be provided for display via the visual display of the user computing device 52.

[0193] The process may be executed iteratively and / or continuously. One or more user inputs to the provided user interface elements may define and / or affect a continuous processing loop.

[0194] The techniques described herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions being performed and the information being sent to and from such systems. The inherent flexibility of computer-based systems enables a wide variety of possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes discussed herein can be implemented using a single device or component, or a composite device or component operating in combination. The database and application can be implemented in a single system or distributed across multiple systems. The distributed components can operate sequentially or in parallel.

[0195] Although the subject matter of the present disclosure has been described in detail with respect to its various specific and exemplary embodiments, each example is provided for illustrative purposes and is not intended to limit the present disclosure. Those skilled in the art, upon understanding the foregoing, can readily make changes and modifications to such embodiments and readily create equivalents. Accordingly, the present disclosure does not exclude including such corrections, changes, and / or additions to the subject matter as would be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment can be used with other embodiments to create still further embodiments. Accordingly, the present disclosure is intended to cover such changes, modifications, and equivalents.

Claims

1. 1. A computing system for determining an output device for providing a query response, comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations; and wherein the operation comprises: obtaining input data, the input data including a query associated with a particular user; acquiring environmental data describing a plurality of computing devices in the user's environment, the plurality of computing devices being associated with a plurality of different output components; generating a prompt based on the input data and the environmental data, the prompt including data describing the query and device information associated with at least a subset of the plurality of computing devices; processing the prompts with a generative model to generate a model-generated output, the model-generated output including a response to the query, the model-generated output being generated to be provided to a particular computing device of the plurality of computing devices; transmitting the model generation output to the particular computing device; Including, the system.

2. the model generated output is generated to be provided to a particular output component of the plurality of different output components; the generative model generates output device instructions, the output device instructions describing a particular computing device of the plurality of computing devices that provides the model-generated output, the particular computing device being associated with the particular output component; The system of claim 1 , wherein the model generation output is sent to the particular computing device based on the output device instructions.

3. Processing the prompts with the generative model to generate the model-generated outputs includes: generating a plurality of model outputs, the plurality of model outputs including a plurality of candidate responses; Transmitting the model generation output to the particular computing device comprises: transmitting a first model output of the plurality of model outputs to a first computing device of the plurality of computing devices; transmitting a second model output of the plurality of model outputs to a second computing device of the plurality of computing devices; The system of claim 1 , comprising:

4. 4. The system of claim 3, wherein the first model output comprises visual data for display via a visual display and the second model output comprises audio data for playback via a speaker component.

5. The system of claim 4 , wherein the first computing device comprises a smart television and the second computing device comprises a smart speaker.

6. Generating the prompt based on the input data and the environmental data includes: determining an environment specific device configuration based on the environmental data; retrieving a prompt template from a prompt library based on the environment specific device configuration; expanding the prompt template based on the input data to generate the prompt; The system of claim 1 , comprising:

7. 7. The system of claim 6, wherein the environment-specific device configuration describes an output type and a respective output quality of each of the plurality of computing devices, and the prompt library includes a plurality of different prompt templates associated with a plurality of different device configurations.

8. 2. The system of claim 1, wherein the plurality of computing devices are connected via a cloud computing system, each of the plurality of computing devices is registered with a platform of the cloud computing system, the environmental data is acquired using the cloud computing system, and the model generation output is transmitted via the cloud computing system.

9. 2. The system of claim 1, wherein the plurality of computing devices are located in close proximity to each other computing device in the plurality of computing devices, the plurality of computing devices are communicatively connected via a local network, and particular computing devices of the plurality of computing devices facilitate obtaining input data and transmitting model generation output.

10. the generative model is communicatively coupled to a search engine via an application programming interface; Processing the prompts with the generative model to generate the model-generated outputs includes: generating an application programming interface call based on the prompt; determining a plurality of search results using the search engine based on the application programming interface calls; processing the plurality of search results with the generative model to generate the model generated output; The system of claim 1 , comprising:

11. 1. A computer-implemented method, comprising: obtaining, by a computing system including one or more processors, input data, the input data including a query associated with a particular user; acquiring, by the computing system, environmental data describing a plurality of computing devices in the user's environment, the plurality of computing devices being associated with a plurality of different output components; generating, by the computing system, a prompt based on the input data and the environmental data, the prompt including data describing the query and device information associated with at least a subset of the plurality of computing devices; processing, by the computing system, the prompt with a generative model to generate a model-generated output and output device instructions, the model-generated output including a response to the query, the model-generated output being generated to be provided with a particular output component of the plurality of different output components, the output device instructions describing a particular computing device of the plurality of computing devices that provides the model-generated output, the particular computing device being associated with the particular output component; transmitting, by the computing system, the model generation output to the particular computing device based on the output device instructions; A method comprising:

12. the plurality of different output components are associated with a plurality of respective output capabilities associated with the plurality of computing devices; The method of claim 11 , wherein each of the plurality of respective output capabilities describes an output type and output quality available via a respective computing device.

13. The method of claim 11 , wherein the output device instructions include an application programming interface call for sending the model generation output to the particular computing device.

14. The method of claim 11 , wherein the plurality of different output components comprises a speaker associated with a first device and a visual display associated with a second device.

15. determining, by the computing system, that the particular output component is associated with an intent of the query; The method of claim 11 , wherein the prompt is generated based on the particular output component associated with the intent of the query.

16. determining, by the computing system, an output hierarchy based on specification information of the plurality of different output components based on the environmental data; The method of claim 11 , wherein the prompt is generated based on the output hierarchy and the query.

17. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: obtaining environmental data, the environmental data describing a plurality of computing devices within an environment associated with a particular user; processing the environmental data to determine a plurality of respective input functions and a plurality of respective output functions associated with the plurality of computing devices, the plurality of respective input functions being associated with candidate input types associated with the plurality of computing devices and the plurality of respective output functions being associated with candidate output types associated with the plurality of computing devices; generating a plurality of respective interfaces for the plurality of computing devices based on the plurality of respective input capabilities and the plurality of respective output capabilities, the plurality of respective interfaces being specialized for the plurality of computing devices based on the plurality of respective input capabilities and the plurality of respective output capabilities; providing said plurality of respective interfaces to said plurality of computing devices; [0023] In one or more non-transitory computer readable media,

18. 20. The one or more non-transitory computer-readable media of claim 17, wherein the plurality of respective interfaces comprises a plurality of device indicators that indicate the plurality of computing devices in the environment associated with the particular user, the plurality of computing devices being configured as a user-specific device ecosystem communicatively connected to receive input and provide output.

19. 20. The one or more non-transitory computer-readable media of claim 17, wherein each of the plurality of respective interfaces is configured to receive a particular input type and provide a particular output type based on a respective input capability and a respective output capability for the particular computing device of the plurality of computing devices.

20. The operation, obtaining user input via a first interface of a first computing device of the plurality of computing devices; processing the user input with a search engine to determine a plurality of search results; processing the plurality of search results with a generative model to generate a model output; providing the model output for display via a second interface of a second computing device of the plurality of computing devices; 20. The one or more non-transitory computer-readable media of claim 17, further comprising:

Citation Information

Patent Citations

  • Information processing system, information processing method and program

    JP2024171311A

  • Method and apparatus for synthesizing adaptive data visualizations

    US20190378506A1

  • Alternate response generation

    US20200184992A1