Environmental multi-device framework for agent companions

CN119493893BActive Publication Date: 2026-08-28GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411625251.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-11-14
Publication Date
2026-08-28
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

无论个人是试图了解他们面前的对象是什么,还是试图确定在其他什么地方可找到该对象,和/或试图确定互联网上的图像是从哪里捕获的,单独进行文本搜索都很困难

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119493893B_ABST
    Figure CN119493893B_ABST
Patent Text Reader

Abstract

This application relates to an environmental multi-device framework for an agent companion. Systems and methods for generating and providing output in a multi-device system can include providing dynamic response generation and display with environment-based prompt generation and generative model response generation. The systems and methods can obtain input data associated with one or more computing devices within an environment, can obtain environmental data describing a plurality of computing devices within the environment, and can generate a prompt based on the input data and the environmental data. The prompt can be processed with a generative model to generate a model-generated output. The model-generated output can then be sent to a particular computing device of the plurality of computing devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority requirements

[0002] This application is based on and claims priority to U.S. nonprovisional application 18 / 390,768, filed on December 20, 2023, which is incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to generating and providing output in a multi-device system. More specifically, this disclosure relates to regulating output generation based on devices in an environment and determining which computing device provides the output. Background Technology

[0004] Computing devices are found to be ubiquitous in users' living rooms, bedrooms, studies, offices, and / or other environments. A multi-device environment provides users with access to multiple computing devices for interaction, including smart TVs, smart speakers, smart appliances, virtual assistant devices, tablets, smart wearables, smartphones, and / or other computing devices available throughout the user's environment. However, due to a lack of interconnectivity between devices, their capabilities may not be fully utilized. For example, a user might be performing a search on their smartphone, which could result in video search results playing on the smartphone, even though the smart TV is just feet away.

[0005] Understanding the world as a whole can be difficult. Whether an individual is trying to understand what an object in front of them is, trying to determine where else to find that object, and / or trying to determine where an image on the internet was captured from, conducting a text search alone is challenging. Specifically, users may struggle to determine which words to use. Furthermore, their vocabulary may not be descriptive or rich enough to produce the desired results. Summary of the Invention

[0006] Various aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.

[0007] One example aspect of this disclosure relates to a computing system for determining output means for providing a query response. The system may include one or more processors and one or more non-transitory computer-readable media that jointly store instructions that, when executed by the one or more processors, cause the computing system to perform operations. These operations may include obtaining input data. The input data may include a query associated with a specific user. The operations may include obtaining environmental data. The environmental data may describe multiple computing devices in the user's environment. In some implementations, the multiple computing devices may be associated with multiple different output components. The operations may include generating a prompt based on the input data and the environmental data. The prompt may include data describing the query and device information associated with a subset of at least a plurality of computing devices. The operations may include processing the prompt with a generative model to generate model-generated output. The model-generated output may include a response to the query. In some implementations, the model-generated output may be generated to be provided by a specific computing device among the multiple computing devices. The operations may include sending the model-generated output to the specific computing device.

[0008] In some implementations, the output of the model generation can be generated to be provided by a specific output component among multiple different output components. The generative model can generate output device instructions. These output device instructions can describe a specific computing device among multiple computing devices used to provide the model-generated output. A specific computing device can be associated with a specific output component. The model-generated output can be sent to the specific computing device based on the output device instructions.

[0009] In some implementations, processing cues with a generative model to generate model-generated outputs may include generating multiple model outputs. These multiple model outputs may include multiple candidate responses. Sending the model-generated outputs to a specific computing device may include: sending a first model output from the multiple model outputs to a first computing device among the multiple computing devices; and sending a second model output from the multiple model outputs to a second computing device among the multiple computing devices. The first model output may include visual data for display via a visual display. The second model output may include audio data for playback via a speaker assembly. In some implementations, the first computing device may include a smart TV, and the second computing device includes a smart speaker.

[0010] In some implementations, generating prompts based on input and environmental data may include: determining an environment-specific device configuration based on the environmental data; obtaining a prompt template from a prompt library based on the environment-specific device configuration; and enhancing the prompt template based on the input data to generate a prompt. The environment-specific device configuration may describe the corresponding output type and corresponding output quality of multiple computing devices. The prompt library may include multiple different prompt templates associated with multiple different device configurations.

[0011] In some implementations, multiple computing devices can be connected via a cloud computing system. Each of the multiple computing devices can register with the cloud computing system's platform. Environmental data can be obtained using the cloud computing system. The output of the model generation can be sent via the cloud computing system. In some implementations, the multiple computing devices can be located close to each other within the multiple computing devices. The multiple computing devices can be communicatively connected via a local network. Specific computing devices among the multiple computing devices can facilitate the acquisition of input data and the transmission of the model generation output.

[0012] In some implementations, the generative model can be communicatively connected to the search engine via an application programming interface (API). Processing prompts with the generative model to generate model-generated output may include: generating API calls based on prompts; determining multiple search results using the search engine based on API calls; and processing multiple search results with the generative model to generate model-generated output.

[0013] Another example aspect of this disclosure relates to a computer-implemented method. The method may include obtaining input data via a computing system comprising one or more processors. The input data may include a query associated with a specific user. The method may include obtaining environmental data via the computing system. The environmental data may describe multiple computing devices in the user's environment. The multiple computing devices may be associated with multiple different output components. The method may include generating a prompt via the computing system based on the input data and the environmental data. The prompt may include data describing the query and device information associated with a subset of at least multiple computing devices. The method may include processing the prompt with a generative model via the computing system to generate model-generated output and output device instructions. The model-generated output may include a response to the query. In some implementations, the model-generated output may be generated to be provided by a specific output component among multiple different output components. The output device instructions may describe a specific computing device among multiple computing devices used to provide the model-generated output. The specific computing device may be associated with a specific output component. The method may include sending the model-generated output to the specific computing device via the computing system based on the output device instructions.

[0014] In some implementations, multiple distinct output components may be associated with multiple corresponding output capabilities linked to multiple computing devices. Each of the multiple corresponding output capabilities may describe the type and quality of output available via the corresponding computing device. Output device instructions may include application programming interface calls for sending the model-generated output to a specific computing device. In some implementations, the multiple distinct output components may include a speaker associated with a first device and a visual display associated with a second device. The method may include determining, via a computing system, the association of a particular output component with the intent of a query. A prompt may be generated based on the association of a particular output component with the intent of the query.

[0015] In some implementations, the method may include determining an output hierarchy based on specification information from multiple different output components by computing the system and using environmental data. Hints may be generated based on the output hierarchy and queries.

[0016] Another exemplary aspect of this disclosure relates to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. These operations may include obtaining environmental data. The environmental data may describe multiple computing devices within an environment associated with a particular user. The operations may include processing the environmental data to determine multiple corresponding input capabilities and multiple corresponding output capabilities associated with the multiple computing devices. The multiple corresponding input capabilities may be associated with candidate input types associated with the multiple computing devices. The multiple corresponding output capabilities may be associated with candidate output types associated with the multiple computing devices. The operations may include generating multiple corresponding interfaces for the multiple computing devices based on the multiple corresponding input capabilities and multiple corresponding output capabilities. Based on the multiple computing devices being based on the multiple corresponding input capabilities and multiple corresponding output capabilities, the multiple corresponding interfaces may be specifically designed for the multiple computing devices. The operations may include providing the multiple corresponding interfaces to the multiple computing devices.

[0017] In some implementations, the multiple corresponding interfaces may include multiple device indicators that indicate multiple computing devices within an environment associated with a specific user. The multiple computing devices may be configured as a user-specific device ecosystem, communicatively connected to receive input and provide output. Each of the multiple corresponding interfaces may be configured to receive a specific input type and provide a specific output type based on the specific input and output capabilities of the specific computing device among the multiple computing devices.

[0018] In some implementations, the operation may include: obtaining user input via a first interface of a first computing device among a plurality of computing devices; processing the user input with a search engine to determine a plurality of search results; processing the plurality of search results with a generative model to generate model output; and providing the model output for display via a second interface of a second computing device among a plurality of computing devices.

[0019] Other aspects of this disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0020] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description

[0021] Referring to the accompanying drawings, a detailed discussion of embodiments is set forth in this specification for those skilled in the art, in which:

[0022] Figure 1 A block diagram of an example generative model system according to an example embodiment of the present disclosure is depicted.

[0023] Figure 2 A block diagram of an example response generation system according to an example embodiment of the present disclosure is depicted.

[0024] Figure 3 A flowchart depicts an example method for generating an execution response according to an example embodiment of the present disclosure.

[0025] Figure 4 A block diagram of an example multi-device management system according to an example embodiment of the present disclosure is depicted.

[0026] Figure 5 A block diagram of an example environment personalization system according to an example embodiment of the present disclosure is depicted.

[0027] Figure 6A An illustration depicts an example interface according to an example embodiment of the present disclosure.

[0028] Figure 6B An illustration depicts an example image capture entry point according to an example embodiment of the present disclosure.

[0029] Figure 6C An illustration depicts an example smart TV entry point according to an example embodiment of the present disclosure.

[0030] Figure 6D An illustration depicts an example schedule interface according to an example embodiment of the present disclosure.

[0031] Figure 6E An illustration depicts an example video conferencing interface according to an example embodiment of the present disclosure.

[0032] Figure 6F An illustration depicts an example email interface according to an example embodiment of the present disclosure.

[0033] Figure 6G An illustration depicts an example video player interface according to an example embodiment of the present disclosure.

[0034] Figure 7 A flowchart is depicted illustrating an example method for generating execution output and selecting routes according to an example embodiment of this disclosure.

[0035] Figure 8 A flowchart depicts an example method for generating an execution interface according to an example embodiment of this disclosure.

[0036] Figure 9A A block diagram of an example computing system for performing multi-device output management according to an example embodiment of the present disclosure is depicted.

[0037] Figure 9B A block diagram of an example computing system for performing multi-device output management according to an example embodiment of the present disclosure is depicted.

[0038] The repeated reference numerals across multiple figures are intended to identify the same features in various implementations. Detailed Implementation

[0039] Generally, this disclosure relates to a multi-device framework for managing input acquisition and output generation for an environment comprising multiple computing devices. Specifically, the systems and methods disclosed herein can utilize environment-based prompt generation and generative modeling processing to generate and / or provide outputs that can be generated based on computing devices available in the environment. For example, the multi-device framework may include a computing system that facilitates the acquisition of input data and environmental data associated with multiple computing devices, and generates one or more model-generated outputs configured to be provided by one or more specific computing devices among the multiple computing devices. The computing system can be used to leverage the diverse input and / or output capabilities of different computing devices within the environment, which may include automatically acquiring different types of input from different devices and / or providing different types of output via different devices (e.g., receiving voice commands via a smartphone, while providing output via a visual display of a television and / or speakers of an audio system).

[0040] Environmental data associated with multiple computing devices within the environment can be used for cue generation, environmental understanding, and / or interface generation. Environmental data can be processed to generate cuees that can adjust the output of a generative model to be compatible with and / or optimized for display and / or playback in the user's environment. Cue generation may include obtaining specific cue templates from a cue library based on the computing devices associated with the environment, and / or may include cue generation by processing input data and / or environmental data using a machine learning model. Additionally and / or optionally, environmental data may be processed to determine which computing devices associated with the environment are used for specific types of input acquisition and / or specific types of output playback. In some implementations, environmental data may be processed to understand the computing devices within the environment and to generate appropriate interfaces for multiple computing devices based on the determined input and / or output capabilities of the computing devices associated with the environment.

[0041] An ambient multi-device framework can be used to provide an immersive virtual assistant that receives input from and outputs to multiple different computing devices. The ambient multi-device framework can obtain queries and / or prompts from the user and provide responses with relevant content types using computing devices within the environment, which can offer relevant content types of higher quality than other devices in the environment. The determination of computing devices can be based on a defined hierarchy, which can be determined based on device specification information (and / or device capability information).

[0042] Computing devices are ubiquitous in users' environments, whether in an office, home, or elsewhere. Smart TVs, smart speakers, smart appliances, virtual assistant devices, tablets, smart wearables, smartphones, and / or other computing devices are available throughout the user's environment; however, the capabilities of these devices may be underutilized due to a lack of interconnectivity and / or collaboration between them. Specifically, computing devices may be unable to determine when and / or how to interact with each other to obtain input and / or provide context-specific output.

[0043] A multi-device framework for an environment may include acquiring and / or generating information describing the computing devices in the environment, including information associated with input and / or output capabilities. This information can then be used to generate prompts for a generative model (e.g., a large language model) that can be tuned to: what data is acquired and / or generated for output and / or what computing devices are used to produce output associated with a user query. In some implementations, for the multi-device framework, the generative model may be fine-tuned (e.g., parametrically efficient fine-tuning and / or soft-cue tuning), which may include tuning the generative model to handle prompts with an context and to generate output based on both the query and the context.

[0044] Specifically, different computing devices may have different components to capture different forms of input data (e.g., text input, voice commands, gesture input, etc.) and / or provide different types of output data (e.g., visual display, audio playback, etc.) to the user. An environmental multi-device framework can leverage capability information associated with computing devices in the environment to provide an immersive and multifaceted computing system that can obtain input and provide output from multiple different computing systems within the environment.

[0045] Smartphones, smartwatches, smart speakers, smart TVs, smart assistant devices, and / or other computing devices are always around the user; however, interconnectivity may be limited and may not be able to efficiently utilize the highest quality input and / or output data. For example, a user may input a query via their smartphone, which may result in video and / or audio being provided as a response on the smartphone, even though a high-quality smart TV and / or smart speaker may be easily accessible and close to the user. Regardless of whether the content is inherently entertainment-oriented, educational, or otherwise intended, a multi-device framework can be utilized to obtain additional forms of input and / or provide additional forms of output of higher quality than a single-device system.

[0046] The systems and methods disclosed herein can be used to manage inputs received from and / or outputs provided via multiple devices. In some implementations, inputs may be received from smartphones and smartwatches, and outputs responding to the inputs may be provided via smart TVs and smart speakers. Additionally and / or alternatively, the system and methods may determine a specific computing device within the environment that has the highest processing power, and then utilize that specific computing device to perform model inference. In some implementations, processing tasks may be split among multiple computing devices within the environment and / or performed by a server computing system.

[0047] The systems and methods disclosed herein offer a variety of technical effects and benefits. As an example, the systems and methods can be used to provide interconnected multi-device ecosystems. Specifically, the systems and methods disclosed herein can obtain input data from one or more devices within the ecosystem. The systems and methods can obtain environmental data describing the devices within the environment, and can then generate prompts based on the input data and environmental data. These prompts can be processed by a generative model to generate responses to input data generated by devices in the environment. The responses can then be sent to one or more specific devices within the environment. For example, text input can be obtained via a tablet computer, and output can be provided via a smart speaker.

[0048] Another example of technical effects and benefits involves improved computational efficiency and enhanced operation of computing systems. For instance, a technical benefit of the systems and methods disclosed herein is the ability to reduce the computational resources required to interact with multiple computing devices within an environment. Specifically, the multi-device framework enables the centralization of data processing and flow, thereby reducing instances of redundant processing among devices within an ecosystem.

[0049] The systems and methods disclosed herein provide a variety of technical effects and benefits. As an example, the systems and methods can provide an interface generation system. The interface generation system can be used to generate environment-specific and device-specific interfaces, which can be specifically generated based on the input and output capabilities available within a multi-device ecosystem.

[0050] Another technical benefit of the systems and methods disclosed herein is the ability to utilize interface generation to provide an immersive multi-device environment. Specifically, interfaces can be generated and provided to offer users the ability to obtain multiple different input types from multiple different devices in an interconnected system and to provide multiple different output types via multiple different devices in the interconnected system.

[0051] Exemplary embodiments of this disclosure will now be discussed in further detail with reference to the accompanying drawings.

[0052] Figure 1 A block diagram of an example generative modeling system 10 according to an exemplary embodiment of the present disclosure is depicted. In some implementations, the generative modeling system 10 is configured to receive and / or obtain a set of input data 12 describing a query and / or scenario, and, as a result of receiving the input data 12, generate, determine, and / or provide a model-generated output 20 describing a response to the query. Thus, in some implementations, the generative modeling system 10 may include a generative model 18 operable to process a prompt 16 to generate a response to a query, the response being configured to be provided by a specific computing device in the environment.

[0053] Specifically, the generative model system 10 can obtain input data 12 and environmental data 14. Input data 12 can describe one or more inputs. Input data 12 can describe queries (e.g., “Who is playing today?”, “How do I make focaccia?”, “What was the song that played at the end of the movie of the year?”). Input data 12 can include direct inputs (e.g., user typing on a graphical keyboard, capturing images, selecting graphical user interface elements, providing voice commands, etc.) and / or contextual inputs (e.g., user search history, current time, user habit data, user browsing history, application activity, currently open applications, temperature, user location, and / or data obtained from other computing devices within the environment).

[0054] Environmental data 14 can describe multiple computing devices within an environment. The environment can be a room, a suite of rooms, a space close to a user, and / or include multiple computing devices that are close to each other. Environmental data 14 can include specification information for each of the multiple computing devices within the environment. Optionally and / or additionally, environmental data 14 can be descriptive identification data for the multiple computing devices (e.g., device name, device serial number, device label, registration data, etc.), capability data for the multiple computing devices (e.g., input capabilities, processing capabilities, and / or output capabilities of the multiple computing devices), and / or other environmental data.

[0055] Generative model system 10 can process input data 12 and environmental data 14 to generate prompts 16. Prompts 16 may include information associated with a query and multiple computing devices. Prompts 16 may include the query and indicators for one or more candidate computing devices used to provide output. Prompts 16 can be generated by querying a prompt library based on environmental data 14 to determine a prompt template. The prompt template can then be populated based on input data 12. Optionally and / or additionally, prompts 16 can be generated by processing input data 12 and environmental data 14 with a prompt generation model. The prompt generation model may include a machine learning model, which may include generative models (e.g., autoregressive language models, diffusion models, visual language models, and / or other generative models).

[0056] The generative model 18 can then process the cues 16 to generate the output 20 produced by the generative model. The generative model 18 may include a generative language model (e.g., a large language model, a visual language model, and / or other language models), an image generation model (e.g., a text-to-image generation model), and / or other generative models. The generative model 18 may include a transformer model, a convolutional model, a feedforward model, a recurrent model, a self-attention model, and / or other models.

[0057] The model-generated output 20 may include a predicted response to a query of the input data 12. The model-generated output 20 may be a novel generated output including multiple predicted features (e.g., predicted text characters, predicted audio signals, predicted pixels, etc.). The model-generated output 20 may be generated to be provided by a specific computing device among multiple computing devices. For example, the model-generated output 20 may be generated with a specific size, quality, and / or content type based on a preferred computing device (with specific capabilities) determined based on prompts and / or processing by the generative model 18.

[0058] The model-generated output 20 can then be provided to a specific computing device within the environment. Providing the model-generated output 20 to the specific computing device may include sending data, which may include generating application programming interface (API) calls using the generative model and executing API calls using one or more APIs. The specific computing device may then provide the model-generated output 20 to a user (e.g., via a visual display, speaker, and / or other output components of the specific computing device).

[0059] Figure 2 A block diagram of an example response generation system 200 according to an example embodiment of the present disclosure is depicted. The response generation system 200 is similar to... Figure 1 The generative model system 10 differs from the response generation system 200 in that it further includes a prompt library 224 and a search engine 226.

[0060] Specifically, the response generation system 200 may acquire input data 212 and environmental data 214. Input data 212 may include text data, image data, audio data, embedded data, signal data, search history data, browsing history data, application interaction data, latent encoded data, multimodal data, global data, and / or other data. Input data 212 may describe one or more inputs, which may include inputs obtained from multiple computing devices within the environment. Input data 212 may describe queries (e.g., “Who is playing today?”, “How do I make focaccia?”, “What was the song that played at the end of the movie of the year?”). Input data 212 may include directly entered inputs (e.g., a user typing on a graphical keyboard, capturing an image, selecting a graphical user interface element, providing a voice command, etc.) and / or contextual inputs (e.g., user search history, current time, user habit data, user browsing history, application activity, currently open applications, temperature, user location, and / or data obtained from other computing devices within the environment). Input data 212 may include contextual data determined to be relevant to the query. Optionally and / or additionally, input data 12 may include predictive queries that can be generated based on predictions of what a user might be interested in, based on one or more defined user contexts (e.g., a defined interest in a particular football team, a possible interest in horror movies determined using a series of searches, and / or a possible interest in purchasing a particular product based on currently purchasing similar products). In some implementations, response generation system 200 may suggest content to a user by generating input data based on personalized and / or contextualized signals.

[0061] Environmental data 214 may describe multiple computing devices 222 within an environment. The environment may be a room, a suite of rooms, proximity to a user, and / or a space comprising multiple computing devices close to each other. Environmental data 214 may include specification information for each of the multiple computing devices 222 within the environment. Optionally and / or additionally, environmental data 214 may be descriptive identification data (e.g., device name, device serial number, device label, registration data, etc.), capability data (e.g., input capabilities, processing capabilities, and / or output capabilities) of the multiple computing devices 222, and / or other environmental data. The multiple computing devices 222 may include smartphones, smart wearable devices (e.g., smartwatches and / or smart glasses), smart speakers, smart TVs, laptops, desktop computers, smart home appliances (e.g., smart refrigerators, smart washing machines, smart dryers, and / or smart dispensers), virtual assistant devices (e.g., smart home panels, room-based assistants, and / or other assistant devices), tablets, and / or other computing devices. The plurality of computing devices 222 may include mobile computing devices and / or fixed computing devices. In some implementations, the plurality of computing devices 222 may include devices that are close to a user, devices that register to a user's mobile device, devices that register to a user's profile, devices that connect to a particular Internet hub, devices that register to a particular assistant device, and / or other related devices.

[0062] The response generation system 200 can process input data 212 and / or environmental data 214 to generate a prompt 216. The prompt 216 may include information associated with a query and multiple computing devices. The prompt 216 may include the query and indicators for one or more candidate computing devices used to provide output. The prompt 216 can be generated by querying a prompt library 224 based on the environmental data 214 to determine a prompt template. Querying the prompt library 224 may include determining prompt templates associated with environmental configurations associated with the user's environment. Optionally and / or additionally, querying the prompt library 224 may include generating an embedding based on the environmental data 214 and then determining a prompt template associated with the embedding based on a nearest neighbor search. The prompt template can then be populated based on the input data 212. Optionally and / or additionally, the prompt 216 can be generated by processing the input data 212 and environmental data 214 using a prompt generation model. The prompt generation model may include a machine learning model, which may include a generative model (e.g., an autoregressive language model, a diffusion model, a visual language model, and / or other generative models). Hint 216 may include text data, image data, audio data, latent encoded data, multimodal data, embedded data, and / or other data. Hint 216 may include hard hints (e.g., text strings and / or other data inputs) and / or soft hints (e.g., learned parameter sets). In some implementations, hint 216 may be a zero-shot prompt and / or a few-shot prompt. Hint 216 may divide the user request into multiple tasks, which may be provided as multiple hints.

[0063] The generative model 218 can then process the cues 216 to generate the output 220 produced by the generative model. The generative model 218 may include generative language models (e.g., large language models, visual language models, and / or other language models), image generation models (e.g., text-to-image generation models), and / or other generative models. The generative model 218 may include transformer models, convolutional models, feedforward models, recurrent models, self-attention models, and / or other models.

[0064] In some implementations, generative model 218 may communicatively connect to one or more processing engines (e.g., search engines, rendering engines, and / or other engines) and / or one or more machine learning models (e.g., classification models, augmentation models, segmentation models, detection models, embedding models, and / or other models). For example, generative model 218 may communicate back and forth with search engine 226 to obtain search results based on input data 212, contextual data 214, and / or inference. Search engine 226 may output search results, which may then be processed by generative model 218 to generate a summary of the search results and / or determine a response to the query based on information from the search results, which may then be used to generate model-generated output 220. Model-generated output 220 may include text data, image data, audio data, latently encoded data, multimodal data, and / or other data.

[0065] The model-generated output 220 may include a predicted response to a query of the input data 212. The model-generated output 220 may be a novel generated output including multiple predicted features (e.g., predicted text characters, predicted audio signals, predicted pixels, etc.). The model-generated output 220 may be generated to be provided by a specific computing device among multiple computing devices. For example, the model-generated output 220 may be generated with a specific size, quality, and / or content type based on a preferred computing device (with specific capabilities) determined based on prompts and / or processing by the generative model 218.

[0066] Generative model 218 can process prompts 216 to generate output device instructions 230. Output device instructions 230 may include instructions for sending the model-generated output 220 to a specific computing device 228 among a plurality of computing devices 222. Output device instructions 230 may include application programming interface calls, registration files, and / or instructions specific to the computing device 228.

[0067] The model-generated output 220 can then be provided to a specific computing device within the environment based on and / or using output device instructions 230. Providing the model-generated output 220 to the specific computing device may include sending data, which may include making application programming interface calls using generative model generation and executing API calls using one or more application programming interfaces. The specific computing device may then provide the model-generated output 220 to a user (e.g., via a visual display, a speaker, and / or other output components of the specific computing device).

[0068] Figure 3 A flowchart depicts an example method performed according to an example embodiment of this disclosure. Although Figure 3The steps performed in a specific order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the specifically illustrated order or arrangement. The various steps of method 300 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.

[0069] At point 302, the computing system may obtain input data. Input data may include queries associated with a specific user. Input data may include text data, audio data, image data, gesture data, latent coding data, multimodal data, and / or other data. Input data may be obtained from and / or generated using one or more computing devices in the environment. Input data may include prompts for generative models. In some implementations, input data may describe a question associated with one or more topics. This question may be associated with content viewed on one or more computing devices and / or features in the environment.

[0070] At point 304, the computing system can obtain environmental data. The environmental data can describe multiple computing devices in the user's environment. These multiple computing devices can be associated with multiple different output components. In some implementations, the multiple computing devices can be located close to each other within the multiple computing devices. The multiple computing devices can be communicatively connected via a network. In some implementations, the environmental data can describe the specifications of the multiple computing devices. For example, the environmental data can indicate components of the computing devices, including input components and / or output components. Additionally and / or optionally, the environmental data can include information describing the input capabilities and / or output capabilities of the multiple different computing devices.

[0071] At point 306, the computing system can generate prompts based on input data and environmental data. Prompts may include data describing the query and device information associated with a subset of at least a plurality of computing devices. Prompts may include text data, image data, audio data, embedded data, statistical data, graphical representation data, latently encoded data, semantic data, multimodal data, and / or other data. Prompts can be generated using machine learning models based on deterministic functions, heuristics, and / or hybrid methods. Prompts can be generated by querying a prompt template library based on environmental data and filling in a selected prompt template based on the input data.

[0072] In some implementations, generating prompts based on input and environmental data may include: determining an environment-specific device configuration based on the environmental data; obtaining a prompt template from a prompt library based on the environment-specific device configuration; and enhancing the prompt template based on the input data to generate a prompt. The environment-specific device configuration may describe the corresponding output type and corresponding output quality of multiple computing devices. The prompt library may include multiple different prompt templates associated with multiple different device configurations.

[0073] At 308, the computing system can use a generative model to process prompts to generate model-generated output. The model-generated output may include a response to a query. In some implementations, the model-generated output may be generated to be provided by a specific computing device among multiple computing devices. The model-generated output may be generated to be provided by a specific output component among multiple different output components. The generative model can generate output device instructions. The output device instructions may describe a specific computing device among multiple computing devices used to provide the model-generated output. A specific computing device may be associated with a specific output component. For example, model-generated output may be generated to be provided to a user via a corresponding output component of a specific computing device (e.g., speakers in a smart surround system and / or a display screen in a smart TV).

[0074] In some implementations, the generative model can be communicatively connected to the search engine via an application programming interface (API). Processing prompts with the generative model to generate model-generated output may include: generating API calls based on prompts; determining multiple search results using the search engine based on API calls; and processing multiple search results with the generative model to generate model-generated output.

[0075] At point 310, the computing system can send the output generated by the model to a specific computing device. The output generated by the model can be sent to the specific computing device based on instructions from the output device. Sending the output generated by the model to the specific computing device may include executing application programming interface calls generated by the generative model. The output generated by the model can be sent to a virtual assistant device in the user environment, and the virtual assistant device can then control the specific computing device to provide the output generated by the model to a specific user.

[0076] In some implementations, multiple computing devices can be connected via a cloud computing system. Each of the multiple computing devices can register with the cloud computing system's platform. Environmental data can be obtained using the cloud computing system. The output generated by the model can be sent via the cloud computing system. Optionally and / or additionally, the multiple computing devices can be communicatively connected via a local network. Specific computing devices among the multiple computing devices can facilitate the acquisition of input data and the transmission of the output generated by the model.

[0077] In some implementations, processing cues with a generative model to generate model-generated outputs may include generating multiple model outputs. These multiple model outputs may include multiple candidate responses. Additionally and / or optionally, sending the model-generated outputs to a specific computing device may include: sending a first model output from the multiple model outputs to a first computing device among the multiple computing devices; and sending a second model output from the multiple model outputs to a second computing device among the multiple computing devices. The first model output may include visual data for display via a visual display. The second model output may include audio data for playback via a speaker assembly. In some implementations, the first computing device may include a smart TV. Additionally and / or optionally, the second computing device may include a smart speaker.

[0078] Figure 4 A block diagram of an example multi-device management system 400 according to an exemplary embodiment of the present disclosure is depicted. Specifically, the multi-device management system 400 may include a computing system comprising a plurality of user computing devices and a server computing system 420. The plurality of user computing devices and the server computing system 420 may be communicatively connected via a network 410.

[0079] Multiple user computing devices may include a first computing device 402, a second computing device 404, a third computing device 406, and / or an nth computing device 408. These multiple user computing devices may include multiple different computing devices, which may have different input, processing, and / or output capabilities. For example, the first computing device 402 may include a smartphone having an image sensor, an audio sensor, a touch sensor, a motion sensor, a speaker, a visual display, haptic components, and / or lights. The second computing device 404 may include a smart wearable device (e.g., a smartwatch) that may include biometric sensors, motion sensors, touch sensors, a visual display, and / or haptic components. The third computing device 406 may include a smart speaker that may include a high-quality speaker and / or a Bluetooth transmitter. The nth computing device 408 may include a smart TV that may include an infrared sensor receiver, a transmitter-receiver, a speaker, a visual display, and / or multiple input ports. The multiple user computing devices may be connected to network 410 via Ethernet, WiFi, and / or Bluetooth connectivity with a companion device.

[0080] Multiple user computing devices can be associated with the environment based on device registration, location, and / or proximity. Environmental data can be generated based on multiple user computing devices, which may include signals obtained from multiple user computing devices.

[0081] Server computing system 420 may obtain input data and / or environmental data from multiple user computing devices via network 410. Server computing system 420 may include multiple processing services for processing the input data and / or environmental data. For example, server computing system 420 may include one or more generative models 422, one or more search engines 424, one or more prompt generation models 426, one or more interface models 428, and / or one or more other models. One or more generative models 422 may be configured, trained, and / or tuned to process prompts associated with the input data and / or environmental data to generate model-generated output that responds to the input data and is configured to have a specific content type based on the environmental data. One or more search engines 424 may communicatively connect to obtain search results that can be used to understand the input data and / or environmental data and / or to respond to the input data and / or environmental data. One or more prompt generation models 426 may be configured, trained, and / or tuned to process the input data and / or environmental data to generate prompts for one or more generative models 422. One or more interface models 428 can be configured, trained, and / or tuned to process environmental data associated with multiple user computing devices and generate multiple corresponding interfaces for multiple user computing devices. Multiple corresponding interfaces can be generated based on the input and / or output capabilities of the devices within the environment.

[0082] Figure 5 A block diagram of an example environment personalization system 500 according to an example embodiment of the present disclosure is depicted. Specifically, the environment personalization system 500 may utilize information from multiple applications and / or platforms to provide personalized responses and / or personalized experiences. For example, data from search assistant 502, document assistant 504, operating system assistant 506, video player assistant 508, browser assistant 510, and / or chat interface assistant 512 are used to determine the user context and / or generate queries and / or suggestions.

[0083] Search Assistant 502 may be associated with a search application (and / or platform). Search Assistant 502 may obtain current query, session state, search history, trend data, and / or other search data.

[0084] Document Assistant 504 may be associated with one or more document applications. Document Assistant 504 may obtain data associated with the currently viewed and / or edited document, drive data (e.g., stored document information), sharing permission data, and / or other document data.

[0085] Operating system assistant 506 may be associated with the operating system of one or more computing devices. Operating system assistant 506 may obtain data associated with the content currently provided for display, application data (e.g., app deep links, application programming interfaces (APIs) and / or usage data) and / or other operational data.

[0086] The Video Player Assistant 508 can be associated with a video player application (and / or platform). The Video Player Assistant 508 can be used to obtain video data of the currently displayed video, video saves, viewing history, followed media providers, subscriptions, comment history, and / or other video player data.

[0087] Browser Assistant 510 can be associated with browser applications. Browser Assistant 510 can obtain data associated with the current page, bookmark data, tab data, browsing history data, and / or other browser data.

[0088] The Chat Interface Assistant 512 can be associated with one or more chatbots. The Chat Interface Assistant can obtain session history data, response history, input history, topics, links, and / or other chatbot data.

[0089] The multi-device management system 514 can obtain data from multiple applications and / or platforms and provide the data to one or more other systems, which may include a core model 516, a grounded service model 518, and / or a personalized model 520.

[0090] For example, core model 516 can be used for summarizing, planning, and reasoning, and / or function calls. Grounding service model 518 can be used to utilize tool libraries, external connector application programming interfaces (APIs), access search results, and / or access context and / or memory. Core model 516 and / or grounding service model 518 can interact with one or more other models and / or services via one or more cloud APIs. The output of core model 516 and / or grounding service model 518 can be fed back to multi-device management system 514, and then to personalized model 520.

[0091] Personalization model 520 can process data to generate personalized output that may include predictive prompts (and / or suggested prompts). Personalized output can be an enhanced prompt response based on user data, interactions, and / or devices associated with user preferences.

[0092] Figure 6AIllustrations depict example interfaces according to exemplary embodiments of the present disclosure. Specifically, the systems and methods disclosed herein can be used to generate and process environmental data associated with computing devices to generate multiple corresponding interfaces for multiple computing devices within the environment. Multiple corresponding interfaces may be considered for multiple other computing devices and may be confused with and / or mixed with pre-existing interfaces. Figure 6A Three example interfaces are depicted that can be associated with different devices in the environment and / or with different environments.

[0093] For example, a first interface 602 may be associated with a first computing device (e.g., a mobile device) within the environment, a second interface 604 may be associated with a second computing device (e.g., a smartwatch) within the environment, and a third interface 606 may be associated with a third computing device (e.g., a smart refrigerator) within the environment. Optionally and / or additionally, the same computing device may have different interfaces based on being in different environments and / or based on different user scenarios.

[0094] Figure 6B An illustration depicts an example image capture entry point according to an exemplary embodiment of this disclosure. Specifically, an entry point may be provided via user interface elements based on a context, utilizing a search interface and / or assistant interface with multi-device input and / or output.

[0095] For example, a user might be capturing image 610 (e.g., capturing an image of a refrigerator). The image can be processed to determine (and / or identify) objects in the image (e.g., classifying the object as a refrigerator of brand X and model Y). Selectable user interface elements can then be provided in viewfinder 612. The image and / or identifiers can then be processed to generate a response that may include search results 614 associated with objects similar to the identified objects. Additional options for interacting with the response may also be provided. These additional options may include an augmented reality experience that could render the object as an image of the user's home 616.

[0096] Figure 6C An illustration depicts an example smart TV entry point according to an exemplary embodiment of the present disclosure. Specifically, an entry point may include a search entry point indicator that provides suggested search entry point indications on one device and can be interacted with on another device.

[0097] For example, the content provided for display on a first computing device 620 (e.g., a smart TV) can be determined to include features that a user might be interested in during a search. Therefore, suggested entry point user interface elements can be rendered over the content. The user can then select a selectable user interface element 622 on a second computing device (e.g., a mobile computing device) to view model-generated output 624. The model-generated output can be provided via the first computing device 620, the second computing device, and / or a third computing device. The model-generated output 624 may include search results, generative model output, one or more renders, one or more suggestions, and / or other options.

[0098] Figure 6D An illustration depicts an example schedule interface according to an exemplary embodiment of the present disclosure. Specifically, Figure 6D The example schedule interface is depicted, which includes a search window that can be adjusted based on schedule data. For example, the search window may include a text input field and two suggested actions. Suggested actions may include finding a meeting room 630 and / or suggesting allocating time 632 for focus time. Suggestions may be generated based on information from multiple applications and / or multiple computing devices. The search window may be available on a device where the schedule application is open and / or on another device.

[0099] Figure 6E An illustration depicts an example video conferencing interface 634 according to an exemplary embodiment of the present disclosure. Specifically, Figure 6E An example video conferencing interface 634 is depicted, which includes a search window 636 that can be adjusted based on meeting data. For example, the search window 636 may include a text input field and one or more suggested actions. Suggested actions may include taking meeting notes (e.g., transcribing the meeting and / or opening a recording application), setting reminders, rescheduling the meeting, and / or obtaining recordings of similar meetings. Suggestions may be generated based on information from multiple applications and / or multiple computing devices. The search window 636 is available on the device where the video conferencing application is open and / or on another device.

[0100] Figure 6F An illustration depicts an example email interface according to an exemplary embodiment of this disclosure. Specifically, Figure 6FAn example email interface 638 is depicted, which includes a search window 640 that can be adjusted based on email data and / or meeting data. For example, the search window 640 may include a text input field and one or more suggested actions. Suggested actions may include summarizing meeting minutes and / or transcripts, scheduling another meeting, obtaining information about the meeting topic, drafting a follow-up email based on the context associated with the meeting minutes, and / or drafting an email based on email data and / or meeting data. Suggestions may be generated based on information from multiple applications and / or multiple computing devices. The search window 640 may be available on a device where the email application is open and / or on another device.

[0101] Figure 6G An illustration depicts an example video player interface according to an exemplary embodiment of the present disclosure. Specifically, Figure 6G An example video player interface 650 is depicted, which includes a search window 654 that can be adjusted based on video data and / or viewing history data. For example, the search window 654 may include a text input field and one or more suggested actions. Suggested actions may include obtaining a product list of similar products shown in the displayed video 652, summarizing the video, obtaining entity tags for the displayed video, searching for similar videos, and / or obtaining additional information associated with the displayed video 652. For example, a user may request further information about a dress shown in the displayed video 652. One or more frames may be segmented from the video and searched, which may include frame cropping. Optionally and / or additionally, entity tags associated with the displayed frames may be obtained and searched. Search results may be provided as a display overlaid on the displayed video in the search window 654. Search results may include product lists, web links, and / or other data. Suggestions may be generated based on information from multiple applications and / or multiple computing devices. The search window 654 may be available on a device where the video player application is open and / or on another device.

[0102] Figure 7 A flowchart depicts an example method performed according to an example embodiment of this disclosure. Although Figure 7 The steps performed in a specific order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the specifically illustrated order or arrangement. The steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.

[0103] At point 702, the computing system can obtain input data. The input data may include queries associated with a specific user. Queries may include questions related to one or more topics, and may request a response to those questions. Input data may include voice commands obtained via a microphone, text strings entered via a graphical keyboard, and / or gestures obtained via a camera, inertial measurement unit, and / or touch sensor. In some implementations, the input data may include data describing inputs obtained from multiple computing devices within the environment (e.g., voice commands obtained via a microphone on a smartphone, gestures obtained via a touch sensor on a smartwatch, image data from a smart refrigerator, and / or viewing history from a smart TV).

[0104] At 704, the computing system can obtain environmental data. The environmental data can describe multiple computing devices in the user's environment. The multiple computing devices can be associated with multiple different output components. The multiple different output components can be associated with multiple corresponding output capabilities associated with the multiple computing devices. In some implementations, each of the multiple corresponding output capabilities can describe the type and quality of output available via the corresponding computing device. The multiple output components may include a speaker associated with a first device and a visual display associated with a second device. The environmental data may include registration data associated with computing devices registered for the environment, WiFi routers, virtual assistant devices, user computing devices, and / or user profiles. In some implementations, the environmental data may include an output hierarchy for multiple candidate output types, which may include a hierarchical representation of the performance capabilities of the multiple computing devices for multiple different output types (e.g., visual display, audio output, haptic feedback, etc.).

[0105] At point 706, the computing system can generate a prompt based on input data and environmental data. The prompt may include data describing the query and device information associated with a subset of at least a plurality of computing devices. The prompt can be generated by processing the input data and environmental data with a prompt generation model. The prompt generation model may include a language model (e.g., a generative language model (e.g., a large language model)). The prompt generation model can be trained and / or tuned to generate prompts based on understanding the intent of the query and determining the output type associated with the intent and / or based on the output type available in the environmental data. The prompt generation model may generate prompt embeddings to modulate the output generation of the generative model.

[0106] In some implementations, the computing system can determine which specific output component is associated with the intent of the query. Hints can be generated based on this association between the specific output component and the query intent (e.g., a request for a song might be associated with a speaker output component, while a request to play a video might be associated with the visual display of a television).

[0107] Additionally and / or optionally, the computing system may determine an output hierarchy based on environmental data and specification information of multiple different output components. Prompts may be generated based on the output hierarchy and queries. For example, prompts may include text generated from the output and / or embeddings tailored to the output capabilities of the computing devices within the environment.

[0108] At 708, the computing system can use generative model processing prompts to generate model-generated output and output device instructions. The model-generated output may include a response to a query. In some implementations, the model-generated output may be generated to be provided by a specific output component among a plurality of different output components. The output device instructions may describe a specific computing device among a plurality of computing devices used to provide the model-generated output. A specific computing device may be associated with a specific output component. In some implementations, the output device instructions may include application programming interface calls for sending the model-generated output to a specific computing device.

[0109] At point 710, the computing system can send the model-generated output to a specific computing device based on instructions from the output device. In some implementations, the transmission can be performed via signal transmission over a network. Optionally and / or additionally, a notification can be provided to the input computing device indicating that an output configured for another device is available, and the user can then interact with this notification before the model-generated output is sent to the specific computing device. In some implementations, the generative model can generate multiple model-generated outputs, and the computing system can send multiple model-generated outputs to multiple different computing devices within the environment (e.g., sending a slideshow to a smart TV for playback, sending a text document to an e-reader or personal computing device (e.g., a smartphone or tablet), sending an audio file to a smart speaker, and / or sending a scheduled set of color and brightness instructions to an RGB smart lighting device).

[0110] Figure 8 A flowchart depicts an example method performed according to an example embodiment of this disclosure. Although Figure 8 The steps performed in a specific order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the specifically illustrated order or arrangement. The steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.

[0111] At point 802, the computing system can obtain environmental data. The environmental data can describe multiple computing devices within the environment associated with a specific user. The environmental data can describe the specification information of the multiple computing devices. In some implementations, the multiple computing devices can be identified based on device registration for a specific network, device registration for a specific user device, device registration for a specific user profile, proximity to the user device, and / or based on signal exchange between the device and the user device. The user device can be a smartphone, tablet computer, smartwatch, smart glasses, and / or other computing device.

[0112] At 804, the computing system can process environmental data to determine multiple corresponding input capabilities and multiple corresponding output capabilities associated with multiple computing devices. The multiple corresponding input capabilities can be associated with candidate input types associated with the multiple computing devices. In some implementations, the multiple corresponding output capabilities can be associated with candidate output types associated with the multiple computing devices. The multiple input capabilities and / or multiple output capabilities can be determined based on specification information (and / or component information) associated with the multiple computing devices. In some implementations, machine learning models can be used to determine the multiple input capabilities and / or multiple output capabilities based on one or more searches, heuristics, and / or one or more other determination techniques. The multiple input capabilities can describe the available input types and the quality range associated with obtaining and / or generating the input types (e.g., microphone decibel range and / or camera resolution). The multiple output capabilities can describe the available output types and the quality range associated with providing the output of that output type (e.g., audio quality (e.g., volume range, frequency range, etc.)).

[0113] At point 806, the computing system can generate multiple corresponding interfaces for multiple computing devices based on multiple corresponding input capabilities and multiple corresponding output capabilities. Based on the multiple computing devices and multiple corresponding input capabilities, the multiple corresponding interfaces can be specifically designed for the multiple computing devices. In some implementations, the multiple corresponding interfaces may include multiple device indicators indicating multiple computing devices within an environment associated with a specific user. The multiple computing devices can be configured as a user-specific device ecosystem, communicatively connected to receive input and provide output. Each of the multiple corresponding interfaces can be configured to receive a specific input type and provide a specific output type based on the corresponding input capabilities and corresponding output capabilities of a specific computing device among the multiple computing devices.

[0114] At point 808, the computing system can provide multiple corresponding interfaces to multiple computing devices. In response to interface generation and / or in response to user interaction with a specific computing device, multiple corresponding interfaces can be sent to multiple computing devices. These multiple corresponding interfaces can be stored by the server computing system and provided to the computing devices when needed. Optionally and / or additionally, the corresponding interface for the corresponding computing device can be downloaded locally and provided during offline and online states.

[0115] In some implementations, the computing system may obtain user input via a first interface of a first computing device among a plurality of computing devices. The computing system may process the user input using a search engine to determine multiple search results, process the multiple search results using a generative model to generate model output, and provide the model output for display via a second interface of a second computing device among the plurality of computing devices.

[0116] More and more users will own multiple devices, which may or may not share operating systems, apps, and / or platforms. Base models (e.g., large base models that may include generative models) can be used as the primary technology by which users interact with their devices. The system can provide a wide range of new experiences because these base models are able to integrate both their inputs and outputs across all devices.

[0117] The multi-device framework utilizing the base model can be used for a variety of different companion tasks, which may include output conditioning, interface generation, input acquisition facilitation, and / or other tasks.

[0118] For example, a multi-device framework can be used for environmental agent companions with fluid interaction mechanisms that depend on the device and surface. Users may own several devices that could be used in a disjointed ecosystem even in the presence of an environmental LLM companion. Agents (e.g., computing systems incorporating a multi-device framework) may include a base model (e.g., a base model with generative and / or cue-generating models) that adapts the input / output interaction model based on device details. For example, a speaker may have a voice-only interface as its primary interaction model, a device with a non-touchscreen may have a hybrid text / voice interface, and / or mobile phones and / or smart wearables (e.g., smartwatches, smart jackets, and / or smart glasses) may have a specially generated user interface based on their shape factors.

[0119] An interface can be generated to blend with existing surfaces (e.g., the user interface can be minimized to one or more app elements (e.g., the agent interface can be text and / or animations within a specific app interface)). The generated interface can have multi-device / surface awareness elements that can highlight which device the environmental companion is actively listening to and / or monitoring based on the user's proximity to the device. When interacting with the companion on one device requires context from other devices and / or requires performing actions on other devices, the user interface can include user interface elements to show that these other devices are being utilized and / or considered.

[0120] Additionally and / or optionally, the system and method may tailor the environmental companion model output based on device shape factors and characteristics. For example, a user in an environment with smart TVs and smart home devices (e.g., virtual assistant devices) can provide input (e.g., via voice commands) that issues a query (e.g., a query similar to “what movies were nominated for the Oscars for their best songs?”). An LLM-enabled search engine / assistant / companion can obtain additional input to the query, where the additional input can describe details of the device (e.g., input specifying the environment includes the screen and speakers and their specific parameters). Higher-quality smart speakers can be provided as part of the LLM prompt. Based on the prompt, a generative model (e.g., LLM) can generate a response for the smart speaker (e.g., “I’ll list out the movies and play samples of their award-winning soundtracks on your speaker”). Meanwhile, another user sending the same query might be in an environment with the same smart TV but with lower speaker quality, so that user could receive the answer via the smart TV in the form of a visual display (for example, the response could include "I'll play parts of the trailers for those movies on your TV. You'll hear fragments of their award-winning soundtracks in these.").

[0121] In some implementations, the system and method may acquire and / or determine the availability and context of nearby devices, and then use this information to modify actions performed on behalf of a user. For example, a user may be in an environment with multiple Internet of Things (IoT) devices, mobile computing devices (e.g., smartphones, tablets, smartwatches, etc.), cars, and / or laptops equipped with an environmental companion. The user may then (via input sensors on the computing devices) inquire with the companion about a trip to a nearby recreational area. The environmental companion may then acquire and / or determine the input in a manner that takes into account all devices in the environment. Input determination and / or providing a response may include responding to the query across more than one device. For example, the companion may identify typical routes and / or destinations and may generate certain visualizations that can be provided within a map application on one or more specific computing devices. A car companion variant of the agent may perform local reasoning and determine that the car will need to be recharged to travel to one or more destinations. On a map app handled by the main agent companion, queries may be processed to output relevant charging stations highlighted along the illustrated route.

[0122] In some implementations, the systems and methods disclosed herein can be used to generate and / or provide interfaces across different computing devices that may be interconnected and / or have similar styles, layouts, and / or semantics, regardless of the manufacturer and / or operating system of the computing device. For example, the systems and methods can generate interfaces that obscure differences between native interfaces and / or operating systems from different computing devices from different manufacturers.

[0123] The system and method may include fluent cross-device representations. Cross-device representations may be provided based on the output of a base model that can tune input / output behavior across multiple devices and surfaces (e.g., apps). Cross-device representations may be implemented via: (1) a centralized base model running in the cloud that is reported directly by device nodes; (2) a hybrid model in which the centralized base model works in conjunction with a distributed local large model that performs inference using device-specific contexts and coordinates with the main model to solve higher-level tasks; and / or (3) a distributed architecture in which each device has its companion version and is grouped together in the same physical space or the same logical unit by a user who owns all of them. A certain level of inference (e.g., >10B parameters) may be required to enable the models to interact with each other.

[0124] Hybrid approaches can be implemented through a variety of different configurations, which may include adaptations for centralized and / or distributed architectures.

[0125] For example, a user may have several registered interconnected devices. The characteristics of these devices may be known and / or determined. The degree of interoperability may vary and may include an API, to which these devices can be fully controlled if plugged in. Device speakers may utilize the API to output audio. The API can be used to directly render elements on the operating system (and / or complex functionality within an app, etc.).

[0126] The API can be displayed in various configurations depending on the presence of a local base model. If no base model exists, the raw functions can be described in some accessible documentation library, and / or optionally, if a local model exists and is provided, a natural language interface may be available. The device may also provide functionality to interface with other external systems, such as sensors that can be used to read ambient temperature, operate blinds, and / or autonomously navigate around the home to perform certain tasks. API examples and information can be fed into the base model, which drives the environmental companion agent in a manner that can be used to adjust response generation and task resolution on behalf of the user.

[0127] Hint generation may include obtaining and / or generating zero-sample or few-sample cues for the base model to process to understand how to use the device's API. If that is insufficient, the device may have a small dataset associated with it (e.g., approximately 1000 examples, which can be used for cue tuning and / or weighting of the base model to understand how to operate the device). Examples may include a task -> decomposition using the API.

[0128] Once surfaced, the devices can be grouped together in a network where each device is a node. The edges of this network can be persistent (e.g., always online) and / or can have associated weights that quantify the correlation between two devices at a given time. For example, two devices might be very close to each other, and the weight between them could be quantified by having smaller / larger values. The two devices can also share certain contexts (e.g., if they interoperate and / or display the same app; for example, state can be dynamically encoded in the graph edges via messaging).

[0129] Available devices can be queried at inference time. Queries can be fed as a list with metadata and / or represented via a graph network, which can also be fed into the environment companion's base model at inference time. Prior to model inference, the base model may have been fine-tuned to work with the device network topology and / or features. The graph network can be a modality that the base model can use to operate. In some implementations, the network can be directly serialized and can be passed as part of a hint to the generative model (e.g., LLM). If that's insufficient, multi-device graph network examples can be used to fine-tune the large base model driving the environment companion (e.g., there are also 1000 or more examples, where each example has a task -> a stepwise breakdown of how to use multiple devices to better solve the task).

[0130] In some implementations, the system and method may include prompting examples to the LLM (or other generative model) (e.g., “[Device Context][Response][Metadata: This response is suitable for a smart speaker at breakfast time]” and / or “[Device Context][Response][Metadata: The response can be visualized by rendering the UI on the phone with three checkboxes for each answer]”). For a device decomposition example, the prompt may include “[Device Context][Response][Metadata: This response should be passed to the smart speaker and should be displayed as a small notification in summary on the phone because the user may not be near the speaker]”. The user can then face one of their devices and decide to interact with the environmental companion.

[0131] Input can be highly dependent on the details of the shape factor. For example, when a user takes out their phone, the companion can activate it via voice alone. When the phone is locked or when the user unlocks it, the companion can allow the UI for keyboard input to render itself. Modalities dependent on device state can be adapted to the context of the user's relationship with nearby devices. For example, if the user's watch is available, voice input can be activated there instead.

[0132] The query issued can be resolved with the help of all devices. For example, if a user issues a statement of the type "I should go to bed soon, when do I need to wake up tomorrow?", the system can trigger a base model to utilize all available scenarios on the device and take action based on the query and / or scenario. Based on the determined action, the response may first process the surfaces that have access to the work and / or schedule scenarios, and respond at a specific suggested time via the speaker.

[0133] The process can continue; for example, in a smart home, the lighting system can determine a user's intention to fall asleep and activate its local reasoning companion to slowly adjust the lighting according to the user's defined bedtime routine. The information can be visually communicated and rendered as UI elements on the user's phone and / or watch, making the decision known to the user.

[0134] Environmental LLM companions may have proactive components. For example, they may interact with devices to determine whether user input is needed. For instance, a support companion with sensors and autonomously operating devices may determine whether to receive input and / or generate prompts based on the user's context.

[0135] In some implementations, the car companion can be built-in. The car companion can be prompted, configured, and / or trained to schedule a 15-minute window before the estimated time the user leaves, allowing the user to check the temperature, rain conditions, etc., and in response to the prompts, configuration, and / or training, the car companion can notify the user to perform an action (e.g., pick up X) and / or not to forget to bring a raincoat.

[0136] The identified context can be received by the environmental companion, which can then decide how to convey that context to the user. For example, the companion could decide to use a smart speaker (e.g., an audio notification: "It's raining and your car can pick you up instead of having to walk to the parking lot"). Optionally and / or additionally, the system can render the same information and / or actions on the user's watch based on the context the device uses.

[0137] A unified behavior for branded elements can exist across all devices in terms of input and / or output. This unified behavior can be achieved through device-specific prompts and available limited-sample examples. The device may be equipped with five voice samples containing ten seconds of audio. The device may also be equipped with five UI examples demonstrating how the assistant can be rendered on that specific device. The underlying environmental model can be tuned against these examples to generate responses and produce a unified interface output from the manufacturer that the user might find familiar.

[0138] Hybrid framework approaches may include devices with a degree of autonomy but typically interoperable via a centralized partner. Variations may include devices lacking the ability to run their local underlying model, with all decisions then made by a centralized partner. Variant approaches may include continuous streaming of information. Another variation may include devices with complete autonomy, without a central partner. In partnerless approaches, the network backbone may be more important and can be guided by simpler signals to determine which devices are interconnected at a given time to resolve a given query.

[0139] The systems and methods disclosed in this paper can use generative artificial intelligence models to connect fragmented systems and unlock new value for users by providing a single service level companion.

[0140] The capabilities of conversational interfaces (chatbots) extend beyond question-and-answer (Q&A) associated with a single device and / or app. For example, a user might be reading a shopping blog and say, "show me the reviews for the products recommended by this article / video." The system might need to understand what the product is, retrieve reviews from the shopping graph, and summarize those reviews in an easily comparable way. In another example, a user might be reading a recipe and say, "add the ingredients to my Shopping App basket." The system might then need to retrieve the ingredients and call the shopping app's API. In yet another example, a user might be watching a travel video blog and say, "show the places mentioned in this article / video on a map." The system might need to understand the mentioned locations, retrieve the addresses, and call the API to create pins on a custom map app. In another example, while viewing a document app, a user might be asking, "Write a 1000-word essay about Lincoln, focusing on the Civil War, to explain his challenges in a way that a 5th grader would understand." Therefore, the system might need to know Lincoln's key facts, his book list, the contents of those books, historian articles, and summarize all this information. In yet another example, while viewing a map app, a user might be zooming in on a map and saying, "Show me a video of things to do around here." The system might need to know the key things to do at that location, the locations mentioned in the video, and retrieve the correct video.In another example, when viewing content on a phone, a user might be on a hiking app screen and say, "show me restaurants near the trailhead." Therefore, the system might need to know the location of the trailhead on the screen being viewed and obtain local results about restaurants near that location.

[0141] The underlying dependencies are that chatbots on these products, regardless of how they are integrated, may rely on common elements such as reasoning, grounding, retrieval, function calls, user state, and personal preferences. Many of these elements come from search and / or knowledge graph services.

[0142] The systems and methods disclosed in this paper can utilize chatbot large language models (LLMs), search understanding large language models and search engines, cloud service large language models, and / or one or more other supporting models. While the architectures of chatbot large language models (LLMs) and search understanding large language models and search engines may appear similar, they can be implemented independently, with differences in each component, such as different RLHF (Reinforcement Learning from Human Feedback) training, different planning algorithms, different sets of third-party plugins, different search backend interfaces, and / or other variations. Different pipelines and / or systems can exchange information for reasoning, understanding, and / or context determination.

[0143] A separate architecture allows for independent and iterative product development. However, while following best practices can ensure the success of individual products, failing to adopt a service-oriented, holistic approach can lead to a disconnect between products.

[0144] Figure 5 An example architecture for building achievable chatbot capabilities is depicted, along with a collection of instantiated chatbot services that provide the appropriate context and personality that may be relevant to the products instantiated within the service.

[0145] For example, there may be common chatbot services built on a robust LLM framework, fine-tuned for common tasks (such as instruction following), with access to search backends, 1p / 3p APIs, and the ability to call functions. Examples of services that can be leveraged to obtain input and / or determine context may include search platforms (e.g., obtaining concise and realistic data by accessing a user's search history), document applications (e.g., obtaining detailed and creative data by accessing document files, emails, currently viewed emails / files, bookmarks, etc.), and / or browser applications (e.g., obtaining concise and realistic data by accessing the current tab, other tabs, bookmarks, history, etc.).

[0146] In some implementations, the system and approach may include fine-tuning the underlying model differently for different products using use case-specific examples and content.

[0147] The systems and methods disclosed herein may include companion models (e.g., the underlying model that may include prompting generative models and / or generative models).

[0148] By integrating AI and LLM capabilities as independent entities within the user interface (UI) (separate from any other user activity), this system and approach ensures persistent availability, allowing users immediate access whenever they need them. This integration allows AI and LLM to remain open and seamlessly connect with user actions as they transition between various applications and tasks.

[0149] This configuration is helpful for users who want quick access to information or to perform tasks without leaving their current app. Additionally, the persistent system is useful for users who want to use the companion app in conjunction with other products. For example, a user can use the companion app to search for information on a topic while simultaneously using a map application for navigation.

[0150] This system and methodology offer two additional benefits. First, it avoids the need to build on existing solutions with potentially incompatible architectures, thus mitigating potential integration challenges. Second, by not altering existing solutions that users are already accustomed to, it minimizes disruption caused by introducing potentially suboptimal major changes.

[0151] If a user is browsing a product in a store, they can use an LLM (Local Management Model) to obtain information about that product (e.g., price, reviews, and specifications). Users can also take photos of the product with a camera and then use the LLM to search for similar products online. This multi-faceted system allows users to access all the information they need about a product without leaving the store.

[0152] Another example is that if a user is viewing a painting in a museum, they can use an LLM (Limited Library Model) to obtain information about the artist, the painting, and its history. The user can also take a photo of the painting and then use the LLM to search for other paintings by the same artist or works on similar themes. This system allows users to learn more about the artwork they are viewing without having to rely on information provided by the museum.

[0153] In some implementations, a user might be looking for a product but unable to find what they're looking for. In such cases, the user can open their phone and describe the product they're looking for, and the phone will search for and find it.

[0154] By allowing users to connect their other devices as companions, this system and method leverages the environmental ecosystem. Connectivity enables users to have a more seamless and integrated experience across their devices. For example, a user can start a task on their phone and then continue it on their laptop without having to re-enter any information. The environmental ecosystem can create a more personalized and convenient experience for users.

[0155] Figure 9A A block diagram of an example computing system 100 performing multi-device output management according to an example embodiment of the present disclosure is depicted. System 100 includes a user computing system 102, a server computing system 130, and / or a third computing system 150 communicatively coupled via a network 180.

[0156] User computing system 102 may include any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0157] User computing system 102 includes one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 may store data 116 and instructions 118 executed by processor 112 to cause user computing system 102 to perform operations.

[0158] In some implementations, the user computing system 102 may store or include one or more machine learning models 120. For example, the machine learning model 120 may be or may otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.

[0159] In some implementations, one or more machine learning models 120 may be received from server computing system 130 via network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing system 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., performing parallel machine learning model processing across multiple instances of input data and / or detected features).

[0160] More specifically, one or more machine learning models 120 may include one or more detection models, one or more classification models, one or more segmentation models, one or more augmentation models, one or more generative models, one or more natural language processing models, one or more optical character recognition models, and / or one or more other machine learning models. One or more machine learning models 120 may include one or more Transformer models. One or more machine learning models 120 may include one or more neural radiation field models, one or more diffusion models, and / or one or more autoregressive language models.

[0161] One or more machine learning models 120 can be used to detect one or more object features. The detected object features can be classified and / or embedded. The classification and / or embedding can then be used to perform a search to determine one or more search results. Optionally and / or additionally, one or more detected features can be used to determine which indicators (e.g., user interface elements indicating detected features) should be provided to indicate that features have been detected. The user can then select an indicator to perform feature classification, embedding, and / or search. In some implementations, classification, embedding, and / or search can be performed before selecting an indicator.

[0162] In some implementations, one or more machine learning models 120 may process image data, text data, audio data, and / or latently encoded data to generate output data, which may include image data, text data, audio data, and / or latently encoded data. One or more machine learning models 120 may perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, context determination, action prediction, image correction, image enhancement, text enhancement, sentiment analysis, object detection, error detection, inpainting, video stabilization, audio correction, audio enhancement, and / or data segmentation (e.g., mask-based segmentation).

[0163] Alternatively or additionally, one or more machine learning models 140 may be included in or otherwise stored and implemented by server computing system 130, which communicates with user computing system 102 according to a client-server relationship. For example, machine learning model 140 may be implemented by server computing system 130 as part of a web service (e.g., viewfinder service, visual search service, image processing service, ambient computing service, and / or overlay application service). Thus, one or more models 120 may be stored and implemented at user computing system 102, and / or one or more models 140 may be stored and implemented at server computing system 130.

[0164] User computing system 102 may also include one or more user input components 122 for receiving user input. For example, user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). Touch-sensitive components can be used to implement a virtual keyboard. Other example user input components include microphones, conventional keyboards, or other devices that a user can use to provide user input.

[0165] In some implementations, the user computing system may store and / or provide one or more user interfaces 124, which may be associated with one or more applications. The one or more user interfaces 124 may be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, augmented reality experiences, virtual reality experiences, and / or other data for display). The user interface 124 may be associated with one or more other computing systems (e.g., server computing system 130 and / or third-party computing system 150). The user interface 124 may include a viewfinder interface, a search interface, a generative model interface, a social media interface, and / or a media content gallery interface.

[0166] User computing system 102 may include one or more sensors 126 and / or receive data from said one or more sensors. The one or more sensors 126 may be housed in a housing assembly that houses one or more processors 112, memory 114, and / or one or more hardware components that may store one or more software packages and / or cause execution of said one or more software packages. The one or more sensors 126 may include one or more image sensors (e.g., cameras), one or more lidar sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biosensors (e.g., heart rate sensors, pulse sensors, retinal sensors, and / or fingerprint sensors), one or more infrared sensors, one or more location sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. The one or more sensors may be used to obtain data associated with the user's environment (e.g., images of the user's environment, records of the environment, and / or the user's location).

[0167] User computing system 102 may include user computing device 104 and / or a portion thereof. User computing device 104 may include mobile computing devices (e.g., smartphones or tablets), desktop computers, laptop computers, smart wearable devices, and / or smart home appliances. Additionally and / or optionally, the user computing system may acquire data from one or more user computing devices 104 and / or generate data using the one or more user computing devices. For example, a smartphone camera may be used to capture image data describing the environment, and / or an overlay application of user computing device 104 may be used to track and / or process data provided to the user. Similarly, one or more sensors associated with a smart wearable device may be used to acquire data about the user and / or about the user's environment (e.g., a camera housed in the user's smart glasses may be used to acquire image data). Additionally and / or optionally, data may be acquired and uploaded from other user devices that may be specifically used for data acquisition or generation.

[0168] Server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 may store data 136 and instructions 138 executed by processor 132 to cause server computing system 130 to perform operations.

[0169] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by said one or more server computing devices. In instances where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0170] As described above, server computing system 130 may store or otherwise include one or more machine learning models 140. For example, model 140 may be, or may otherwise include, various machine learning models. Example machine learning models include neural networks or other multi-layered nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Reference Figure 9B Discuss example model 140.

[0171] Additionally and / or optionally, the server computing system 130 may include a search engine 142 and / or be communicatively connected to such a search engine, which may be used to crawl one or more databases (and / or resources). The search engine 142 may process data from the user computing system 102, the server computing system 130, and / or a third-party computing system 150 to determine one or more search results associated with input data. The search engine 142 may perform term-based searches, tag-based searches, Boolean-based searches, image searches, embedding-based searches (e.g., nearest neighbor searches), multimodal searches, and / or one or more other search techniques.

[0172] Server computing system 130 may store and / or provide one or more user interfaces 144 for obtaining input data and / or providing output data to one or more users. The one or more user interfaces 144 may include one or more user interface elements, which may include input fields, navigation tools, content tiles, selectable tiles, widgets, data display carousels, dynamic animations, information pop-ups, image enhancement, text-to-speech, speech-to-text, augmented reality, virtual reality, feedback loops, and / or other interface elements.

[0173] User computing system 102 and / or server computing system 130 may train models 120 and / or 140 via interaction with a third-party computing system 150 communicatively coupled via network 180. The third-party computing system 150 may be separate from or part of the server computing system 130. Optionally and / or additionally, the third-party computing system 150 may be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more scenarios.

[0174] The third-party computing system 150 may include one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 executed by the processor 152 to cause the third-party computing system 150 to perform operations. In some implementations, the third-party computing system 150 includes one or more server computing devices or is otherwise implemented by said one or more server computing devices.

[0175] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. Generally, communication conducted through network 180 may use a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL) via any type of wired and / or wireless connection.

[0176] The machine learning models described in this specification can be used for a variety of tasks, applications, and / or use cases.

[0177] In some implementations, the input to the machine learning model of this disclosure may be image data. The machine learning model can process the image data to generate output. As an example, the machine learning model can process image data to generate image recognition output (e.g., image data identification, latent embedding of image data, encoded representation of image data, hashing of image data, etc.). As another example, the machine learning model can process image data to generate image segmentation output. As another example, the machine learning model can process image data to generate image classification output. As another example, the machine learning model can process image data to generate image data modification output (e.g., image data alteration, etc.). As another example, the machine learning model can process image data to generate encoded image data output (e.g., encoded and / or compressed representation of image data, etc.). As another example, the machine learning model can process image data to generate magnified image data output. As another example, the machine learning model can process image data to generate prediction output.

[0178] In some implementations, the input to the machine learning model of this disclosure can be text or natural language data. The machine learning model can process the text or natural language data to generate output. As an example, the machine learning model can process natural language data to generate a language-encoded output. As another example, the machine learning model can process text or natural language data to generate a latent text embedding output. As another example, the machine learning model can process text or natural language data to generate a translation output. As another example, the machine learning model can process text or natural language data to generate a classification output. As another example, the machine learning model can process text or natural language data to generate a text segmentation output. As another example, the machine learning model can process text or natural language data to generate a semantic intent output. As another example, the machine learning model can process text or natural language data to generate an amplified text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine learning model can process text or natural language data to generate a predictive output.

[0179] In some implementations, the input to the machine learning model of this disclosure may be speech data. The machine learning model may process the speech data to generate an output. As an example, the machine learning model may process speech data to generate a speech recognition output. As another example, the machine learning model may process speech data to generate a speech translation output. As another example, the machine learning model may process speech data to generate a latent embedding output. As another example, the machine learning model may process speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of speech data, etc.). As another example, the machine learning model may process speech data to generate an amplified speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine learning model may process speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model may process speech data to generate a predicted output.

[0180] In some implementations, the input to the machine learning model of this disclosure may be sensor data. The machine learning model may process the sensor data to generate output. As an example, the machine learning model may process sensor data to generate identification output. As another example, the machine learning model may process sensor data to generate prediction output. As another example, the machine learning model may process sensor data to generate classification output. As another example, the machine learning model may process sensor data to generate segmentation output. As another example, the machine learning model may process sensor data to generate segmentation output. As another example, the machine learning model may process sensor data to generate visualization output. As another example, the machine learning model may process sensor data to generate diagnostic output. As another example, the machine learning model may process sensor data to generate detection output.

[0181] In some cases, the input includes visual data, and the task is a computer vision task. In other cases, the input includes pixel data for one or more images, and the task is an image processing task. For example, an image processing task could be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the probability that one or more images depict an object belonging to that object class. An image processing task could be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, identifies the probability that that region depicts an object of interest. As another example, an image processing task could be image segmentation, where the image processing output defines a corresponding probability for each category in a predetermined set of categories for each pixel in one or more images. For example, the set of categories could be foreground and background. As another example, the set of categories could be object classes. As another example, an image processing task could be depth estimation, where the image processing output defines a corresponding depth value for each pixel in one or more images. As another example, an image processing task could be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at that pixel between the images in the network input for each pixel in one of the input images.

[0182] A user computing system may include multiple applications (e.g., applications 1 to N). Each application may include its own corresponding machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.

[0183] Each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a field manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In other implementations, the API used by each application is application-specific.

[0184] User computing system 102 may include multiple applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).

[0185] The central intelligence layer may include multiple machine learning models. For example, a corresponding machine learning model (e.g., a model) may be provided for each application, and this machine learning model is managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligence layer is included within the operating system of the computing system 100 or otherwise implemented by the operating system.

[0186] The central intelligence layer can communicate with the central device data layer. The central device data layer can serve as a centralized data repository for the computing system 100. The central device data layer can communicate with many other components of the computing device, such as one or more sensors, a field manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a proprietary API).

[0187] Figure 9B A block diagram of an example computing system 50 performing multi-device output management according to an exemplary embodiment of the present disclosure is depicted. Specifically, the example computing system 50 may include one or more computing devices 52, which can be used to acquire and / or generate one or more datasets, which may be processed by a sensor processing system 60 and / or an output determination system 80 to provide feedback to a user, who can provide information about features in one or more acquired datasets. The one or more datasets may include image data, text data, audio data, multimodal data, latently encoded data, etc. The one or more datasets may be acquired via one or more sensors (e.g., one or more sensors within the computing device 52) associated with the one or more computing devices 52. Additionally and / or optionally, the one or more datasets may be stored data and / or retrieved data (e.g., data retrieved from web resources). For example, a user may interact with images, text, and / or other content items. One or more determinations may then be generated using the interaction with the content items.

[0188] One or more computing devices 52 may acquire and / or generate one or more datasets based on image capture, sensor tracking, data storage retrieval, content download (e.g., downloading images or other content items from web resources via the Internet), and / or via one or more other technologies. Sensor processing system 60 may be used to process one or more datasets. Sensor processing system 60 may use one or more machine learning models, one or more search engines, and / or one or more other processing technologies to perform one or more processing techniques. One or more processing techniques may be performed in any combination and / or individually. One or more processing techniques may be performed serially and / or in parallel. Specifically, context determination block 62 may be used to process one or more datasets, which determines the context associated with one or more content items. Context determination block 62 may identify and / or process metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data to determine the specific context associated with a user. A context can be associated with an event, an identified trend, a specific action, a specific type of data, a specific environment, and / or another context associated with a user and / or data retrieved or obtained.

[0189] The sensor processing system 60 may include an image preprocessing block 64. The image preprocessing block 64 can be used to adjust one or more values ​​of the acquired and / or received image to prepare the image for processing by one or more machine learning models and / or one or more search engines 74. The image preprocessing block 64 can resize the image, adjust saturation values, adjust resolution, strip and / or add metadata, and / or perform one or more other operations.

[0190] In some implementations, the sensor processing system 60 may include one or more machine learning models, which may include a detection model 66, a segmentation model 68, a classification model 70, an embedding model 72, and / or one or more other machine learning models. For example, the sensor processing system 60 may include one or more detection models 66 that can be used to detect specific features in a processed dataset. Specifically, one or more detection models 66 may be used to process one or more images to generate one or more bounding boxes associated with the detected features in the one or more images.

[0191] Additionally and / or optionally, one or more segmentation models 68 may be used to segment one or more portions of a dataset from one or more datasets. For example, one or more segmentation models 68 may utilize one or more segmentation masks (e.g., manually generated and / or generated based on one or more bounding boxes) to segment a portion of an image, a portion of an audio file, and / or a portion of text. Segmentation may include isolating one or more detected objects and / or removing one or more detected objects from an image.

[0192] One or more classification models 70 can be used to process image data, text data, audio data, latently encoded data, multimodal data, and / or other data to generate one or more classifications. The one or more classification models 70 may include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. The one or more classification models 70 can process data to determine one or more classifications.

[0193] In some implementations, one or more embedding models 72 may be used to process data to generate one or more embeddings. For example, one or more embedding models 72 may be used to process one or more images to generate one or more image embeddings in an embedding space. The one or more image embeddings may be associated with one or more image features of one or more images. In some implementations, one or more embedding models 72 may be configured to process multimodal data to generate multimodal embeddings. The one or more embeddings may be used for classification, searching, and / or learning the embedding space distribution.

[0194] The sensor processing system 60 may include one or more search engines 74, which can be used to perform one or more searches. The one or more search engines 74 may crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more dedicated databases, and / or one or more general databases) to determine one or more search results. The one or more search engines 74 may perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), metadata-based search, multimodal search, web resource search, image search, text search, and / or application search.

[0195] Additionally and / or optionally, the sensor processing system 60 may include one or more multimodal processing blocks 76, which may be used to assist in processing multimodal data. The one or more multimodal processing blocks 76 may include generating multimodal queries and / or multimodal embeddings for processing by one or more machine learning models and / or one or more search engines 74.

[0196] The output determination system 80 can then be used to process the output of the sensor processing system 60 to determine one or more outputs to be provided to the user. The output determination system 80 may include heuristic-based determination, machine learning model-based determination, user-selection-based determination, and / or context-based determination.

[0197] Output determination system 80 can determine how and / or where to provide one or more search results in search results interface 82. Additionally and / or optionally, output determination system 80 can determine how and / or where to provide one or more machine learning model outputs in machine learning model output interface 84. In some implementations, one or more search results and / or one or more machine learning model outputs can be provided for display via one or more user interface elements. One or more user interface elements can be overlaid on the displayed data. For example, one or more detection indicators can be overlaid on detected objects in the viewfinder. One or more user interface elements can be selected to perform one or more additional search and / or one or more additional machine learning model processes. In some implementations, user interface elements can be provided as application-specific user interface elements and / or uniformly across different applications. One or more user interface elements may include pop-up displays, interface overlays, interface tiles and / or small pieces, carousels, audio feedback, animations, interactive widgets, and / or other user interface elements.

[0198] Additionally and / or optionally, data associated with the output of the sensor processing system 60 may be used to generate and / or provide augmented reality and / or virtual reality experiences 86. For example, one or more acquired datasets may be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which can then be used to provide the augmented reality and / or virtual reality experience 86 to a user. The augmented reality experience may render information associated with the environment into the appropriate environment. Optionally and / or additionally, objects associated with the processed dataset may be rendered into the user environment and / or virtual environment. Rendering dataset generation may include training one or more neural radiation field models to learn a three-dimensional representation of one or more objects.

[0199] In some implementations, one or more action prompts 88 may be determined based on the output of the sensor processing system 60. For example, a search prompt, purchase prompt, generate prompt, reservation prompt, call prompt, redirection prompt, and / or one or more other prompts may be determined to be associated with the output of the sensor processing system 60. The one or more action prompts 88 may then be provided to the user via one or more selectable user interface elements. In response to the selection of one or more selectable user interface elements, a corresponding action of the action prompt may be performed (e.g., performing a search, utilizing the purchase application programming interface, and / or opening another application).

[0200] In some implementations, one or more generative models 90 may be used to process one or more datasets and / or the output of the sensor processing system 60 to generate model-generated content items, which can then be provided to a user. This generation may be based on user selection and / or may be performed automatically (e.g., automatically based on one or more conditions, which may be associated with a threshold amount of unrepresented search results).

[0201] One or more generative models 90 may include language models (e.g., large language models and / or visual language models), image generation models (e.g., text-to-image generation models and / or image enhancement models), audio generation models, video generation models, graphics generation models, and / or other data generation models (e.g., other content generation models). One or more generative models 90 may include one or more Transformer models, one or more convolutional neural networks, one or more recurrent neural networks, one or more feedforward neural networks, one or more generative adversarial networks, one or more self-attention models, one or more embedding models, one or more encoders, one or more decoders, and / or one or more other models. In some implementations, one or more generative models 90 may include one or more autoregressive models (e.g., machine learning models trained to generate predicted values ​​based on previously generated behavioral data) and / or one or more diffusion models (e.g., machine learning models trained to generate predicted data based on generating and processing distributional data associated with input data).

[0202] One or more generative models 90 can be trained to process input data and generate model-generated content items, which may include multiple predicted words, pixels, signals, and / or other data. The model-generated content items may include novel content items that differ from any existing work. One or more generative models 90 may utilize learned representations, sequences, and / or probability distributions to generate content items, which may include phrases, storylines, settings, objects, characters, beats, lyrics, and / or other aspects not included in existing content items.

[0203] One or more generative models 90 may include a visual language model. The visual language model may be trained, tuned, and / or configured to process image data and / or text data to generate natural language output. The visual language model may leverage a pre-trained large language model (e.g., a large autoregressive language model) and one or more encoders (e.g., one or more image encoders and / or one or more text encoders) to provide detailed natural language output that simulates human-written natural language.

[0204] Visual language models can be used for zero-shot image classification, few-shot image classification, image caption addition, multimodal query extraction, multimodal question answering, and / or can be tuned and / or trained for multiple different tasks. Visual language models can perform visual question answering, image caption generation, feature detection (e.g., content monitoring (e.g., for inappropriate content)), object detection, scene recognition, and / or other tasks.

[0205] Visual language models can leverage pre-trained language models and then tune them for multimodality. Training and / or tuning of visual language models can include image-text matching, masked language modeling, multimodal fusion with cross-attention, contrastive learning, prefix language model training, and / or other training techniques. For example, a visual language model can be trained to process images to generate predicted text similar to ground-value text data (e.g., ground-value captions for an image). In some implementations, a visual language model can be trained to replace masked lexical units of a natural language template with textual lexical units describing features depicted in an input image. Optionally and / or additionally, training, tuning, and / or model inference can include multi-layered connections of visual and textual embedding features. In some implementations, a visual language model can be trained and / or tuned via jointly learning image embeddings and textual embedding generation, which can include training and / or tuning a system to map embeddings to a joint feature embedding space that maps textual and image features to a shared embedding space. Joint training can include image-text pair parallel embeddings and / or can include triple training. In some implementations, images can be used and / or processed as prefixes for language models.

[0206] The output determination system 80 can use the data augmentation block 92 to process one or more datasets and / or the output of the sensor processing system 60 to generate augmented data. For example, the data augmentation block 92 can be used to process one or more images to generate one or more augmented images. Data augmentation may include data correction, data cropping, removal of one or more features, addition of one or more features, resolution adjustment, lighting adjustment, saturation adjustment, and / or other enhancements.

[0207] In some implementations, one or more datasets and / or the output of the sensor processing system 60 may be stored based on the determination of the data storage block 94.

[0208] The output of the output determination system 80 can then be provided to the user via one or more output components of the user computing device 52. For example, one or more user interface elements associated with one or more outputs can be provided for display via the visual display of the user computing device 52.

[0209] The process can be performed iteratively and / or continuously. One or more user inputs to the provided user interface elements can be adjusted and / or affect the continuous processing loop.

[0210] This paper discusses technologies related to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0211] While the subject matter has been described in detail with respect to various specific example embodiments, each example is provided by way of explanation and not as a limitation of this disclosure. Modifications, alterations, and equivalents of such embodiments will be readily apparent to those skilled in the art upon understanding the foregoing. Therefore, this disclosure does not exclude such modifications, alterations, and / or additions to the subject matter that will be readily understood by those of ordinary skill in the art. For example, features shown or described as part of an embodiment may be used with another embodiment to produce yet another embodiment. Therefore, this disclosure is intended to cover such modifications, alterations, and equivalents.

Claims

1. A computing system for determining an output device for providing a query response, the system comprising: One or more processors; as well as One or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions, the instructions causing the computing system to perform operations when executed by the one or more processors, the operations including: Obtain input data, wherein the input data includes queries associated with a specific user; Obtain environmental data, wherein the environmental data describes multiple computing devices in the user's environment, wherein the multiple computing devices are associated with multiple different output components; A prompt is generated based on the input data and the environment data, wherein the prompt includes data describing the query and device information associated with a subset of at least the plurality of computing devices, wherein generating the prompt includes: The output capabilities of the plurality of computing devices are determined based on the environmental data. Generate a representation of the performance capabilities of the plurality of computing devices; and The prompt is generated based on the representation of the performance capabilities of the plurality of computing devices and the input data; Based on the data describing the query and the device information, the prompt is processed using a generative model to generate a model-generated output, wherein processing the prompt using the generative model to generate the model-generated output includes: A response to the query is generated based on the data describing the query processed using the generative model; and Based on the device information, the generative model is used to generate a model-generated output based on the representation and the response; the model-generated output will be provided to a specific computing device among the plurality of computing devices; and The output generated by the model is sent to the specific computing device.

2. The system of claim 1, wherein the output of the model generation is generated to be provided by a specific output component among the plurality of different output components; The generative model generation output device instructions describe a specific computing device among the plurality of computing devices used to provide the output of the model generation, wherein the specific computing device is associated with the specific output component; and The output generated by the model is sent to the specific computing device based on instructions from the output device.

3. The system of claim 1, wherein processing the prompt with the generative model to generate the output generated by the model comprises: Generate multiple model outputs, wherein the multiple model outputs include multiple candidate responses; and Sending the output generated by the model to the specific computing device includes: The first model output from the plurality of model outputs is sent to the first computing device among the plurality of computing devices; as well as The second model output from the plurality of model outputs is sent to the second computing device among the plurality of computing devices.

4. The system of claim 3, wherein the first model output includes visual data for display via a visual display, and wherein the second model output includes audio data for playback via a speaker assembly.

5. The system of claim 4, wherein the first computing device includes a smart TV, and wherein the second computing device includes a smart speaker.

6. The system of claim 1, wherein generating the prompt based on the input data and the environmental data comprises: Based on the environmental data, determine the environment-specific device configuration; Based on the specific device configuration of the environment, obtain prompt templates from the prompt library; as well as The prompt template is enhanced based on the input data to generate the prompt.

7. The system of claim 6, wherein the environment-specific device configuration describes the corresponding output types and corresponding output qualities of the plurality of computing devices, and wherein the prompt library includes a plurality of different prompt templates associated with a plurality of different device configurations.

8. The system of claim 1, wherein the plurality of computing devices are connected via a cloud computing system, each of the plurality of computing devices registers with the platform of the cloud computing system, wherein the environmental data is obtained using the cloud computing system, and wherein the output generated by the model is sent via the cloud computing system.

9. The system of claim 1, wherein the plurality of computing devices are located in proximity to each other computing device within the plurality of computing devices, wherein the plurality of computing devices are communicatively connected via a local network, and wherein a particular computing device among the plurality of computing devices facilitates the transmission of input data acquisition and model generation outputs.

10. The system of claim 1, wherein the generative model is communicatively connected to a search engine via an application programming interface, and wherein processing the prompts with the generative model to generate output generated by the model includes: Generate an application programming interface call based on the aforementioned prompts; The search engine is used to determine multiple search results based on the application programming interface call; as well as The generative model is used to process the multiple search results to generate the output generated by the model.

11. A computer-implemented method, the method comprising: Input data is obtained through a computing system including one or more processors, wherein the input data includes queries associated with a specific user; Environmental data is obtained through the computing system, wherein the environmental data describes multiple computing devices in the user's environment, wherein the multiple computing devices are associated with multiple different output components; The computing system generates a prompt based on the input data and the environmental data, wherein the prompt includes data describing the query and device information associated with a subset of at least the plurality of computing devices, and generating the prompt includes: The output capabilities of the plurality of computing devices are determined based on the environmental data. Generate a representation of the performance capabilities of the plurality of computing devices; and The prompt is generated based on the representation of the performance capabilities of the plurality of computing devices and the input data; Based on the data describing the query and the device information, the computational system processes the prompt using a generative model to generate model-generated output and output device instructions, wherein processing the prompt using the generative model to generate the model-generated output and output device instructions includes: A response to the query is generated based on the data describing the query processed using the generative model; and Based on the device information, the generative model is used to generate model-generated outputs and output device instructions based on the representations and responses. The model-generated outputs will be provided to a specific computing device among the plurality of computing devices, wherein the output device instructions include instructions for providing the model-generated outputs to the specific computing device among the plurality of computing devices; and The computing system sends the output generated by the model to the specific computing device based on the instructions of the output device.

12. The method of claim 11, wherein the plurality of different output components are associated with a plurality of corresponding output capabilities associated with the plurality of computing devices, and wherein each of the plurality of corresponding output capabilities describes the type and quality of output available via the corresponding computing device.

13. The method of claim 11, wherein the output device instructions include an application programming interface call for sending the output generated by the model to the particular computing device.

14. The method of claim 11, wherein the plurality of different output components include a speaker associated with the first device and a visual display associated with the second device.

15. The method of claim 11, wherein, The output generated by the model is generated to be provided to a specific output component among the plurality of different output components, and the method further includes: The computing system determines that the specific output component is associated with the intent of the query; and The prompts are generated based on the association between the specific output component and the intent of the query.

16. The method of claim 11, further comprising: The computing system determines the output hierarchy based on the specification information of the multiple different output components using the environmental data. and The prompts are generated based on the output hierarchy and the query.

17. One or more non-transitory computer-readable media, the non-transitory computer-readable media collectively storing instructions, the instructions causing the one or more computing devices to perform operations when executed by the one or more computing devices, the operations including: Obtain environmental data, wherein the environmental data describes multiple computing devices within an environment associated with a specific user; The environmental data is processed to determine a plurality of corresponding input capabilities and a plurality of corresponding output capabilities associated with the plurality of computing devices, wherein the plurality of corresponding input capabilities are associated with candidate input types associated with the plurality of computing devices, and wherein the plurality of corresponding output capabilities are associated with candidate output types associated with the plurality of computing devices; Multiple corresponding interfaces are generated for the multiple computing devices based on the multiple corresponding input capabilities and the multiple corresponding output capabilities, wherein the multiple computing devices are based on the multiple corresponding input capabilities and the multiple corresponding output capabilities, and the multiple corresponding interfaces are specifically designed for the multiple computing devices; Provide the plurality of corresponding interfaces to the plurality of computing devices; A query is obtained from a first computing device among the plurality of computing devices using one or more of the plurality of corresponding interfaces; Processing the query and the environment data to generate a prompt, wherein generating the prompt includes: To obtain the determined multiple corresponding input capabilities and multiple corresponding output capabilities; Generate a representation of the performance capabilities of the plurality of computing devices; and The prompt is generated based on the representation of the performance capabilities of the plurality of computing devices and the query; and Processing the prompt with a generative model to generate model-generated output and output device instructions, wherein processing the prompt with the generative model to generate model-generated output and output device instructions includes: A response to the query is generated based on processing the query using the generative model; and The generative model is used to generate model-generated outputs and output device instructions based on the various corresponding input capabilities, the various corresponding output capabilities, and the response. The model-generated outputs will be provided to a second computing device among the plurality of computing devices, wherein the output device instructions include instructions for providing the model-generated outputs to the second computing device among the plurality of computing devices; and The output generated by the model is sent to the second computing device based on the instructions from the output device.

18. The one or more non-transitory computer-readable media of claim 17, wherein the plurality of respective interfaces include a plurality of device indicators indicating the plurality of computing devices within the environment associated with the particular user, and wherein the plurality of computing devices are configured as a user-specific device ecosystem, the user-specific device ecosystem being communicatively connected to receive input and provide output.

19. The one or more non-transitory computer-readable media of claim 17, wherein each of the plurality of respective interfaces is configured to receive a particular input type and provide a particular output type based on the respective input capabilities and respective output capabilities of the plurality of computing devices.

20. One or more non-transitory computer-readable media as claimed in claim 17, wherein the operation further comprises: User input is obtained via a first interface of the first computing device among the plurality of computing devices; The user input is processed by a search engine to determine multiple search results; The generative model is used to process the multiple search results to generate model output; as well as The model output is provided for display via a second interface of the second computing device among the plurality of computing devices.

Citation Information

Patent Citations

  • Method and system for presenting privacy-friendly query activity based on environmental signals

    CN115735201A

  • Medical service method, apparatus and device based on LLM intelligent agent architecture, and medium

    CN117112759A