Generating and / or adapting the content of an automated assistant according to the distance between a user and an automated assistant interface
By generating and adapting automated assistant content based on user distance, the inefficiencies of resource consumption and output perception issues are addressed, enhancing user interaction and optimizing resource use.
Patent Information
- Application Number
- JP2022084510
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2038-05-04
AI Technical Summary
Automated assistants render user interface output regardless of user distance, leading to inefficient resource consumption and difficulty in perceiving output, especially for less dexterous users, and may waste computational resources rendering output that is not perceived by the user.
Generate and adapt automated assistant content based on user distance using sensors like vision cameras or distance sensors, selecting a subset of agent data tailored to the user's proximity, reducing resource usage by only transmitting and rendering necessary content.
Improves user interaction by ensuring content is easily perceived, optimizes resource use by avoiding unnecessary computation and network transmission, and adapts output as the user moves.
Smart Images

Figure 0007715681000001 
Figure 0007715681000002 
Figure 0007715681000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods, apparatuses, systems, and computer-readable media for generating and / or adapting the content of an automated assistant according to the distance of a user from an automated assistant interface that renders the content of the automated assistant.
Background Art
[0002] A person may engage in a conversation between a human and a computer using an interactive software application (also referred to as a "digital agent", "chatbot", "interactive personal assistant", "intelligent personal assistant", "assistant application", "conversation agent", etc.) called an "automated assistant" herein. For example, a person (sometimes called a "user" when interacting with an automated assistant) may use oral natural language input (i.e., utterances) that may be converted to text and then processed, and / or provide text (e.g., typed) natural language input to give commands and / or requests to an automated assistant. The automated assistant responds to the request by providing a response user interface output that may include an audible and / or visual user interface output.
Summary of the Invention
Means for Solving the Problems
[0003] The applicant has generally recognized that when an automated assistant causes a user interface output to be rendered for presentation to a user (e.g., in response to a request from the user), the user interface output is generally rendered regardless of the distance of the user from the user interface output device that renders the content. As a result, the user may not be able to perceive the user interface output from the user's current location (e.g., the displayed output may be too small and / or the volume of the audible output may be too low). This may require the user to move and provide a user interface input that requests the output to be rendered again. Processing such user interface inputs and / or rendering the content again can cause unnecessary consumption of computing resources and / or network resources. Further, with respect to less dexterous users, the user may have difficulty moving to a location where the user interface input can be perceived. Additionally or alternatively, the user interface output from the automated assistant may be rendered in a computationally expensive manner as a result of rendering the output regardless of the user's distance. For example, the audible output may be rendered at a volume that is louder than necessary, and / or the displayed output may be displayed by multiple frames for a longer duration than would be the case if the content of multiple frames were instead displayed by a single frame.
[0004] Furthermore, the applicant recognized that when a user interface output is being rendered, the user's movement can potentially interfere with the user's ability to perceive further output from the automated assistant. Additionally, when the automated assistant is causing output to be provided to a particular client device, the user may desire to perceive more output due to the user's proximity to the particular interface of the client device as the user approaches the client device. However, since many automated assistants do not typically know the user's distance, those automated assistants may waste computational resources rendering output that may not be perceived by the user. Moreover, considering the number of ways in which a user can perceive output, computational resources may not be used efficiently when the rendered output is not adapted for nearby users.
[0005] The implementations disclosed herein are directed to methods, apparatus, and (transitory and non-transitory) computer-readable media for generating and / or adapting the content of an automated assistant according to the distance of at least one user to an automated assistant interface that renders the content of the automated assistant. Some implementations for generating the content of an automated assistant according to the distance of at least one user generate the content based on generating a request for an agent that includes distance measurement criteria based on the currently determined distance of the user. The current distance of the user can be determined based on signals from one or more sensors such as vision sensors (e.g., monographic camera, stereographic camera), dedicated distance sensors (e.g., laser rangefinder), microphones (e.g., using beamforming and / or other techniques). Further, the request for the agent is sent to the corresponding agent, and the corresponding agent responds to the request for the agent with agent data adapted to the distance measurement criteria. The automated assistant can then provide the agent data (or a transformation thereof) as content for rendering to the user. Since the content is adapted to the distance measurement criteria and can be easily perceived by the user at the user's current distance, the interaction of the user with the automated assistant is improved. Further, the agent data can be a subset of the agent data of candidate agents available for the request, and the subset is selected by the agent based on a subset fit with the distance measurement criteria of the request for the agent. In these and other ways, only a subset of the agent data of candidate agents is provided by the agent instead of the entire agent data of candidate agents (which requires more network resources to transmit). Further, the client device of the automated assistant that renders the content can receive only a subset (or a transformation thereof) of the agent data instead of the entire agent data (or a transformation thereof) of candidate agents.The specific nature of the content adapted to the distance measurement criteria may ensure the efficient use of computing and other hardware resources in a computing device such as a user device that executes an automated assistant. This is because, at least, the implementation of potentially computationally expensive capabilities of the assistant that cannot be perceived by the user can be avoided. For example, network resources (such as when sending only a subset to the client device), memory resources of the client device (such as when buffering only a subset in the client device), and / or savings in the processor and / or power resources of the client device (such as when rendering only a part or all of the subset) may be possible.
[0006] As one non-limiting example of generating automated assistant content according to a user's distance, assume that the user is 7 feet away from the display of a client device having an assistant interface. Further assume that the user gives the spoken utterance "local weather forecast". The user's estimated distance can be determined based on signals from sensors of the client device and / or other sensors in the vicinity of the client device. The spoken utterance can be processed to generate a request of an agent (e.g., specifying the intent of "weather forecast" and a location value corresponding to the location of the client device), and a distance metric based on the user's estimated distance can be included in the agent's request. The agent's request can be sent to the corresponding agent, and in response, the corresponding agent can return graphical content including only a graphical representation of the 3-day weather forecast for that location. The graphical representation of the 3-day weather forecast can be sent to the client device and rendered graphically via the display of the client device. The corresponding agent can select the graphical representation of the 3-day weather forecast based on the correspondence of the distance metric with the graphical representation of the 3-day weather forecast (e.g., instead of a 1-day, 5-day, or other different weather forecast).
[0007] As a variation of the example, instead, assume that the user is 20 feet away from the display and gives the same verbal utterance "local weather forecast". In such a variation, the distance measurement criteria included in the agent's request reflects an estimated distance of 20 feet (instead of an estimated distance of 7 feet), and as a result, the content returned by the agent in response to the request may include text or audible content that conveys the weather forecast for the location for three days -- and may exclude any graphical content. The audible content (or text content, or audio that is a text-to-speech conversion of the text content) can be sent to the client device to be rendered audible by the speaker of the client device without any weather-related graphical content being visually rendered. The corresponding agent may select a three-day text or audible weather forecast based on the correspondence of the distance measurement criteria with the three-day text or audible weather forecast (e.g., instead of a graphical representation of the weather forecast).
[0008] As a further variation of the example, instead, assume that the user is 12 feet away from the display and gives the same verbal utterance "local weather forecast". In such a further variation, the distance measurement criteria included in the agent's request reflect an estimate of the 12-foot distance, and as a result, the content returned by the agent in response to the request may include text or audible content that conveys the weather forecast for the location for three days - and may also include graphical content that conveys only the forecast for one day (i.e., for that day) of the location. The audible content (or text content, or audio that is a text-to-speech conversion of the text content) can be sent to the client device to be rendered audible by the speaker of the client device, and the graphical content of the weather for one day can also be sent to be graphically rendered by the display of the client device. Also in this case, the corresponding agent can select the content to be returned based on the correspondence of the distance measurement criteria with the content to be returned.
[0009] In some implementations, the content of the automated assistant rendered by the client device can additionally or alternatively be adapted according to the user's distance. For example, when the automated assistant is performing a particular automated assistant action, the automated assistant can "switch" the rendering of different subsets of the content of the automated assistant, such as a subset of the content of the automated assistant that can be utilized locally at the client device that is rendering the content (e.g., the content of the automated assistant stored in the local memory of the client device). The automated assistant can use a measurement of distance at a given time to select a subset of the content of the automated assistant to be used for rendering at the client device at the given time. The content of the automated assistant can be provided to the client device from a remote device, for example, in response to the remote device receiving a request related to the action of the automated assistant. The content provided can correspond to the content of the automated assistant that can be adapted by the client device for multiple different positions and / or distances of the user. In this way, as long as the user is changing position according to the corresponding position and / or distance, the automated assistant can adapt the rendering or presentation of the content of the automated assistant according to the change in the user's position and / or the user's distance. When the user changes position to a position and / or location that does not accommodate any suitable adaptation of the rendered content (and / or changes position near such a position and / or location), the automated assistant can cause the client device to request additional content of the automated assistant for that position and / or location. And the additional content of the automated assistant can be used to render more suitable content at the client device.
[0010] As an example, an automated assistant may execute a routine that includes a plurality of different actions. The automated assistant may execute the routine in response to a user command (e.g., a verbal utterance, a tap of a user interface element) and / or in response to the occurrence of one or more conditions (e.g., based on detecting the presence of the user, based on a particular time, based on an alarm being dismissed by the user). In some implementations, one of the actions among the plurality of different actions of the routine may include rendering content corresponding to a podcast. The content can be rendered using data that can be utilized locally on the client device and can be adapted according to the distance of the user from the client device. For example, when the user is at a first distance from the client device, the automated assistant may cause a portion of the available data to be rendered as content restricted to audible content. Further, when the user moves to a second distance that is shorter than the first distance, the automated assistant can adapt the content to be rendered to include video content and / or cause the audible content to be rendered at a louder volume. For example, the video content can correspond to a video recording of an interview from which the audio content is derived. The data providing the basis for the audible content and the video content can be sent to the client device in response to the initialization of the routine (e.g., by a remote automated assistant component) and / or can be pre-downloaded by the client device in anticipation prior to the initialization of the routine (e.g., according to subscription data indicating that the user desires such content to be automatically downloaded or an instruction of the automated assistant according to the user's preferences).
[0011] In these and other ways, the rendered content can adapt to changes in the user's position and / or the user's location without necessarily requiring additional data each time the user moves. This can reduce latency in adapting the rendered content. If the user changes position to a location not corresponding to data available locally, the automated assistant can cause the client device to request additional data, and / or the automated assistant can generate a request for additional data. Optionally, the request can include information based on a measurement of distance. When the client device receives additional data accordingly (e.g., from a server hosting podcast data), the automated assistant can cause the client device to render the content using the additional data and based on the distance data.
[0012] In some implementations, the client device can anticipate and / or buffer the content in anticipation of the user moving to a position corresponding to the particular rendered content. For example, the client device can have data available locally corresponding to when the user is between 5 feet and 10 feet from the client device. While the user is within 5 to 10 feet of the client device but still moving towards the client device, the client device can render the data available locally and request additional data in anticipation. The additional data can correspond to a distance between 2 feet and 5 feet from the client device, such that when the user enters the area between 2 and 5, the client device can render the additional data. This can reduce latency when switching subsets of data to be rendered as the user moves towards or away from the client device.
[0013] As an example, a user can provide an oral utterance such as "Assistant, play my song." In response, the client device can request data associated with various distances and determine the content to be rendered based on the detected distance of the user from the client device. For example, when the user is 20 feet away from the client device, the client device can render content limited to audio and pre-load album art configured to be rendered when the user is closer than 20 feet but farther than 12 feet. In some implementations, when the user moves to a location at a distance between 20 feet and 12 feet, the album art can replace any previous graphical content (e.g., lyrics) on the client device. Alternatively or additionally, when the user is closer than 12 feet but farther than 6 feet, the client device can render a video and synchronize it with any audio being rendered. In some implementations, the rendered video can be based on data requested by the client device in response to determining that the user is on a trajectory towards the client device, although the video may not have been locally available when the user was 20 feet away. In this way, the requested data is mutually exclusive with the data provided as a basis for the rendered audio data, and the rendered video replaces any graphical content that was being rendered before the user reaches a distance between 12 feet and 6 feet from the client device. Further, when the user is closer than 6 feet, the client device can continue to render the video but can also further visually render touchable media controls (e.g., interactive control elements for rewinding, pausing, and / or fast-forwarding), although those controls were not rendered before the user got closer than 6 feet.
[0014] In some implementations, multiple users can be in an environment shared by client devices that can access an automated assistant. Thus, determining a distance measurement can depend on at least one user who is "active" or otherwise directly or indirectly interacting with the automated assistant. For example, one or more sensors communicating with the client device can be used to detect whether a user is an active user among a group of multiple people. For example, data generated from the output of a visual sensor (e.g., a camera of the client device) can be processed to determine an active user among multiple users based on, for example, the posture, gaze, and / or mouth movement of the active user. As one specific example, a single user can be determined to be the active user based on the user's posture and gaze being directly oriented towards the client device and based on the postures and gazes of other users not being directly oriented towards the client device. In a specific example, the distance measurement can be based on the determined distance of a single user (which can be determined based on the output from the visual sensor and / or the output from other sensors). As another specific example, two users can be determined to be active users based on the postures and gazes of the two users being directly oriented towards the client device. In such another specific example, the distance measurement can be based on the determined distances of the two users (e.g., the average of the two distances).
[0015] Alternatively or additionally, audible data generated from the output of a transducer (e.g., a microphone of a client device) can be processed using beamforming, voice recognition, and / or other techniques to identify an active user among a plurality of users. For example, spoken utterances can be processed using beamforming to estimate the distance of the user giving the spoken utterance, the user giving the spoken utterance is considered to be the active user, and the estimated distance is utilized as the distance of the active user. Also, for example, voice recognition of a spoken utterance can be utilized to identify a user profile that matches the spoken utterance, and the active user in the captured image can be determined based on the facial features and / or other features of the active user that match the corresponding features of the user profile. As yet another example, spoken utterances can be processed using beamforming to estimate the direction of the user giving the spoken utterance, and the active user who gave the spoken utterance is determined based on the fact that the active user is in that direction within the captured image and / or other sensor data. In these and other ways, it is possible to identify the active user among a plurality of users within the environment of the client device, and content can be generated and / or adapted to that active user instead of other users within the environment. Such information can then be used as a basis for generating and / or adapting content for the user. In other implementations, a user's voice signature or voice identifier (ID) can be detected, and the voice signature and / or voice ID can be processed in combination with one or more images from a camera to identify the user's status. For example, audio data collected based on the output of a microphone can be processed to detect features of the voice and compare the features of the voice to one or more profiles that can be accessed by an automated assistant.A profile that has the highest correlation with the characteristics of the voice can be used to determine how to generate and / or adapt content for the user.
[0016] The above description has been provided as an overview of some implementations of the present disclosure. Further descriptions of those implementations and other implementations are shown in more detail below.
[0017] In some implementations, provided is a method executed by one or more processors, including the step of receiving a request to initiate execution of an action by an automated assistant. The automated assistant is accessible via an automated assistant interface of a client device communicating with a display device and a sensor, and the sensor provides an output indicative of a distance of a user relative to the display device. The method further includes the step of determining a measured value of a distance corresponding to an estimated distance of the user relative to the display device based on the output of the sensor. The method further includes the step of identifying an agent for completing the action based on the received request. The agent is accessible by the automated assistant and is configured to provide data for the client device based on the estimated distance of the user relative to the display device. The method further includes the step of generating an agent request to cause the identified agent to provide a content item to facilitate the action in response to receiving the request and identifying the agent based on the received request. The agent request identifies the determined measured value of the distance. The method further includes the step of sending the agent request to the agent to cause the agent to select a subset of candidate content items for the action based on a correspondence between the subset of candidate content items and the measured value of the distance included in the agent request, wherein the subset of candidate content items is configured to be uniquely rendered on the client device as compared to other content items excluded from the subset of content items. The method further includes the step of causing the client device to render the selected subset of candidate content items.
[0018] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0019] In some implementations, a subset of content items includes a first subset corresponding to a first range of distances that encompasses the user's estimated distance, and a second subset corresponding to a second range of distances. The second range of distances excludes the estimated distance and has a common boundary with the first range of distances. In those implementations, the step of causing the client device to render a selected subset of candidate content items includes causing the client device to first render only the first subset, and buffering the second subset at the client device and then causing the second subset to be rendered in response to determining that the user has moved to a new distance within the second range of distances. In some versions of those implementations, the step of causing the second subset to be rendered at the client device includes causing the first subset to be replaced with the second subset at the client device in response to determining that the user has moved to a new distance. In some of those versions, the second subset can optionally have no content that is also included in the first subset. In some other versions of those implementations, the first subset includes audio data, the second subset includes graphical content, the step of causing the client device to first render only the first subset includes causing the audio data to be rendered audibly at the client device, and the step of causing the second subset to be rendered at the client device includes causing the graphical content to be rendered at the client device along with an audible rendering of the audio data. In some of those other versions, the graphical content is an image or the graphical content is video that is rendered in synchronization with the audio data.In some additional or alternative versions, the agent selects a first subset based on a first subset corresponding to a first range of distances that includes the user's estimated distance corresponding to the measured value of the distance, and the agent selects a second subset based on the user's estimated distance being within a threshold distance of a second range of distances corresponding to the second subset. In still other additional or alternative versions, the method further includes determining, based on the output from the sensor, an estimated rate of change of the estimated distance, and including an indication of the estimated rate of change in the agent request. In those other additional or alternative versions, the agent selects a first subset based on a first subset corresponding to a first range of distances that includes the user's estimated distance corresponding to the measured value of the distance, and the agent selects a second subset based on an indication of the estimated rate of change.
[0020] In some implementations, the user and one or more additional users are within the environment of the client device, and the method further includes determining that the user is the currently active user of the automated assistant. In those implementations, the step of determining a measured value of the distance corresponding to the user's estimated distance includes determining the measured value of the user's distance instead of one or more additional users in response to determining that the user is the currently active user of the automated assistant. In some of those implementations, the step of determining that the user is the active user is based on one or both of the output from the sensor and additional output from at least one additional sensor. For example, the sensor or additional sensor can include a camera, the output or additional output can include one or more images, and the step of determining that the user is the active user can be based on one or both of the user's pose determined based on one or more images and the user's gaze determined based on one or more images.
[0021] In some implementations, the method includes sending an agent request to an agent and causing a client device to render a selected subset of candidate content items, and then determining a distinct distance measurement, where the distinct distance measurement indicates that the user's distance from the display device has changed; generating a distinct agent request for a specified agent in response to determining the distinct distance measurement, where the distinct agent request includes the distinct distance measurement; sending the distinct agent request to the agent to cause the agent to select a distinct subset of candidate content items for an action based on a correspondence between the distinct subset of candidate content items and the distinct distance measurement included in the agent request; and causing the client device to render the selected distinct subset of candidate content items.
[0022] In some implementations, the received request is based on an oral utterance received at an automated assistant interface and includes audio data embodying the user's voiceprint, and the method further includes selecting a user profile indicating user preferences for content adapted to proximity based on the user's voiceprint. In those implementations, the subset of content items is selected based on the user preferences.
[0023] In some implementations, the distance measurement is embodied in the received request or is received separately from the received request.
[0024] In some implementations, the client device generates a distance measurement from the output of the sensor, sends the distance measurement in a request or an additional transmission, and the step of determining the distance measurement is performed at the server device. For example, the server device can determine the distance measurement based on the distance measurement being included in a request or an additional transmission and can determine the distance measurement without directly accessing the output of the sensor.
[0025] In some implementations, a method is provided that is executed by one or more processors and includes rendering first content to facilitate an action already requested by a user during an interaction between the user and an automated assistant. The automated assistant is accessible via an automated assistant interface of a client device, and the first content is rendered based on a first subset of content items stored locally on the client device. The method further includes determining, based on an output of a sensor connected to the client device, that the user's position has changed from a first position to a second position while the client device is rendering the first content. The method further includes identifying a second subset of content items for rendering second content therefrom to facilitate the action based on the output of the sensor. The second subset of content items includes data exclusive from the first subset of content items and is stored locally on the client device. The method further includes rendering the second content based on the identified second subset of content items. The method further includes monitoring a subsequent output of the sensor while the client device is rendering the second content, and when the subsequent output of the sensor indicates that the user has moved to a third position different from the first and second positions, determining that a third subset of content items for rendering third content when the user is at the third position is not available locally on the client device, and generating a request to receive the third subset of content items from a remote server device accessible by the automated assistant.
[0026] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0027] In some implementations, the third position is closer to the client device than the first and second positions, and the method further includes receiving a third subset of content items and rendering third content based on the third subset of content items. In some of those implementations, the first content is rendered by a first modality, and the third content is rendered by a second modality different from the first modality. In some versions, the step of rendering the third content includes replacing the second content with the third content, and / or the first modality is an audio modality, the first content is rendered by one or more speakers connected to the client device, the second modality is a display modality, and the third content is rendered by a display device connected to the client device.
[0028] In some implementations, the method further includes receiving an oral utterance at an automated assistant interface of the client device, and the sensor is integral to the automated assistant interface and includes one or more microphones configured to respond to audible input from the user. In some of those implementations, the method further includes determining a target application for performing an action and a user orientation with respect to the client device based on audio data corresponding to the received oral utterance.
[0029] In some implementations, the sensor includes a camera, and the method further includes determining whether the user is an active user based on one or more images captured by the camera when subsequent output of the sensor indicates that the user has moved to a third position, based on the user's pose determined based on the processing of one or more images, the direction of the user's gaze determined based on the processing of one or more images, the movement of the user's mouth determined based on the processing of one or more images, and one or more of the user's gestures detected based on the processing of one or more images.
[0030] In some implementations, a method is provided that is executed by one or more processors and includes receiving, at a remote automated assistant system, an automated assistant request sent by a client device that includes a display device. The method further includes determining, by the remote automated assistant system, based on the content of the automated assistant request, an automated assistant agent for the automated assistant request and a measurement of the user's distance indicating the current distance between the client device and the user within the client device's environment. The method further includes sending, by the remote automated assistant system, an agent request that includes the measurement of the user's distance to the determined automated assistant agent for the automated assistant request. The method further includes receiving, by the remote automated assistant system, from the automated assistant agent in response to the agent request, a content item adapted to the measurement of the user's distance. The method further includes sending, to the client device from the remote automated assistant in response to the automated assistant request, the content item adapted to the measurement of the user's distance. Sending the response content causes the client device to render the response content by the client device's display device.
[0031] These and other implementations of the technologies disclosed in this specification may include one or more of the following features.
[0032] In some implementations, the step of determining a measurement of the user's distance includes determining that the measurement of the user's distance meets a first distance threshold and a second distance threshold, and the content item includes a first subset of content items adapted to the first distance threshold and a second subset of content items adapted to the second distance threshold. In some versions of those implementations, the client device is configured to determine a measurement of the user's distance and select data for rendering response content from one of the first subset of content items and the second subset of content items. In some of those versions, the client device is further configured to render response content based on the first subset of content items when the measurement of the user's distance meets only the first distance threshold, and to render response content based on the second subset of content items when the measurement of the user's distance meets only the second distance threshold. The first subset of content items may include data embodying a data format that is excluded from the second subset of content items.
[0033] In some implementations, a method is provided that is executed by one or more processors and includes determining that a given user among a plurality of users in an environment is the currently active user of an automated assistant that can be accessed by the given user via a client device based on output from one or more sensors associated with the client device in the environment. The method further includes determining a distance measurement corresponding to the distance of a given user to the client device based on the output from one or more sensors and / or based on additional output from (one or more sensors and / or other sensors). The method may further include causing the client device to render content that is tailored to the distance of the given user. The content is tailored to the distance of the given user instead of other users in the environment based on determining that the given user is the currently active user of the automated assistant.
[0034] These and other implementations of the technology may optionally include one or more of the following features.
[0035] In some implementations, the method can further include generating content that is tailored to the distance of a given user, and generating the content that is tailored to the distance of the given user is based on determining that the given user is the currently active user of the automated assistant. In some of those implementations, generating the content that is tailored to the distance of the given user is sending an agent request to a given agent, the agent request including the distance measurement, and receiving content from the given agent in response to sending the agent request.
[0036] In some implementations, the method may further include, during the rendering of the content, determining that a given user has moved and is at a new estimated distance from the client device. In some of those implementations, the method may further include, based on a given user being the currently active user, rendering by the client device a second content adjusted to the new estimated distance in response to determining that the given user has moved and is at a new estimated distance from the client device. In some versions of those implementations, the step of rendering the second content by the client device may include causing the client device to replace the content with the second content. In other versions of those implementations, the content may be capable of including only audible content, the second content may be capable of including graphical content, and the step of rendering the second content by the client device may include rendering the second content together with the content.
[0037] In some implementations, the step of rendering by the client device content adjusted to a given user's distance may include selecting the content in place of other candidate content based on the selected content corresponding to a measurement of the distance and other candidate content not corresponding to the measurement of the distance.
[0038] Other implementations may include a non - transient computer - readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform methods such as one or more of the methods described above and / or elsewhere in this specification. Still other implementations may include one or more computers and / or one or more robotic systems including one or more processors operable to execute instructions stored to perform methods such as one or more of the methods described above and / or elsewhere in this specification.
[0039] It should be understood that all combinations of the above - described concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are considered to be part of the subject matter disclosed herein.
Brief Description of the Drawings
[0040]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0041] FIG. 1 shows FIG. 100 which provides an example of adapting response content according to the distance of user 102 from client device 104 and / or an automated assistant interface. Implementations contemplated herein relate to generating and / or adapting content in response to changes in the position of user 102 who is attempting to access content directly or indirectly via an automated assistant. Generally, computing devices can adapt content according to the distance of the user from the computing device, but such content may be limited to that which can be utilized locally on the computing device. Further, restricting the availability of data to that which is locally accessible can suppress the efficiency of the computing device when more suitable data can be retrieved quickly from an external source such as a remote server. For example, a user closer to a computing device having a display panel and speakers may more easily or quickly understand data such as a weekly weather forecast presented on the display panel rather than being output as audio via the speakers. Thus, by adapting such data according to the proximity of the user to the computing device, the computing device can reduce the amount of time at which a particular output is presented at an interface such as the speakers of the computing device.
[0042] In some implementations contemplated herein, user 102 can request that an action be performed or initiated by an automated assistant, and in response, any data provided to perform the request can be adapted according to the location or change in location of user 102 relative to client device 104. The automated assistant can be, but is not limited to, a tablet computing device or the like that includes one or more sensors that can operate as an automated assistant interface and / or provide an output for determining the distance of the user from client device 104, and can be accessed by the user via the automated assistant interface of client device 104. To invoke the automated assistant, the user can provide a verbal utterance such as, for example, "Assistant, what's the weather today." In response, client device 104 can convert the verbal utterance into audio data, which can be processed at client device 104 and / or transmitted to a remote device 116 (e.g., a remote server) for processing. Further, in response to receiving the verbal utterance, client device 104 can determine a measurement of distance corresponding to the distance between user 102 and client device 104, or the distance between user 102 and a peripheral device communicating with client device 104. Client device 104 can transmit the audio data and the measurement of distance to remote device 116 for the remote device 116 to determine the action and / or application that the user is attempting to initiate via the automated assistant.
[0043] In various implementations, the client device 104 determines local distance measurements based on the output from one or more sensors. For example, the client device 104 can process an image captured by a monographic camera of the client device 104 to estimate the distance of a user. For example, the client device 104 can estimate the distance of a user by processing the image (e.g., using one or more local machine learning models) to classify regions of the image as likely to contain a person's head, and estimating the distance of the user based on the size of the user's head in the image (e.g., based on the size of the region). As another example, the client device 104 can estimate the distance of a user based on the output from a stereographic camera of the client device 104, such as a stereographic image including a depth channel. For example, the client device 104 can process the image (e.g., using one or more local machine learning models) to classify regions of the stereographic image as likely to contain a person, and estimate the distance of that person based on the depth value of that region (e.g., based on the average, median, or other statistical measure of multiple depth values). As yet another example, the client device 104 can estimate the distance of a user based on the output from a microphone of the client device 104. For example, the client device can analyze audio data corresponding to a user's spoken utterance using beamforming and / or other techniques to estimate the distance of the user. As yet another example, the client device 104 can estimate distance based on the output from a combination of sensors, such as based on the output from a visual sensor and based on the output from a microphone. Additional and / or alternative sensors, such as dedicated distance sensors, light detection and ranging (LIDAR) sensors, etc., can be utilized. Also, in some implementations, the client device 104 is external to the client device 104 but communicates with the client device 104 It can rely on the output from one or more sensors. Further, in various implementations, the client device 104 can optionally provide the output from the sensor (and / or its conversion) to a remote device 116, and the remote device 116 can optionally determine a distance measurement based on such provided data.
[0044] The determined action can be associated with content items that can be provided by an automated assistant, applications that can be accessed by the client device, and / or third-party (or first-party) agents hosted on a separate remote device 116. In some implementations, the remote device 116 can compare the received distance measurement to one or more distance thresholds (e.g., thresholds corresponding to a first distance threshold 108 and a second distance threshold 110) to determine a suitable subset of content items that can be used to render content for the user. Alternatively or additionally, the remote device 116 can provide the distance measurement to an application tasked with providing content items, and the application can perform a comparison of the distance measurement to one or more distance thresholds to identify a suitable subset of content items. Alternatively or additionally, the application can receive the distance measurement and provide the distance measurement as an input to a model configured to provide one or more values that can form a basis for generating and / or selecting a subset of content items.
[0045] A suitable subset of content items for a particular distance measurement value can be a subset that is used to render content that is more easily perceivable by user 102 compared to content rendered based on other distance measurement values. For example, when the position of user 102 corresponds to a first distance threshold 108 (e.g., where N can be any distance that can define the limit of the device's visible range, such as the distance N from the client device) that is closest to the client device or within the visible range of the client device, the subset of content items selected to render the first content 112 can include video data (e.g., an image or video presenting a weather forecast). Further, when the position of user 102 corresponds to a second distance threshold 110 that is close to the non - visible range of the client device (e.g., between N and N + m, where m is any positive real number), the subset of content items selected to render the second content 114 can include image data and / or video data of lower quality compared to the above - mentioned video data (e.g., an image or video containing larger and fewer graphical elements than the above - mentioned image or video). Further, when the position of the user corresponds to a third distance threshold 118 within the non - visible range of the client device (e.g., between N + m and N + p, where p is any positive real number larger than m), the subset of content items selected to render the content can include audio data (e.g., a recording of a person's voice providing a weather forecast). When the subset of content items is selected based on the distance measurement value, the subset of content items can be transmitted from the remote device 116 to the client device 104 so that the client device 104 can use the selected subset of content items to render the content. The labels "A" and "B" indicate the correlation between the respective distance thresholds (i.e., the first distance threshold 108 and the second distance threshold 110) and the respective content to be rendered (i.e., the first content 112 and the second content 114). Note that.
[0046] In some implementations, while rendering the first content 112 using a selected subset of content items, the user 102 may move from the first position 120 to the second position 122. A change in the distance of the user 102 or the most recent distance of the user 102 can be detected at the client device 104, and additional distance measurements can be generated at the client device 104 and / or a remote device 116. And the additional distance measurements can be used to select an additional subset of content items for rendering additional content for the user 102 while the user remains at the second position 122. For example, the client device 104 is rendering visual content corresponding to a weather forecast (i.e., the first content 112), and the user is within the visible range of the client device 104 (e.g., an area corresponding to the first threshold 108), but the user 102 can move from the first position 120 to a second position 122 (e.g., a position corresponding to the third threshold 118) that does not allow the user to see the client device 104. Accordingly, the client device 104 can generate a measured value of the detected or estimated distance, and this measured value of the distance can be provided to the remote device 116.
[0047] A remote device 116 may enable an application that has already selected a subset of content items to select an additional subset of content items for rendering further content for user 102. The additional subset of content items may be, for example, audio data that can be used to render further content that can be perceived by user 102 despite user 102's movement to a second location 122. In this way, client device 104 is not strictly limited to local data to adapt to changes in the user's distance; rather, it can use remote services and / or applications to identify more suitable data for rendering content. Further, this enables an automated assistant to replace data during the execution of an action so that any rendered content is adapted to a user who changes their relative position.
[0048] In some implementations, user 102 may move to positions corresponding to a tolerance range or an overlapping range of values corresponding to multiple distance thresholds. As a result, an application or device tasked with selecting a subset of content items for rendering the content may select multiple subsets of content items. In this way, if the user moves from a first position that satisfies a first distance threshold to a second position that satisfies a second distance threshold, client device 104 can locally adapt any rendered content in response to the change in the user's position. In some implementations, the user's trajectory and / or speed may also be used to select multiple different subsets of content items for rendering the content in order to adapt the content being rendered in real time while the user is in motion 106. For example, user 102 can request their automated assistant to queue a song while the user walks towards their television or display projector (e.g., "Assistant, queue my favorite song on my TV"), and in response, a first subset of content items and a second subset of content items can be selected for rendering the content on the television or display projector. The first subset of content items can correspond to audio data that can be rendered by the television or display projector when the user is further away from the television or display projector during the user's movement, and the second subset of content items can correspond to audio-visual data that can be rendered by the television or display projector when the user is closest to the television or display projector. In this way, the second subset of content items supplements the first subset of content items with some amount of mutually exclusive data since the first subset of content items did not include video data.In some implementations, the rate of change of the user's position or location and / or the user's trajectory can be determined by the client device and / or an automated assistant in addition to or instead of determining the distance. In this way, content can be requested and / or buffered in anticipation based on the rate of change of the user's position or location and / or the user's trajectory. For example, when the user is determined to be moving at a rate of change that meets a particular rate-of-change threshold and / or exhibits a trajectory that is at least partially towards or away from the client device, the client device can render different content in response to the determination, and / or request additional data for which other content can be rendered when the user moves towards or away from the client device.
[0049] Figures 2A-2C illustrate diagrams providing examples of content being rendered based on the distance of a user to a client device 204. In particular, FIG. 2A shows a diagram 200 of a user 208 approaching a client device 210 within an environment 218 such as a kitchen. User 208 can approach client device 210 after initializing an automated assistant to perform one or more actions. For example, user 208 can trigger a sensor that user 208 has installed in their kitchen, and in response to the sensor being triggered, the automated assistant can initialize the performance of an action. Alternatively, user 208 can call the automated assistant via the automated assistant interface of client device 210, as contemplated herein.
[0050] Actions that are performed in response to initializing an automated assistant may include presenting media content, such as music, for user 208. Initially, the automated assistant may cause client device 210 to render first content 212 that provides audible content with little or no graphical content. This can conserve computational resources and / or network resources in consideration of the possibility that user 208 is far from client device 210 and thus may not be able to perceive graphical content.
[0051] As shown in FIG. 204, when user 208 moves closer to client device 210, client device 210 can receive and / or process one or more signals capable of providing an indication that user 208 has moved closer to client device 210 compared to FIG. 2A. In response, the automated assistant can receive some amount of data based on the one or more signals and cause client device 210 to render second content 214 at client device 210. Second content 214 can provide graphical content that is more graphical than first content 212, content that is greater in amount compared to first content 212, and / or content that has a higher bitrate compared to first content 212. In some implementations, second content 214 can include at least some amount of content that is exclusive to first content 212. Alternatively or additionally, second content 214 may not be locally available on client device 210 when user 208 was at the location corresponding to FIG. 2A, but rather may be rendered based on data retrieved in response to user 208 moving to the location corresponding to FIG. 2B.
[0052] Further, FIG. 206 shows how a third content 216 can be rendered on the client device 210 when the user 208 moves to a position closer to the client device 210 compared to the user 208 of FIGS. 2A and 2B. In particular, the third content 216 may include content that is adjusted according to the user closest to the client device 210. For example, an automated assistant may determine that the user 208 is even closer to the client device 210 compared to FIGS. 2A and 2B, and cause the client device 210 to render text content (e.g., "[CONTENT]"). The data providing the basis for the text content may be locally available when the user 208 is further away, or may be requested by the client device 210 from a remote device depending on the trajectory of the user 208 towards the client device 210. In this way, the content provided under the instructions of the automated assistant can be dynamic according to the distance of the user 208 from the client device 210. Further, the client device 210 can render unique content according to where the user 208 is located relative to the client device 210.
[0053] Figure 3 shows a method 300 for rendering the content of an automated assistant according to the distance between a user and an automated assistant interface. Method 300 can be executed by one or more computing devices, applications, and / or any other device or module capable of interacting with the automated assistant. Method 300 can include an operation 302 of receiving a request to initiate the execution of an action on the automated assistant. The automated assistant can be accessed via the automated assistant interface of the client device, and the client device may include or communicate with a display device and sensors. The sensor can provide an output by which the distance of the user with respect to the display device can be determined. For example, the sensor can be a camera that provides an output by which an image can be generated to determine the distance between the user and the display device. Alternatively or additionally, the client device can include one or more acoustic sensors, and the output from the acoustic sensors can be analyzed (e.g., using beamforming techniques) to identify the position of the user with respect to the client device.
[0054] Method 300 may further include an operation 304 of identifying an agent to complete an action based on a received request. The agent can be one or more applications or modules associated with a third party that is separate from an entity that manages an automated assistant and can be accessed by the automated assistant. Further, the agent can be configured to provide data for a client device based on a distance of the user from a display device. In some implementations, the agent can be one of a plurality of different agents that are called by the automated assistant to facilitate one or more actions performed based on a direct request from the user (e.g., "Assistant, perform [action].") or an indirect request (e.g., an action performed as part of a learned user schedule).
[0055] Method 300 may also include an operation 306 of determining a measurement of a distance corresponding to an estimated distance of the user from the client device. The measurement of the distance can be determined based on data provided by the client device. For example, a sensor of the client device can provide an output embodying information regarding the position of the user relative to the sensor. The output can be processed and embodied in a request to initialize the automated assistant to perform an action. In some implementations, the measurement of the distance can correspond to various data for which the position and / or location characteristics of the user can be determined. For example, the measurement of the distance can also indicate the distance between the user and the client device and the orientation of the user relative to the client device (e.g., whether the user is facing the client device or not).
[0056] Method 300 may further include an operation 308 of generating an agent request that causes an identified agent to provide a content item based on a determined distance measurement to facilitate an action to be performed. The agent request may be generated by an automated assistant and may include one or more slot values to be processed by the identified agent. For example, the slot values of the agent request can identify a distance measurement, contextual data related to a received request such as a time, user preferences, historical data based on previous agent requests, and / or any other data that can be processed by the agent application.
[0057] Method 300 may also include an operation 310 of sending a request to an agent to cause the agent to select a subset of content items based on the request and the determined distance measurement. The subset of content selected by the agent may correspond to a distance threshold of the user. Further, the subset of content may be rendered by the client device uniquely compared to other content items based on a correspondence between the user's distance threshold and the distance measurement. In other words, the agent may select a subset of content items from a group of content items, but the selected subset is adjusted according to the determined distance measurement. Thus, if different distance measurements are determined, different subsets of content items are selected and the client device renders different content based on the different subsets of content items.
[0058] Method 300 may further include an operation 312 of causing a client device to render a selected subset of content items. The selected subset of content items may be rendered as content presented on a display device of the client device. However, in some implementations, the selected subset of content items may be rendered as audible content, video content, audiovisual content, still images, tactile feedback content, control signals, and / or any other output that can be perceived by a person. In some implementations, an agent generated and / or adapted the selected subset of content items, but the client device can further adapt the selected subset of content items according to context data that can be utilized by the client device. For example, the client device can further adapt the subset of content items according to the user's location, the user's expression, the time, the occupancy of the environment in which the client device and the user participate, the geolocation of the client device, the user's schedule, and / or any other information that can indicate the context in which the user is interacting with an automated assistant. For example, -- since the user is within the audible range of the client device, the actions performed can include rendering audible content, and the selected subset of content items can include audio data, but -- the client device can dynamically adapt the volume of any rendered audio according to the presence of others in the environment and / or whether the user is on a call or using the audio subsystem of the client device for a separate action.Alternatively or additionally, since the user is within the audible range of the client device, the action to be performed can include rendering audible content, and the selected subset of content items can include audio data. However, the client device can cause a different client device to render audio data when the context data indicates that the user has moved closer to a different client device (i.e., a separate distance greater than the distance previously indicated by the distance measurement).
[0059] Figure 4 shows a method 400 for adapting the content of an automated assistant based on the position of a user relative to an automated assistant interface. Method 400 may be performed by one or more computing devices, applications, and / or any other device or module capable of interacting with the automated assistant. Method 400 may include an operation 402 of rendering first content to facilitate an action already requested by the user during an interaction between the user and the automated assistant. The first content may be rendered by a client device through one or more different modalities such as, but not limited to, a touch display panel, a speaker, a haptic feedback device, and / or any other interface that may be used by the computing device. Further, the first content may be rendered based on a first subset of content items that may be available locally on the client device. For example, the first content may be a subset of content items retrieved from a remote server device in response to a routine initialized by the user at the instruction of the automated assistant. For example, the routine may be a "morning" routine that is initialized in response to the user entering the user's kitchen in the morning and a sensor connected to a client device in the kitchen indicating the presence of the user. As part of the "morning" routine, the automated assistant may download content items corresponding to the user's schedule. Thus, the first content item may be associated with the user's schedule, and the first content to be rendered may correspond to a graphical user interface (GUI) having k display elements, where k is any positive integer.
[0060] Method 400 may further include an operation 404 of determining that a user's proximity has changed from a first position to a second position while the client device is rendering first content, based on the output of one or more sensors connected to the client device. For example, the sensors may include multiple microphones for using beamforming techniques to identify the user's position. Alternatively or additionally, the sensors may also include a camera by which the user's orientation, gaze, and / or position can be determined. Using the information from the sensors, an automated assistant can identify a subset of one or more active users from among the multiple users in the environment to generate and / or adapt content for the active users. For example, the content may be generated based on the distance of the active users, independent of the distance of users not included in the subset determined to be active users. Further, the information from the sensors may be used to determine the distance of the user from an automated assistant interface, the client device, and / or any other device that may be communicating with the client device. For example, while the user is viewing the rendered first content, the user may move towards or away from the display panel on which the first content is being rendered.
[0061] Method 400 may also include operation 406 of identifying a second subset of content items for rendering second content therefrom to facilitate an action. For example, when the action is related to a "morning" routine and the content items are associated with the user's schedule, the second subset of content items may be selected according to the user's ability to perceive the second subset of content items. More specifically, if the second location is closer to an automated assistant interface (e.g., a display panel) that is more automated than the first location, the second subset of content items may include additional graphical elements that allow the user to perceive more information. As a result, the user can gather more details about the user's schedule as the user moves closer to the automated assistant interface. Further, the computing resources used for rendering additional graphical elements that may be triggered in response to the second location being closer to the interface than the first location are used in an efficient manner in line with the above considerations.
[0062] Method 400 may further include operation 408 of rendering second content based on the identified second subset of content items. The second content to be rendered may be capable of corresponding to a GUI having l display elements, where l is any positive integer greater than or less than k. For example, the first content to be rendered may include k display elements corresponding to the user's schedule over a few hours. Further, the second content to be rendered may include l display elements corresponding to the user's schedule for the entire day. In this way, the second subset of content items has one or more content items that are mutually exclusive with the first subset of content items. As a result, the user sees different graphical elements as the user changes position closer to the display panel.
[0063] Method 400 may further include operation 410 of monitoring subsequent outputs of the sensor while the client device is rendering the second content. In some implementations, the automated assistant may monitor the output of the sensor under the user's permission to determine whether the user has moved further away from or closer to the automated assistant interface. In this way, the automated assistant can further adapt the content being rendered so that the content is more efficiently perceived by the user. In operation 412 of method 400, a determination is made as to whether the user has moved to a third position different from the first and second positions. If the user has not moved to the third position, the automated assistant can continue to monitor the output of the sensor at least in accordance with operation 410. If the user has moved to the third position, method 400 can proceed to operation 414.
[0064] In operation 414 of method 400, a determination is made as to whether third content can be utilized locally on the client device. The third content may correspond to a third subset of content items that, if the third content were rendered on the client device, would provide the user with additional information about the user's schedule. For example, the third subset of content items may include information about the user's schedule that was not included in the first subset of content items and / or the second subset of content items. In particular, the third subset of content items may include at least some amount of data that is mutually exclusive with respect to the first subset of content items and the second subset of content items. For example, the third subset of content items may include different types of data such as images and / or videos that were not included in the first subset of content items and / or the second subset of content items. The third subset of content items can include data related to the user's schedule for the next week or month, thereby enabling the user to perceive additional information about the user's schedule as the user moves closer to the automated assistant interface.
[0065] When a third subset of content items cannot be utilized locally on a client device, method 400 can proceed to operation 416, which can include generating a request to receive the third subset of content items. The request can be transmitted over a network, such as the Internet, to a remote server device to receive the third subset of content items. For example, the remote server device can host an agent associated with a scheduling application that can be accessed by an automated assistant. The agent can receive the request and identify additional content items related to the request. The agent can then transmit the additional content items as the third subset of content items to the automated assistant and / or the client device. Thereafter, method 400 can proceed to operation 418, which can include rendering third content based on the third subset of content items. Alternatively, when the third subset of content items can be utilized locally on the client device, operation 416 can be skipped, and method 400 can proceed from operation 414 to operation 418.
[0066] FIG. 5 shows a system 500 for adapting response content according to the distance of a user from client device 516 and / or automated assistant interface 518. The automated assistant interface 518 can enable a user to communicate with automated assistant 504, and automated assistant 504 can operate as part of an assistant application provided on one or more computing devices such as client device 516 (e.g., a tablet device, a stand-alone speaker device, and / or any other computing device), and / or remote computing device 512 such as server device 502. The assistant interface 518 can include one or more of a microphone, a camera, a touch screen display, a user interface, and / or any other device or combination of devices that can provide an interface between the user and the application. For example, a user can initialize automated assistant 504 by providing verbal, text, and / or graphical input to the assistant interface to cause automated assistant 504 to perform functions (e.g., provide data, control a peripheral device, access an agent or third-party application, etc.). The client device 516 can include a display device that can be a display panel including a touch interface for receiving touch input and / or gestures to enable a user to control an application on the client device 516 via the touch interface.
[0067] Client device 516 may communicate with a remote computing device 512 via a network 514 such as the Internet. The client device 516 can offload computing tasks to the remote computing device 512 to conserve computing resources in the client device 516. For example, the remote computing device 512 can host an automated assistant 504, and the client device 516 can send inputs received at one or more assistant interfaces 518 to the remote computing device 512. However, in some implementations, the automated assistant 504 can be hosted on the client device 516. In various implementations, all or some aspects of the automated assistant 504 can be implemented on the client device 516. In some of those implementations, aspects of the automated assistant 504 are implemented by a local assistant application on the client device 516, and the local assistant application can interface with the remote computing device 512 to implement other aspects of the automated assistant 504. The remote computing device 512 can optionally provide services to multiple users and their associated assistant applications via multiple threads. In some implementations where all or some aspects of the automated assistant 504 are implemented by a local assistant application on the client device 516, the local assistant application can be an application that is separate from the operating system of the client device 516 (e.g., installed "on top of" the operating system) -- or alternatively, can be implemented directly in the operating system of the first client device 516 (e.g., an application of the operating system but considered an application integrated with the operating system).
[0068] In some implementations, the remote computing device 512 can include a voice-to-text engine 506 that can process audio data received at the assistant interface to identify text embodied within the audio data. The process for converting the audio data to text can include a speech recognition algorithm that can use a neural network, a word2vec algorithm, and / or a statistical model to identify groups of audio data corresponding to words or phrases. The text converted from the audio data can be parsed by a text parser engine 508 and made available to the automated assistant 504 as text data that can be used to generate and / or identify command phrases from the user.
[0069] In some implementations, the automated assistant 504 can adapt content for the client device 516 and the agent 532 that can be accessed by the automated assistant 504. During an interaction between the user and the automated assistant 504, user data 506 and / or context data 522 can be collected at the client device 516, the server device 502, and / or any other device that can be associated with the user. The user data 506 and / or context data 522 can be used under the user's permission by one or more applications or devices that are integral with or accessible by the client device 516. For example, the context data 522 can include data corresponding to time data, location data, event data, media data, and / or any other data that can be related to the interaction between the user and the automated assistant 504. Additionally, the user data 506 can include account information, message information, calendar information, user preferences, historical interaction data between the user and the automated assistant 504, content items related to applications and / or agents that can be accessed by the client device 516, and / or any other data that can be associated with the user.
[0070] To enable the automated assistant 504 to adapt content for the user, the automated assistant 504 can interact with an agent 532, which can provide agent data 536 (i.e., content items) to a remote device 512 and / or a client device 516 for rendering the content at the automated assistant interface 518. As used herein, an "agent" refers to one or more computing devices and / or software that is separate from the automated assistant. In some cases, the agent may be a third-party (3P) agent as it may be managed by someone other than the person managing the automated assistant. In some implementations, the automated assistant 504 may use an agent selection engine 528 to select an agent from a plurality of different agents to perform a particular action in response to a direct or indirect request from the user. The selected agent may be configured to receive requests from the automated assistant (e.g., via a network and / or an API). In response to receiving a request, the agent generates response content based on the request and transmits the response content to provide an output based on the response content. For example, the agent 532 may transmit the response content to the automated assistant 504 for providing an output based on the response content by the automated assistant 504 and / or the client device 516. As another example, the agent 538 itself may provide the output. For example, the user can interact with the automated assistant 504 via the client device 516 (e.g., the automated assistant is implemented on the client device and / or can communicate with the client device over a network), and the agent 538 is an application installed on the client device 516 or is remotely executable on the client device 516 but "streams It is possible for the application to be an application that can be "rung". When the application is called, the application can be executed by the client device 516 and / or presented in front of the client device (for example, the content of the application can take over the display of the client device).
[0071] Calling an agent can include sending a request (e.g., using an application programming interface (API)) to cause the agent to generate content for presentation to the user, including values for call parameters (e.g., values for intent parameters, values for intent slot parameters, and / or values for other parameters), via one or more user interface output devices (e.g., one or more of the user interface output devices used in interaction with an automated assistant). The response content generated by the agent can be adjusted according to the parameters of the request. For example, the automated assistant 504 can use data generated based on the output from one or more sensors in the client device 516 to generate one or more distance measurements. The distance measurements can be embodied as parameters of the request to the agent 538 such that the agent data 536 (i.e., the response content) can be generated, selected, and / or otherwise adapted based on the distance measurements. In some implementations, the agent 538 can include an agent data selection engine 534 that generates, selects, and / or adapts the agent data 536 based at least on the parameters of the request received from the remote device 512 and / or the client device 516. In this way, the client device 516 can render content for the user based at least on a subset of the agent data 536 provided by the agent 532 in response to the distance measurements corresponding to the user.
[0072] FIG. 6 is a block diagram of an exemplary computer system 610. Generally, computer system 610 includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices can include, for example, a storage subsystem 624 that includes a memory 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with computer system 610. Network interface subsystem 616 provides an interface to an external network and is coupled to a corresponding interface device of other computer systems.
[0073] User interface input device 622 can include a keyboard, a mouse, a trackball, a pointing device such as a touchpad or a graphics tablet, a scanner, a touch screen incorporated in a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. Generally, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into computer system 610 or a communication network.
[0074] The user interface output device 620 may include non-visual displays such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for generating a visual image. The display subsystem may also provide a non-visual display such as an audio output device. Generally, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 610 to a user or another machine or computer system.
[0075] The storage subsystem 624 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 624 may include logic for performing selected aspects of methods 300, 400, and / or for implementing one or more of client device 104, remote device 116, client device 516, server device 502, remote device 512, remote device 530, automated assistant 504, agent 532, and / or any other device or operation contemplated herein.
[0076] These software modules are generally executed by processor 614 alone or by processor 614 in combination with other processors. Memory 625 used in storage subsystem 624 may include several memories, such as main random access memory (RAM) 630 for storing instructions and data during program execution and read-only memory (ROM) 632 in which certain instructions are stored. File storage subsystem 626 can provide persistent storage for programs and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functions of a particular implementation may be stored by file storage subsystem 626 within storage subsystem 624 or in other machines accessible by processor 614.
[0077] Bus subsystem 612 provides a mechanism for enabling the various components and subsystems of computer system 610 to communicate with each other as intended. Bus subsystem 612 is shown schematically as a single bus, but alternative implementations of the bus subsystem may use multiple buses.
[0078] Computer system 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 610 shown in FIG. 6 is intended only as a specific example for the purpose of showing some implementations. Many other configurations of computer system 610 with more or fewer components than the computer system shown in FIG. 6 are possible.
[0079] In situations where the systems described in this specification may collect or use personal information about a user (or often referred to as a "participant" in this specification), the user may be given the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences, or the user's current geographical location), or to control whether and / or how the user should receive content from a content server that may be relevant to the user. Also, certain data may be processed in one or more ways before being stored or used so that information that can identify an individual is removed. For example, the user's identity may be anonymized such that information that can identify the individual cannot be determined about the user, or in the case where geographical location information is obtained, the user's geographical location may be generalized (such as to the city, zip code, or state level), and thus processed so that the user's specific geographical location cannot be determined. Thus, the user may have the ability to control how information about the user is collected and / or used.
[0080] Although several implementations are described and illustrated herein, various other means and / or structures may be utilized to perform the functions described herein and / or to obtain one or more of the results and / or advantages, and each such change and / or modification is to be regarded as being within the scope of the implementations described herein. More broadly, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend upon the specific one or more applications for which the teachings are used. One of ordinary skill in the art will recognize or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the above-described implementations are presented by way of example only, and it should be understood that within the scope of the appended claims and their equivalents, implementations may be practiced in a manner other than specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
Explanation of Signs
[0081] 100 Figure 102 User 104 Client Device 106 Movement 108 First Distance Threshold 110 Second Distance Threshold 112 First Content 114 Second Content 116 Remote Device 118 Third Distance Threshold 120 First Position 122 Second Position 200 Figure 204 Client Device Figure 206 User 208 Client Device 210 First Content 212 Second Content 214 Third Content 216 Environment 218 Method 300 Method 400 System 500 Server Device 502 Automated Assistant 504 Voice-to-Text Engine, User Data 506 Text Parser Engine 508 Remote Computing Device 512 Client Device 516 Automated Assistant Interface 518 Context Data 522 Agent Selection Engine 528 Remote Device 530 Agent 532 Agent Data Selection Engine 534 Agent Data 536 Agent 538 Computer System 610 Bus Subsystem 612 Processor 614 Network Interface Subsystem 616 User Interface Output Device 620 User Interface Input Device 622 Storage Subsystem 624 Memory 625 File Storage Subsystem 626 Main Random Access Memory (RAM) 630 Read Only Memory (ROM) 632
Claims
Claim 1 A method executed by one or more processors, comprising: Receiving a request to perform one or more actions using an automated assistant of a client device; Determining that a user among a plurality of users in the environment of the client device provided the request and that the user is a currently active user of the client device, wherein there are multiple currently active users; Determining a distance of the user relative to a display of the client device in response to determining that the user is a currently active user of the automated assistant of the client device, wherein the distance is an average of the respective distances determined for multiple currently active users; Selecting a subset of the plurality of content items based on the determined distance of the user relative to the display of the client device from among the plurality of content items corresponding to the one or more actions, wherein one or more additional content items among the plurality of content items are excluded from the subset; Causing output to be rendered on the display of the client device based on the selected subset of content items corresponding to the one or more actions, and when movement of the user is detected during the rendering, detecting a most recent distance of the user and generating an additional distance measurement value at the client device, and causing output of additional content items for the user to be rendered using the additional distance measurement value; Receiving an additional request to perform the one or more additional actions using the automated assistant of the client device in response to the user moving from an initial position when making the request to perform the one or more actions to another position when making an additional request to perform one or more additional actions; Selecting an additional subset of the plurality of additional content items based on an additional distance of the user from the display of the client device from the plurality of additional content items corresponding to the one or more additional actions, wherein one or more further additional content items among the plurality of additional content items are excluded from the additional subset; Rendering an output to the client device on the display based on the selected additional subset of the content items corresponding to the one or more additional actions; comprising a method. [
2. ] The step of determining the distance of the user from the display of the client device includes determining the distance of the user instead of one or more further users among the plurality of users in the environment. The method according to claim 1. [
3. ] Determining that the user among the plurality of users in the environment of the client device provided the request and determining that the user is the currently active user of the client device is based on one or both of the user's posture determined based on one or more instances of visual data and the user's gaze determined based on one or more instances of visual data includes determining whether the user is the currently active user. The method according to claim 2. [
4. ] Determining that the user provided the additional request and determining that the user is also the currently active user with respect to the additional request; Determining an additional distance of the user from the display of the client device in response to determining that the user is the currently active user with respect to the additional request; further comprising The method according to claim 1. [
5. ] The user moves from an initial position when making the request to perform the one or more actions to another position when making the additional request to perform the one or more additional actions. the determined distance of the user with respect to the display of the client device is based on the initial position of the user, the determined additional distance of the user with respect to the display of the client device is based on another position of the user, The method according to claim 4.
6. when making the request to perform the one or more actions, the user is at a predetermined position, and when making the additional request to perform the one or more additional actions, the user remains at the predetermined position, the determined distance of the user with respect to the display of the client device is based on the predetermined position of the user, the determined additional distance of the user with respect to the display of the client device is based on the predetermined position of the user, The method according to claim 4.
7. a system comprising one or more processors and a memory, wherein the memory, when executed by the one or more processors, causes the one or more processors to receive a request to perform one or more actions using an automated assistant of a client device, determine that a user among a plurality of users in the environment of the client device provided the request, and determine that the user is a currently active user of the client device, where there are a plurality of the currently active users, in response to determining that the user is a currently active user of the automated assistant of the client device, determine the distance of the user with respect to the display of the client device, where the distance is an average of the respective distances determined for the plurality of the currently active users, select a subset of the plurality of content items based on the determined distance of the user with respect to the display of the client device from among the plurality of content items corresponding to the one or more actions, where one or more further content items among the plurality of content items are excluded from the subset, Rendering an output to the client device on the display based on the selected subset of content items corresponding to the one or more actions, and when detecting the user's movement during the rendering, detecting the user's latest distance and generating a measurement of an additional distance on the client device, and using the measurement of the additional distance to render an output of additional content items for the user. Receiving, using the automated assistant of the client device, an additional request to perform the one or more additional actions in response to the user moving from an initial position when making the request to perform the one or more actions to another position when making an additional request to perform one or more additional actions. Selecting an additional subset of the plurality of additional content items based on an additional distance of the user with respect to the display of the client device from among the plurality of additional content items corresponding to the one or more additional actions, wherein one or more further additional content items among the plurality of additional content items are excluded from the additional subset. Rendering an output to the client device on the display based on the selected additional subset of content items corresponding to the one or more additional actions. Configured to store instructions for performing operations including. System. Claim 8 Determining the distance of the user with respect to the display of the client device includes determining the distance of the user instead of one or more additional users among the plurality of users in the environment. The system according to claim 7. Claim 9 Determining that the user among the plurality of users in the environment of the client device provided the request and determining that the user is the currently active user of the client device is. Based on one or more instances of visual data, the posture of the user and the user's fixation determined based on one or more instances of the visual data based on one or both of determining whether the user is the currently active user, including The system according to claim 8.
10. wherein the operation determining that the user provided the additional request and that the user is also the currently active user with respect to the additional request; determining an additional distance of the user relative to the display of the client device in response to determining that the user is the currently active user with respect to the additional request; further comprising The system according to claim 7.
11. when the user moves from an initial position when making the request to perform the one or more actions to another position when making the additional request to perform the one or more additional actions, the determined distance of the user relative to the display of the client device is based on the initial position of the user, the determined additional distance of the user relative to the display of the client device is based on the other position of the user, The system according to claim 10.
12. when making the request to perform the one or more actions, the user is at a predetermined position, and when making the additional request to perform the one or more additional actions, the user remains at the predetermined position, the determined distance of the user relative to the display of the client device is based on the predetermined position of the user, the determined additional distance of the user relative to the display of the client device is based on the predetermined position of the user, The system according to claim 10.
13. when executed by one or more processors, cause the one or more processors to receive a request to perform one or more actions using an automated assistant of a client device Determining that a user among a plurality of users within the environment of the client device provided the request and that the user is a currently active user of the client device, where there are a plurality of the currently active users; Determining a distance of the user from a display of the client device in response to determining that the user is a currently active user of the automated assistant of the client device, where the distance is an average of the respective distances determined for a plurality of the currently active users; Selecting a subset of the plurality of content items from the plurality of content items corresponding to the one or more actions based on the determined distance of the user from the display of the client device, where one or more additional content items among the plurality of content items are excluded from the subset; Causing output to be rendered on the display of the client device based on the selected subset of the content items corresponding to the one or more actions, and when movement of the user is detected during the rendering, detecting a most recent distance of the user and generating a measurement of an additional distance at the client device and using the measurement of the additional distance to cause output of further content items for the user to be rendered; Receiving, using the automated assistant of the client device, an additional request to perform one or more additional actions in response to the user having moved from an initial position when making the request to perform the one or more actions to another position when making an additional request to perform one or more additional actions; Selecting an additional subset of the plurality of additional content items based on an additional distance of the user from the display of the client device from the plurality of additional content items corresponding to the one or more additional actions, wherein one or more further additional content items among the plurality of additional content items are excluded from the additional subset; Causing the client device to render an output on the display based on the selected additional subset of the content items corresponding to the one or more additional actions; A non-transitory computer-readable storage medium configured to store instructions for performing operations including the above.
14. Determining the distance of the user from the display of the client device includes determining the distance of the user instead of one or more further users among the plurality of users in the environment. The non-transitory computer-readable storage medium according to claim 13.
15. Determining that the user among the plurality of users in the environment of the client device provided the request and determining that the user is the currently active user of the client device, Based on one or both of the user's posture determined based on one or more instances of visual data and The user's gaze determined based on the one or more instances of visual data, Including determining whether the user is the currently active user. The non-transitory computer-readable storage medium according to claim 14.
16. The operations include Determining that the user provided the additional request and determining that the user is also the currently active user with respect to the additional request, Determining an additional distance of the user from the display of the client device in response to determining that the user is the currently active user with respect to the additional request, And further including. The non-transitory computer-readable storage medium according to claim 13.
17. when the user makes the additional request to perform the one or more additional actions from an initial position when making the request to perform the one or more actions, wherein the determined distance of the user relative to the display of the client device is based on the initial position of the user, wherein the determined additional distance of the user relative to the display of the client device is based on the another position of the user, A non-transitory computer-readable storage medium according to claim 16. **Claim 18** when making the request to perform the one or more actions, the user is at a predetermined position, and when making the additional request to perform the one or more additional actions, the user remains at the predetermined position, wherein the determined distance of the user relative to the display of the client device is based on the predetermined position of the user, wherein the determined additional distance of the user relative to the display of the client device is based on the predetermined position of the user, A non-transitory computer-readable storage medium according to claim 16.
Citation Information
Patent Citations
Information processing device, information processing method and program
JP2017144521A
Digital Assistant Experience based on Presence Detection
US20170289766A1