Dynamic assistant suggestions during assistant browsing
By providing dynamic update suggestions while rendering content on the display interface, the automation assistant solves the problem of low efficiency in web page interaction, achieving more efficient resource utilization and user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2026-06-23
AI Technical Summary
Existing automated assistants cannot effectively render website content when interacting with web pages, resulting in excessive use of computing resources and low interaction efficiency, especially on dedicated display devices where they cannot achieve complete interaction with the website.
The automated assistant renders content on the display interface while providing dynamically updated assistant suggestions. It generates concise navigation suggestions based on user interaction history and current content, reducing the user's scrolling on the webpage and allowing interaction with content via voice or touch input.
It improves the efficiency of web page interaction and the utilization of computing resources, reduces manual operations by users on web pages, and enhances the ability to interact with automated assistants.
Smart Images

Figure CN117099097B_ABST
Abstract
Description
Background Technology
[0001] Humans can engage in human-computer dialogue with interactive software applications referred to in this paper as “automated assistants” (also known as “digital agents,” “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “assistant applications,” “conversational agents,” etc.). For example, humans (who may be referred to as “users” when interacting with automated assistants) can provide commands / requests to automated assistants using spoken natural language input (i.e., utterances) (which in some cases may be converted to text and then processed) and / or by providing textual (e.g., typed) natural language input.
[0002] While some automation assistants can allow users to access web page data, many web pages are programmed for desktop computers and portable devices such as cellular phones and tablets. Therefore, although an automation assistant may be able to open a website on the specified device, further user interaction with the website may not be possible via the automation assistant, and may be limited to interactions via certain input modalities, and / or limited to certain restricted interactions via those input modalities. For example, if a user expects to participate in the website's graphical rendering, the user may have to use a cursor, keyboard, and / or touch interface—and may not be able to interact with the rendering via voice, and / or may be limited to only certain voice interactions (e.g., limited to "scroll up" and "scroll down"). This can lead to inefficient interaction with the website, the inability to perform certain interactions with the website, and / or several other drawbacks, each of which can lead to excessive use of the specified device's computing resources.
[0003] As an example, when a user invokes the automation assistant to access a website via a standalone display device, the automation assistant may not be able to fully (or at all) render the website, thus recommending that the user access the website via their cellular phone. Therefore, because the automation assistant cannot facilitate interaction with the website on a standalone display device, power and processing bandwidth will be consumed across multiple devices. In some instances, this lack of universality on certain websites may be due to limitations in stylesheets and / or other specifications stored in association with the website. Websites designed for mobile and desktop browsers may not be suitable for automation assistant access, except for retrieving data fragments from the website. If a user expects to further browse the website or view other related websites, they may necessarily need to access a separate application with web browser functionality, even if that separate application may not simultaneously offer one or more useful utilities of the automation assistant. Summary of the Invention
[0004] Some implementations described herein relate to an automated assistant that can be invoked to visually render content at a display interface, and simultaneously render (e.g., visually) assistant suggestions, which may optionally be dynamically updated by the automated assistant as the user interacts with the content. In some implementations, the assistant suggestions rendered concurrently with certain content may be based on various different data associated with the user, data associated with the automated assistant, and / or information associated with the content itself.
[0005] As an example of some of these implementations, a user can invoke an automation assistant to view a recipe. After ordering groceries using their automation assistant, the user can invoke the assistant by providing a spoken phrase such as “Assistant, show me a recipe for chow mein.” In response, the automation assistant can access webpage data from the recipe webpage and render the webpage data on the display interface of the computing device. Optionally, and as described herein, the automation assistant can initially respond to the spoken phrase to perform a search for “recipe for chow mein,” identify multiple resources (including the recipe webpage) in response to the search, and present the corresponding search results for each resource (e.g., each search result may include fragments of content(s) from the corresponding resource). The user can then select (e.g., by clicking or by voice selection) the search results corresponding to the recipe webpage, which can cause the webpage data to be rendered. While the webpage data is being rendered, the user can scroll through the content rendered on the display interface by providing another input to the automation assistant—such as touch input and / or another spoken phrase (e.g., “Scroll down.”). The automation assistant can display assistant suggestions along with the rendered content (e.g., in a separate pane to the left or right of the pane rendering the web page content), and one or more topics among the assistant suggestions displayed at a given time can be selected by the automation assistant based on a portion of the rendered content currently visible to the user at that time.
[0006] For example, when an automation assistant scrolls through a webpage (such as a stir-fry recipe) in response to user input (or even automatically), the display might eventually render a portion of the webpage data that provides a list of the recipe's ingredients. Because the user had previously queried the automation assistant about ordering groceries before visiting the webpage, the automation assistant can prioritize any suggestions associated with the food order as lower than other suggestions. Therefore, while the ingredient list might lack any specific food order suggestions when rendered, the assistant's suggestions could include navigation suggestions, such as those that can be selected or spoken individually to navigate and display the corresponding sections of the webpage. Continuing with this example, when the automation assistant scrolls to the section of the webpage detailing oven preheating instructions, in response, it could render an assistant suggestion for controlling the smart oven (which wasn't previously displayed). When the smart oven assistant suggestion is based on user input selection, the automation assistant can preheat the smart oven in the user's home to the temperature specified on the webpage. Alternatively, if the assistant suggestion is not selected and the automation assistant continues to scroll the webpage, the assistant suggestion can be replaced by a different assistant suggestion based on other data in the webpage data and / or other data (e.g., for setting a cooking timer based on recipe instructions).
[0007] In some implementations, multiple assistant suggestions may be rendered near content (e.g., web page data and / or other application data) to provide a unique list of suggestions generated based on the content. For example, each assistant suggestion may include natural language content summarizing a portion of the content, and when a user selects an assistant suggestion (e.g., through touch or verbal input matching the suggestion), the automated assistant may scroll to the corresponding section of the content. For example, according to the example above, the rendered assistant suggestions may include navigation suggestions that include content summarizing a portion of the recipe (e.g., the navigation suggestion could be "Serving Instructions"). In response to the user selecting the rendered assistant suggestion, the automated assistant may scroll to a portion of the recipe provided under the title "Tips for Beautifully Plating This Masterpiece," and the automated assistant may also replace the rendered assistant suggestion. For example, replacing an assistant suggestion could refer to an assistant action that might be helpful for the visible portion of the recipe. Assistant actions can be, but are not limited to, text-to-speech actions to read "tips for serving" to the user, scrolling actions to further scroll through content, playing videos that may be included in the recipe page, and / or actions that cause the automated assistant to "pivot" to another page (e.g., different recipe articles based on other web page data and / or other application data).
[0008] Continuing with the example from the previous paragraph, note that the navigation suggestion "Serving Instructions" differs from the webpage's title "Tips for Beautifully Plating This Masterpiece," which the automation assistant navigates to when a suggestion is selected. For example, the navigation suggestion does not include any of the same terms as the title and is more concise (i.e., fewer characters) than the title—but is still semantically aligned with the title and the content immediately following it. In some implementations, navigation suggestions are generated to be more concise (i.e., fewer characters) to allow selection of navigation suggestions with more concise spoken language, resulting in reduced computational resources used when performing speech recognition of spoken language and also leading to a shorter duration of user interaction with the webpage. In some implementations, navigation suggestions are generated to be more concise (i.e., fewer characters) so that they can be rendered within the display constraints of the assistant device's display. For example, the assistant suggestion portion of the displayed interface may be limited to a small fraction of the available display (e.g., 25% or less, 20% or less) to ensure efficient viewing of the webpage within the content viewing portion of the interface, and navigation suggestions are generated with maximum character limits to ensure they can be rendered within these constraints.
[0009] Various techniques can be used to generate more concise navigation suggestions while ensuring that the suggestions are semantically aligned with the linked content. As an example, navigation suggestions can be generated using text summarization techniques with maximum character constraints (e.g., using a trained text summarization machine learning model). For instance, a title can be processed using text summarization to generate a summary text used as a navigation suggestion. As another example, navigation suggestions can be generated based on anchor text (i.e., hyperlinks to specific parts of a resource, such as specific XML or HTML tags) that jump to a portion of the content. For example, certain anchor texts can be selected based on whether they satisfy the maximum character constraint, and optionally, based on whether they are the most frequently occurring anchor texts that satisfy the maximum character constraint. Alternatively, if no anchor text satisfies the character constraint, navigation suggestions can be generated by applying text summarization techniques to the anchor texts (e.g., the most frequently occurring anchor texts) and using the summary text as the navigation suggestion. As another example, navigation suggestions can be generated based on the embeddings that generated the content (e.g., by processing the text using Word2Vec, BERT, or other trained models) and by identifying words or phrases that satisfy character constraints and have alternating embeddings in the embedding space (i.e., embeddings that are "close to" the generated ones). For example, the title "Tips for Beautifully Plating This Masterpiece" and the text immediately following it can be processed to generate a first word embedding. Further, based on determining that it satisfies character constraints and that it has a second embedding in the embedding space that is closest to the first embedding among candidate navigation suggestions that satisfy character constraints, "Serving Instructions" can be selected as the navigation suggestion for that title.
[0010] In some implementations, when accessing similar content, assistant suggestions rendered adjacent to other content can be based on historical interactions between the user and the automated assistant, and / or between other users and their automated assistants. For example, a user providing spoken phrases such as "Assistant, how is hail formed?" can cause the automated assistant to render web page data detailing how hail forms in the atmosphere. When a user is viewing a specific section of web page data detailing the potential damage hail can cause, the automated assistant can generate supplementary content based on that specific section of the web page data. For example, when a user or another user views that section of web page data and causes the automated assistant to perform a specific action, the automated assistant can generate supplementary content based on previous instances.
[0011] For example, historical interaction data can indicate that other users who viewed content related to hail damage also used their automated assistant to call an insurance company. Based on this historical interaction data, and in response to a user viewing that section of webpage data, an assistant suggestion for calling the insurance company can be rendered for the automated assistant. This assistant suggestion can be rendered simultaneously with the rendering of that section of webpage data, and in response to the user selecting the assistant suggestion, the automated assistant can cause the computing device or another computing device to call the insurance company. As the user scrolls through the section of webpage data related to hail damage, the automated assistant can replace that assistant suggestion (e.g., "Call my insurance provider") with a different assistant suggestion based on another section of subsequently rendered webpage data.
[0012] In some implementations, one or more assistant suggestions rendered with search result data and / or other application data may correspond to actions that modify and / or enhance the search result data and / or other application data. For example, when a user is viewing the "hail damage" section of webpage data, the user can select an assistant suggestion to have other webpage data rendered. This other webpage data may correspond to a form for receiving information about hail damage insurance. When the automation assistant determines that the user is viewing the "address" section of the form, it may render another assistant suggestion. This other assistant suggestion may correspond to an action to have the user's address data populated into the form with the user's prior permission. When the user clicks on the other assistant suggestion, the automation assistant may have the address data placed in the form, and then the alternative suggestion may be rendered on the display screen.
[0013] In some implementations, search result data can be dynamically adapted based on the distance between the user and the display interface rendering the search result data. This can facilitate time-efficiency reviews of search results based on the user's distance from the display interface. Furthermore, search result data can be dynamically adapted based on any size limitations of the display interface and / or any other limitations on the user's access to the automation assistant's device. For example, when the user is outside a threshold distance from the display interface rendering the search result, the automation assistant can render a single search result at the display interface, making the search result visible from the user's distance. When the user is outside the threshold distance, the automation assistant can receive input from the user for navigating to different search results (e.g., swipe gestures, spoken words, etc. at a non-zero distance from the display interface). When the user moves within the threshold distance, the automation assistant can render additional details about the currently rendered search result. Alternatively or additionally, when the user is within the threshold distance, the automation assistant can render one or more additional search results from the search result data, allowing the user to identify the appropriate search result more quickly. This can reduce search navigation time and thus conserve power and other resources consumed while the user continues navigating the assistant's search results.
[0014] The above description is provided as an overview of some embodiments of this disclosure. Further descriptions of these and other embodiments are given below in more detail.
[0015] Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., multiple central processing units (CPUs), multiple graphics processing units (GPUs), and / or multiple tensor processing units (TPUs)) to perform one or more methods such as those described above and / or elsewhere herein. Further embodiments may include a system of one or more computers including one or more processors operable to execute stored instructions to perform one or more methods such as those described above and / or elsewhere herein.
[0016] It should be understood that all combinations of the foregoing and additional concepts described in more detail herein are contemplated as part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as part of the subject matter disclosed herein. Attached Figure Description
[0017] Figure 1A and Figure 1B The illustration depicts a view of a user interacting with an automated assistant to reveal assistant suggestions that may accompany at least a portion of the content.
[0018] Figure 2A , Figure 2B and Figure 2C The illustration shows a view of a user interacting with an automation assistant to view content, which can provide a basis for certain assistant suggestions that can be rendered in certain scenarios.
[0019] Figure 3 The illustration shows a system that provides an automated assistant that can offer suggestions based on content rendered on the display interface and can adapt the suggestions as the user and / or the automated assistant navigates the content.
[0020] Figure 4 The illustration shows a method for providing assistant suggestions that can help users navigate content on a computing device's display interface and can dynamically adapt as the user navigates the content.
[0021] Figure 5 This is a block diagram of an example computer system. Detailed Implementation
[0022] Figure 1A and Figure 1B The illustrations depict views 100 and 140 where user 102 interacts with an automation assistant to reveal assistant suggestions that may accompany at least a portion of the content. For example, user 102 may provide verbal utterances 106, such as “Assistant, show my deep-dish pizza recipe,” which may cause the automation assistant to render a specific recipe website 126 at computing device 104. A portion of the content of the recipe website 126 rendered at the display interface 110 of computing device 104 may be processed by the automation assistant. The content may be processed to generate one or more assistant suggestions based on the content and / or any other relevant data.
[0023] In some implementations, the automation assistant may process additional content accessible via a website and / or application to supplement the content rendered at display interface 110. For example, recipe website 126 may include a background section, ingredient section, instruction section, and media section where the user can watch a video of the recipe being prepared. Based on these additional sections of content, the automation assistant may generate a first assistant suggestion 114, which, when selected, causes display interface 110 to render the ingredient section of recipe website 126. In some implementations, the assistant suggestion may be initialized in response to verbal input, touch gestures, motion gestures at a non-zero distance from display interface 110, and / or any other input that may be provided to the automation assistant.
[0024] In some implementations, the automated assistant may generate one or more assistant suggestions based on historical interaction data and / or contextual data associated with user 102. For example, with prior permission from user 102, the automated assistant may access historical interaction data indicating that user 102 interacted with an app for ordering flour several months ago. Based on this determination, and if user 102 views a website that mentions flour in the ingredients section, the automated assistant may generate a second assistant suggestion 116 corresponding to the flour ordering action. For example, in response to user 102 selecting the second assistant suggestion 116, the automated assistant may access an app for ordering groceries and select "flour" as the item to purchase.
[0025] Alternatively or additionally, based on the recipe website 126 with media sections, the automation assistant can generate a third assistant suggestion 118, which corresponds to a "shortcut" to another part of the content (e.g., a recipe video). In some implementations, the third assistant suggestion 118 can be generated based on historical interaction data indicating that a user or another user, with or without their automation assistant, accessed the recipe website 126 and viewed the recipe videos included in the recipe website 126. In this way, user 102 can be provided with assistant suggestions that can save them time by not requiring them to scroll through the entire recipe website 126 to ultimately watch the video. Moreover, user 102 may be unaware of the existence of the recipe videos; therefore, with prior permission from other users, the automation assistant can learn from past interactions with other users to simplify the interaction between user 102 and their computing device 104.
[0026] In some implementations, the automation assistant may generate additional assistant suggestions based on the content of the display interface 110 and / or one or more previous requests from user 102. For example, based on user 102 requesting a "deep dish pizza" recipe and viewing the content, the automation assistant may generate a fourth assistant suggestion 120 corresponding to a different website and / or application. For example, the fourth assistant suggestion 120 may correspond to navigating to another website that provides pizza dough recipes instead of Chicago-style pizza.
[0027] When user 102 provides additional spoken words 112, such as "See ingredients section," the automation assistant can process audio data corresponding to the additional spoken words 112. The automation assistant can determine, based on the audio data, that user 102 is selecting a first assistant suggestion 114. In some implementations, the speech processing may be biased based on the content rendered at the display interface 110 and / or the natural language content associated with the assistant suggestion. In response to receiving additional spoken words, the automation assistant can cause computing device 104 to render another part of the recipe website 126, such as... Figure 1B As illustrated in view 140.
[0028] For example, the automation assistant can render the "ingredients" section of recipe website 126 on display interface 110. This conserves computing resources and power that might otherwise be consumed if user 102 is prompted to scroll to the ingredients section by manually clicking and dragging scroll element 128. In some implementations, the automation assistant can render one or more additional assistant suggestions based on the content rendered at display interface 110. The additional assistant suggestions can be based on updated historical interaction data that indicates, for example, that user 102 has previously viewed the "background" section of recipe website 126. Based on this updated historical interaction data, the automation assistant can omit the assistant suggestion corresponding to the "background" section. Additionally or alternatively, the automation assistant can determine, based on updated historical interaction data, that user 102 has not viewed the recipe video of recipe website 126. Based on this determination, the automation assistant can cause display interface 110 to continue rendering a third assistant suggestion 118, even if the user is currently viewing a different part of recipe website 126.
[0029] Alternatively or additionally, the automation assistant can determine that user 102 interacts with the smart oven within a threshold timeframe while viewing recipe website 126. Based on this determination, the automation assistant can cause display interface 110 to render assistant suggestions 144 for accessing the smart oven. Assistant suggestions 144 can be selected by user 102 using verbal phrases such as "Show me the oven," which can initialize applications associated with the smart oven. In some embodiments, the assistant suggestions rendered after user 102 navigates web page data and / or application data can be based on how other users interact with the automation assistant while viewing web page data and / or application data. For example, one or more users viewing recipe website 126 might invoke their automation assistant to play music while they are preparing a recipe. Based on this information, when user 102 is viewing the "ingredients" section of recipe website 126, the automation assistant can render assistant suggestions 148 for playing music. Alternatively or additionally, one or more other users viewing recipe website 126 may have invoked its automation assistant to view other similar recipes, such as those prepared by the same chef who wrote recipe website 126. Based on this determination, the automation assistant may cause display interface 110 to render assistant suggestions 142 to direct the user to other recipes from this chef (e.g., "See more recipes from this chef").
[0030] In some implementations, while the automation assistant is reading content from user 102, the settings can be modified. Figure 1A The assistant suggestions rendered at point 110 in the display interface and in Figure 1B Other assistant suggestions are rendered at the display interface 110. For example, when user 102 has requested the automation assistant to retrieve website content and / or application content for user 102, the automation assistant can scroll the content along the display interface 110 when the content is specified. As the content scrolls along the display interface 110, the automation assistant can dynamically render assistant suggestions at the display interface 110. One or more assistant suggestions rendered at the display interface 110 can be generated based on a portion of the content currently being rendered at the display interface 110. Therefore, various different assistant suggestions can be rendered at the display interface 110 at multiple different times when the automation assistant audibly specifies natural language content to be rendered at the display interface 110.
[0031] Figure 2A , Figure 2B and Figure 2CThe illustration shows views 200, 240, and 260 where user 102 interacts with an automation assistant to view content, which can provide a basis for certain assistant suggestions that can be rendered in certain scenarios. For example, user 202 can provide the automation assistant with a verbal phrase 212, such as “Assistant, control my lights,” which is accessible via computing device 204. In response to receiving the verbal phrase 212, the automation assistant can cause the display interface 210 of computing device 204 to render content based on the request implemented in the verbal phrase 212. For example, the content may include an Internet of Things (IoT) home application 208 that allows user 202 to control the kitchen lights in their home.
[0032] The content of display interface 210 can be processed by an automation assistant and / or another application to generate assistant suggestions that can be simultaneously rendered to the interface of the IoT home application 208 being rendered. Assistant suggestions may include assistant suggestion 216 for accessing user 202's utility bills, assistant suggestion 218 for purchasing new lights, assistant suggestion 220 for controlling other devices in user 202's home, and / or assistant suggestion 222 for communicating with specific content (e.g., "Call mom"). Each assistant suggestion may be generated based on the content of display interface 210, the content of spoken words 212, historical interaction data, context data, and / or any other data accessible to the automation assistant.
[0033] User 202 can select a specific assistant suggestion by providing another verbal phrase 214, such as “See my electric bill,” which may refer to assistant suggestion 216 for accessing the utility website. In response to receiving the other verbal phrase 214, the automated assistant can cause display interface 210 to render utility website 242, which may provide details about utility usage in user 202's household with prior permission from user 202. The automated assistant can process the content of display interface 210 to render additional assistant suggestions. In some implementations, the automated assistant can process additional content of utility website 242 that may not currently be rendered at display interface 210 to generate suggestion data characterizing assistant suggestion 246 (e.g., “View usage stats”). Assistant suggestion 246 may be a navigation suggestion that implements more concise content than the content linked to by assistant suggestion 246. For example, assistant suggestion 246 may link to a portion of utility website 242 that includes monthly usage data with charts and schedules. This part of utility website 242 can be processed by an automation assistant to generate content “View usage stats”, which can be a text summary and / or concise representation of this part of utility website 242.
[0034] Alternatively or additionally, the automation assistant may identify one or more actions that user 202 can perform to interact with content at display interface 210. For example, the automation assistant may provide an assistant suggestion 248 such as "Read this to me," which, when selected, causes the automation assistant to read natural language content from utility website 242 to user 202. Alternatively or additionally, the automation assistant may provide assistant suggestions 244 and / or assistant suggestions 222 based on historical interaction data that may instruct other users on how to interact with the automation assistant when viewing content similar to utility website 242.
[0035] In some implementations, with prior permission from the content author, the automated assistant can adapt and / or enhance the content of the display interface 210 based on how the user 202 interacts with the content. For example, when the user 202 is located at a distance from... Figure 2BWhen the suggestion element 250 is located at a first distance beyond a threshold distance from the display interface 210, the automation assistant can render the suggestion element 250 in a stacked arrangement. The stacked arrangement can be an arrangement of suggestion elements that can be modified by a gesture, causing the suggestion element (e.g., utility website 242) to be removed from the foreground of the display interface 210. When a suggestion element in the foreground of the display interface 210 is discarded, another alternative suggestion in the stacked arrangement can be rendered in the foreground of the display interface 210.
[0036] When user 202 repositions to the second distance within the threshold distance of the display interface 210, such as Figure 2C As illustrated in view 260, the automation assistant may cause the assistant suggestions to no longer appear in the stacked arrangement. Instead, in some embodiments, the automation assistant may cause the assistant suggestions to appear on the display interface 210 in a "carousel-like" arrangement and / or coupled arrangement. When the user is within a threshold distance, the user 202 may provide an input gesture to the computing device 204, which causes the content element 262 (i.e., the selectable content element) to move simultaneously on the display interface 210. For example, when a left swipe gesture is performed at the display interface 210, the content element of the utility website 242 may move to the left and / or at least partially remove from the display interface 210. Simultaneously, and in response to the left swipe gesture, content elements 264, 266, and / or 270 may be further rendered to the left side of the display interface 210, while also revealing another content element(s) not previously displayed at the display interface 210.
[0037] If user 202 returns to a distance beyond the threshold with prior permission from user 202, as detected by computing device 204 and / or the automation assistant, then content element 262 may return to the specified distance. Figure 2B The display interface 210 can be arranged in a stacked layout. In this way, when the user 202 is further away from the computing device 204, the display interface 210 can render fewer suggestion elements compared to when the user 202 is closer to the computing device 204. Alternatively or additionally, when the user 202 is closer to the computing device 204 (e.g., ... Figure 2C Compared to the previous method, when the user 202 is farther away from the computing device 204, each content element 262 can be rendered over a larger area (e.g., ...). Figure 2B middle).
[0038] Figure 3The illustration depicts a system 300 providing an automation assistant 304, which can offer assistance suggestions based on content rendered at a display interface and adapt these suggestions as the user and / or the automation assistant 304 browses the content. The automation assistant 304 can operate as part of an assistant application provided on one or more computing devices, such as computing device 302 and / or server devices. Users can interact with the automation assistant 304 via assistant interfaces(s) 320, which can be a microphone, camera, touchscreen display, user interface, and / or any other device capable of providing an interface between the user and the application. For example, a user can initialize the automation assistant 304 by providing language, text, and / or graphical input to the assistant interface 320, causing the automation assistant 304 to initiate one or more actions (e.g., providing data, controlling peripheral devices, accessing agents, generating input and / or output, etc.).
[0039] Alternatively, the automation assistant 304 may be initialized based on the processing of context data 336 using one or more trained machine learning models. Context data 336 may characterize one or more features of the environment accessible to the automation assistant 304 and / or be predicted to be one or more features of a user intended to interact with the automation assistant 304. The computing device 302 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures to allow the user to control applications 334 of the computing device 302 via the touch interface. In some embodiments, the computing device 302 may lack a display device, thereby providing audible user interface output instead of graphical user interface output. Furthermore, the computing device 302 may provide a user interface, such as a microphone, for receiving spoken natural language input from the user. In some embodiments, the computing device 302 may include a touch interface and may not have a camera, but may optionally include one or more other sensors.
[0040] Computing device 302 and / or other third-party client devices can communicate with the server device via a network (such as the Internet). Additionally, computing device 302 and any other computing devices can communicate with each other via a local area network (LAN) (such as a Wi-Fi network). Computing device 302 can offload computational tasks to the server device to conserve computing resources at computing device 302. For example, the server device can host an automation assistant 304, and / or computing device 302 can transmit input received at one or more assistant interfaces 320 to the server device. However, in some implementations, the automation assistant 304 can be hosted at computing device 302, and various processes associated with the operation of the automation assistant can be executed at computing device 302.
[0041] In various implementations, all or fewer aspects of the automation assistant 304 may be implemented on the computing device 302. In some of these implementations, aspects of the automation assistant 304 may be implemented via the computing device 302 and may interface with a server device that may implement other aspects of the automation assistant 304. The server device may optionally serve multiple users and their associated assistant applications via multiple threads. In implementations where all or fewer aspects of the automation assistant 304 are implemented via the computing device 302, the automation assistant 304 may be an application separate from the operating system of the computing device 302 (e.g., installed "on top" of the operating system), or alternatively, it may be implemented directly by the operating system of the computing device 302 (e.g., considered an application of the operating system, but integrated with the operating system).
[0042] In some implementations, the automation assistant 304 may include an input processing engine 306, which may employ multiple different modules to process inputs and / or outputs from the computing device 302 and / or the server device. For example, the input processing engine 306 may include a voice processing engine 308, which can process audio data received at the assistant interface 320 to identify text implemented within the audio data. The audio data may be transferred from, for example, the computing device 302 to the server device to conserve computing resources at the computing device 302. Additionally or alternatively, the audio data may be processed exclusively at the computing device 302.
[0043] The process of converting audio data into text may include a speech recognition algorithm that employs neural networks and / or statistical models to identify groups of audio data corresponding to words or phrases. The text converted from the audio data may be parsed by a data parsing engine 310 and used as text data by an automation assistant 304. This text data may be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, the output data provided by the data parsing engine 310 may be provided to a parameter engine 312 to determine whether the user has provided input corresponding to a specific intent, action, and / or routine that can be performed by the automation assistant 304 and / or by an application or agent accessible via the automation assistant 304.
[0044] For example, assistant data 338 may be stored on a server device and / or computing device 302, and may include data defining one or more actions that can be performed by the automation assistant 304, as well as parameters necessary to perform these actions. Parameter engine 312 may generate one or more parameters of intent, action, and / or slot values, and provide one or more parameters to output generation engine 314. Output generation engine 314 may use one or more parameters to communicate with automation assistant 320 to provide output to the user, and / or to communicate with one or more applications 334 to provide output to one or more applications 334.
[0045] In some implementations, the automation assistant 304 may be an application that can be installed "on top" of the operating system of the computing device 302, and / or may form part (or all) of the operating system of the computing device 302 itself. The automation assistant application includes and / or accesses on-device speech recognition, on-device natural language understanding, and on-device performance. For example, on-device speech recognition can be performed using an on-device speech recognition module that processes audio data (detected by microphone(s)) using an end-to-end speech recognition machine learning model locally stored at the computing device 302. On-device speech recognition generates recognized text of spoken utterances (if any) present in the audio data. Furthermore, on-device natural language understanding (NLU) can be performed, for example, using an on-device NLU module that processes the recognized text generated using on-device speech recognition, along with optional contextual data, to generate NLU data.
[0046] NLU data may include multiple intents corresponding to a spoken utterance and optionally multiple parameters (e.g., slot values) of those intents. On-device fulfillment may be performed using an on-device fulfillment module that utilizes NLU data (from on-device NLU) and optionally other local data to determine multiple actions to be taken to resolve the multiple intents (and optionally multiple parameters) of the spoken utterance. This may include determining local and / or remote responses (e.g., answers) to the spoken utterance, interactions with multiple locally installed applications performed based on the spoken utterance, multiple commands transmitted to multiple Internet of Things (IoT) devices (directly or via corresponding remote systems) based on the spoken utterance, and / or multiple other resolution actions performed based on the spoken utterance. On-device fulfillment may then initiate local and / or remote execution / enforcement of the determined multiple actions to resolve the spoken utterance.
[0047] In various implementations, remote speech processing, remote NLU, and / or remote execution can be utilized at least selectively. For example, recognized text can be transmitted at least selectively to at least one of the remote automation assistant components for remote NLU and / or remote execution. For example, recognized text can be selectively transmitted in parallel with on-device execution for remote execution, or in response to a failure of on-device NLU and / or on-device execution. However, on-device speech processing, on-device NLU, on-device execution, and / or on-device execution can be prioritized at least due to the reduced latency they provide when resolving spoken utterances (since no client-server round trips are required to resolve spoken utterances). Furthermore, on-device functionality may be the only functionality available in the absence of or with limited network connectivity.
[0048] In some implementations, computing device 302 may include one or more applications 334, which may be provided by a third-party entity different from the entity providing computing device 302 and / or automation assistant 304. The application state engine of automation assistant 304 and / or computing device 302 may access application data 330 to determine one or more actions that can be performed by the one or more applications 334, as well as the state of each application in the one or more applications 334 and / or the state of the corresponding device associated with computing device 302. The device state engine of automation assistant 304 and / or computing device 302 may access device data 332 to determine one or more actions that can be performed by computing device 302 and / or one or more devices associated with computing device 302. In addition, application data 330 and / or any other data (such as device data 332) can be accessed by the automation assistant 304 to generate context data 336, which can characterize the context in which application 334 and / or device are being executed and / or the context in which user accesses computing device 302, accesses application 334 and / or any other device or module.
[0049] When one or more applications 334 are executed at computing device 302, device data 332 can characterize the current operating state of each application 334 executed at computing device 302. Furthermore, application data 330 can characterize one or more features of the executing application 334, such as the content of one or more graphical user interfaces rendered in the direction of the one or more applications 334. Alternatively or additionally, application data 330 can characterize action patterns that can be updated by the respective application and / or automation assistant 304 based on the current operating state of the respective application. Alternatively or additionally, one or more action patterns of one or more applications 334 can remain static but can be accessed by the application state engine to determine the appropriate action to be initialized via automation assistant 304.
[0050] In some implementations, system 300 may include a prior interaction engine 324 that processes data associated with prior interactions between one or more users and one or more of their respective automation assistants. The prior interaction engine 324 may process this data to identify actions that may have been initiated by the user when the user viewed certain content at computing device 302. Subsequently, when the user or another user is viewing the specific content, the prior interaction engine 324 may identify one or more actions that may be associated with that user or the other user.
[0051] In some implementations, system 300 may include an assistant suggestion engine 316 that processes data accessible to system 300 to generate one or more assistant suggestions to recommend to a user via assistant interface 320 of computing device 302. For example, a prior interaction engine 324 may determine that the user is viewing something associated with a previous interaction between the user and the automation assistant 304. Based on this determination, assistant suggestion engine 316 may generate suggestion data characterizing one or more alternative suggestions based on the previous interaction and what the user is currently viewing. For example, during a previous interaction, the user may have viewed recipe ingredients and then added them to a website checkout page using a grocery app on their phone. When the user subsequently views the recipe ingredients, prior interaction engine 324 may query the grocery app with prior permission from the user to determine if the user has purchased the ingredients. Based on this determination, assistant suggestion engine 316 may generate suggestion data that advises the user to complete their purchase of the ingredients or view another part of the recipe corresponding to the recipe ingredients (e.g., view ingredient preparation instructions).
[0052] In some implementations, system 300 may include an element placement engine 326 that can place content elements and / or assistant suggestions based on data accessible to system 300. For example, when a user provides a query to automation assistant 304 about viewing certain content, automation assistant 304 may identify multiple different instances of content to be presented to the user. When the user is within a threshold distance of computing device 302, element placement engine 326 may arrange multiple content elements in a coupled arrangement on the display interface of computing device 302. Alternatively, when the user is within a threshold distance of computing device 302, element placement engine 326 may arrange multiple content elements in a stacked arrangement. When content elements are in a stacked arrangement, the area of the foreground content element may be larger than one or more of the content elements; otherwise, these content elements will be rendered in a coupled arrangement (i.e., a carousel arrangement).
[0053] In some implementations, system 300 may include an entry point engine 318 that can determine how a user arrived at a specific page of content, with prior permission from the user. This determination may be based on data accessible to system 300 and may be used by an assistant suggestion engine 316 and / or an element placement engine 326 to generate and / or place assistant suggestions and / or content elements. For example, entry point engine 318 may determine that a first user arrived at the recipe page via a link sent to them via text message by a friend (e.g., Matthew). Based on this determination, assistant suggestion engine 316 may generate assistant suggestions corresponding to actions for sending text messages (e.g., “Send a message to Matthew”). Entry point engine 318 may also determine that a second user arrived at the recipe page via a link in a video. Based on this determination, assistant suggestion engine 316 may generate assistant suggestions corresponding to actions for returning to view the video and / or for viewing another video based on content selections on the recipe page.
[0054] In some implementations, element placement engine 326 can modify and / or enhance the content of content elements rendered at computing device 302 with prior permission from the content author. For example, element placement engine 326 can determine, based on data available to system 300, that content rendered at computing device 302 can be modified and / or enhanced based on information associated with the user. In some instances, based on application data 330, device data 332, and / or context data 336, the assistant suggests that a portion of the content can be rendered above and / or in place of a portion of the content. For example, when a user is searching for information about a food allergy and then visits a website for viewing recipes, element placement engine 326 can modify and / or supplement a portion of the recipe content to include food allergy information previously viewed by the user. For example, a portion of the interface relating to "wheat flour" could be supplemented with a content element relating to "chickpea flour," which the user may have viewed after requesting automation assistant 304 to "show information about wheat allergies."
[0055] Figure 4The illustration depicts a method 400 for providing assistant suggestions that can assist a user in navigating content on a computing device's display interface and can dynamically adapt as the user navigates the content. Method 400 can be performed by one or more devices, applications, and / or any other application or module capable of interacting with an automation assistant. Method 400 may include actions to determine whether the automation assistant has received a request to render content on the computing device's display interface. This request can be implemented in a spoken statement or other input that the automation assistant can respond to. For example, the request can be implemented in a spoken statement such as, "Assistant, how much can a solar panel reduce my electric bill?" In response to receiving the spoken statement, the automation assistant can determine that the user is requesting the automation assistant to render content related to a solar panel.
[0056] When it is determined that the request for content was provided by a user, method 400 can proceed from operation 402 to operation 404, which may include causing the display interface of the computing device to render a first portion of the content (e.g., a solar panel review article) for the user. According to the example above, the automation assistant may render content obtained from an application and / or website in response to receiving spoken words. The content may include text, graphics, video, and / or any other content accessible to the automation assistant. Due to the limited size of the display interface of the computing device, the automation assistant may render a first portion of the content, and optionally, preload a second portion of the content for rendering.
[0057] Method 400 can proceed from operation 404 to operation 406, which may include generating a first assistant suggestion based on a first portion of content rendered at the display interface of the computing device. For example, the first portion of the content may include details about how to reduce electricity bills. Based on this first portion of the content, the automation assistant may generate the first assistant suggestion and identify an action for the automation assistant to reduce the output of one or more lights in the user's home. For example, the first assistant suggestion may include natural language content such as "Turn off my basement lights," and alternatively or additionally, the first assistant suggestion may correspond to an action by the automation assistant scrolling from the first portion of the content to a second portion of the content. For example, the first assistant suggestion may include natural language content such as "Go to 'SolarPanel Prices'," which may be a reference to the second portion of the content, detailing the prices of certain solar panels.
[0058] In some implementations, the first assistant suggestion can be generated through one or more interactions between the user and the automated assistant before the user provides verbal input. For example, the interaction between the user and the automated assistant could involve the user requesting the automated assistant to modify the settings of one or more lights in the user's home. Alternatively or additionally, with prior permission from other users, the first assistant suggestion can be generated based on how other users interact with their respective automated assistants while viewing solar panel review articles. For example, historical interaction data could indicate that one or more other users have quickly scrolled to solar panel prices. Based on this historical interaction data, the automated assistant can generate a first assistant suggestion corresponding to the action of scrolling to the "Solar Panel Prices" section of a solar panel review article.
[0059] Method 400 can proceed from operation 406 to operation 408, which may include rendering the first assistant suggestion on the display interface of the computing device. When the first assistant suggestion is rendered on the display interface, method 400 may proceed to operation 410 to determine whether the user has rendered the second part of the content on the display interface. When the automation assistant determines that the user has not rendered the second part of the content on the display interface, method 400 may optionally return to operation 402. However, when the automation determines that the user has rendered the second part of the content on the display interface, method 400 may proceed to operation 412.
[0060] Operation 412 may include enabling the automation assistant to generate a second assistant suggestion based on a second part of the content. The second assistant suggestion may include natural language content that summarizes another part of the content accessed by the user in response to spoken utterance received in operation 402. For example, the automation assistant may generate a second assistant suggestion to draw the user's attention to another part of the content, which may include information that the user might be interested in, thus saving the user time and effort from scrolling through the entire article. For example, based on the type of smart device installed in the user's home (e.g., a smart thermostat), the automation assistant may identify another part of the article related to the smart device (e.g., "Subtitle: Tips for Scheduling a Smart Thermostat"). The automation assistant may use natural language understanding and / or one or more other natural language processing techniques to provide a summary of the other part of the article (e.g., "Use Geofencing feature of your thermostat to reduce energy").
[0061] Alternatively or additionally, when the second part of the content discusses calling an entity that might help fulfill a request from the user (such as a power company), the automation assistant may generate a second assistant suggestion for controlling the communication operations of the automation assistant. For example, the second part of the content may discuss calling the power company to discuss net metering options for solar energy, and based on this information, the automation assistant may identify the power company used by the user (e.g., based on application data and / or other contextual data). The second assistant suggestion may then include natural language content such as "Call my utility company," and, in response to the user selecting the second assistant suggestion, the automation assistant may make a call to the utility company.
[0062] When the second assistant suggestion has been generated, method 400 may proceed to operation 414, which renders the second assistant suggestion along with a second portion of the content on the display interface of the computing device. In some embodiments, the second assistant suggestion and / or the second portion of the content may be arranged according to the user's distance from the display interface. For example, when the user is within a threshold distance from the computing device, the content element corresponding to that content may be displayed along with other content elements identified in response to a request from the user. Each content element may be coupled such that when the user provides a specific input gesture (e.g., a swipe gesture) to the automation assistant, the content element may move in the direction of the gesture (e.g., in a turntable manner) to reveal other content elements (e.g., other solar panel review articles). Alternatively, when the user is outside the threshold distance, the content element may be arranged in a stacked arrangement on top of other content elements. When the user provides a specific input gesture to the automation assistant, the content element corresponding to the first portion of the content may move to reveal another content element identified in response to a request. Method 400 may optionally return to operation 402 for further detection of another request from the user.
[0063] Figure 5 This is a block diagram 500 of an example computer system 510. The computer system 510 typically includes at least one processor 514 that communicates with a plurality of peripheral devices via a bus subsystem 512. These peripheral devices may include a storage subsystem 524 (including, for example, memory 525 and file storage subsystem 526), a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices allow users to interact with the computer system 510. The network interface subsystem 516 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.
[0064] User interface input device 522 may include a keyboard, pointing devices (such as a mouse, trackball, touchpad, or drawing tablet), scanner, touchscreen incorporated into a display, audio input devices (such as a voice recognition system, microphone), and / or other types of input devices. Generally, the term "input device" is used to encompass all possible types of devices and methods for inputting information into computer system 510 or into a communication network.
[0065] User interface output device 520 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or other mechanisms for creating visible images. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and methods for outputting information from computer system 510 to a user or to another machine or computer system.
[0066] Storage subsystem 524 stores the programming and data construction functionality of some or all of the modules described herein. For example, storage subsystem 524 may include logic for performing selected aspects of method 400 and / or implementing system 300, computing device 104, computing device 204, automation assistant, and / or any other application, device, apparatus, and / or module discussed herein.
[0067] These software modules are typically executed by processor 514 alone or in combination with other processors. Memory 525 in storage subsystem 524 may include several memories, including main random access memory (RAM) 530 for storing instructions and data during program execution and read-only memory (ROM) 532 where fixed instructions are stored. File storage subsystem 526 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives with associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Functional modules implementing the specific embodiments may be stored by file storage subsystem 526 in storage subsystem 524 or in other machines accessible by processor(s) 514.
[0068] Bus subsystem 512 provides mechanisms for allowing various components and subsystems of computer system 510 to communicate with each other as intended. Although bus subsystem 512 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0069] Computer system 510 can be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, for the purposes of illustrating some embodiments, Figure 5 The description of the computer system 510 depicted is intended only as a concrete example. Figure 5 Compared to the computer system described in the diagram, computer system 510 has many other configurations with more or fewer components.
[0070] In situations where the system described herein collects or can utilize personal information about a user (or generally referred to herein as a "participant"), the user may be provided with the opportunity to: control whether the program or feature collects user information (e.g., information about the user's social networks, social actions or activities, occupation, user preferences, or the user's current geographic location) or to control whether and / or how content that may be more relevant to the user is received from the content server. Furthermore, before specific data is stored or used, it may be disposed of in one or more ways to remove personally identifiable information. For example, a user's identifier may be disposed of such that the user's personally identifiable information cannot be determined, or the user's geographic location (from which geographic location information (such as city, zip code, or state) is obtained) may be generalized such that the user's specific geographic location cannot be determined. Therefore, the user can control how information about them is collected and / or used.
[0071] While several embodiments have been described and illustrated herein, various other components and / or structures may be utilized for performing functions and / or obtaining results and / or one or more of the advantages described herein, and each of such variations and / or modifications is considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications in which the teachings are used. Those skilled in the art will recognize, or can determine, many equivalents of the specific embodiments described herein using only conventional experimentation. Therefore, it is to be understood that the foregoing embodiments are presented by way of example only, and that embodiments may be practiced in ways different from the specific descriptions and claims within the scope of the appended claims and their equivalents. Embodiments of this disclosure relate to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of this disclosure if such features, systems, articles, materials, kits, and / or methods do not contradict each other.
[0072] In some embodiments, a method implemented by a processor(s) is provided, and includes receiving at a computing device a request for an automation assistant to render content at a display interface of the computing device. The request is implemented in verbal utterance provided by a user to the computing device. The method also includes, in response to receiving the request from the automation assistant, causing the display interface of the computing device to render a first portion of the content. The method further includes having the automation assistant process content data representing the first portion of the content rendered at the display interface of the computing device. The content data is processed to facilitate the generation of a first assistant suggestion to be rendered at the display interface. The method also includes, based on the processed content data, causing the display interface of the computing device to render the first assistant suggestion and the first portion of the content. The method further includes, when the first portion of the content is rendered at the display interface: by the automation assistant, determining that the user has provided input to facilitate the rendering of a second portion of the content, wherein the second portion of the content differs from the first portion of the content; in response to the input, processing other content data representing the second portion of the content, wherein the other content data is processed to facilitate the generation of a second assistant suggestion to be rendered at the display interface; and, based on other supplementary content, causing the display interface of the computing device to render the second assistant suggestion and the second portion of the content, wherein the second assistant suggestion differs from the first assistant suggestion.
[0073] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0074] In some implementations, the input from the user relates to a first assistant suggestion, and causes a second portion of the content to be rendered on the display interface of the computing device, wherein the first assistant suggestion includes natural language content identifying the second portion of the content. In some of these implementations, the input is additional spoken language provided by the user to the automated assistant.
[0075] In some implementations, the second assistant suggests an operation performed by an automation assistant at a separate computing device, which is different from the computing device itself.
[0076] In some implementations, processing content data includes generating a first assistant suggestion based on historical interaction data. For example, historical interaction data may characterize previous interactions in which a user or another user caused a computing device or another computing device to perform an action, while some content was rendered on the computing device or another computing device. For example, some content may include a first portion of the content.
[0077] In some implementations, when the second assistant suggestion is rendered on the display screen, the first assistant suggestion is omitted from the display screen. In some of these implementations, when the second assistant suggestion is rendered on the display screen, the first portion of the content is omitted from the display screen.
[0078] In some implementations, the second assistant suggestion includes natural language content based on a second part of the content, and the selection of the second assistant suggestion causes the automated assistant to initialize a separate application that is different from the automated assistant.
[0079] In some implementations, a method implemented by processor(s) is provided, and includes receiving at a computing device a request for an automated assistant to render content at a display interface of the computing device. The request is implemented in verbal utterance provided by a user to the computing device. The method also includes causing the display interface to render the content in response to the request. The method further includes processing interaction data in response to the request, the interaction data representing a previous interaction by the user that caused the content to be rendered at the display interface of the computing device. During the previous interaction, a first assistant suggestion was rendered at the display interface along with the content. The method also includes generating suggestion data based on the processed interaction data, the suggestion data representing a second assistant suggestion different from the first assistant suggestion. The second assistant suggestion is further generated based on the content rendered at the display interface. The method also includes causing the display interface to render the second assistant suggestion and the content based on the suggestion data. The second assistant suggestion is selectable via subsequent verbal utterance from the user to the automated assistant.
[0080] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0081] In some implementations, the second assistant suggestion is rendered along with specific natural language content, and the method further includes biasing one or more speech recognition processes toward the specific natural language content of the second assistant suggestion when the second assistant suggestion is rendered at a display device.
[0082] In some implementations, the first assistant suggestion is rendered along with the specific natural language content, and the method further includes causing one or more speech recognition processes to deviate from the specific natural language content when the second assistant suggestion is rendered at the display device.
[0083] In some implementations, the second assistant suggestion corresponds to an operation initialized via an automation assistant. In some of these implementations, the method further includes, based on the second assistant suggestion, enabling the computing device to access operation data corresponding to the operation before the user provides subsequent verbal utterances to the automation assistant.
[0084] In some implementations, receiving a request for the automation assistant to render content on the display interface of the computing device includes receiving a selection of search results from a list of search results rendered on the display interface of the computing device. In some versions of these implementations, the second assistant suggests generating content further based on search results in the search results list. In some additional or alternative versions of these implementations, the second assistant suggests generating content further based on one or more other search results in the search results list.
[0085] In some implementations, a method implemented by processor(s) is provided, and includes receiving at a computing device a request for an automation assistant to render content at a display interface of the computing device. The request is made in a verbal utterance provided by a user located at a first distance from the display interface. The method also includes, in response to the request, causing the display interface of the computing device to render selectable content elements in a coupled arrangement extending across the display interface. Input gestures receivable by the automation assistant cause multiple selectable content elements to move simultaneously across the display interface. The method further includes determining, by the automation assistant, that the user has repositioned from the first distance to a second distance from the display interface. The method also includes, based on the user's repositioning to the second distance from the display interface, causing the selectable content elements to be rendered in a stacked arrangement. In this stacked arrangement, foreground content elements among the selectable content elements are rendered on top of other content elements. The method also includes receiving, by the automation assistant, another request for the automation assistant to reveal a specific content element among the other content elements. The specific content element is distinct from the foreground content element among the selectable content elements. The method also includes responding to other requests by replacing the foreground content element in the selectable content elements with a specific content element in other content elements in the area of the display interface.
[0086] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0087] In some implementations, the method further includes rendering one or more assistant suggestions along with the selectable content elements at the display interface when the selectable content elements are rendered in a coupled arrangement. The one or more assistant suggestions are based on specific content associated with the selectable content elements. In some versions of these implementations, the method further includes rendering one or more additional assistant suggestions along with the selectable content elements at the display interface when the selectable content elements are rendered in a stacked arrangement. The one or more additional assistant suggestions are based on other specific content corresponding to the foreground content elements, and the one or more assistant suggestions are different from the one or more additional assistant suggestions. In some versions of these implementations, when the one or more additional assistant suggestions are rendered at the display interface, the one or more assistant suggestions are omitted from the display interface.
[0088] In some implementations, one or more assistants suggest identifying operations initiated by the automation assistant at a separate computing device, which is different from the computing device.
[0089] In some implementations, one or more assistants suggest actions based on historical interaction data accessible to the automation assistant. Optionally, the historical interaction data characterizes previous interactions in which the user or another user caused the computing device or another computing device to perform an action, and specifically indicates that content was rendered on the computing device or another computing device.
Claims
1. A method implemented by one or more processors, the method comprising: Receive a request at the computing device for the automation assistant to render content on the display interface of the computing device. The request is made in verbal speech provided by the user to the computing device; In response to receiving the request from the automation assistant, the display interface of the computing device renders a first portion of the content, wherein, based on the limited size of the display interface, the first portion of the content is rendered independently of a second portion that also displays the content; The automated assistant processes the content data, which represents the first portion of the content rendered at the display interface of the computing device. The content data is processed to facilitate the generation of a first assistant suggestion to be rendered at the display interface; Based on processing the content data, the display interface of the computing device renders the first assistant suggestion and the first portion of the content; When the first portion of the content is rendered at the display interface: The automation assistant determines that the user has provided input to facilitate the rendering of the second portion of the content by the display interface. Wherein, the second part of the content is different from the first part of the content, and The second part of the content is preloaded when the first part of the content is rendered; In response to the input, other content data representing the second portion of the content is processed. The other content data is processed to facilitate the generation of a second assistant suggestion to be rendered at the display interface; and Based on other supplementary information, the display interface of the computing device renders the second assistant suggestion and the second portion of the content. The second assistant's suggestion differs from the first assistant's suggestion.
2. The method according to claim 1, in, The input from the user relates to the first assistant's suggestion, and causes the second portion of the content to be rendered on the display interface of the computing device. The first assistant suggestion includes natural language content that identifies the second part of the content.
3. The method according to claim 2, wherein, The input is additional verbal utterances provided by the user to the automated assistant.
4. The method according to claim 1, wherein, The second assistant suggestion corresponds to an operation performed by the automation assistant at a separate computing device, which is different from the computing device.
5. The method according to claim 1, wherein, Processing the content data includes: The first assistant suggestion is generated based on historical interaction data. The historical interaction data represents previous interactions in which the user or another user caused the computing device or another computing device to perform an action, while specific content was rendered on the computing device or the other computing device.
6. The method according to claim 5, wherein, The specific content includes the first part of the content.
7. The method according to claim 1, in, When the second assistant suggestion is rendered at the display interface, the first assistant suggestion is omitted from the display interface, and Specifically, when the second assistant suggestion is rendered at the display interface, the first part of the content is omitted from the display interface.
8. The method according to claim 1, in, The second assistant's suggestion includes natural language content based on the second part of the content, and The selection of the second assistant suggestion causes the automation assistant to initialize a separate application, which is different from the automation assistant.
9. A method implemented by one or more processors, the method comprising: Receive a request at the computing device for the automation assistant to render content on the display interface of the computing device. The request is made in verbal speech provided by the user to the computing device; In response to the request, the display interface renders the content; In response to the request, interaction data is processed, the interaction data representing a previous interaction by the user that caused the content to be rendered at the display interface of the computing device. During the previous interaction, the first assistant suggested that the content be rendered on the display interface together with the content. Based on the processed interaction data, suggestion data is generated, which represents a second assistant suggestion that differs from the first assistant suggestion. The second assistant suggests further generating the content based on the content rendered at the display interface; and Based on the suggested data, the display interface renders the second assistant's suggestion and the content. Wherein, the second assistant suggests that subsequent spoken utterances from the user to the automated assistant are selectable and rendered with specific natural language content; and When the second assistant suggestion is rendered at the display interface and based on the second assistant suggestion rendered at the display interface, one or more speech recognition processes are biased toward the specific natural language content of the second assistant suggestion.
10. The method according to claim 9, wherein, The first assistant suggests rendering with specific natural language content, and the method further includes: When the second assistant suggests that the speech recognition process deviate from the specific natural language content when it is rendered on the display interface, one or more speech recognition processes may be made to deviate from the specific natural language content.
11. The method according to claim 9, in, The second assistant suggestion corresponds to the operation initialized via the automated assistant, and The method further includes: Based on the second assistant's suggestion, the operation data corresponding to the operation is accessed by the computing device before the user provides the subsequent verbal utterance to the automation assistant.
12. The method according to claim 9, wherein, Receiving the request for the automation assistant to render content at the display interface of the computing device includes: Receive a selection of a search result from a list of search results rendered at the display interface of the computing device. The second assistant suggests generating the solution based on the search results in the search results list.
13. The method according to claim 9, wherein, Receiving the request for the automation assistant to render content at the display interface of the computing device includes: Receive a selection of a search result from a list of search results rendered at the display interface of the computing device. The second assistant suggests generating the search results based on one or more other search results in the search results list.
14. A system comprising: One or more processors; as well as A memory for storing instructions, which, when executed, cause the one or more processors to perform the operations of any one of claims 1 to 13.
15. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the operations of any one of claims 1 to 13.
Citation Information
Patent Citations
Intelligent automated assistant for TV user interactions
US20150382047A1