Automated assistant for introducing or controlling search filter parameters in individual applications

CN116724304BActive Publication Date: 2026-08-18GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180088467.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-20
Filing Date
2021-12-15
Publication Date
2026-08-18
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

以这种方式在界面之间的切换可能跨越计算设备的许多方面而消耗资源,并且可能增加将由应用和/或自动化助理提供不准确的搜索结果的可能性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116724304B_ABST
    Figure CN116724304B_ABST
Patent Text Reader

Abstract

Implementations set forth herein relate to an automated assistant that can operate as an interface between a user and a separate application to search application content of the separate application. The automated assistant can interact with existing search filter features of the other application and can also adapt in cases where certain filter parameters are not directly controllable at a search interface of the application. For example, when a user requests to perform a search operation using certain items, those items can refer to content filters that can not be available at the search interface of the application. However, the automated assistant can generate assistant input based on those content filters in order to ensure that any resulting search results will be filtered accordingly. The assistant input can then be submitted into a search field of the application and the search operation can be performed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Humans can engage in human-computer dialogue using interactive software applications referred to herein as “automated assistants” (also known as “digital agents,” “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “assistant applications,” “conversational agents,” etc.). For example, humans (who may be referred to as “users” when interacting with automated assistants) can provide commands and / or requests to automated assistants using spoken natural language input (i.e., utterances) and / or by providing text-based (e.g., typed) natural language input, which in some cases can be converted into text and then processed.

[0002] For example, a user who invokes an automation assistant to perform a search operation via a specific application may be limited by whether the application has enabled features for interfacing with the automation assistant. Depending on whether the automation assistant can control certain features of the application, it may only complete a limited number of requests from the user. In these instances, the user may have to be entrusted with individually identifying performed and unperformed requests and then subsequently interacting with the touch interface to manually complete any unperformed requests. Switching between interfaces in this way can be resource-intensive across many aspects of the computing device and may increase the likelihood of inaccurate search results being provided by the application and / or automation assistant. Summary of the Invention

[0003] The embodiments described herein relate to an automation assistant that allows a user to search and / or filter application content (e.g., a website, client application, server application, browser, etc.) by providing verbal commands to the automation assistant without requiring direct input from the user to the application. A search operation can be initiated when the user requests the automation assistant to access the application and search for application content. In response to such a request from the user, the automation assistant can determine whether the application, as identified by the user, provides any features other than the search field for filtering search results. The automation assistant can determine whether the application's search interface includes one or more optional graphical user interface (GUI) elements for restricting the types of content to be included in the search results. When the automation assistant determines that the application interface includes one or more optional filter elements corresponding to one or more items in the assistant's input from the user, the automation assistant can adjust one or more filter elements based on one or more items. The automation assistant can then populate the search field of the application interface with one or more other items identified in the assistant's input and initiate the search operation.

[0004] When initiating a search operation, the application can search for application content related to one or more items in the search fields. As a result, the user receives search results from the application without directly interacting with it to search for application content. Instead, the user relies on an automated assistant to perform Natural Language Understanding (NLU) and / or speech-to-text processing to interact with the application based on the user's requests. In this way, the user can reduce the amount of time spent trying to identify certain filter elements at the application interface and / or manually typing search terms into the application's search fields.

[0005] In some implementations, search results provided by the application can be further filtered by the automation assistant in response to another request from the user to the automation assistant. For example, after the user interacts with the application to provide search results, the user can provide additional verbal utterances to the automation assistant. These verbal utterances can identify one or more additional search terms that the automation assistant can use to filter search results and / or additionally select a subset of search results. For example, in response to receiving additional verbal utterances, the automation assistant can determine whether any additional search terms embodied in the additional verbal utterances correspond to one or more optional filter elements rendered at the application's search results interface. When the automation assistant determines that an additional search term does not correspond to one or more optional filter elements of the application, the automation assistant can generate a search command to be executed by the application. The search command can be generated to ensure that the application provides a subset of search results, rather than resetting any established search parameters used to generate the search results, and / or not starting a new search from an empty state.

[0006] In some implementations, a user can provide an automation assistant with a search request that includes parameters regarding when any search results content should be provided to the user. For example, a user could provide a search request for content from a search app (e.g., a news app) and also specify a subsequent time when the user wants the search results to be provided (e.g., "Assistant, search for blockchain articles from yesterday in my news app and read them to me at 10 AM"). In this way, as the automation assistant operates within the device's ecosystem, it can search and / or download any relevant search results available to the user on the device at the specified time. This also allows the automation assistant to select a reliable network for retrieving search results, rather than downloading content from any network available when the user requests to receive the results content.

[0007] The above description serves as an overview of some embodiments of this disclosure. These and other embodiments are further described in more detail below.

[0008] Other embodiments may include a non-transitory computer-readable storage medium storing instructions that can be executed by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform methods, such as one or more of the methods described above and / or elsewhere herein. Other embodiments may include a system of one or more computers including one or more processors operable to execute the stored instructions to perform methods, such as one or more of the methods described above and / or elsewhere herein.

[0009] It should be understood that all combinations of the foregoing concepts and the additional concepts described in more detail herein are considered part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered part of the subject matter disclosed herein. Attached Figure Description

[0010] Figure 1A , Figure 1B , Figure 1C and Figure 1D The illustration shows a view of a user interacting with an automation assistant to control the app's search operations.

[0011] Figure 2 The diagram illustrates a system that provides an automated assistant for controlling search operations of individual applications, allowing the implementation of search filters regardless of whether the individual application provides quick control over the search filters.

[0012] Figure 3 The diagram illustrates a method for operating an automation assistant to interface with individual applications in order to search and / or filter certain application content.

[0013] Figure 4 This is a block diagram of an example computer system. Detailed Implementation

[0014] Figure 1A , Figure 1B , Figure 1C and Figure 1DViews 100, 120, 140, and 160 illustrate user 102 interacting with an automation assistant to control search operations of application 132. Application 132 may be separate from the automation assistant, but the automation assistant may be allowed to control certain operations of application 132. For example, application 132 could be a hardware shopping application that user 102 could use to purchase computer parts. To invoke the automation assistant to control application 132, user 102 can provide spoken words 106 to the audio interface of computing device 104. Spoken words 106 could be, for example, "Assistant, search the hardware shopping application for RAM," etc. Figure 1A As shown in view 100.

[0015] In response to spoken words, the automation assistant can initialize application 132 and cause application 132 to perform a search operation based on the natural language content of the spoken words. For example, the item "RAM" can be incorporated by the automation assistant into the search field 122 of application 132 and cause application 132 to perform a search based on the item "RAM". As a result, the automation assistant can cause the search results interface 138 to be rendered on the display interface of computing device 104 in response to spoken words 106. The search results interface 138 at application 132 may include a list of search results 126, one or more optional GUI elements 124 for controlling one or more filters, one or more optional checkboxes 130, one or more image results, and / or one or more other optional elements 136 for controlling application 132.

[0016] Although user 102 can manually interact with the display interface at computing device 104 to further refine search results 126, user 102 can alternatively continue to interact with the automation assistant to refine search results 126. For example, with prior permission from user 102, the automation assistant can continue to detect whether user 102 has provided any input for controlling application 132. For example, user 102 can provide another verbal statement 134 for filtering search results 126 based on the content of another verbal statement 134. The other verbal statement 134 could be, for example, “SODIMM,” which could refer to a category of a subset of items listed in search results 126. In response to the other verbal statement 134, the automation assistant can determine whether the content of the other verbal statement 134 includes one or more additional search terms and / or one or more filter parameters.

[0017] For example, the automation assistant can determine that the user has identified filter parameters for one or more filters available at the search results interface 138. In some implementations, the automation assistant can identify the relevance between the content of spoken words and the application's filter parameters based on available assistant data. The assistant data can characterize one or more trial-and-error processes and / or one or more trained machine learning models based on various application interfaces and / or application metadata (e.g., HTML, XML, PHP, etc.) associated with application 132 and / or different applications. Therefore, when the user identifies a specific filter parameter, the assistant data can be used to generate one or more commands for modifying one or more optional GUI elements according to the specific filter parameter.

[0018] In response to another spoken utterance 134, the automation assistant can interact with one or more optional GUI elements 124 to enable filters based on the content of the spoken utterance 134. For example, as Figure 1C As shown in view 140, the automation assistant can generate commands that allow selection of a "Type" dropdown menu and selection of the RAM type "SODIMM". In response to modification of the optional GUI element 124 controlling the RAM type, application 132 can subsequently render another list of search results 142, and may also include instructions (e.g., SODIMM) from filters 146 activated by the automation assistant.

[0019] In some implementations, the application 132 that user 102 is interacting with may not include certain filters for controlling one or more filtering operations to filter search results. In any case, the automation assistant may be invoked by user 102 to further refine the search results based on one or more search parameters, which may not correspond to any optional GUI elements provided by application 132. For example, user 102 may provide additional verbal utterances 144, such as “non-ECC,” which could refer to another category of memory chips. In response to receiving additional verbal utterances 144, the automation assistant may determine whether the content of the additional verbal utterances 144 includes one or more items associated with any selection elements available at the search results interface 138.

[0020] For example, the automation assistant can determine whether any content in the search results interface 138 and / or metadata associated with application 132 corresponds to an optional filter in search results 142 that can filter search results 142 based on the item "non-ECC". When the automation assistant determines that one or more items in another utterance 144 are not associated with any available optional filters in application 132, the automation assistant can generate search terms 148 to replace application 132, which does not provide appropriate user-optional filters. In some implementations, when the automation assistant determines that one or more items in a spoken utterance are not associated with any available content filters, the automation assistant can identify one or more alphanumeric characters and / or non-alphanumeric characters to incorporate into the search command.

[0021] In response to additional verbal utterance 144, the automation assistant can incorporate one or more search terms 148 into the search field 122 and initiate a search operation via application 132. Search terms 148 may include one or more search terms from a previous search requested by user 102 (e.g., “RAM”) and one or more additional search terms from a recent search requested by user 102 (e.g., “non-ECC”). As a result, application 132 can render filtered search results 162 at the search results interface 138. In some implementations, search terms 148 may be retained in the search field 122 to notify user 102 of any assistant filters already applied by the automation assistant, such as… Figure 1D The view provided in 160.

[0022] In some implementations, when user 102 has already provided another input to further refine the currently available search results (e.g., ... Figure 1CAs shown, the automation assistant can identify the corresponding state of each of one or more filters available at application 132. For example, in response to user 102 providing utterance 144, the automation assistant can identify the settings of the "type" filter and generate command data that can be submitted to application 132 along with the assistant input to ensure that the "type" filter has the same state when performing subsequent searches. For example, the automation assistant can generate a search command provided in search field 148 based on utterance 144, and when executing the search command (e.g., "non-ECCRAM"), the automation assistant can check to determine whether the "type" filter is restricted to "SODIMM". When the automation assistant determines that the "type" filter remains unchanged, the automation assistant may not submit another command. However, when the "type" filter is reset after executing a search command based on utterance 144, the automation assistant can modify the "type" filter to be restricted to "SODIMM". In some implementations, the automation assistant can determine whether user 102 is requesting further refinement of the current set of search results or starting a new search from an empty state (e.g., resetting all filters). In some implementations, this determination can be based on the content of the assistant's input (e.g., whether the spoken words include or omit the term "search") and / or any context that can be associated with a new search or a more refined search.

[0023] In some implementations, user 102 may select a specific search result without explicitly identifying the search result and / or by describing another filter parameter imposed on recent search results 162. For example, a specific search result may be identified based on features compared to other search results. In some implementations, user 102 may specify visual features and / or natural language content features of a specific search result to select a search result from a list of search results. For example, visual features may correspond to one or more images 128 rendered by application 132 in association with certain search results. Alternatively or additionally, alphanumeric and / or non-alphanumeric characters in a specific search result may be interpreted by an automated assistant to identify any relevance between the specific search result and the content of the assistant input provided by user 102.

[0024] For example, user 102 could provide a verbal statement such as “The biggest one,” which could refer to the search result with the largest memory size (e.g., 16GB) among the most recent search results 162. In some implementations, each search result could be processed by an automation assistant to generate a corresponding embedding that can be mapped into a latent space. Subsequently, in response to subsequent auxiliary input, the automation assistant can compare the embedding input by the assistant with each corresponding embedding mapped into the latent space. Alternatively or additionally, a heuristic process can be performed to identify the specific search result that the user is referring to. For example, verbal statement 164 could cause the automation assistant to compare one or more items in each corresponding search result with each other to identify items that can indicate the size of each search result. When the automation assistant selects the search result that user 102 is referring to (e.g., by selecting checkbox 130 for 16GB RAM), user 102 can provide a subsequent verbal statement 166 (e.g., “checkout”) to continue using the automation assistant as the interface between user 102 and application 132.

[0025] Figure 2 The illustration depicts a system 200 that provides search operations for controlling individual applications, allowing for the implementation of search filters regardless of whether the individual application provides quick control over the search filters. The automation assistant 204 may operate as part of an assistant application provided on one or more computing devices, such as computing device 202 and / or server devices. A user can interact with the automation assistant 204 via an assistant interface 220, which may be a microphone, camera, touchscreen display, user interface, and / or any other device capable of providing an interface between the user and the application. For example, a user can initialize the automation assistant 204 by providing verbal, textual, and / or graphical input to the assistant interface 220 to cause the automation assistant 204 to initialize one or more actions (e.g., providing data, controlling peripheral devices, accessing agents, generating inputs and / or outputs, etc.).

[0026] Alternatively, the automation assistant 204 can be initialized based on the processing of the context data 236 using one or more trained machine learning models. The context data 236 may characterize one or more features of the environment in which the automation assistant 204 can access, and / or predict one or more features of a user intending to interact with the automation assistant 204. The computing device 202 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures to allow the user to control the application 234 of the computing device 202 via the touch interface. In some embodiments, the computing device 202 may be without a display device, thus providing audible user interface output instead of a graphical user interface output. Furthermore, the computing device 202 may provide a user interface, such as a microphone, for receiving spoken natural language input from a user. In some embodiments, the computing device 202 may include a touch interface and may be without a camera, but may optionally include one or more other sensors.

[0027] Computing device 202 and / or other third-party client devices can communicate with the server device via a network such as the Internet. Additionally, computing device 202 and any other computing devices can communicate with each other via a local area network (LAN) such as a Wi-Fi network. Computing device 202 can offload computing tasks to the server device to save computing resources at computing device 202. For example, the server device can host automation assistant 204, and / or computing device 202 can send input received at one or more assistant interfaces 220 to the server device. However, in some embodiments, automation assistant 204 can be hosted at computing device 202, and various processes that can be associated with the operation of the automation assistant can be performed at computing device 202.

[0028] In various implementations, all or fewer aspects of the automation assistant 204 may be implemented on the computing device 202. In some of these implementations, aspects of the automation assistant 204 are implemented via the computing device 202 and may interface with a server device, which may implement other aspects of the automation assistant 204. The server device may optionally serve multiple users and their associated assistant applications via multithreading. In implementations where all or fewer aspects of the automation assistant 204 are implemented via the computing device 202, the automation assistant 204 may be an application separate from the operating system of the computing device 202 (e.g., installed "on top" of the operating system) – or it may alternatively be implemented directly by the operating system of the computing device 202 (e.g., considered an application of the operating system, but integrated with the operating system).

[0029] In some implementations, the automation assistant 204 may include an input processing engine 206, which may employ multiple different modules to process inputs and / or outputs from the computing device 202 and / or the server device. For example, the input processing engine 206 may include a speech processing engine 208, which can process audio data received at the assistant interface 220 to recognize text contained within the audio data. The audio data may be transferred from, for example, the computing device 202 to the server device to conserve computing resources at the computing device 202. Alternatively or additionally, the audio data may be processed specifically at the computing device 202.

[0030] The process for converting audio data into text may include: a speech recognition algorithm, which may employ a neural network, and / or a statistical model for recognizing audio data sets corresponding to words or phrases. The text converted from the audio data may be parsed by data parsing engine 210 and made available as text data to automation assistant 204. This text data may be used to generate and / or recognize command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, the output data provided by data parsing engine 210 may be provided to parameter engine 212 to determine whether the user has provided input corresponding to a specific intent, action, and / or routine that can be performed by automation assistant 204, and / or an application or agent that can be accessed via automation assistant 204. For example, assistant data 238 may be stored at server device and / or computing device 202 and may include data defining one or more actions that can be performed by automation assistant 204, as well as parameters required to perform the actions. Parameter engine 212 may generate one or more parameters for intents, actions, and / or slot values ​​and provide one or more parameters to output generation engine 214. The output generation engine 214 can communicate with the assistant interface 220 using one or more parameters to provide output to the user, and / or communicate with one or more applications 234 to provide output to one or more applications 234.

[0031] In some implementations, the automation assistant 204 may be an application that can be installed "on top" of the operating system of the computing device 202 and / or may itself form part (or all) of the operating system of the computing device 202. The automation assistant application includes, and / or can access, on-device speech recognition, on-device natural language understanding, and on-device implementation. For example, on-device speech recognition can be performed using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model locally stored at the computing device 202. On-device speech recognition generates recognized text for spoken utterances (if any) present in the audio data. Furthermore, for example, on-device natural language understanding (NLU) can be performed using an on-device NLU module that processes the recognized text generated using on-device speech recognition, along with optional context data, to generate NLU data.

[0032] NLU data can include the intent corresponding to the spoken utterance and optional parameters (e.g., slot values) for that intent. On-device execution can be performed using an on-device execution module that utilizes the NLU data (from on-device NLU) and optional additional local data to determine the action to be taken to resolve the intent of the spoken utterance (and optional parameters for that intent). This can include determining local and / or remote responses to the spoken utterance (e.g., answers), interactions with locally installed applications performed based on the spoken utterance, commands transmitted to Internet of Things (IoT) devices (directly or via corresponding remote systems) based on the spoken utterance, and / or other parsed actions to be performed based on the spoken utterance. On-device execution can then initiate local and / or remote execution / implementation of the determined actions to resolve the spoken utterance.

[0033] In various implementations, remote speech processing, remote NLU, and / or remote execution can be utilized at least selectively. For example, identified text can be selectively transmitted to a remote automation assistant component for remote NLU and / or remote execution. For example, identified text can be selectively transmitted for remote execution in parallel with on-device execution, or transmitted in response to failure of on-device NLU and / or on-device execution. However, on-device speech processing, on-device NLU, on-device execution, and / or on-device execution can be preferred, at least due to the reduced latency they provide when parsing spoken utterance (due to the absence of the (multiple) client-server round trips required for parsing spoken utterance). Furthermore, on-device functionality may be the only available functionality in the absence of network connectivity or with limited network connectivity.

[0034] In some implementations, computing device 202 may include one or more applications 234, which may be provided by a third-party entity different from the entity providing computing device 202 and / or automation assistant 204. The application state engine of automation assistant 204 and / or computing device 202 may access application data 230 to determine one or more actions that can be performed by one or more applications 234, and the state of each of the one or more applications 234 and / or the state of the corresponding device associated with computing device 202. The device state engine of automation assistant 204 and / or computing device 202 may access device data 232 to determine one or more actions that can be performed by computing device 202 and / or one or more devices associated with computing device 202. Furthermore, application data 230 and / or any other data (e.g., device data 232) may be accessed by automation assistant 204 to generate context data 236, which may characterize the context in which a particular application 234 and / or device is performing, and / or the context in which a particular user is accessing computing device 202, accessing application 234 and / or any other device or module.

[0035] When one or more applications 234 are executed at computing device 202, device data 232 can characterize the current operating state of each application 234 executed at computing device 202. Furthermore, application data 230 can characterize one or more features of the executing application 234, such as the content of one or more graphical user interfaces rendered in the direction of the one or more applications 234. Alternatively or additionally, application data 230 can characterize action patterns, which can be updated by the respective application and / or by the automation assistant 204 based on the current operating state of the respective application. Alternatively or additionally, one or more action patterns for one or more applications 234 can remain static but can be accessed by the application state engine to determine the appropriate action initialized via the automation assistant 204.

[0036] The computing device 202 may also include an assistant invocation engine 222, which can use one or more trained machine learning models to process application data 230, device data 232, context data 236, and / or any other data accessible to the computing device 202. The assistant invocation engine 222 can process this data to determine whether to wait for the user to explicitly utter an invocation phrase to invoke the automation assistant 204, or to consider the data as indicating the user's intention to invoke the automation assistant, rather than requiring the user to explicitly utter an invocation phrase. For example, one or more trained machine learning models can be trained using instances of training data based on scenarios in which the user is in an environment where multiple devices and / or applications are exhibiting various operational states. Instances of training data can be generated to capture training data characterizing scenarios where the user invokes the automation assistant and other scenarios where the user does not invoke the automation assistant.

[0037] When one or more trained machine learning models are trained based on these instances of training data, the assistant invocation engine 222 can enable the automated assistant 204 to detect or restrict verbal invocation phrases from the user based on features of the context and / or environment. Alternatively or additionally, the assistant invocation engine 222 can enable the automated assistant 204 to detect or restrict the detection of one or more assistant commands from the user based on features of the context and / or environment. In some implementations, the assistant invocation engine 222 can be disabled or restricted based on the computing device 202 detecting assistant suppression output from another computing device. In this way, when the computing device 202 detects assistant suppression output, the automated assistant 204 will not be invoked based on context data 236; otherwise, if no assistant suppression output is detected, the automated assistant 204 will be invoked based on context data 236.

[0038] In some implementations, system 200 may include a filter identification engine 216, whose processing may include application data 230, device data 232, context data 236, and / or any other data, and auxiliary data 238, to determine whether an application provides access to one or more filter features and / or how to control one or more filter features. The filter identification engine 216 may process the auxiliary data 238 to determine whether an application 234 executing at computing device 202 is rendering one or more optional GUI elements for controlling one or more filter features. In some implementations, the auxiliary data 238 may be processed based on one or more heuristic processes and / or using one or more trained machine learning models. When one or more filter features are identified for a particular application, the filter identification engine 216 may communicate with the input engine 218 of system 200 to determine whether input from a user is associated with one or more filter features.

[0039] For example, filter recognition engine 216 can generate data characterizing one or more filter features of the application the user is visiting, and automation assistant 204 can compare this data with input from the user. The user can provide input such as spoken words, including one or more items and / or a request for the automation assistant to perform a search for certain application content. Automation assistant 204 can compare the natural language content of this input with the data from filter recognition engine 216 to determine whether the content of the input is associated with any of the one or more filter features. When automation assistant 204 determines that the user has provided a search request identifying one or more filter features, the automation assistant can generate command data to communicate to the application. The command data received by the application can modify one or more filter parameters of one or more filter features based on the input from the user and cause the application to perform a search operation.

[0040] In some implementations, user input may include items that may not be associated with any filter features of the application, but may still be items the user intends to use to filter search results. As a result, the automation assistant 204 may employ a search input engine 226 to determine whether any items in the user input can be used as the basis for generating a search command (i.e., application input) that can be incorporated into the application's search field. For example, when the user includes items for filtering search results (e.g., "search for RAM manufactured this year"), but the application does not have a corresponding filter feature (e.g., no slider for limiting the year of manufacture), the search input engine 226 may generate part of a search command (e.g., "MFR>=2021") to be incorporated into the application's search field when the search operation is performed. Alternatively or additionally, when the input item engine 218 determines that certain items in the input can be incorporated as search items (e.g., "RAM", "laptop memory", etc.) into the search command, the search input engine 226 may combine such search items with any other identified filter parameters to incorporate them into the search command (e.g., search field: "RAM laptop memory MFR>=2021").

[0041] When a search is performed at an application via an automated assistant, the application may render some content as search results. Search results may be rendered at the application's search results interface, and may be processed by the search results engine 224 of system 224. The search results engine 224 may generate further data based on the search results to determine whether any subsequent input (e.g., user input provided while rendering the search results in the foreground of computing device 202) is relevant to the search results. The search results engine 224 may process data including, but not limited to, screenshots, metadata, source code, and / or any other data that may be associated with the search results interface. In some implementations, the search results engine 224 may generate training data for further training one or more trained machine learning models to render more accurate search results in response to a user's request to perform a search. For example, the weighting of terms and / or embeddings may be modified to make a particular trained machine learning model more reliable when processing search terms and / or filter parameters specified by the user. For example, the weighting of terms for a first application may differ from another weighting of terms for a second application, at least based on how reliable the generated results are in relation to the user's search request.

[0042] Figure 3 The illustration depicts a method 300 for operating an automation assistant to interface with a separate application for searching and / or filtering certain application content. Method 300 can be performed by one or more applications, devices, and / or any other means or modules capable of interacting with the automation assistant. Method 300 may include an operation 302 determining whether the automation assistant has received a spoken utterance. For example, the spoken utterance could be a request from the automation assistant to access an encyclopedia application to identify certain articles (e.g., “Assistant, search for cryptography articles written this year in my encyclopedia application”). In response to receiving the spoken utterance, the automation assistant can process audio data to identify one or more requests embodied in the spoken utterance.

[0043] Method 300 can proceed from operation 302 to operation 304, which may include determining whether the user is requesting the initialization of a search operation in another application. Otherwise, if no verbal utterance is received, the automation assistant can continue to detect assistant input. The request for the search operation to be initialized may specify the application that the user wishes to use in conjunction with the automation assistant to search for certain application content. Alternatively or additionally, the request for the search operation may identify one or more items that should be used to identify specific application content. When it is determined that the user has already requested the initialization of a search operation in another application, method 300 can proceed from operation 304 to optional operation 306. Otherwise, method 300 may return to operation 302 for detecting assistant input from one or more users.

[0044] Optional operation 306 may include processing automation assistant data based on one or more search interfaces of one or more applications. For example, assistant data may characterize one or more trial-and-error processes and / or one or more machine learning models, which can be used to process data based on the application's search interface. In response to spoken utterances, the automation assistant may initialize the application to render the application's search interface at a display on a computing device. Assistant data may be processed to determine whether certain features of the search interface can be controlled by the automation assistant. For example, the search interface may include search fields for providing search terms and / or other characters for defining search actions to be performed by the application. Alternatively or additionally, the search interface may include one or more optional GUI elements for establishing filter settings for search actions.

[0045] Method 300 can proceed from optional operation 306 to operation 308, which determines whether the spoken utterance recognizes one or more filter parameters associated with the application. For example, the assistant data can be processed using input audio data to determine if there is any correlation between the natural language content of the spoken utterance and one or more filter parameters associated with the application. Following a previous example, the automation assistant can determine that the application's search interface includes filter parameters for filtering encyclopedia articles published before a specific date. When the automation assistant determines that the spoken utterance recognizes one or more filter parameters, method 300 can proceed from operation 308 to operation 310. Otherwise, method 300 can proceed from operation 308 to operation 314.

[0046] Operation 310 may include modifying one or more filter settings based on verbal utterances. For example, verbal utterances may reflect one or more filter parameters specified by the user, allowing the automation assistant to modify one or more filters in the application accordingly. In some instances, the user may identify one or more filter parameters corresponding to one or more optional GUI elements (e.g., one or more checkboxes and / or one or more dials). The automation assistant may determine how to modify one or more optional GUI elements based on the one or more filter parameters identified by the user to perform a search operation according to a request from the user. For example, when a user requests the automation assistant to search for articles published after a specific year in an encyclopedia application, the automation assistant may adjust the optional GUI element on the dial that controls the publication date filter. Method 300 may then proceed from operation 310 to operation 312, which may include having the application initialize a search operation based on verbal utterances. As a result, the performed search operation can be initialized using filter parameters specified by the user and implemented by the automation assistant, without requiring the user to manually interact with the display interface to activate certain filters. Alternatively or additionally, the search operation may be initialized using one or more search terms identified in the verbal utterances and incorporated into the search field of the application's search interface.

[0047] In some implementations, when spoken words identify one or more filter parameters that may not be associated with the application or are otherwise available at the application's search interface, method 300 can proceed from operation 308 to operation 314. Operation 314 may include generating application input based on one or more filter parameters. The application input may be, for example, a search command including alphanumeric characters and / or non-alphanumeric characters, which may be provided to the application's search field to perform a search operation. In some implementations, when filter parameters are identified by the user but are not available at the search interface, the automation assistant may identify one or more special characters (e.g., non-alphanumeric characters). For example, when the search interface does not include optional GUI elements for limiting search results associated with a specific time, the automation assistant may identify one or more special characters and / or equations (e.g., >, <, >=, <=, etc.) that can be used to filter out certain search results that may not be associated with a specific time range (e.g., "cryptography <= 1 year"). In this way, the user can perform such a search as a single input to the automation assistant, rather than waiting for search results to appear and subsequently adjusting any filters that may or may not be available at the search results interface.

[0048] Figure 4This is a block diagram 400 of an example computer system 410. The computing device 410 typically includes at least one processor 414 that communicates with a plurality of peripheral devices via a bus subsystem 412. These peripheral devices may include a storage subsystem 424 (including, for example, a memory subsystem 424 and a file storage subsystem 426), a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices allow users to interact with the computing device 410. The network interface subsystem 416 provides an interface to an external network and couples to corresponding interface devices in other computing devices.

[0049] User interface input device 422 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen integrated into a display, a voice recognition system, a microphone, and / or other types of input devices. Generally, the term "input device" is used to encompass all possible types of devices and methods for inputting information onto computing device 410 or a communication network.

[0050] User interface output device 420 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating visual images. The display subsystem may also provide non-visual displays, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and methods for outputting information from computing device 410 to a user or another machine or computing device.

[0051] Storage subsystem 424 stores the program and data constructs that provide functionality for some or all of the modules described herein. For example, storage subsystem 424 may include selected aspects of performing method 300, and / or logic implementing system 200, computing device 104, and / or any other application, device, apparatus, and / or module discussed herein.

[0052] These software modules are typically executed by processor 414 alone or in combination with other processors. Memory 425 used in storage subsystem 424 may include multiple memories, including main random access memory (RAM) 430 for storing instructions and data during program execution and read-only memory (ROM) 432 for storing fixed instructions. File storage subsystem 426 can provide permanent storage for program and data files and may include hard disk drives, floppy disk drives with associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored by file storage subsystem 426 within storage subsystem 424 or on other machines accessible to processor 414.

[0053] Bus subsystem 412 provides a mechanism for allowing various components and subsystems of computing device 410 to communicate with each other as intended. Although bus subsystem 412 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0054] The computing device 410 can be of various types, including workstations, servers, computing clusters, blade servers, server groups, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, in order to illustrate some implementation methods, Figure 4 The description of the computing device 410 depicted is intended only as a specific example. Many other configurations of the computing device 410 may have... Figure 4 The computing device depicted in the text has more or fewer components.

[0055] In situations where the systems described herein collect or may use personal information about users (or, generally referred to herein as "participants"), users may be given the opportunity to control whether a program or function collects user information (e.g., information about a user's social networks, social behaviors or activities, occupation, user preferences, or the user's current geographic location), or to control whether and / or how content that may be more relevant to the user is received from a content server. Furthermore, certain data may be processed in one or more ways before being stored or used, thereby removing personally identifiable information. For example, a user's identity may be processed so that the user's personally identifiable information cannot be determined, or the user's geographic location may be generalized (e.g., generalized to a city, zip code, or state level) when geographic location information is obtained, making it impossible to determine the user's specific geographic location. Therefore, users can control how information about themselves is collected and / or used.

[0056] While several embodiments have been described and illustrated herein, various other means and / or structures may be utilized for performing functions and / or obtaining results and / or one or more advantages described herein, and each such variation and / or modification is considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications to which the teachings are addressed. Those skilled in the art will recognize, or may determine, many equivalents of the particular embodiments described herein using only conventional experimentation. Therefore, it should be understood that the foregoing embodiments are presented as examples only, and embodiments may be implemented in ways different from those specifically described and claimed within the scope of the appended claims and their equivalents. Embodiments of this disclosure pertain to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of this disclosure if such features, systems, articles, materials, kits, and / or methods do not contradict each other.

[0057] In some implementations, a method implemented by one or more processors is described as including operations such as: receiving, at a computing device and from a user, a spoken utterance directed to an automation assistant accessible via the computing device, wherein the computing device includes a display interface that, upon receiving the spoken utterance, renders an application's search results interface. The method may further include operations such as: determining, based on the spoken utterance, whether the user identifies a specific filter setting not included in one or more filter settings of the search results interface. The method may further include operations such as: when the automation assistant determines that the user identifies a specific filter setting not included in one or more filter settings, generating application input based on application content from the search results interface and the spoken utterance, and causing the application to initialize the execution of a search operation based on the application input.

[0058] In some implementations, application input identifies specific filter settings and one or more search terms corresponding to application content rendered at the search results interface. In some implementations, application input includes one or more items contained in spoken language, as well as one or more search terms previously submitted to the application to render application content at the search results interface. In some implementations, the method may further include modifying the application's specific filter settings based on spoken language from the user when the automation assistant determines that the user has identified specific filter settings included in one or more filter settings, wherein modifying the specific filter settings causes different application content to be rendered at the search results interface. In some implementations, initiating the execution of a search operation based on application input includes incorporating the application input into the search fields of the search results interface.

[0059] In some implementations, determining whether a user identifies a specific filter setting not included in one or more filter settings of the search results interface includes processing assistant data based on one or more search interfaces previously rendered by the application or different applications. In some implementations, the method may further include determining a corresponding state for each of the one or more filter settings of the search results interface when the automation assistant determines that the user has identified a specific filter setting not included in one or more filter settings, wherein the application input is also based on each corresponding state of each of the one or more filter settings of the search results interface. In some implementations, causing the application to initialize the execution of a search operation based on application input includes causing the application to render a subset of application content that has been filtered according to each corresponding state of each of the one or more filter settings.

[0060] In other embodiments, a method implemented by one or more processors is described as including operations such as: receiving verbal utterances from a user at a computing device to further prompt an automation assistant to initiate a search operation using an application separate from the automation assistant, wherein the verbal utterances identify one or more items. The method may further include determining, based on the verbal utterances, whether one or more items of the verbal utterances are associated with one or more optional graphical user interface (GUI) elements rendered at the interface of the application, wherein the one or more optional GUI elements control one or more filter parameters of the application's search features. The method may further include operations such as: when one or more items of the verbal utterances correspond to one or more optional GUI elements rendered at the interface of the application, causing one or more specific optional GUI elements to control one or more specific filter parameters, and causing the application to initialize the search operation based on the one or more specific filter parameters.

[0061] In some implementations, initiating a search operation with an application includes: including one or more items identified in the spoken utterance in the search field if the search field does not include one or more specific filter parameters. In some implementations, the method may further include: generating application input representing one or more specific filter parameters when one or more items in the spoken utterance do not correspond to one or more optional GUI elements rendered in the application's interface, and initiating the search operation with the application input. In some implementations, initiating the search operation with the application input includes: including the application input in the application's search field, wherein the application input identifies one or more specific filter parameters.

[0062] In some implementations, application input includes non-alphanumeric characters selected based on one or more specific filter parameters. In some implementations, spoken utterances include requests to an automation assistant to: search for specific content using the application and subsequently render that specific content to the user. In some implementations, the method may further include: accessing specific application content that satisfies one or more specific filter parameters when one or more items of the spoken utterance do not correspond to one or more optional GUI elements rendered in the application's interface, and, after accessing the specific application content, causing the automation assistant to render audible content based on the specific application content.

[0063] In some embodiments, the method may further include the following operations: rendering multiple different search results at another interface of the application based on a search operation when one or more items of a spoken word correspond to one or more optional GUI elements rendered on an interface of the application; receiving additional spoken words from the user after rendering the multiple different search results, wherein the additional spoken words identify one or more additional items for identifying a subset of the multiple different search results; and filtering the multiple different search results according to one or more additional items in response to receiving the additional spoken words. In some embodiments, filtering the multiple different search results according to one or more additional items includes: determining that one or more other optional GUI elements rendered on another interface correspond to one or more additional items; and selecting one or more other optional GUI elements according to one or more additional items.

[0064] In some implementations, a method implemented by one or more processors is described as including operations such as: receiving a verbal utterance at a computing device, the verbal utterance including a request for an automation assistant to perform a search for application content accessible via an application interface, wherein the verbal utterance identifies one or more items associated with the application content to be searched. The method may further include operations based on the verbal utterance to determine whether the application provides one or more filtering features for filtering the application content based on one or more items identified in the verbal utterance. The method may further include operations such as: when it is determined that the application does not provide one or more filtering features, identifying one or more filter parameters for submission to the application based on one or more items to induce the execution of a search for application content, causing the automation assistant to provide application input to the application, wherein the application input identifies one or more filter parameters, and causing the application to render search results based on the application input, wherein the search results include a subset of application content that satisfies one or more filter parameters.

[0065] In some implementations, one or more filtering features include one or more optional graphical user interface (GUI) elements that control one or more filtering operations of the application. In some implementations, the method may also include the following operation: when it is determined that the application does not provide one or more filtering features: identifying one or more non-alphanumeric characters selected based on one or more filter parameters, wherein the application input identifies one or more non-alphanumeric characters.

Claims

1. A method implemented by one or more processors, the method comprising: The user receives spoken commands from the computing device and directs them to an automated assistant accessible via the computing device. The computing device includes a display interface that renders a search results interface of an application when the spoken words are received from the user, and the display interface includes one or more optional graphical user interface (GUI) elements that establish one or more filter settings for the search results interface. Determine whether the user recognizes a specific filter setting that is not included in one or more filter settings in the search results interface based on the spoken utterance; and When the automated assistant determines, based on the spoken words, that the user has identified a specific filter setting that is not included in one or more filter settings on the search results interface: Generate application input, which is based on the spoken words and the application content of the search results interface, and The application initializes the execution of the search operation based on the application input.

2. The method according to claim 1, wherein, The application input identifies the specific filter settings and one or more search terms corresponding to the application content rendered at the search results interface.

3. The method according to claim 1 or claim 2, wherein, The application input includes one or more items contained in the spoken utterance and one or more search items previously submitted to the application to render the application content at the search results interface.

4. The method according to any one of the preceding claims further includes: When the automation assistant determines, based on the spoken words, that the user has identified a specific filter setting included in one or more of the filter settings: The application's specific filter settings are modified based on the user's spoken words. Specifically, modifying the specific filter settings allows different application content to be rendered on the search results interface.

5. The method according to any one of the preceding claims, wherein, The process of enabling the application to initialize the execution of the search operation based on the application input includes: The application input is incorporated into the search field of the search results interface.

6. The method according to any one of the preceding claims, wherein, Determining whether the user recognizes a specific filter setting that is not included in the one or more filter settings on the search results interface includes: Process assistant data, which is based on one or more search interfaces previously rendered by the application or different applications.

7. The method according to any one of the preceding claims further comprises: When the automation assistant determines that the user has identified a specific filter setting that is not included in the one or more filter settings: Determine the appropriate state for each filter setting in the one or more filter settings of the search results interface. The application input is further based on each corresponding state of each filter setting in the one or more filter settings of the search results interface.

8. The method according to claim 7, wherein, The process of enabling the application to initialize the execution of the search operation based on the application input includes: The application renders a subset of application content, which is filtered according to each corresponding state of each of the one or more filter settings.

9. A method implemented by one or more processors, the method comprising: The system receives verbal commands from the user at the computing device to further prompt the automation assistant to initiate the search operation using an application separate from the automation assistant. The spoken utterance identifies one or more items; Based on the spoken utterance, determine whether one or more items of the spoken utterance are associated with one or more optional graphical user interface (GUI) elements rendered at the interface of the application. Wherein, the one or more optional GUI elements control one or more filter parameters of the search features of the application; and When one or more items of the spoken utterance correspond to one or more optional GUI elements rendered at the interface of the application: To enable one or more specific optional GUI elements among the one or more optional GUI elements to control one or more specific filter parameters, and The application initializes the search operation based on the one or more specific filter parameters.

10. The method according to claim 9, wherein, Initializing the search operation for the application includes: If the search field does not include the one or more specific filter parameters, the application's search field will include the one or more items identified in the spoken utterance.

11. The method of claim 9, further comprising: When one or more items of the spoken utterance do not correspond to one or more optional GUI elements rendered at the interface of the application: Generate application inputs characterizing the one or more specific filter parameters, and The application is made to initialize the search operation using the application input.

12. The method according to claim 11, wherein, Making the application use the application input to initialize the search operation includes: Make the application's search field include the application's input. The application input identifies one or more specific filter parameters.

13. The method according to claim 12, wherein, The application input includes non-alphanumeric characters selected based on one or more specific filter parameters.

14. The method according to any one of claims 9 to 13, wherein, The spoken utterances include requests to the automated assistant for: searching for specific content using the application; and subsequently rendering the specific content for the user.

15. The method of claim 14, further comprising: When one or more items of the spoken utterance do not correspond to one or more optional GUI elements rendered at the interface of the application: Access specific application content that satisfies one or more of the specific filter parameters, and After accessing the specific application content, the automation assistant renders audible content based on the specific application content.

16. The method according to any one of claims 9 to 15, further comprising: When one or more items of the spoken utterance correspond to one or more optional GUI elements rendered at the interface of the application: Based on the search operation, multiple different search results are rendered on another interface of the application. After rendering the multiple different search results, additional spoken words are received from the user. Wherein, the additional spoken utterances identify one or more additional items, the one or more additional items being used to identify a subset of the multiple different search results, and In response to receiving the additional verbal utterance, the multiple different search results are filtered based on the one or more additional items.

17. The method according to claim 16, wherein, Filtering the multiple different search results based on the one or more additional items includes: Determine that one or more other optional GUI elements rendered at the other interface correspond to the one or more additional items, and Select one or more other optional GUI elements based on the one or more additional items.

18. A method implemented by one or more processors, the method comprising: The system receives spoken commands at a computing device, including requests for the automated assistant to perform searches for application content accessible via the application's interface. Wherein, the spoken language recognition is associated with one or more items of the application content to be searched; Determine the filtering features available for the application to filter the application content; Based on the spoken utterance, determine whether the filtering features available to the application include one or more specific filtering features for filtering the application content based on the one or more items identified in the spoken utterance; When it is determined that the application does not provide the one or more specific filtering features: Based on the one or more items, identify one or more filter parameters to be submitted to the application to prompt a search for the application's content. The automation assistant provides application input to the application. The application input identifies the one or more filter parameters, and The application renders search results based on the input from the application. The search results include a subset of the application content that satisfies one or more filter parameters.

19. The method according to claim 18, wherein, The one or more specific filtering features include one or more optional graphical user interface (GUI) elements for controlling one or more filtering operations of the application.

20. The method according to claim 18 or claim 19, further comprising: When it is determined that the application provides one or more of the specific filtering features: Select one or more optional GUI elements to control one or more filtering operations of the application.

21. The method according to claim 18 or claim 19, further comprising: When the application is determined to not provide one or more of the filtering features: One or more non-alphanumeric characters are identified based on the one or more filter parameters, wherein the one or more non-alphanumeric characters are selected based on the one or more filter parameters. The application input identifies one or more non-alphanumeric characters.

22. A computer program comprising instructions that, when executed by one or more processors of a computing system, cause the computing system to perform the method according to any one of claims 1-21.

23. One or more computing devices configured to perform the method according to any one of claims 1-21.

Citation Information

Patent Citations

  • Initializing conversation with automated agent via selectable graphical element

    CN110622136A

  • Conversational bot to navigate upwards in the funnel

    US10558693B1