Automated assistant for introducing or controlling search filter parameters in separate applications
The automated assistant addresses the limitations of application features by enabling verbal search and filter control, reducing manual interaction and ensuring accurate results through automated filtering.
Patent Information
- Application Number
- JP2023537160
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-20
- Filing Date
- 2021-12-15
- Publication Date
- 2025-08-21
- Estimated Expiration
- 2041-12-15
AI Technical Summary
Users are limited by the features of an application when interacting with automated assistants, requiring manual interaction to fulfill unfulfilled requests and leading to resource consumption and inaccurate search results.
An automated assistant that enables users to search and filter application content through verbal utterances, adjusting application interface elements and search fields without direct user input, and refining results based on additional verbal commands.
Reduces the time spent on manual interaction and ensures accurate search results by automating the filtering process, allowing users to specify search parameters and refine results verbally.
Smart Images

Figure 0007727734000001 
Figure 0007727734000002 
Figure 0007727734000003
Abstract
Description
[Background technology]
[0001] Humans can engage in human-computer interactions using interactive software applications referred to herein as "automated assistants" (also referred to as "digital agents," "chatbots," "interactive personal assistants," "intelligent personal assistants," "assistant applications," "conversational agents," etc.). For example, humans (who may be referred to as "users" when interacting with an automated assistant) can provide commands and / or requests to the automated assistant using verbal natural language input (i.e., utterances) and / or by providing textual (e.g., typed) natural language input, which may optionally be converted to text and then processed.
[0002] For example, a user who invokes their automated assistant to perform a search operation through a particular application may be limited by whether the particular application has enabled features for interfacing with the automated assistant. Depending on whether the automated assistant has control over particular features of the application, the automated assistant may fulfill only a limited number of requests from the user. In such cases, the user may be forced to individually identify fulfilled and unfulfilled requests and then subsequently interact with a touch interface to manually complete any unfulfilled requests. This switching between interfaces may consume resources across many aspects of the computing device and may increase the likelihood that the application and / or automated assistant will provide inaccurate search results. Summary of the Invention [Means for solving the problem]
[0003] Implementations described herein relate to an automated assistant that enables a user to search and / or filter application content of an application (e.g., a website, a client application, a server application, a browser, etc.) by providing verbal utterances to the automated assistant and without the user providing direct input to the application. A search operation can be initiated when a user requests the automated assistant to access the application and search the application content. In response to such a request from the user, the automated assistant can determine whether the application identified by the user provides any features for filtering search results aside from a search field. The automated assistant can determine whether the application's search interface includes one or more selectable graphical user interface (GUI) elements for limiting the type of content included in the search results. When the automated assistant determines that the application interface includes one or more selectable filter elements that correspond to one or more words in the assistant input from the user, the automated assistant can adjust the one or more filter elements according to the one or more words. The automated assistant can then fill in a search field in the application interface with one or more other words identified in the assistant input and initiate a search operation.
[0004] When a search operation is initiated, the application can search application content related to one or more terms in the search field. As a result, the user receives search results from the application without directly interacting with the application to search for application content. Rather than directly interacting with the application, the user relies on an automated assistant performing natural language understanding (NLU) and / or speech-to-text processing to interact with the application according to a request from the user. In this way, the user can reduce the amount of time they spend attempting to identify specific filter elements in the application interface and / or manually typing search terms into the search field of the application interface.
[0005] In some implementations, the search results provided by the application can be further filtered by the automated assistant in response to another request from the user to the automated assistant. For example, a user can provide an additional verbal utterance to the automated assistant after interacting with the application to cause the automated assistant to provide search results. The verbal utterance can identify one or more additional search terms that can be used by the automated assistant to filter the search results and / or otherwise select a subset of the search results. For example, in response to receiving the additional verbal utterance, the automated assistant can determine whether any additional search terms embodied in the additional verbal utterance correspond to one or more selectable filter elements rendered in the application's search result interface. When the automated assistant determines that the additional search terms do not correspond to one or more selectable filter elements of the application, it can generate a search command to be executed by the application. The search command can be generated to ensure that the application provides a subset of the search results rather than resetting any established search parameters used to generate the search results and / or starting a new search from a null state.
[0006] In some implementations, a user can provide an automated assistant with a search request that includes parameters regarding when any search result content should be provided to the user. For example, a user can provide a search request seeking content from an application (e.g., a news application) and subsequently specify the time at which the user wants the search results provided to them (e.g., "Assistant, search my news application for blockchain articles from yesterday and read them to me at 10:00 AM"). In this way, when operating on an ecosystem of devices, the automated assistant can search for and / or download any relevant search results on devices the user may have access to at the specified time. This can also enable the automated assistant to select a trusted network to retrieve search results from, rather than immediately downloading content from whatever network is available at the time the user requests to receive the resulting content.
[0007] The above description is provided as a summary of some implementations of the present disclosure. Further description of these and other implementations is provided in more detail below.
[0008] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) for performing a method, such as one or more of the methods described above and / or elsewhere herein. Still other implementations may include a system of one or more computers including one or more processors operable to execute stored instructions for performing a method, such as one or more of the methods described above and / or elsewhere herein.
[0009] It should be understood that any combination of the foregoing concepts, and additional concepts described in more detail herein, is contemplated as being part of the present subject matter disclosed herein. For example, any combination of claimed subject matter appearing at the end of this disclosure is contemplated as being part of the present subject matter disclosed herein. [Brief explanation of the drawings]
[0010] [Figure 1A] FIG. 1 illustrates a view in which a user interacts with an automated assistant to control the search behavior of an application. [Figure 1B] FIG. 1 illustrates a view in which a user interacts with an automated assistant to control the search behavior of an application. [Figure 1C] FIG. 1 illustrates a view in which a user interacts with an automated assistant to control the search behavior of an application. [Figure 1D] FIG. 1 illustrates a view in which a user interacts with an automated assistant to control the search behavior of an application. [Figure 2] FIG. 1 illustrates a system that provides an automated assistant for controlling the search operations of a separate application in a manner that allows search filters to be implemented regardless of whether the separate application provides explicit control over the search filters. [Figure 3] FIG. 1 illustrates a method for operating an automated assistant to interface with a separate application to search and / or filter specific application content. [Figure 4] FIG. 1 is a block diagram of an exemplary computer system. DETAILED DESCRIPTION OF THE INVENTION
[0011] 1A, 1B, 1C, and 1D show views 100, 120, 140, and 160 in which a user 102 interacts with an automated assistant to control the search behavior of an application 132. The application 132 can be separate from the automated assistant but can allow the automated assistant to control certain operations of the application 132. For example, the application 132 can be a hardware shopping application that the user 102 can use to purchase computer parts. To invoke the automated assistant to control the application 132, the user 102 can provide a spoken utterance 106 to an audio interface of the computing device 104. The spoken utterance 106 can be, for example, "Assistant, search the hardware shopping application for RAM," as shown in view 100 of FIG. 1A.
[0012] In response to the verbal utterance, the automated assistant can initialize an application 132 and cause the application 132 to perform a search operation based on the natural language content of the verbal utterance. For example, the word “RAM” can be embedded in the search field 122 of the application 132 by the automated assistant, causing the application 132 to perform a search based on the word “RAM.” As a result, in response to the verbal utterance 106, the automated assistant can cause a search results interface 138 to be rendered on the display interface of the computing device 104. The search results interface 138 in the application 132 can include a list of search results 126, one or more selectable GUI elements 124 for controlling one or more filters, one or more selectable checkboxes 130, one or more image results, and / or one or more other selectable elements 136 for controlling the application 132.
[0013] The user 102 can manually interact with a display interface at the computing device 104 to further refine the search results 126, or alternatively, can continue to interact with the automated assistant to refine the search results 126. For example, with prior permission from the user 102, the automated assistant can continue to detect whether the user 102 has provided any input to control the application 132. As an example, the user 102 can provide another spoken utterance 134 to filter the search results 126 according to the content of the another spoken utterance 134. The another spoken utterance 134 can be, for example, "SODIMM," which can refer to a classification of a subset of items listed in the search results 126. In response to the another spoken utterance 134, the automated assistant can determine whether the another spoken utterance 134 includes one or more additional search terms and / or one or more filter parameters.
[0014] The automated assistant can determine, for example, that the user has specified filter parameters for one or more filters available in search results interface 138. In some implementations, the automated assistant can identify a correlation between the content of the verbal utterance and the filter parameters of the application based on available assistant data. The assistant data can characterize one or more heuristic processes and / or one or more trained machine learning models based on various application interfaces and / or application metadata (e.g., HTML, XML, PHP, etc.) associated with application 132 and / or another application. Thus, when a particular filter parameter is specified by the user, the assistant data can be used to generate one or more commands to modify one or more selectable GUI elements according to the particular filter parameter.
[0015] In response to this other verbal utterance 134, the automated assistant can interact with one or more selectable GUI elements 124 to enable a filter based on the content of the verbal utterance 134. For example, the automated assistant can generate a command to select a "Type" drop-down menu and select the RAM type "SODIMM" as shown in view 140 of FIG. 1C. In response to changing the selectable GUI element 124 controlling the RAM type, the application 132 can then render another list of search results 142, which can also include an indication of the filter 146 (e.g., SODIMM) that was activated via the automated assistant.
[0016] In some implementations, the application 132 with which the user 102 is interacting may not include a specific filter to control one or more filtering operations that filter the search results. Nevertheless, an automated assistant can be invoked by the user 102 to further refine the search results according to one or more search parameters that may not correspond to any selectable GUI elements provided by the application 132. For example, the user 102 may provide an additional verbal utterance 144, such as "non-ECC," which may refer to a different classification of memory chips. In response to receiving the additional verbal utterance 144, the automated assistant can determine whether the content of the additional verbal utterance 144 includes one or more words associated with any selection elements available in the search results interface 138.
[0017] For example, the automated assistant can determine whether any content in the search results interface 138 and / or metadata associated with the application 132 corresponds to a selectable filter that can filter the search results 142 according to the term "non-ECC." When the automated assistant determines that one or more words in the additional spoken utterance 144 are not associated with any available selectable filters of the application 132, the automated assistant can generate search terms 148 on behalf of the application 132 that does not provide an appropriate filter selectable by the user. In some implementations, when the automated assistant determines that one or more words in the spoken utterance are not associated with any available content filters, the automated assistant can identify one or more alphanumeric and / or non-alphanumeric characters for incorporation into the search command.
[0018] In response to the additional verbal utterance 144, the automated assistant can incorporate one or more search terms 148 into the search field 122 and initiate a search operation via the application 132. The search terms 148 can include one or more search terms from a prior search requested by the user 102 (e.g., "RAM") and one or more additional search terms from the most recent search requested by the user 102 (e.g., "non-ECC"). As a result, the application 132 can render filtered search results 162 in the search result interface 138. In some implementations, the search terms 148 can remain within the search field 122 to inform the user 102 of any assistant filters used by the automated assistant, as presented in view 160 of FIG. 1D .
[0019] In some implementations, when user 102 issues another input to further refine currently available search results (e.g., as shown in FIG. 1C ), the automated assistant can identify the respective state of each filter of one or more filters available in application 132. For example, in response to user 102 providing verbal utterance 144, the automated assistant can identify the setting of the “Type” filter and generate command data that can be submitted to application 132 with the assistant input to ensure that the “Type” filter has the same state when a subsequent search is performed. As an example, the automated assistant can generate a search command to provide in search field 148 based on verbal utterance 144; when this search command (e.g., “Non-ECC RAM”) is executed, the automated assistant can check to determine whether the “Type” filter is limited to “SODIMM.” When the automated assistant determines that the “Type” filter remains unchanged, it may not submit another command. However, when the "type" filter has been reset after executing a search command based on the verbal utterance 144, the automated assistant may change the "type" filter to be limited to "SODIMM." In some implementations, the automated assistant may determine whether the user 102 is requesting to further refine the current set of search results or to start a new search from a null state (e.g., with all filters reset). In some implementations, this determination may be based on the content of the assistant input (e.g., whether the verbal utterance includes or omits the word "search") and / or any context that may be relevant to the new search or search refinement.
[0020] In some implementations, the user 102 can select a particular search result without explicitly identifying the search result and / or by indicating another filter parameter to impose on the most recent search results 162. For example, a particular search result can be identified based on the characteristics of the particular search result relative to other search results. In some implementations, the user 102 can specify visual features and / or natural language content features of a particular search result to select the search result for the list of search results. The visual features can correspond, for example, to one or more images 128 rendered by the application 132 in association with some search results. Alternatively or additionally, the alphanumeric and / or non-alphanumeric characters of a particular search result can be interpreted by the automated assistant to identify any correlation between the particular search result and the content of the assistant input provided by the user 102.
[0021] For example, user 102 can provide a verbal utterance such as “biggest,” which can refer to the search result with the largest memory capacity (e.g., 16 GB) among the most recent search results 162. In some implementations, each search result can be processed by the automated assistant to generate a respective embedding that can be mapped into a latent space. Then, in response to a subsequent assistant input, the automated assistant can compare the embedding for the assistant input to the respective embeddings mapped into the latent space. Alternatively or additionally, a heuristic process can be performed to identify the specific search result the user may be referring to. For example, given verbal utterance 164, the automated assistant can compare one or more words in each search result to each other to identify words that can indicate the size of each search result. When the automated assistant selects the search result to which user 102 is referring (e.g., by selecting checkbox 130 for 16 GB RAM), user 102 can provide a subsequent verbal utterance 166 (e.g., “check out”) to continue using the automated assistant, which serves as an interface between user 102 and application 132.
[0022] 2 illustrates a system 200 that provides an automated assistant for controlling the search operations of a separate application to allow search filters to be implemented regardless of whether the separate application provides explicit control over the search filters. Automated assistant 204 can operate as part of an assistant application provided on one or more computing devices, such as computing device 202 and / or a server device. A user can interact with automated assistant 204 through assistant interface 220, which can be a microphone, a camera, a touchscreen display, a user interface, and / or any other device capable of providing an interface between a user and an application. By way of example, a user can initialize automated assistant 204 by providing verbal, textual, and / or graphical input to assistant interface 220 to cause automated assistant 204 to initiate one or more actions (e.g., providing data, controlling a peripheral device, accessing an agent, generating input and / or output, etc.).
[0023] Alternatively, the automated assistant 204 may be initialized based on processing the context data 236 using one or more trained machine learning models. The context data 236 may characterize one or more features of an environment accessible to the automated assistant 204 and / or one or more features of a user predicted to be interacting with the automated assistant 204. The computing device 202 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures that enable a user to control the application 234 of the computing device 202 via the touch interface. In some implementations, the computing device 202 may lack a display device and thus provide an audible user interface output without providing a graphical user interface output. Additionally, the computing device 202 may provide a user interface such as a microphone for receiving verbal natural language input from a user. In some implementations, the computing device 202 may include a touch interface, and the computing device 202 may lack a camera, although the computing device 202 may optionally include one or more other sensors.
[0024] The computing device 202 and / or other third-party client devices can be in communication with a server device via a network, such as the Internet. In addition, the computing device 202 and any other computing devices can be in communication with each other via a local area network (LAN), such as a Wi-Fi network. The computing device 202 can offload computational tasks to the server device to conserve computational resources at the computing device 202. As an example, the server device can host the automated assistant 204 and / or the computing device 202 can send inputs received at one or more assistant interfaces 220 to the server device. However, in some implementations, the automated assistant 204 can be hosted at the computing device 202, and various processes that can be associated with the operation of the automated assistant can be performed at the computing device 202.
[0025] In various implementations, all or some aspects of the automated assistant 204 can be implemented on the computing device 202. In some of these implementations, aspects of the automated assistant 204 can be implemented by the computing device 202 and interface with a server device that can implement other aspects of the automated assistant 204. The server device can optionally serve multiple users and their associated assistant applications via multiple threads. In implementations in which all or some aspects of the automated assistant 204 are implemented by the computing device 202, the automated assistant 204 can be an application separate from (e.g., installed “on top of”) the operating system of the computing device 202, or alternatively, can be implemented directly by (e.g., integral with but considered an operating system application of) the operating system of the computing device 202.
[0026] In some implementations, the automated assistant 204 can include an input processing engine 206, which can process input and / or output of the computing device 202 and / or the server device using multiple different modules. As an example, the input processing engine 206 can include a speech processing engine 208, which can process audio data received at the assistant interface 220 to identify text embodied in the audio data. To conserve computational resources at the computing device 202, the audio data can be transmitted from the computing device 202 to a server device, for example. Additionally or alternatively, the audio data can be processed solely at the computing device 202.
[0027] The process for converting audio data to text can include a speech recognition algorithm, which can use neural networks and / or statistical models to identify groups of audio data that correspond to words or phrases. The text converted from the audio data can be parsed by a data parsing engine 210 and made available to the automated assistant 204 as text data that can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, the output data provided by the data parsing engine 210 can be provided to a parameter engine 212 to determine whether the user has provided input that corresponds to a particular intent, action, and / or routine that can be performed by the automated assistant 204 and / or an application or agent that can be accessed via the automated assistant 204. For example, assistant data 238 can be stored on the server device and / or the computing device 202, and can include data defining one or more actions that can be performed by the automated assistant 204, as well as parameters necessary to perform the action. The parameter engine 212 can generate one or more parameters for the intent, action, and / or slot value and provide the one or more parameters to the output generation engine 214. The output generation engine 214 can use the one or more parameters to communicate with the assistant interface 220 to provide output to the user and / or to communicate with one or more applications 234 to provide output to the one or more applications 234.
[0028] In some implementations, the automated assistant 204 can be an application that can be installed “on top of” the operating system of the computing device 202 and / or can itself form part of (or the entire) the operating system of the computing device 202. The automated assistant application can include and / or access on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition can be performed using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model stored locally on the computing device 202. The on-device speech recognition generates recognized text for the spoken utterances present in the audio data, if any. Also, for example, on-device natural language understanding (NLU) can be performed using an on-device NLU module that processes the recognized text generated using the on-device speech recognition and, optionally, contextual data to generate NLU data.
[0029] The NLU data can include an intent corresponding to the verbal utterance and, optionally, parameters related to the intent (e.g., slot values). On-device fulfillment can be implemented using an on-device fulfillment module that utilizes the NLU data (from the on-device NLU) and, optionally, other local data, to determine actions to take to resolve the intent of the verbal utterance (and, optionally, parameters related to the intent). This can include determining local and / or remote responses (e.g., answers) to the verbal utterance, interactions with locally installed applications to perform based on the verbal utterance, commands to send to Internet of Things (IoT) devices (directly or via corresponding remote systems) based on the verbal utterance, and / or other resolution actions to perform based on the verbal utterance. The on-device fulfillment can then initiate local and / or remote implementation / execution of the determined actions to resolve the verbal utterance.
[0030] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment can be at least selectively utilized. For example, recognized text can at least selectively be sent to a remote automated assistant component for remote NLU and / or remote fulfillment. By way of example, recognized text can optionally be sent for remote fulfillment in parallel with on-device fulfillment or in response to failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution can be prioritized at least due to the reduced latency they provide when resolving spoken utterances (because client-server round trips are not required to resolve the spoken utterance). Furthermore, on-device functionality can be the only functionality available in situations without network connectivity or with limited network connectivity.
[0031] In some implementations, the computing device 202 may include one or more applications 234 that may be provided by a third-party entity different from the entity that provided the computing device 202 and / or the automated assistant 204. The application state engine of the automated assistant 204 and / or the computing device 202 may access the application data 230 to determine one or more actions that may be performed by the one or more applications 234, as well as the state of each application of the one or more applications 234 and / or the state of each device associated with the computing device 202. The device state engine of the automated assistant 204 and / or the computing device 202 may access the device data 232 to determine one or more actions that may be performed by the computing device 202 and / or one or more devices associated with the computing device 202. Additionally, the application data 230 and / or any other data (e.g., device data 232) can be accessed by the automated assistant 204 to generate context data 236, which can characterize the context in which a particular application 234 and / or device is running and / or the context in which a particular user is accessing the computing device 202, the application 234, and / or any other device or module.
[0032] While one or more applications 234 are executing on the computing device 202, the device data 232 may characterize the current operational state of each application 234 executing on the computing device 202. Additionally, the application data 230 may characterize one or more features of the executing applications 234, such as the content of one or more graphical user interfaces being rendered at the direction of the one or more applications 234. Alternatively or additionally, the application data 230 may characterize action schemas that may be updated by the respective applications and / or by the automated assistant 204 based on the respective applications' current operational state. Alternatively or additionally, the one or more action schemas for one or more applications 234 may remain fixed but may be accessed by the application state engine to determine appropriate actions to initiate via the automated assistant 204.
[0033] Computing device 202 can further include assistant invocation engine 222, which can use one or more trained machine learning models to process application data 230, device data 232, context data 236, and / or any other data accessible to computing device 202. Assistant invocation engine 222 can process this data to determine whether the data should be considered indicative of a user's intent to invoke an automated assistant, instead of waiting for or requesting the user to explicitly speak an invocation phrase to invoke automated assistant 204. For example, one or more trained machine learning models can be trained using training data instances based on a scenario in which a user is present in an environment in which multiple devices and / or applications exhibit various operating states. Training data instances can be generated to capture training data characterizing contexts in which a user invokes an automated assistant and other contexts in which the user does not invoke an automated assistant.
[0034] When one or more trained machine learning models are trained according to these training data instances, assistant invocation engine 222 can cause automated assistant 204 to detect or limit the detection of a spoken invocation phrase from a user based on context and / or environmental features. Additionally or alternatively, assistant invocation engine 222 can cause automated assistant 204 to detect or limit the detection of one or more assistant commands from a user based on context and / or environmental features. In some implementations, assistant invocation engine 222 can be disabled or limited based on computing device 202 detecting an assistant suppression output from another computing device. In this manner, when computing device 202 detects an assistant suppression output, automated assistant 204 will not be invoked based on context data 236 that would otherwise cause automated assistant 204 to be invoked if the assistant suppression output had not been detected.
[0035] In some implementations, system 200 can include a filter identification engine 216 that processes assistant data 238, which can include application data 230, device data 232, context data 236, and / or any other data, to determine whether an application provides access to one or more filter features and / or to determine how to control one or more filter features. Filter identification engine 216 can process assistant data 238 to determine whether an application 234 running on computing device 202 renders one or more selectable GUI elements for controlling one or more filter features. In some implementations, assistant data 238 can be processed according to one or more heuristic processes and / or using one or more trained machine learning models. When one or more filter features are identified for a particular application, filter identification engine 216 can communicate with input word engine 218 of system 200 to determine whether input from a user is associated with the one or more filter features.
[0036] For example, filter identification engine 216 can generate data characterizing one or more filter features of an application being accessed by a user, and automated assistant 204 can compare the data with input from the user. The user can provide input, such as a verbal utterance, including one or more words and / or a request to the automated assistant to have the application conduct a search for specific application content. Automated assistant 204 can compare the natural language content of the input with data from filter identification engine 216 to determine whether the content of the input is relevant to any of the one or more filter features. When automated assistant 204 determines that the user has provided a search request that identifies one or more of the filter features, the automated assistant can generate command data to communicate to the application. The command data received by the application can modify one or more filter parameters of one or more filter features according to the input from the user and cause the application to perform a search operation.
[0037] In some implementations, input from a user can include words that may not be relevant to any filter features of the application but may still be intended by the user to filter search results. As a result, the automated assistant 204 can use the search input engine 226 to determine whether any words in the user input can be used as a basis for generating a search command (i.e., application input) that can be incorporated into the application's search field. For example, when a user includes a word to filter search results (e.g., "search for RAM manufactured this year") but the application does not have a corresponding filtering feature (e.g., no slider bar to limit the manufacturing year), the search input engine 226 can generate a portion of the search command (e.g., "MFR>=2021") to be incorporated into the application's search field when the search operation is performed. Alternatively or additionally, when the input term engine 218 determines that some words of the input can be incorporated into a search command as search terms (e.g., "RAM," "laptop memory," etc.), the search input engine 226 can incorporate such search terms into a search command in combination with any other identified filter parameters (e.g., Search Field: "RAM laptop memory MFR>=2021").
[0038] When a search operation is performed in an application via an automated assistant, the application can render certain content as search results. The search results can be rendered in a search result interface of the application, and the search results can be processed by the search result engine 224 of the system 200. The search result engine 224 can generate additional data based on the search results to determine whether any subsequent input (e.g., user input provided while the search results are being rendered in the foreground of the computing device 202) is relevant to the search results. The search result engine 224 can process data including, but not limited to, screenshots, metadata, source code, and / or any other data that can be associated with the search result interface. In some implementations, the search result engine 224 can generate training data to further train one or more trained machine learning models to render more accurate search results in response to a user request to perform a search operation. For example, term and / or embedding weightings can be changed to modify a particular trained machine learning model to be more reliable when used to process search terms and / or filter parameters specified by a user. For example, the weighting of terms for a first application may differ from a different weighting of those terms for a second application based at least on how reliably each term produces results relevant to a search request from a user.
[0039] 3 illustrates a method 300 for operating an automated assistant to interface with a separate application to search and / or filter specific application content. Method 300 can be performed by one or more applications, devices, and / or any other apparatus or module capable of interacting with the automated assistant. Method 300 can include an operation 302 of determining whether a spoken utterance has been received by the automated assistant. For example, the spoken utterance can be a request to the automated assistant to access an encyclopedia application to identify a particular article (e.g., "Assistant, search my encyclopedia application for articles on encryption technology written this year"). In response to receiving the spoken utterance, the automated assistant can process the audio data to identify one or more requests embodied in the spoken utterance.
[0040] Method 300 can move from operation 302 to operation 304, which can include determining whether a user has requested that a search operation be initiated in another application. Alternatively, when no verbal utterance is received, the automated assistant can continue to detect assistant input. The request for a search operation to be initiated can specify an application that the user wants to use in combination with the automated assistant to search for specific application content. Alternatively or additionally, the request for a search operation can specify one or more words to be used to identify specific application content. When it is determined that the user has requested that a search operation be initiated in another application, method 300 can move from operation 304 to optional operation 306. Otherwise, method 300 can return to operation 302 to detect assistant input from one or more users.
[0041] Optional operation 306 can include processing automated assistant data based on one or more search interfaces of one or more applications. For example, the assistant data can characterize one or more heuristic processes and / or one or more machine learning models that can be used to process data based on the application's search interface. In response to the verbal utterance, the automated assistant can initialize the application so that the application's search interface is rendered on the display interface of the computing device. The assistant data can be processed to determine whether particular features of the search interface can be controlled by the automated assistant. For example, the search interface can include a search field for input of search terms and / or other characters to define a search operation to be performed by the application. Alternatively or additionally, the search interface can include one or more selectable GUI elements for establishing filter settings for the search operation.
[0042] Method 300 can move from optional operation 306 to operation 308 for determining whether the spoken utterance identifies one or more filter parameters associated with the application. For example, the assistant data can be processed along with the input audio data to determine whether there is any correlation between the natural language content of the spoken utterance and one or more filter parameters associated with the application. Following the previous example, the automated assistant can determine that the application's search interface includes filter parameters for filtering encyclopedia articles published before a particular date. When the automated assistant determines that the spoken utterance identifies one or more filter parameters, method 300 can move from operation 308 to operation 310. Otherwise, method 300 can move from operation 308 to operation 314.
[0043] Operation 310 may include modifying one or more filter settings based on the verbal utterance. For example, the verbal utterance may embody one or more filter parameters specified by a user to enable the automated assistant to modify one or more filters of the application accordingly. In some cases, the user may identify one or more filter parameters corresponding to one or more selectable GUI elements, such as one or more checkboxes and / or one or more dials. The automated assistant may determine, based on the one or more filter parameters specified by the user, how to modify the one or more selectable GUI elements to perform a search operation according to the user's request. For example, when a user requests the automated assistant to search an encyclopedia application for articles published after a particular year, the automated assistant may adjust a dial-type selectable GUI element that controls a publication date filter. Method 300 may then move from operation 310 to operation 312, which may include having the application initiate a search operation based on the verbal utterance. As a result, the search operation performed can be initiated using filter parameters specified by the user and performed by the automated assistant without the user having to manually interact with the display interface to activate a particular filter. Alternatively or additionally, the search operation can be initiated using one or more search terms specified within a spoken utterance and incorporated into a search field of the application's search interface.
[0044] In some implementations, when the verbal utterance identifies one or more filter parameters that may not be relevant to the application or otherwise available in the application's search interface, method 300 may move from operation 308 to operation 314. Operation 314 may include generating an application input based on the one or more filter parameters. The application input may be, for example, a search command including alphanumeric characters and / or non-alphanumeric characters that may be provided in a search field of the application to perform a search operation. In some implementations, when a filter parameter is specified by the user but is not available in the search interface, the automated assistant may identify one or more special characters (e.g., non-alphanumeric characters). For example, when the search interface does not include a selectable GUI element for limiting search results with respect to a particular time, the automated assistant may identify one or more special characters and / or expressions (e.g., >, <, >=, <=, etc.), which may be used to filter out some search results that may not be relevant to a particular time range (e.g., "encryption technology <= 1 year"). In this way, users can perform such searches as a single input to an automated assistant, rather than waiting for search results to be displayed and then adjusting any filters that may or may not be available in the search results interface.
[0045] 4 is a block diagram 400 of an exemplary computer system 410. The computer system 410 typically includes at least one processor 414 that communicates with several peripheral devices via a bus subsystem 412. These peripheral devices may include, for example, a storage subsystem 424 including memory 425 and a file storage subsystem 426, a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices enable user interaction with the computer system 410. The network interface subsystem 416 provides an interface to external networks and is coupled to corresponding interface devices in other computer systems.
[0046] The user interface input devices 422 may include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen integrated into a display, a voice recognition system, an audio input device such as a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into the computer system 410 or onto a communications network.
[0047] The user interface output devices 420 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and ways for outputting information from the computer system 410 to a user or to another machine or computer system.
[0048] Storage subsystem 424 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 424 may include logic for performing selected aspects of method 300 and / or for implementing one or more of system 200, computing device 104, and / or any other applications, devices, apparatuses, and / or modules discussed herein.
[0049] These software modules are typically executed by the processor 414 alone or in combination with other processors. The memory 425 used within the storage subsystem 424 may include several memories, including a main random access memory (RAM) 430 for storing instructions and data during program execution and a read-only memory (ROM) 432 in which fixed instructions are stored. The file storage subsystem 426 may provide persistent storage for program files and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation may be stored within the storage subsystem 424 or within another machine accessible by the processor 414 via the file storage subsystem 426.
[0050] Bus subsystem 412 provides a mechanism for allowing the various components and subsystems of computer system 410 to communicate with each other as intended. Although bus subsystem 412 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0051] Computer system 410 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 410 depicted in Figure 4 is intended only as a specific example intended to illustrate some implementations. Many other configurations of computer system 410 are possible, having more or fewer components than the computer system depicted in Figure 4.
[0052] In situations where the systems described herein collect or may use personal information about users (or, as often referred to herein, “participants”), users may be given the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how to receive content from content servers that may be more relevant to the user. Also, certain data may be handled in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identification information may be handled so that personally identifiable information cannot be identified for that user, or if geographic location information is obtained, the user's geographic location may be generalized (e.g., to the city level, zip code level, or state level) so that the user's specific geographic location cannot be identified. Thus, users may have control over how information is collected and / or used about them.
[0053] While several implementations have been described and illustrated herein, various other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein can be utilized, and each such variation and / or modification is considered to be within the scope of the implementations described herein. More generally, any parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications in which the present teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, it should be understood that the foregoing implementations are presented by way of example only, and that, within the scope of the appended claims and their equivalents, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0054] In some implementations, a method implemented by one or more processors is described as including an operation such as receiving, at a computing device, a verbal utterance from a user directed to an automated assistant accessible through the computing device, the computing device including a display interface rendering a search results interface of an application when the verbal utterance is received. The method may further include an operation of determining, based on the verbal utterance, whether the user has specified a particular filter setting that is not included in one or more filter settings of the search results interface. When the automated assistant determines that the user has specified a particular filter setting that is not included in the one or more filter settings, the method may further include an operation of generating application content for the search results interface and application input based on the verbal utterance, and causing the application to initiate performance of a search operation based on the application input.
[0055] In some implementations, the application input identifies one or more search terms that correspond to specific filter settings and application content that is rendered in the search results interface. In some implementations, the application input includes one or more words included in a spoken utterance and one or more search terms that were previously submitted to the application to cause the application content to be rendered in the search results interface. In some implementations, when the automated assistant determines that the user has specified a specific filter setting included in the one or more filter settings, the method can further include an operation of modifying the specific filter setting of the application according to the spoken utterance from the user, wherein modifying the specific filter setting causes different application content to be rendered in the search results interface. In some implementations, causing the application to initiate performance of a search operation based on the application input includes incorporating the application input into a search field of the search results interface.
[0056] In some implementations, determining whether the user has specified a particular filter setting that is not included in the one or more filter settings of the search result interface includes processing assistant data based on one or more search interfaces previously rendered by the application or another application. In some implementations, when the automated assistant determines that the user has specified a particular filter setting that is not included in the one or more filter settings, the method can further include the operation of determining a respective state of each filter setting of the one or more filter settings of the search result interface, where the application input is further based on the respective state of each filter setting of the one or more filter settings of the search result interface. In some implementations, causing the application to initialize performance of the search operation based on the application input includes causing the application to render a filtered subset of the application content according to the respective state of each filter setting of the one or more filter settings.
[0057] In another implementation, a method implemented by one or more processors is described as including an operation, such as receiving, from a user, a verbal utterance at a computing device to prompt an automated assistant to initiate a search operation using an application separate from the automated assistant, the verbal utterance identifying one or more words. The method may further include an operation of determining, based on the verbal utterance, whether one or more words of the verbal utterance are associated with one or more selectable graphical user interface (GUI) elements rendered in an interface of the application, the one or more selectable GUI elements controlling one or more filter parameters of a search feature of the application. When the one or more words of the verbal utterance correspond to one or more selectable GUI elements rendered in the interface of the application, the method may further include an operation of causing one or more specific selectable GUI elements of the one or more selectable GUI elements to control one or more specific filter parameters and causing the application to initiate a search operation according to the one or more specific filter parameters.
[0058] In some implementations, causing the application to initiate a search operation includes causing a search field of the application to include one or more words identified in the spoken utterance without the search field including one or more particular filter parameters. In some implementations, the method may further include, when the one or more words of the spoken utterance do not correspond to one or more selectable GUI elements rendered in an interface of the application, the operation of generating application input characterizing one or more particular filter parameters and causing the application to initialize a search operation using the application input. In some implementations, causing the application to initialize a search operation using the application input includes causing a search field of the application to include the application input, where the application input specifies one or more particular filter parameters.
[0059] In some implementations, the application input includes non-alphanumeric characters selected based on one or more specific filter parameters. In some implementations, the spoken utterance includes a request to the automated assistant to search for specific content using the application and then render the specific content for the user. In some implementations, when one or more words of the spoken utterance do not correspond to one or more selectable GUI elements rendered in the interface of the application, the method can further include the operations of accessing specific application content that satisfies the one or more specific filter parameters, and, after accessing the specific application content, causing the automated assistant to render audible content based on the specific application content.
[0060] In some implementations, the method may further include the operations of: causing a plurality of different search results to be rendered in another interface of the application based on the search operation when one or more words of the spoken utterance correspond to one or more selectable GUI elements rendered in the interface of the application; after rendering the plurality of different search results, receiving an additional spoken utterance from the user, the additional spoken utterance specifying one or more additional words for identifying a subset of the plurality of different search results; and, in response to receiving the additional spoken utterance, causing the plurality of different search results to be filtered according to the one or more additional words. In some implementations, causing the plurality of different search results to be filtered according to the one or more additional words includes determining that one or more other selectable GUI elements rendered in the other interface correspond to the one or more additional words and selecting the one or more other selectable GUI elements according to the one or more additional words.
[0061] In some implementations, a method implemented by one or more processors is described in a computing device as including operations such as receiving a spoken utterance including a request to an automated assistant to conduct a search of application content accessible through an interface of the application, the spoken utterance identifying one or more words related to the application content to be searched. The method may further include an operation of determining, based on the spoken utterance, whether the application provides one or more filtering features for filtering the application content according to the one or more words identified in the spoken utterance. When it is determined that the application does not provide one or more filtering features, the method may further include operations of identifying, based on the one or more words, one or more filter parameters to submit to the application to facilitate conducting a search of the application content; causing the automated assistant to provide application input to the application, the application input identifying the one or more filter parameters; and causing the application to render search results based on the application input, the search results including a subset of the application content that satisfies the one or more filter parameters.
[0062] In some implementations, the one or more filtering features include one or more selectable graphical user interface (GUI) elements that control one or more filtering operations of the application. In some implementations, when it is determined that the application does not provide the one or more filtering features, the method may further include an operation of identifying, based on the one or more filter parameters, one or more non-alphanumeric characters selected based on the one or more filter parameters, wherein the application input identifies the one or more non-alphanumeric characters. [Explanation of symbols]
[0063] 100 views 102 users 104 Computing Devices 106 Oral Speech 120 Views 122 search fields 124 Selectable GUI Elements 126 results 128 images 130 selectable checkboxes 132 Applications 134 Alternative oral utterances 136 other selectable elements 138 Search Results Interface 140 Views 142 results 144 additional verbal utterances 146 filters 148 search terms, search fields 160 Views 162 Filtered Search Results, Most Recent Search Results 164 Oral Speech 166 Subsequent oral utterances 200 systems 202 Computing Devices 204 Automated Assistant 206 Input Processing Engine 208 Audio Processing Engine 210 Data Parsing Engine 212 Parameter Engine 214 Output Generation Engine 216 Filter Identification Engine 218 Input Word Engine 220 Assistant Interface 222 Assistant Call Engine 224 Search Results Engine 226 Search Input Engine 230 Application Data 232 Device Data 234 Applications 236 Context Data 238 Assistant Data 300 ways 400 Block Diagram 410 Computer Systems 412 Bus Subsystem 414 processor 416 Network Interface Subsystem 420 User Interface Output Device 422 User Interface Input Devices 424 Storage Subsystem 425 memory 426 File Storage Subsystem 430 Main Random Access Memory (RAM) 432 Read-Only Memory (ROM)
Claims
1. 1. A method implemented by one or more processors, comprising: receiving, at a computing device, a verbal utterance from a user directed to an automated assistant accessible via the computing device, the computing device includes a display interface rendering an application search result interface when the spoken utterance is received, and the automated assistant has access to a plurality of external applications; determining, based on assistant data, whether the application is rendering one or more selectable graphical user interface (GUI) elements; and, when the automated assistant determines that the application is rendering the one or more selectable GUI elements, determining, based on the verbal utterance, whether the user has specified a particular filter setting that is not included in one or more filter settings associated with the one or more selectable GUI elements of the search result interface; When the automated assistant determines that the user has specified the particular filter setting that is not included in the one or more filter settings, generating application input based on application content of the search result interface and the spoken utterance; causing the application to initiate performance of a search operation based on the application input; A method comprising:
2. The method of claim 1 , wherein the application input identifies one or more search terms that correspond to the particular filter settings and the application content being rendered in the search results interface.
3. 3. The method of claim 1 or claim 2, wherein the application input includes one or more words contained within the spoken utterance and one or more search terms previously submitted to the application to cause the application content to be rendered in the search results interface.
4. When the automated assistant determines that the user has specified the particular filter setting included in the one or more filter settings, changing the particular filter setting of the application in accordance with the verbal utterance from the user, The method of claim 1 , further comprising the step of causing different application content to be rendered in the search result interface by changing the particular filter setting.
5. causing the application to initiate performance of the search operation based on the application input; The method of claim 1 , further comprising incorporating the application input into a search field of the search results interface.
6. determining whether the user has identified the particular filter setting that is not included in the one or more filter settings of the search result interface, 6. The method of claim 1, including processing assistant data based on one or more search interfaces previously rendered by the application or another application.
7. When the automated assistant determines that the user has specified the particular filter setting that is not included in the one or more filter settings, determining a respective state of each filter setting of the one or more filter settings of the search result interface; The method of claim 1 , further comprising the step of: the application input being further based on a respective state of each filter setting of the one or more filter settings of the search result interface.
8. causing the application to initiate performance of the search operation based on the application input; 8. The method of claim 7, comprising causing the application to render a filtered subset of application content according to a respective state of each filter setting of the one or more filter settings.
9. A computer program comprising instructions which, when executed by one or more processors of a computing system, cause the computing system to perform the method of any one of claims 1 to 8.
10. One or more computing devices configured to perform the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Voice recognition apparatus, voice recognition method and voice recognition program
JP2008089625A
Data search device, navigation apparatus and data search method
JP2009266129A
Information processing system
JP2014106927A
A Method for Adaptive Conversational State Management with Filtering Operators Applied Dynamically as Part of a Conversational Interface
JP2016502696A
Information processing device
JP2017058599A