Determine the response content of a composite query based on a set of generated subqueries
By generating a subquery set of composite queries and distributing them to the agent to generate response content, the problem of low composite queries processing efficiency in the prior art is solved, and more efficient user interaction and computing resource utilization is achieved.
Patent Information
- Application Number
- CN201880094249.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-05-07
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2038-11-11
AI Technical Summary
Existing computing systems are difficult to effectively handle composite queries, resulting in users needing to split queries into multiple separate queries, increasing computing resource consumption and interaction complexity.
By generating a subquery set of composite queries, the quality score is determined and distributed to the corresponding agent to generate the response content and finally rendered to the user.
Simplify user input, reduce the number of interactions, improve computing efficiency, and reduce network and computing resource consumption.
Smart Images

Figure CN112236765B_ABST
Abstract
Description
Background Art
[0001] Computing systems can process various queries from humans. For example, a human can engage in a human - machine conversation with an interactive software application referred to herein as an "automatic assistant" (also known as a "chatbot", "interactive personal assistant", "intelligent personal assistant", "personal voice assistant", "conversation agent", etc.). For example, a human (who can be referred to as a "user" when interacting with the computing system) can provide commands, queries, and / or requests (collectively referred to herein as "queries") using free - form natural language input and / or free - form natural language input typed in, which can be voice utterances that are converted into text and then processed.
[0002] The computing system uses response content to respond to the query. For example, in response to a query of "What time is it in Paris, France" submitted to a search system, the search system can provide a response reflecting the current time in Paris, France. Similarly, for example, in response to the utterance "Turn off the couch light" provided to an automatic assistant, the automatic assistant can turn off the "couch light" and provide a response confirming that the couch light has been turned off. Although various computing systems can appropriately respond to many queries, many systems are unable to execute compound queries. For example, many computing systems fail due to the user's utterance "What time is it in Paris, France, what's the population, and what's the average high temperature in July". For example, many computing systems fail by providing an incorrect response, not determining an appropriate response, but providing a default error response, etc. As a result, users are often forced to split compound queries into multiple different types of queries, which may collectively require more computing and / or network resources to process. In addition, the interaction between the user and the computing system is inhibited. Summary of the Invention
[0003] The implementations described herein involve determining that a set of multiple sub-queries is a proper interpretation of a composite query based on the query submitted as the composite query. Those implementations further involve, in response to such determination, providing a corresponding command for each sub-query of the determined set. Each command is sent to a corresponding agent (in one or more agents) and causes the agent to generate and provide corresponding response content. Additionally, those implementations further involve causing content to be rendered in response to the submitted query, where the content is based on the corresponding response content received in response to the commands.
[0004] Accordingly, the implementations disclosed herein enable a user to submit a composite query and to generate and provide a proper response to those composite queries. For example, this can improve human-machine interaction by reducing, e.g., the amount of input that needs to be provided by the user (e.g., the input provided by the user to an automated assistant interface). For example, the user can provide a single spoken utterance that includes the composite query and has a first duration, instead of needing to provide multiple different types of utterances that together have a second duration that is longer than the first duration. In addition to improving human-machine interaction (e.g., human-automated assistant interaction), such implementations can also directly result in various computer and / or network efficiencies as compared to, e.g., techniques that require the user to break a composite query into multiple different types of queries. For example, a single spoken utterance that includes the composite query can have a shorter duration and / or include fewer words as compared to multiple spoken utterances that break the composite query into multiple different types of queries. Thus, the speech-to-text and / or other processing of such a single spoken utterance can be more efficient than the speech-to-text processing of multiple spoken utterances. As another example, a single client-to-server transmission (or its client-side speech-to-text conversion) of a single spoken utterance can be utilized to convey the composite query, instead of multiple (overall higher data) transmissions that may be required for multiple spoken utterances. As yet another example, a single server-to-client transmission can be utilized to provide content in response to the composite query, where the size of the content can be less than the size of multiple content instances in response to multiple spoken utterances, and / or the content can be rendered more efficiently (e.g., in less time) as compared to multiple content instances in response to multiple spoken utterances.
[0005] In various implementations, a user submits a compound query by speaking a single spoken utterance. For example, speech-to-text conversion of the single spoken utterance can be performed to generate text corresponding to the compound query. For example, in response to the user providing the corresponding spoken utterance via an assistant interface of a client device, text such as "How far is Paris and what's the weather there" can be generated. The assistant interface can be implemented via an automatic assistant client of the client device and enables the user to interact with the automatic assistant, which can be implemented via the automatic assistant client and / or via one or more remote computing devices (e.g., cloud-based automatic assistant components).
[0006] The implementations described herein generate a set of one or more subqueries, where each subquery in each set collectively represents an interpretation of the compound query. As an example, and continuing with the previous example, a set of subqueries can be generated that includes subqueries such as "How far is Paris" and "What's the weather in Paris".
[0007] In addition, a quality score can be determined for each set of subqueries. If the quality score of one of the sets meets a threshold, that set can be selected as the set to be used to obtain and / or provide content for rendering in response to the composite query. In some implementations, determining that the quality score of a given set meets the threshold can include: determining that the quality score of the given set meets the threshold relative to the quality scores of other sets (thereby indicating that the subqueries of the given set are a more appropriate interpretation compared to the subqueries of other sets); determining that the quality score of the given set meets the threshold relative to a composite query quality score, which is optionally determined for the composite query without splitting it into subqueries (thereby indicating that splitting the composite query into subqueries is more appropriate compared to not splitting the composite query); and / or determining that the quality score of the given set meets a fixed threshold (thereby indicating that the subqueries of the given set are a consistent interpretation). In some implementations, when determining the quality score of a set of subqueries, an individual quality score can be determined for each subquery of the set, and then the quality score of the subqueries of the set can be determined based on the individual quality scores of the subqueries of the set. The quality score of a given set can indicate, for example, the likelihood that the set correctly represents the intent of the individual queries included in the text. For example, for the composite query "What's the weather in Turks and Caicos and Haiti", a first set of subqueries can be generated, which includes "What's the weather in Turks", "What's the weather in Caicos", and "What's the weather in Haiti" (i.e., a set that does not correctly reflect the user's intent since "Turks and Caicos" is a country). In addition, a second set of subqueries can be generated, which includes "What's the weather in Turks and Caicos" and "What's the weather in Haiti", correctly separating the individual countries. According to the techniques described herein, a first quality score can be determined for the first set, and a second quality score can be determined for the second set, and the second set (which can appropriately reflect the user's intent) can be selected over the first set based on the first and second quality scores.
[0008] As described above, if the quality score of one of the sets meets the threshold, it can be selected as the set that will be used to obtain and / or provide content for rendering in response to the composite query. For each sub-query of the selected set, a command is generated and provided to the corresponding agent, the response content is received from the corresponding agent, and the content is provided for rendering, where the content is based on the response content received from the agent. The agent to which the command is provided can be, for example, a general search system or a specific agent, such as the first-party (IP) or third-party (BP) assistant agent described herein. As an example, for the sub-query "What's the weather in Paris", a structured command can be provided to a weather agent that requests the current weather in "Paris". As another example, for the sub-query "Turn on the couch light", a structured command can be provided to a 3P lighting agent, causing the 3P agent to transition the "couch light" (mapped to the client device via which the composite query was submitted) to the "on" state. As yet another example, for the sub-query "can an ant jump", a command that conforms to the query can be submitted to a search system. In some implementations, multiple commands generated from the set of sub-queries can be provided to a single agent, while in other implementations, a first command is provided to a first agent, a second command is provided to a second agent, and so on.
[0009] In response to being provided with a command, a corresponding agent provides response content. Then, the content based on the response content can be rendered to the user. The rendered content can be the response content itself, or content determined based on the response content but not strictly conforming to the response content. For example, a first sub-query in response to a set response content from a weather agent can include the text "80 degrees and sunny", and a second sub-query in response to a set response content from a lighting agent can include "confirmation" that implements a lighting command. The content to be rendered can include an audible "ding" indicating that the lighting command has been implemented, and a spoken output of "80 degrees and sunny", i.e., a text-to-speech conversion of the text "80 degrees and sunny". In some implementations, one or more terms from the original sub-query can be utilized to render the response content. For example, for the "Paris" command sent to a weather agent generated from a sub-query based on "What's the weather in Paris", the agent can provide the result "68", and the term "Paris" identified from the sub-query can be utilized to render the response content to the user via an audible response "The weather in Paris is 68 degrees Fahrenheit right now". In some implementations, the response content from an assistant agent can be used to generate additional sub-queries, and the response content from the additional sub-queries is rendered to the user.
[0010] In some implementations, one or more entities can be identified in the query text and can be used to generate a set of sub-queries. For example, for the text "Tell me how long it would take to bike 25 miles and run 2 miles", aliases of entities associated with "biking" and "running" can be identified. Sub-queries can be generated that include patterns commonly used as search queries, such as "How long would it take to travel <distance>by <mode of transportations (The mode of "How long would it take to travel <distance> by <mode of transportation>"). Since the query includes pairs of distances and modes of transportation, a set of two sub-queries is generated: "How long would it take to travel 25 miles by bike" and "How long would it take to travel 2 miles by running". In some cases, multiple modes can be identified in a single text (e.g., "I want to go to Paris, show me restaurants and hotels there").
[0011] As another example, a mode can include one or more terms and / or entities, such as "What's the weather in <city> ( <city>How about the weather in)”, where <city>Placeholder for the alias of a city entity. Thus, for a query like "What's the weather in Paris and Rome", subqueries like "What's the weather in Paris" and "What's the weather in Rome" can be generated based on the similarity between the terms in the text and the terms in the pattern. In some cases, commonly recognized patterns can be associated with available proxies, and these common patterns can be used to determine subqueries. For example, a weather proxy can be associated with "weather in <city> ( <city>"weather in <city>on <date> ( <date>When <city>"weather in" and / or "temperature in" <city> ( <city>is associated with a common pattern of "temperature in)". Then, the associated pattern can be utilized to generate one or more subqueries based on entities in the text and the similarity between the pattern and the text.
[0012] In some implementations, one or more subqueries can be generated based on prefixes and / or suffixes in the text (i.e., one or more terms at the beginning and end of the text, respectively). For example, the text "When were Romania and Poland founded" includes the prefix "When were" and the suffix "founded". A set of subqueries can be generated that includes "When was Romania founded" and "When was Poland founded", both of which include the prefix of the original text (with a re-conjugation of the verb) and the suffix.
[0013] As another example, one or more subqueries can be generated based on text that includes only a prefix or only a suffix. For example, the text "turn on the kitchen lights and the TV in the living room" includes the prefix "turn on", which applies to two different actions (i.e., turning on the kitchen lighting and turning on the TV in the living room). By appending the prefix "turn on" to each subquery, a set of subqueries can be generated that includes "turn on the kitchen lights" and "turn on the living room TV".
[0014] The above description is provided as an overview of the various implementations disclosed herein. Those various implementations, as well as additional implementations, are described in more detail herein.
[0015] On the one hand, a method implemented by one or more processors is provided. The method includes: receiving, at an assistant interface of a user's client device, text generated in response to detecting a single spoken utterance of the user, the text corresponding to the single spoken utterance and generated based on a speech-to-text conversion of the single spoken utterance; generating a set of subqueries based on the text, the set of subqueries including at least a first subquery and a second subquery, wherein the set of subqueries together define a candidate interpretation of the text; determining a quality score of the set of subqueries; and determining that the quality score of the set of subqueries meets at least one threshold. In response to determining that the quality score meets the threshold, the method further includes: providing a corresponding command to a corresponding assistant agent for each subquery of the set, the corresponding command including at least a first command based on the first subquery and a second command based on the second subquery; receiving corresponding response content in response to providing the corresponding command, the corresponding response content including at least a first response content in response to the first command and a second response content in response to the second command; and causing the client device to render content based on the corresponding response content to the user.
[0016] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0017] In some implementations, causing the client device to render content to the user may include causing the client device to render a first content based on the first response content, and may further include causing the client device to render a second content based on the second response content.
[0018] In some implementations, the first command is provided to a first assistant agent, and the first assistant agent is a third-party agent controlled by a third party different from the party controlling the assistant interface. In those implementations, the first command may be a structured command generated based on the first subquery that causes the first assistant agent to change the state of a first connected device, and the first response content may include an acknowledgement of the change in state in response to the structured command.
[0019] In some implementations, the method may further include: generating a second set of second subqueries based on the text, the second subqueries of the second set being unique compared to the subqueries of the set and together defining an additional candidate interpretation of the text; and determining an additional quality score of the second set of second subqueries. Additionally, it may be determined that the quality score of the set of subqueries meets the threshold based on a comparison of the quality score of the set with the additional quality score of the second set.
[0020] In some implementations, determining a quality score for a set may include: determining a first quality score for a first subquery; determining a second quality score for a second subquery; and determining the quality score for the set based on the first quality score and the second quality score. In some of those implementations, determining the first quality score for the first subquery may include determining whether the first subquery conforms to one or more identified commands of any assistant agent, and determining the second quality score for the second subquery may include determining whether the second subquery conforms to one or more identified commands of any assistant agent.
[0021] In some implementations, the set may include at least three subqueries.
[0022] In some implementations, generating the set may include generating a first portion of text and a second portion of text, where the first portion is unique compared to the second portion, and generating the first subquery based on the first portion. Additionally, in those implementations, the second subquery may be generated based on the second portion and based on one or more terms in the text that are not included in the second portion of the text. In some versions of those implementations, the first subquery may further be generated based on one or more terms in the text that are not included in the first portion of the text. In some versions of those implementations, one or more terms that are not included in the second portion of the text and are used in generating the second subquery may include at least one term from the first portion of the text. In some of those versions, one or more terms that are not included in the second portion of the text and are used in generating the second subquery may include at least one term that occurs before the first portion of the text. In some of those versions, one or more terms that are not included in the second portion of the text and are used in generating the second subquery may include at least one term that occurs after the second portion of the text.
[0023] In some implementations, generating the set may include: identifying aliases for one or more entities included in the text; identifying a first pattern based on the text, where the first pattern includes one or more terms and an entity type; generating the first portion based on the first pattern; identifying a second pattern based on the text, where the second pattern includes one or more second terms and an entity type or an additional entity type; generating the second portion based on the second pattern; generating the first subquery by replacing the entity type in the first portion with a corresponding alias; and generating the second subquery by replacing the entity type or the additional entity type in the second portion with a corresponding alias. In some of those versions, the first pattern and the second pattern are unique, and the alias in the first subquery may be the same as the alias in the second subquery. In some of those implementations, the first pattern and the second pattern may be the same, and the alias in the first subquery may be different from the alias in the second subquery.
[0024] In some implementations, the method may further include: generating an aggregate query based on the response content; providing an aggregate command to one or more additional assistant agents based on the aggregate query; and receiving, in response to providing the aggregate command, aggregate response content from the one or more additional assistant agents based on the aggregate response content.
[0025] In some implementations, the method may further include determining a text quality score for the text, wherein determining that the quality score of the set of subqueries meets a threshold may be based on a comparison of the quality score with the text quality score. Additionally, in response to the quality score of the set of subqueries meeting the threshold, the method may further include generating a command based on the set of subqueries instead of generating a command based on the text.
[0026] On the other hand, provided is another method implemented by one or more processors, the method including: receiving, at an interface of a user's client device, text generated in response to a user input; generating a set of subqueries based on the text including at least a first subquery and a second subquery, wherein the set of subqueries collectively define a candidate interpretation of the text; determining a set quality score for the set of subqueries; determining a text quality score for the text; and determining that the quality score of the set of subqueries meets at least one threshold, including a threshold relative to the text quality score of the text. In response to determining that the quality score meets the threshold, the method further includes: providing a corresponding command to a corresponding agent for each subquery of the set, the corresponding command including at least a first command based on the first subquery and a second command based on the second subquery; receiving corresponding response content in response to providing the corresponding command, including at least a first response content in response to the first command and a second response content in response to the second command; and causing the client device to render content to the user based on the corresponding response content.
[0027] On the other hand, another method implemented by one or more processors is provided, the method comprising: receiving, at an interface of a user's client device, text generated in response to a user input; generating a set of subqueries based on the text, the set of subqueries including at least a first subquery and a second subquery, the set of subqueries jointly defining a candidate interpretation of the text; determining a set quality score for the set of subqueries; and determining that the quality score for the set of subqueries meets at least one threshold, including a threshold relative to a text quality score of the text. In response to determining that the quality score meets the threshold, the method further comprises: for each subquery of the set, providing a corresponding command to a corresponding agent, the corresponding command including at least a first command based on the first subquery and a second command based on the second subquery; receiving corresponding response content in response to providing the corresponding command, the corresponding response content including at least a first response content in response to the first command and a second response content in response to the second command; and causing the client device to render a first content based on the first response content to the user and render a second content based on the second response content to the user.
[0028] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0029] In some implementations, the method further comprises determining that the text meets one or more conditions. Execution of generating the set of subqueries may be performed only after determining that the text meets one or more conditions.
[0030] On the other hand, a method implemented by one or more processors is provided, the method comprising: receiving, at an assistant interface of a user's client device, text generated in response to detecting a single spoken utterance of the user, the text corresponding to the single spoken utterance and generated based on a speech-to-text conversion of the single spoken utterance; generating a first portion of the text and a second portion of the text, wherein the first portion is unique compared to the second portion; generating a first subquery based on the first portion; generating a second subquery based on the second portion and based on one or more terms of the text not included in the second portion of the text; for each subquery of the set, providing a corresponding command to a corresponding assistant agent, the corresponding command including at least a first command based on the first subquery and a second command based on the second subquery; receiving corresponding response content in response to providing the corresponding command, the corresponding response content including at least a first response content in response to the first command and a second response content in response to the second command; and causing the client device to render content based on the corresponding response content to the user.
[0031] Additionally, some implementations include one or more processors of one or more computing devices, where the one or more processors are operable to execute instructions stored in an associated memory, and where the instructions are configured to cause the execution of any of the methods described herein. Some implementations also include one or more non-transitory computer-readable storage media that store computer instructions executable by one or more processors to execute any of the methods described herein.
[0032] It should be understood that all combinations of the foregoing concepts and additional concepts described in greater detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 Depicts an example process flow according to various implementations disclosed herein, which example process flow illustrates generating a set of multiple subqueries from a single utterance and generating example response content for each subquery in the set.
[0034] Figure 2 Depicts an example process flow according to various implementations disclosed herein, which example process flow illustrates an example of generating a set of subqueries from text including a compound query.
[0035] Figure 3 Illustrates an example process flow according to various implementations disclosed herein, which example process flow illustrates generating response content for multiple subqueries generated based on a compound query and utilizing the response content in generating content for rendering in response to the compound query.
[0036] Figure 4 Is a block diagram of an example environment in which response content can be generated in response to an utterance including a compound query.
[0037] Figure 5 Illustrates a flowchart of an example method according to various implementations disclosed herein.
[0038] Figure 6 Illustrates an example architecture of a computing device. DETAILED DESCRIPTION
[0039] First, referring to Figure 1 , a single spoken utterance 180 of a user is received. For example, the single spoken utterance 180 can be generated based on output from one or more microphones of a client device and can be embodied in audio data generated in response to the user speaking the spoken utterance by one or more microphones. The single spoken utterance 180 can be a continuous utterance of the user, such as a single utterance determined by a voice activity detector based on analysis of the audio data.
[0040] The spoken utterance 180 is provided to the speech-to-text module 162, which converts the spoken utterance 180 into text 181. In some cases, a single spoken utterance 180 may include a compound query as it includes multiple sub-queries that are combined into a single utterance. For example, the text 181 "I want to go to Paris, show me hotels and the weather" includes two independent but related queries. One sub-query is related to hotels in Paris and the second sub-query is related to the current weather in Paris.
[0041] Although Figure 1 the techniques shown in start with an input of a single spoken utterance 180, the text 181 can be received from the user via one or more alternative sources. For example, the user can directly submit the text 181, e.g., via a virtual keyboard (e.g., via a touch-sensitive display) and / or via another input device. Thus, although described in as originating from a single spoken utterance, the text 181 can be generated based on any single input from the user. Figure 1 as being derived from a single spoken utterance, the text 181 can be generated based on any single input from the user.
[0042] As described in more detail below, the set generator 121 can generate a set of one or more sub-queries (e.g., set 182) based on the text 181, and can select a given set of sub-queries as the one for which response content should be generated and render a corresponding set of content in response to the single utterance 180. For example, a given set can be selected based on a determination that the quality score of the given set by the scorer 122 meets at least one threshold. In some of those implementations, the scorer 122 can also determine the quality score of the text 181 itself and select the given set based on a determination that its quality score meets a threshold relative to the quality score for the text 181 itself. In other words, in those implementations, a given set of sub-queries is only selected for use instead of the text 181 itself if the quality score of the selected given set meets a threshold relative to the quality score for the text 181 (and optionally if it meets one or more other thresholds).
[0043] In some implementations, the set generator 121 always generates one or more sets and compares the quality score of the generated sets with the quality score of the text 181. In some implementations, the generation of sets by the set generator 121 depends on the text 181 meeting one or more criteria. For example, one or more criteria may include that the quality score of the text 181 does not meet a threshold, thereby indicating that the response content of the text 181 (if any) is of low quality. Similarly, for example, one or more criteria may additionally or alternatively include that the text 181 does not match the required input for any single agent, the text 181 is unable to return any "non-error" response content from one or more (e.g., any) agents, and so on. Similarly, for example, one or more criteria may additionally or alternatively include that the text 181 includes one or more terms indicating a possible compound query (e.g., the text 181 includes the term "and").
[0044] The set generator 121 generates a set of one or more subqueries, such as a set 182 that includes two or more subqueries 183 and 184. Although the set 182 is shown as including two subqueries, the set generator 121 may generate a set that includes any number of subqueries, as further described herein. For example, the text 181 "What are the populations of Paris and London, and who are their mayors" may cause the set 182 to include the subqueries "What is the population of Paris", "What is the population of London", "Who is the mayor of Paris", and "Who is the mayor of London". When generating the set 182, the set generator 121 may utilize one or more of a variety of techniques, such as those further described herein.
[0045] In some implementations, the set generator 121 may generate the subqueries 183 and 184 based on a prefix that occurs at the beginning of the text 181 and / or a suffix that occurs at the end of the text 181. For example, and referring to Figure 2 , the text "When were Poland and Romania founded" 205 includes the prefix "When were" 210 and the suffix 215 "founded". The set generator 121 can determine that the non-prefix and non-suffix portion of the text 205 "Poland and Romania" can be divided into two sub-queries based on (for example) two terms separated by "and" for the two countries. Thus, the set generator 121 divides the middle portion into two parts, "Poland" 220 and "Romania" 225. The prefix 210 and suffix 215 can then be attached to the portion 220 and the portion 225, resulting in sub-queries "When was Poland founded" 230 and "When was Romania founded" 235. As shown in this example, when prefixes and / or suffixes are attached to the portions to allow for subject consistency, some terms of the prefix and / or suffix may be slightly changed, such as verb conjugation. In addition, although Figure 2 Examples of include both prefixes and suffixes, but in some cases, the set generator 121 may append only prefixes or only suffixes, depending on the initial text 181. For example, for the text "Turn on the kitchen lightsand the living room TV", the prefix "Turn on" may be used to generate subqueries "Turn on the kitchenlights" and "Turn on the living room TV".
[0046] In some implementations, the set generator 121 can generate sub-queries 183 by identifying one or more patterns in the text 181. Patterns can be identified from a database that includes frequent search query patterns, frequent voice commands, and / or other frequent queries. Frequent patterns can be identified based on a log of previous commands. Also, for example, one or more agents can be associated with patterns that are commonly identified as input for a particular agent, and the set generator 121 can examine the text 181 for terms that are close to or common to available agent commands. For example, an agent that provides the temperature at a location can be associated with a common pattern of asking about the weather, such as "What is the current temperature in a certain location?" <city> ( <city>What's the current temperature in) and / or "What's the weather in <city> ( <city>What's the weather like)". Similarly, for example, an agent controlling one or more lighting devices can have associated patterns "Turn on the lights" and / or "Turn off the <room>lights(turn off the lights in the <room>).
[0047] In some implementations, text 181, which has been tagged with voice, entity, and / or other background information, can be provided to the set generator 121. The tags of the text 181 can be used to identify potential patterns that can be provided to the agent. For example, the set generator 121 can receive text with tagged terms such as "I want to go to Paris, show me bars and what's the weather there" (I want to go to Paris, tell me about the bars and the weather there), including being tagged as <city>The entity "Paris". Based on the terms in the text, the set generator 121 can identify the common query "bars in <city>” and "weather in <city>”. Based on the identified patterns, the set generator 121 can generate the first sub-query "bars in Paris" and the second query "weather in Paris" by replacing the entity placeholder with the cities tagged in text 181. As another example, the text "I want to go Paris or Rome. What's the weather there (I want to go to Paris or Rome. What's the weather like there)" can result in the identified pattern "weather in <city>" and generates the subqueries "weather in Paris" and "weather in Rome".
[0048] In some implementations, the set generator 121 can generate subqueries 183 by context expansion of the text 181. The context expansion includes identifying a list of entities in the text 181, and further identifying one or more terms and / or patterns in the text 181. Thus, the entity list can be used to expand the text into multiple subqueries for each entity in the list. For example, for the text "What are the populations of Rome, Paris, and London, and who are their mayors", identify the subqueries including "Rome", "Paris", and "London". <city>List of entities. Two further patterns were identified in the text: "Population of <city>” and "Mayor of <city>". Then, the set generator 121 can generate a set including the following subqueries "population of Rome", "population of Paris", "population of London", "mayor of Rome", "mayor of Paris", and "mayor of London" by applying each entity to each identified pattern.
[0049] Once the set 182 of subqueries 183 and 184 has been generated, the set 182 is provided to the set scorer 122 to determine the quality score of the set 182. In some implementations, the scorer 122 determines a quality score indicative of the likelihood that a user submitting the text 181 may obtain results from the subqueries of the set 181. For example, for a set 182 including the subqueries "Weather in Paris" and "hotels in Paris", the scorer 122 can determine a quality score indicative of the likelihood that the queries included in the original text 181 "I want to go to Paris, show me the weather and hotels there" are represented by the subqueries. In some implementations, the quality score can indicate the likelihood that one or more available agents will receive the subqueries as input. For example, the set generator 121 can generate a set including the subqueries "What's the weather in Paris" and "What's the weather in London", and determine a quality score indicative of the likelihood that a weather agent will provide meaningful response content based on the subqueries.
[0050] In some implementations, in addition to the set 182, the original text 181 can also be provided to the set scorer 122. A quality score can be determined for the set 182, which indicates whether the set 182 would elicit meaningful content from available agents if provided to one or more agents. For example, for the text "What's the weather in Trinidad and Tobago", the set generator 121 can generate subqueries "weather in Trinidad" and "weather in Tobago". However, since "Trinidad and Tobago" is a single country, the original text "What's the weather in Trinidad and Tobago" can provide a more meaningful result than the subqueries. Thus, the quality score of the original text can be more indicative of quality than the quality score of the set. Thus, in those implementations, the original text 181 can be used to generate commands to be provided to the agent, and the response content from the agent can be used to render content to the user. Thus, the user interaction with the automated assistant can be improved by providing commands based on the set of subqueries only when the set is more likely to provide meaningful response content compared to the original text.
[0051] In some implementations, the quality score can be determined based on the quality scores of each subquery of the set 182. For example, the set 182 can be provided to the scorer 122, and then the scorer 122 can determine the quality score for each of the subqueries 182 and 183. These quality scores can be based on, for example, the likelihood that each subquery is a valid input to the agent and / or based on whether the subquery is a good candidate for representing the text 181. In some implementations, the scorer 122 can access one or more databases (not shown) that include, for example, multiple mappings between grammar and response actions (or more generally, intents), visual cues and response actions, and / or touch inputs and response actions. For example, the grammar included in the mappings can be selected and / or learned over time and can represent common intents of the user. For example, one grammar "turnoff the <room>lights(Off <room>The "lamp)" can be mapped to an intent that invokes a response action that causes <room>One or more of the lights are turned off. Another syntax, "[weather|forecast]today (today's [weather|forecast])", can match a user query such as "what's the weather today" and "what's the forecast for today?". In some implementations, the quality score of a subquery can be based on the frequency of the subquery that appears in one or more previous query logs. In some implementations, the quality score of a subquery can be determined based on the number of search results (and / or the quality of the search results) in response to the subquery. In some implementations, the quality score can be determined based on whether the subquery corresponds to a voice action operation. In some implementations, the quality score of a subquery can be based on one or more of multiple terms in the subquery that are not in the text, or multiple terms in the text that are not in any of the subqueries in the set. The quality score of the subquery can then be used to generate an overall set quality score for set 182.
[0052] If the quality score of set 182 meets the threshold, set 181 is provided to the agent engine 123. For example, the quality score determined by the scorer 122 can be in the range of 0 to 1.0, where 1.0 is used to indicate an ideal candidate for text 181 and / or an exact match to one or more patterns associated with the available agents. For example, if the set quality score exceeds a threshold quality, such as 0.8, the subqueries 182 and 183 can be provided to the agent engine 123. In some implementations, if the set quality score does not meet the threshold, text 181 can be provided to the agent engine 123 instead. In some implementations, multiple sets 182 can be generated, scored by the scorer 122, and only the set 182 with the quality score that best indicates quality can be provided to the agent engine 123. For example, a first set can be generated and the determined quality score can be 0.6, while a second set can be generated with the determined quality score of 0.8. Only the set with the higher quality score can be provided to the agent engine 123.
[0053] As an example, the set generator 121 generates a first set and a second set based on text. Then, each set can be scored by the scorer 122. In addition, the original text is scored by the scorer 122. For example, if it is determined that the quality score of the first set is 0.8, the quality score of the second set is 0.7, and the quality score of the text is 0.6, then the first set can be provided to the proxy engine 123 because the quality score of the first set is higher than both the quality score of the second set and the quality score of the text. Similarly, for example, a set (or text) may have to meet a basic threshold before being provided to the proxy engine 123. In the previous example, the basic threshold can be 0.75, and the quality score of the first set is the only quality score that meets the basic threshold. Thus, the first set is provided to the proxy engine 123.
[0054] The proxy engine 123 determines commands to provide to the proxy for each subquery of the set. The commands include one or more terms provided directly to the proxy that convey the necessary information from the subquery. For example, a user may provide a query of "What's the weather in Paris", which includes one or more terms that are unnecessary for providing to the proxy to receive valid response content. Instead, the proxy engine 123 can generate a "Paris" command and provide the command to the weather proxy. Additionally, for example, for a query of "What is the weather in Paris tomorrow", the proxy engine 123 can generate a "Paris, tomorrow" command to the weather proxy.
[0055] For each of the subqueries 183 and 184, the proxy engine 123 generates commands 185 and 186 to provide to the proxies 191 and 192. The "proxy" as used herein refers to one or more computing devices and / or software employed by an automated assistant. In some cases, the proxy can be separate from the automated assistant and / or can communicate with the automated assistant via one or more communication channels. In some of those cases, the automated assistant can transmit data (e.g., proxy commands) from a first network node to a second network node that implements all or aspects of the proxy's functionality. In some cases, the proxy can be a third-party (3P) proxy because it is managed by a party separate from the party that manages the automated assistant. In some other cases, the proxy can be a first-party (IP) proxy because it is managed by the same party that manages the automated assistant.
[0056] The agent is configured to receive call requests and / or other agent commands from the automated assistant (e.g., via a network and / or via an API). In response to receiving an agent command, the agent generates response content based on the agent command and sends the response content to provide a user interface output based on the response content. For example, the agent can transmit the response content to the automated assistant for the automated assistant to provide an output based on the response content. As another example, the agent itself can provide the output. For example, a user can interact with the automated assistant via a client device (e.g., the automated assistant can be implemented on the client device and / or in a network communicating with the client device). The agent can be an application installed on the client device or an executable application remote from the client device but "streamable" on the client device. When the application is invoked, it can be executed by the client device and / or brought to the forefront of the client device (e.g., its content can replace the display of the client device).
[0057] In some cases, in response to an invocation of a particular agent according to the techniques disclosed herein, the human-automated assistant conversation can be transferred to the particular agent at least temporarily (in fact or effectively). For example, an output based on the response content of the particular agent can be provided to the user to facilitate the conversation, and further user input can be received in response to the output. The further user input (or its transformation) can be provided to the particular agent. The particular agent can utilize its own semantic engine and / or other components when generating further response content, which can be used to generate further output to facilitate the conversation. This general process can continue until, for example, the particular agent provides response content that terminates the particular agent conversation (e.g., an answer or solution rather than a prompt), additional user interface input from the user terminates the particular agent conversation (e.g., instead of invoking a response from the automated assistant or other agent), and so on.
[0058] In some cases, when the conversation is effectively transferred to a particular agent, the automated assistant can still act as an intermediary. For example, in the case of acting as an intermediary for the user's natural language input as voice input, the automated assistant can convert the voice input to text, provide the text (and optional text annotations) to the particular agent, receive response content from the particular agent, and provide an output based on the particular response content for rendering to the user. Similarly, for example, when acting as an intermediary, the automated assistant can analyze the user input and / or response content of the particular agent to determine whether the conversation with the particular agent should be terminated, whether the user should be transferred to an alternative agent, whether global parameter values should be updated based on the particular agent conversation, etc. In some cases, the conversation can be actually transferred to the particular agent (once transferred, without the automated assistant acting as an intermediary), and when one or more conditions occur, such as when terminated by the particular agent (e.g., in response to an intent being completed via the particular agent), it can optionally be transferred back to the automated assistant.
[0059] As shown, the flowchart includes only two subqueries in set 182; however, set 182 can include any number of subqueries, and the agent engine 123 can generate commands for each of the multiple subqueries. In some implementations, the commands can be based on one or more keywords included in the corresponding subquery. For example, the set generator 121 can generate a first subquery 183 "What's the weather in Paris” and a second subquery 184 "Show me restaurants in Paris". Once the set has been scored and the quality score of 182 meets the threshold, the agent engine 123 identifies a first agent 191 that is a weather agent (i.e., returns meaningful response content based on the first subquery 183) and a second agent 192 that provides entertainment venues. The agent engine 123 generates a first command 185 "Paris” to provide to the first agent 191, and generates a second command 186 "Paris, restaurants” to provide to the second agent 192.
[0060] In some implementations, multiple commands from subqueries can be provided to the same agent. For example, the set generator 121 can generate a set 181 that includes the subqueries "What's the weather in London” and "What's the weather in Paris”. Both subqueries can produce commands that can both be provided to the same agent (e.g., "London” and "Paris” provided to the weather agent).
[0061] In response to a command being provided, agents 191 and 192 provide response content 172 and 173. The response content can include, for example, written text, images, audio, and / or other media that can be provided to the user. For example, a "weather" agent can receive a command that includes a location and can provide text describing the current weather at that location, an image showing the weather (e.g., an image of rain, temperature), and / or an audio description of the weather. Also, for example, an agent associated with the control of one or more lighting devices of a user can receive a command indicating a specific light and the user's intent to change the light settings (e.g., "livingroom, dim to 50%"), and can provide response content that includes an acknowledgement that the action has been successfully performed. In some implementations, the agent command can be a structured command that identifies an "intent" and values for one or more time slots, where the "intent" and values are generated based on a sub-query. For example, for a sub-query of "turn on couch light", the intent of "turn on" can be paired with the value of "couch light" (or some other associated identifier of the specific light), and a second time slot can include an identifier of the user and / or the client device so that the agent can determine who is making the request.
[0062] Based on receiving response content 172 and 173, the content can then be provided to the user's device to render the corresponding content to the user. Rendering the content can include, for example, providing a visual output to the user, providing a text output to the user, and / or providing an audio output to the user. The form of the rendered content can vary based on the hardware and / or software capabilities of the user device. For example, the rendered content can be provided to a smart phone, and the rendered content is provided as audio and images. Also, for example, the rendered content can be provided to an assistant device that only includes a speaker, and it is only rendered as audio. In some implementations, the content rendered by the client device may be different from response content 172 and 173. For example, the agent can provide the response content as text, and the text can be used to generate speech, which is then rendered by the client device.
[0063] In some implementations, the response content 172 and 173 can be rendered with additional terms from subqueries 183 and 184. For example, for the subquery "What's the weather in Paris", the command "Paris" can be provided to the weather agent. The agent can provide the response content "72, sunny, with a chance of rain later". However, when rendering the response content, the term "the weather in Paris" can be recognized from the subquery and further used to render the content, such as "The weather in Paris is 72, sunny, with a chance of rain later", so that the user can have the context of what query the response content is answering. Similarly, for example, for the query "Turn off the kitchen lights", the response content can only include an indication of successfully turning off the lights. Thus, to provide context, the response content can be rendered as "Kitchen lights are now off" based on the initial query (and / or the command provided to the agent).
[0064] In some implementations, the response content 172 and 173 are not immediately rendered by the client device, but can be used by the proxy engine to submit new commands to the proxy. Subsequently, the subsequent response content can then be rendered to the user. For example, referring to Figure 3 , an example process flow is provided that demonstrates generating response content for multiple sub - queries generated based on a compound query and adopting an example of the response content when generating content for rendering in response to the compound query. For the example shown, the text 305 "Divide the population of California by the population of Ohio" results in commands for "California" 310 and "Ohio" 315. These commands can be generated by the proxy engine 123 based on the sub - queries "What's the population of California" and "What's the population of Ohio", each sub - query being generated by the set generator 121. The commands can be provided to the assistant agent 320, which accepts the commands for the states and provides the population of the states. Then the response content is received from the proxy, which includes "California population" 325 and "Ohio population 330", each number representing the population of the respective state. Then a new aggregation command 335 can be generated, which is an aggregation of the response content 325 and 330. Then the aggregation command 335 is provided to the arithmetic agent 340, which accepts the two numbers and the operation. Subsequently, the subsequent response content 350 (a number) can be used to cause the client device to render the content.
[0065] Reference Figure 4 , a block diagram of an example environment that can generate response content in response to an utterance including a compound query is provided. The example environment includes a client computing device 106 that executes an instance of an auto - assistant client 107. One or more cloud - based auto - assistant components 160 can be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) that are communicatively coupled to the client device 106 via one or more local area networks and / or wide area networks (e.g., the Internet), typically indicated at 110.
[0066] An instance of the Auto-Assistant Client 107 can, through its interaction with one or more cloud-based Auto-Assistant components 160, form a logical instance of an Auto-Assistant 140 that, from the user's perspective, appears to be an Auto-Assistant with which the user can engage in a human-machine conversation. Thus, it should be understood that in some implementations, a user interacting with the Auto-Assistant Client 107 executing on the client device 106 is actually interacting with a logical instance of his / her own Auto-Assistant 140. For the sake of brevity, the term "Auto-Assistant" as used herein to refer to serving a particular user is generally meant to refer to the combination of the Auto-Assistant Client 107 executing on the client device 106 operated by the user and one or more cloud-based Auto-Assistant components 160 (which may be shared among multiple Auto-Assistant Clients of multiple client computing devices). It should also be understood that in some implementations, the Auto-Assistant 140 can respond to requests from any user, regardless of whether the user is actually "served" by the particular instance of the Auto-Assistant 140.
[0067] The client computing device 106 can be, for example: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance device such as a smart TV, and / or a wearable device of the user that includes a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices can be provided. In various implementations, the client computing device 106 can optionally operate one or more other applications in addition to the Auto-Assistant Client 107, such as, a messaging client (e.g., SMS, MMS, online chat), a browser, etc. In some of these various implementations, one or more of the other applications can optionally interface with the Auto-Assistant 107 (e.g., via an application programming interface), or include an instance of its own Auto-Assistant application (which can also interface with the cloud-based Auto-Assistant components 160).
[0068] The automated assistant 140 conducts a human-machine conversation session with the user via the user interface input and output devices of the client device 106. To protect user privacy and / or save resources, in many cases, the user typically has to explicitly invoke the automated assistant 140 before the automated assistant will fully process an uttered phrase. The explicit invocation of the automated assistant 140 can occur in response to a certain user interface input received at the client device 106. For example, the user interface input that can be used to invoke the automated assistant 140 via the client device 106 can optionally include the actuation of a hardware and / or virtual button of the client device 106. Additionally, the automated assistant client can include one or more local engines 108, such as an invocation engine operable to detect the presence of one or more uttered invocation phrases. The invocation engine can invoke the automated assistant 140 in response to detecting one of the uttered invocation phrases. For example, the invocation engine can invoke the automated assistant 140 in response to detecting an uttered invocation phrase such as "Hey Assistant", "OK Assistant", and / or "Assistant". The invocation engine can continuously process (e.g., if not in an "inactive" mode) a stream of audio data frames based on the output from one or more microphones of the client device 106 to monitor for the occurrence of an uttered invocation phrase. When monitoring for the occurrence of an uttered invocation phrase, the invocation engine discards (e.g., after temporarily storing in a buffer) any audio data frames that do not include an uttered invocation phrase. However, when the invocation engine detects the occurrence of an uttered invocation phrase in the processed audio data frames, the invocation engine can invoke the automated assistant 140. As used herein, "invoking" the automated assistant 140 can include activating one or more previously inactive functions of the automated assistant 140. For example, invoking the automated assistant 140 can include causing one or more local engines 108 and / or cloud-based automated assistant components 130 to further process the audio data frame based on which the invocation phrase was detected, and / or one or more subsequent audio data frames (however, no further processing of the audio data frames occurred prior to the invocation). For example, in response to the invocation of the automated assistant 140, local and / or cloud-based components can perform a process of streaming audio to text to generate a set of sub-queries.
[0069] One or more local engines 108 of the automatic assistant 140 are optional and may include, for example, the aforementioned invocation engine, a local speech-to-text ("STT") engine (which converts a user's single spoken utterance into text), a local text-to-speech ("TTS") engine (which converts text into speech), a local natural language processor (which determines the semantics of the audio and / or the text converted from the audio), and / or other local components. Since the client device 106 is relatively limited in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the local engine 108 may have limited functionality relative to any counterparts included in the cloud-based automatic assistant component 160.
[0070] The cloud-based automatic assistant component 160 utilizes the almost infinite resources of the cloud to perform more robust and / or more precise processing of audio data and / or other user interface inputs relative to any counterparts of the local engine 108. Similarly, in various implementations, the client device 106 may provide audio data and / or other data to the cloud-based automatic assistant component 160 in response to the invocation engine detecting the utterance of an invocation phrase or detecting some other explicit invocation of the automatic assistant 140.
[0071] The illustrated cloud-based automatic assistant component 160 includes a cloud-based TTS module 161, a cloud-based STT module 162, a natural language processor 163, a dialogue state tracker 164, and a dialogue manager 165. The illustrated cloud-based automatic assistant component 160 also includes a query engine 120 that generates a set of subqueries based on the audio data received from the client device 106. Additionally, the cloud-based automatic assistant component 160 includes a proxy engine 122 that may provide commands to one or more proxies based on the generated set of subqueries.
[0072] In some implementations, one or more of the engines and / or modules of the automatic assistant 140 may be omitted, combined, and / or implemented in a component separate from the automatic assistant 140. For example, in some implementations, the query engine 120 and / or the proxy engine 123 may be wholly or partially implemented on the client device 106. Additionally, in some implementations, the automatic assistant 140 may include additional and / or alternative engines and / or modules.
[0073] The cloud-based STT module 162 may convert the audio data into text, and then may provide the text to the natural language processor 163. The text may include, for example Figure 1 The text 181 and can be based on a single spoken utterance of the user. In various implementations, the cloud-based STT module 162 can convert the audio data into text at least in part based on the speaker label and the assignment indication provided by the assignment engine 124. The cloud-based TTS module 161 can convert the text data (e.g., the natural language response formulated by the automatic assistant 140 and / or one or more assistant agents 160) into a computer-generated speech output. In some implementations, the TTS module 161 can provide the computer-generated speech output to the client device 106 for direct output, for example, using one or more speakers. In other implementations, the text data (e.g., the natural language response) generated by the automatic assistant 140 can be provided to one of the local engines 108, and one of the local engines can then convert the text data into a computer-generated speech for local output.
[0074] The natural language processor 163 of the automatic assistant 140 processes the free-form natural language input and generates an annotated output for use by one or more other components of the automatic assistant 140. For example, the natural language processor 163 can process the free-form natural language input, which is a text input that is the conversion of the audio data provided by the user via the client device 106 by the STT module 162. The generated annotated output can include one or more annotations of the natural language input and optionally one or more (e.g., all) terms of the natural language input.
[0075] In some implementations, the natural language processor 163 is configured to identify and annotate various types of syntactic information in the natural language input. For example, the natural language processor 163 can include a part-of-speech tagger (not described) that is configured to annotate terms with the syntactic role of the terms. Moreover, for example, in some implementations, the natural language processor 163 can additionally and / or alternatively include a dependency parser (not depicted) that is configured to determine the syntactic relationships between the terms in the natural language input.
[0076] In some implementations, the natural language processor 163 may additionally and / or alternatively include an entity tagger (not depicted) configured to annotate entity references in one or more segments, such as references to persons (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and virtual), and the like. The entity tagger of the natural language processor 163 may annotate references to entities at a relatively high granularity level (e.g., such that all references to an entity category such as a person can be identified) and / or at a relatively low granularity level (e.g., such that all references to a specific entity such as a specific person can be identified). The entity tagger may rely on the content of the natural language input to resolve specific entities and / or may optionally communicate with a knowledge graph or other entity database to resolve specific entities. The identified entities may be utilized to identify patterns in the text as described herein.
[0077] In some implementations, the natural language processor 163 may additionally and / or alternatively include a coreference resolver (not depicted) configured to group or "cluster" references to the same entity based on one or more context clues. For example, the coreference resolver may be used to resolve the term "there" in the natural language input "I liked Hypothetical Café last time we ate there" to "Hypothetical Café".
[0078] In some implementations, one or more components of the natural language processor 163 may rely on annotations from one or more other components of the natural language processor 163. For example, in some implementations, the aforementioned entity tagger may rely on annotations from the coreference resolver and / or the dependency parser when annotating all mentions of a specific entity. Similarly, for example, in some implementations, the coreference resolver may rely on annotations from the dependency parser when clustering references to the same entity. In some implementations, when processing a specific natural language input, one or more components of the natural language processor 163 may use relevant previous inputs and / or other relevant data outside of the specific natural language input to determine one or more annotations.
[0079] In some implementations, the dialogue state tracker 164 can be configured to track a "dialogue state", which includes, for example, the belief state of one or more user goals (or "intentions") during a human-machine dialogue session and / or across multiple dialogue sessions. When determining the dialogue state, some dialogue state trackers may attempt to determine the most likely values of the slots instantiated in the dialogue based on the user and system utterances in the dialogue session. Some techniques utilize a fixed ontology that defines a set of slots and a set of values associated with those slots. Some techniques may alternatively or additionally be customized for individual slots and / or domains. For example, some techniques may require training a model for each slot type in each domain.
[0080] The dialogue manager 165 can be configured to map the current dialogue state provided, for example, by the dialogue state tracker 164 to one or more "response actions" among a plurality of candidate response actions that are then performed by the automated assistant 140. Depending on the current session state, the response actions may take various forms. For example, the initial and in-flow dialogue states corresponding to the rounds of the dialogue session that occurred before the last round (e.g., when performing the task desired by the end user) can be mapped to various response actions that include the automated assistant 140 outputting additional natural language dialogue. This response dialogue can include, for example, requesting that the user provide parameters (i.e., fill the slots) for certain actions that the dialogue state tracker 164 believes the user intends to perform. In some implementations, the response actions can include, such as "request" (e.g., looking for parameters for slot filling), "provide" (e.g., suggesting an action or course of action to the user), "select", "notify" (e.g., providing the requested information to the user), "mismatch" (e.g., informing the user that the user's last input is not understood), commands to peripheral devices (e.g., turning off a light bulb), and so on.
[0081] The client device 106 submits a request via one or more local area networks and / or wide area networks (e.g., the Internet), typically indicated at 110. The request can be a single spoken utterance of the user, as received via the microphone of the client device 106. The single spoken utterance can be converted to text by the STT module 162, processed by the natural language processor 163 to identify the entity parts in the speech and / or text, and the text can be provided to the query engine 120. As previously described with respect to Figures 1 to 3 As described, the set generator 121 can generate a set of subqueries based on the transformed text and the tokenized terms and / or entities in the text. The set can be scored by a scorer 122, and if the resulting quality score meets a threshold, the query engine 120 can provide the set of subqueries to the proxy engine 123. The proxy engine 123 generates commands for each subquery in the set and provides the commands to one or more agents 190, which then provide response content in response to receiving the commands. The response content can then be utilized to provide content to the user via the client device 106. For example, the response content can be converted to speech via the TTS module 161 and provided to the user as a turn of a conversation by the dialogue manager 165, as described herein.
[0082] Reference Figure 5 , a flowchart illustrating an example method according to various implementations described herein. For convenience, reference is made to the operations of the system performing the operations Figure 5 of the flowchart. The system can include various components of various computer systems. Additionally, although the operations of the Figure 5 method are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.
[0083] At block 505, the system receives text generated in response to a single spoken utterance of a user. The text can be based on the speech-to-text conversion performed by the STT module 162 and also on the audio received from the client device 106. For example, the client device 106 can execute an assistant client 107 that can provide a single spoken utterance to the assistant 140. In some cases, the text can be a compound query. For example, the text "What's the weather in Paris and show me restaurants near me” includes a request for the weather in Paris and a request to provide restaurants near the user.
[0084] At block 510, the system generates a set of subqueries based on the text received at block 505. The set can include any number of subqueries based on the text, and the set collectively defines a candidate interpretation of the text. As described herein, the subqueries can be generated based on multiple techniques. For example, the subqueries can be generated based on identifying patterns in the text, appending prefixes and / or suffixes to the text, context expansion of the text, and / or rewriting the text as a set of subqueries that retain a similar meaning to the text. In some implementations, multiple sets can be generated for the text.
[0085] At block 515, the system determines a quality score for a set of subqueries. The quality score can indicate, for example, the likelihood that the set correctly represents the intent of the individual queries included in the text. In some implementations, the quality score can indicate the similarity between one or more identified patterns and the subqueries. For example, each subquery can be scored, and a quality score for the set can be generated that is a function of the scores of the individual subqueries. In cases where multiple sets can be generated for the text, a score can be determined for each set.
[0086] At block 520, the system determines that the quality score for the set meets a threshold. Determining that the quality score meets a threshold can include determining that the quality score meets the threshold relative to the quality scores of other sets (thereby indicating that the subqueries of the given set are a more appropriate interpretation than the subqueries of other sets); determining that it meets a threshold relative to a quality score for the composite query, which is optionally determined for the composite query without splitting it into subqueries (thereby indicating that splitting the composite query into subqueries is more appropriate than not splitting the composite query); and / or determining that it meets a fixed threshold (thereby indicating that the subqueries of the given set are a consistent interpretation).
[0087] At block 525, in response to determining that the quality score for the set meets the threshold, the system provides commands to one or more corresponding assistant agents. For each subquery, the system generates a corresponding command based on the corresponding subquery and provides the corresponding command to the corresponding assistant agent. For example, for a subquery of "What's the weather in Paris", a structured command can be generated with an intent of "current weather" and a location value of "Paris", and this command can be provided to the assistant agent. The commands can be generated and provided to the assistant agents by a component that shares one or more characteristics with the agent engine 123.
[0088] At block 530, the system receives response content from one or more assistant agents in response to providing the commands to the agents. The response content can be, for example, text, images, audio, video, and / or other content that can be rendered to the user and / or that can be utilized to render content to the user. For example, a "Paris" command provided to a weather assistant may result in response content of "68, sunny", which can be provided as text, audio, and / or an image.
[0089] At block 535, the system utilizes the response content from the assistant agent to cause the user's client device 106 to render content based on the response content. Rendering the content can include providing audio to the user based on the response content, providing the response content as text, and / or providing the response content visually. In some implementations, as described herein, the response content can be utilized to generate an aggregated command provided to the assistant agent, and the response content from additional agents is utilized to render content to the client device 106.
[0090] Figure 6 is a block diagram of an example computing device 610 that can optionally be used to perform one or more aspects of the techniques described herein. For example, client device 106 can include one or more components of example computing device 610 and / or one or more server devices that implement cloud-based assistant component 110.
[0091] Computing device 610 generally includes at least one processor 614 that communicates with a plurality of peripheral devices via bus subsystem 612. These peripheral devices can include storage subsystem 624, including, for example, memory subsystem 625 and file storage subsystem 626; user interface output device 620; user interface input device 622; and network interface subsystem 616. The input and output devices allow a user to interact with computing device 610. Network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0092] User interface input device 622 can include a keyboard; a pointing device such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touchscreen incorporated into the display; an audio input device such as a speech recognition system; a microphone; and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into computing device 610 or a communication network.
[0093] User interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or other mechanisms for creating visual images. The display subsystem can also provide a non-visual display such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways to output information from computing device 610 to a user or another machine or computing device.
[0094] Storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, storage subsystem 624 can include executing Figures 1 - 3 Example process flows, Figure 5 selected aspects of the methods depicted in Figures 1 - 5 and the logic of the various components depicted in
[0095] These software modules are typically executed by the processor 614, either alone or in combination with other processors. The memory 625 used in the storage subsystem 624 can include multiple memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 for storing fixed instructions. The file storage subsystem 626 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical disk drive, or a removable media cartridge. Modules implementing the functionality of certain implementations can be stored by the file storage subsystem 626 in the storage subsystem 624 or in other machines accessible by the processor 614.
[0096] The bus subsystem 612 provides a mechanism for enabling the various components and subsystems of the computing device 610 to communicate with each other as expected. Although the bus subsystem 612 is schematically shown as a single bus, alternative implementations of the bus subsystem can use multiple buses.
[0097] The computing device 610 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 6 the description of the computing device 610 depicted in Figure 6 is only intended as a specific example for the purpose of illustrating some implementations. Many other configurations of the computing device 610 may have more or fewer components than the computing device depicted in
[0098] In the case where certain implementations discussed herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, and the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use a user's personal information only after receiving explicit authorization to do so from the relevant user.
[0099] For example, provide the user with control over whether a program or feature collects user information about that particular user or other users related to the program or feature. Present one or more options to each user for whom personal information is to be collected to allow control over the information collection related to that user, to provide permission or authorization regarding whether information is to be collected and which parts of the information are to be collected. For example, one or more such control options can be provided to the user via a communication network. Additionally, before storing or using certain data, certain data may be processed in one or more ways such that personally identifiable information is removed. As an example, a user's identity may be processed such that the personally identifiable information cannot be determined. As another example, a user's geographical location may be generalized to a larger area such that the user's specific location cannot be determined.< / room> < / room> < / room> < / city> < / city> < / city> < / city> < / city> < / city> < / city> < / room> < / city> < / city> < / city> < / city> < / city> < / city> < / city> < / date> < / date> < / city> < / city> < / city> < / city> < / city> < / city> < / distance>
Claims
1. A method implemented by one or more processors, the method comprising: Receiving, at an assistant interface of a user's client device, text generated in response to detecting a single spoken utterance of the user, the text corresponding to the single spoken utterance and generated based on a speech-to-text conversion of the single spoken utterance; Generating a set of subqueries based on the text, the set of subqueries including at least a first subquery and a second subquery, wherein the set of subqueries jointly define a candidate interpretation of the text, and wherein generating the set of subqueries includes: Identifying aliases of one or more entities included in the text; Identifying a first pattern based on the text, wherein the first pattern includes one or more terms and entity types; Generating a first part based on the first pattern; and Generating a first subquery by replacing the entity type in the first part with a corresponding alias in the aliases; Determining a quality score for the set of subqueries, wherein the quality score of the set indicates the likelihood that the subqueries of the set represent the intent of the text; Determining that the quality score of the set of subqueries meets at least one threshold; In response to determining that the quality score meets the threshold: Providing corresponding commands to corresponding assistant agents for each subquery of the set, wherein providing the corresponding commands includes providing at least: A first command based on the first subquery, and A second command based on the second subquery, the second command being a structured command to change the state of a connected device; Receiving corresponding response content in response to providing the corresponding commands, the corresponding response content including at least first response content in response to the first command; and Causing the client device to render content based on the corresponding response content to the user.
2. The method according to claim 1, wherein, The response content further includes second response content in response to the second command, and wherein causing the client device to render content to the user includes: Causing the client device to render first content based on the first response content; and Causing the client device to render second content based on the second response content.
3. The method according to claim 1, wherein, The second command is provided to a third-party assistant agent, and wherein the third-party assistant agent is a third-party agent controlled by a third party, the third party being different from the party controlling the assistant interface.
4. The method according to claim 1, further comprising: Generating a second set of second subqueries based on the text, wherein the second subqueries of the second set are unique compared to the subqueries of the set, and wherein the second subqueries of the second set jointly define an additional candidate interpretation of the text; and Determining an additional quality score for the second set of second subqueries, wherein determining that the quality score of the set of subqueries meets the threshold is based on a comparison of the quality score of the set with the additional quality score of the second set.
5. The method according to claim 1, wherein, Determining a quality score for the set includes: Determining a first quality score for the first subquery; Determining a second quality score for the second subquery; and Determining the quality score for the set based on the first quality score and the second quality score.
6. The method according to claim 5, wherein, Determining the first quality score for the first subquery includes determining whether the first subquery conforms to one or more identified commands of any assistant agent, and wherein determining the second quality score for the second subquery includes determining whether the second subquery conforms to one or more identified commands of any assistant agent.
7. The method according to claim 1, wherein, The set includes at least three subqueries.
8. The method according to claim 1, wherein, Generating the set includes: Generating a first portion of the text and a second portion of the text, wherein the first portion is unique compared to the second portion; and Generating the first subquery based on the first portion and generating the second subquery based on the second portion and based on one or more terms in the text that are not included in the second portion of the text.
9. The method according to claim 8, wherein, Generating the first subquery is further based on one or more terms in the text that are not included in the first portion of the text.
10. The method according to claim 8, wherein, One or more terms that are not included in the second portion of the text and are used in generating the second subquery include at least one term from the first portion of the text.
11. The method according to claim 8, wherein, One or more terms that are not included in the second portion of the text and are used in generating the second subquery include at least one term that occurs before the first portion of the text.
12. The method according to claim 8, wherein, One or more terms that are not included in the second portion of the text and are used in generating the second subquery include at least one term that occurs after the second portion of the text.
13. The method according to claim 1, wherein, Generating the set further includes: Identifying a second pattern based on the text, wherein the second pattern includes one or more second terms and the entity type or an additional entity type; Generating a second portion based on the second pattern; and Replacing the entity type or the additional entity type in the second portion with the corresponding alias in the aliases to generate the second subquery.
14. The method according to claim 13, wherein, The first pattern and the second pattern are unique, and wherein the alias in the first subquery is the same as the alias in the second subquery.
15. The method according to claim 13, wherein, The first pattern is the same as the second pattern, and wherein the alias in the first subquery is different from the alias in the second subquery.
16. The method according to claim 1, further including: Generating an aggregate query based on the response content; Provide an aggregation command to one or more additional assistant agents based on the aggregation query; And In response to providing the aggregation command, receive aggregation response content from the one or more additional assistant agents, wherein the rendered content is based on the aggregation response content.
17. The method according to any one of claims 1-16, further comprising: Determine a text quality score for the text, wherein determining that the quality score of the set of subqueries meets the threshold is based on a comparison of the quality score with the text quality score; and In response to the quality score of the set of subqueries meeting the threshold: Generate the command based on the set of subqueries, rather than based on the text.
18. A method implemented by one or more processors, the method comprising: Receive text generated in response to user input at an interface of a user's client device; Generate a set of subqueries based on the text, the set of subqueries including at least a first subquery and a second subquery, wherein the set of subqueries jointly define a candidate interpretation of the text, wherein generating the set of subqueries includes: Identify aliases of one or more entities included in the text; Identify a pattern based on the text, wherein the pattern includes one or more terms and entity types; Generate a first part based on the pattern; and Generate a first subquery by replacing the entity type in the first part with a corresponding alias in the aliases; Determine a set quality score for the set of subqueries, wherein the quality score of the set indicates the likelihood that the set of subqueries represents the intent of the text; Determine a text quality score for the text, wherein the text quality score indicates the likelihood that the text is not a compound query; Determine that the quality score of the set of subqueries meets a threshold relative to the text quality score of the text; In response to determining that the quality score meets the threshold: Provide a corresponding command for each subquery of the set to a corresponding agent in a set of assistant agents, the corresponding command including at least a first command based on the first subquery and a second command based on the second subquery; Receive corresponding response content in response to providing the corresponding command, the corresponding response content including at least a first response content in response to the first command and a second response content in response to the second command; and Cause the client device to render content based on the corresponding response content to the user.
19. The method according to claim 18, wherein, Causing the client device to render content to the user includes: Causing the client device to render a first content based on the first response content; and Causing the client device to render a second content based on the second response content.
20. The method according to claim 18, wherein, The first command is provided to a first agent, and the second command is provided to the first agent, and wherein the first agent is a search agent.
21. The method according to any one of claims 19-20, wherein, Generating the set includes: Generating a first part of the text and a second part of the text, where the first part is unique compared to the second part; and Generating the first subquery based on the first part and generating the second subquery based on the second part and based on one or more terms in the text that are not included in the second part of the text.
22. The method according to any one of claims 19 to 20, wherein, Generating the set further includes: Identifying a second pattern based on the text, where the second pattern includes one or more second terms and an entity type or additional entity types; Generating a second part based on the second pattern; Generating a first subquery by replacing the entity type in the first part with a corresponding alias; and Generating a second subquery by replacing the entity type or additional entity types in the second part with a corresponding alias.
23. A method implemented by one or more processors, the method includes: Receiving, at an interface of a user's client device, text generated in response to a user input; Generating a set of subqueries based on the text, the subqueries of the set including at least a first subquery and a second subquery, where the subqueries of the set jointly define a candidate interpretation of the text, and where generating the set of subqueries includes: Identifying aliases for one or more entities included in the text; Identifying a pattern based on the text, where the pattern includes one or more terms and an entity type; Generating a first part based on the pattern; and Generating a first subquery by replacing the entity type in the first part with the corresponding alias in the aliases; Determining a set quality score for the set of subqueries, where the quality score of the set indicates the likelihood that the subqueries of the set represent the intent of the text; Determining that the quality score of the set of subqueries meets at least one threshold; In response to determining that the quality score meets the threshold: For each subquery of the set, providing a corresponding command to a corresponding agent in a set of assistant agents, the corresponding command including at least a first command based on the first subquery and a second command based on the second subquery; Receiving corresponding response content in response to providing the corresponding command, the corresponding response content including at least a first response content in response to the first command and a second response content in response to the second command; and Causing the client device to render a first content based on the first response content to the user and render a second content based on the second response content to the user.
24. The method according to claim 23, further includes: Determining that the text meets one or more conditions; where generating the set of subqueries is only performed after determining that the text meets one or more conditions.
25. A non-transitory computer-readable storage medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 24.
26. A system comprising one or more processors for performing the method according to any one of claims 1 to 24.
Citation Information
Patent Citations
Multi-command single utterance input method
CN106471570A