Selectively invoking an automated assistant based on detected environmental conditions in the case of voice-based calls that do not require the automated assistant
Through the machine learning model, the automatic assistant directly responds to user commands when meeting specific conditions, solving the problem of explicit calls prolonged interaction and resource waste, achieving more efficient interaction.
Patent Information
- Application Number
- CN202080093495.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-01-17
AI Technical Summary
In the prior art, the process of users requiring explicit call to the automatic assistant leads to interaction extension and waste of resources, especially in environments where the user interacts frequently with the automatic assistant.
The machine learning model is used to process environmental status data, bypass explicit invocation phrases, learn user intentions through training data, and the automatic assistant responds directly to user commands when meeting specific environmental conditions.
It reduces the time and consumption of computing resources for users to interact with automatic assistants, improves interaction efficiency, and saves power and processing bandwidth.
Smart Images

Figure CN114981772B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Humans can use an interactive software application, referred to herein as an "automatic assistant" (also known as a "digital agent", "chatbot", "interactive personal assistant", "intelligent personal assistant", "conversation agent", etc.), to engage in a human-machine conversation. For example, a human (who may be referred to as a "user" when interacting with the automatic assistant) can provide commands and / or requests using spoken natural language input (i.e., utterances) that can be converted into text and then processed in some cases and / or by providing text (e.g., typed) natural language input.
[0002] In some instances, the responsiveness of an automatic assistant can be limited to scenarios in which the user explicitly invokes the automatic assistant. For example, the user must often explicitly invoke the automatic assistant before the automatic assistant will fully process a spoken utterance. Some user interface inputs that can be used to invoke the automatic assistant via a client device can include a hardware button and / or a virtual button (e.g., a tap on the hardware button, a selection of a graphical interface element displayed by the client device) for invoking the automatic assistant at the client device. Many automatic assistants can be additionally or alternatively invoked in response to one or more specific spoken invocation phrases, which are also referred to as "hot words / phrases" or "trigger words / phrases" (e.g., an invocation phrase such as "Hey, assistant"). As a result of the explicit invocation, the user typically spends time invoking their automatic assistant before instructing their automatic assistant to assist with a particular task. This can cause the interaction between the user and the automatic assistant to be unnecessarily prolonged and can result in a corresponding prolonged use of various computing and / or network resources. SUMMARY OF THE INVENTION
[0003] The implementations described herein relate to the training and / or implementation of one or more machine learning models that can be used to at least selectively bypass an explicit invocation of an assistant, which otherwise may be required before invoking the assistant to perform various tasks. In other words, the output generated using the machine learning model can be used to determine when the assistant should respond to a spoken utterance and when the spoken utterance does not precede an explicit invocation of the assistant. To determine whether to invoke the assistant based on environmental conditions, a trained machine learning model can be employed when processing various different signals to generate an output indicating whether an explicit invocation of the assistant should be bypassed. For example, the trained machine learning model can be used to process data representing an environment in which a user can interact with the assistant. In some implementations, a signal vector can be generated to represent the operating states of various different devices within the environment. These operating states can indicate the user's intent to invoke the assistant and can thus effectively replace a spoken invocation phrase. In other words, when the user is in a particular environment where the user would typically require the assistant to perform a specific action, the trained machine learning model can be used to process context data representing the environment, the user, the time of day, the location, and / or any other characteristics associated with the environment and / or the user. Processing of the context data can produce an output (e.g., a probability) indicating whether the user will request the performance of an assistant action. This probability can be used to cause the assistant to require or bypass requiring the user to provide an invocation phrase (or other explicit invocation) before responding to an assistant command.
[0004] As a result of invoking the assistant without initially detecting an invocation phrase (or other explicit invocation), various computational and power resources can be saved. For example, a computing device that requires an explicit spoken invocation phrase before each assistant command can consume more resources than another computing device that does not require an explicit spoken invocation phrase before each assistant command. Resources such as power and processing bandwidth can be saved when the computing device no longer continuously monitors for an invocation but instead processes the context signals that are already available. Further resources such as processing bandwidth and client device power resources can be saved when the interaction between the user and the assistant is shortened because it is no longer necessary for the user to provide an invocation phrase before most assistant commands. For example, the interaction between the user and the client device that includes the assistant can be shorter in duration because the user at least selectively does not need to use a spoken invocation phrase or other explicit invocation input as a start to an assistant command.
[0005] Examples of training data for training a machine learning model can be based on interactions between one or more users and one or more automated assistants. For example, in at least one interaction, a user can provide a call phrase (e.g., "Hey, Assistant...") followed by an assistant command (e.g., "Protect my alarm system."), and another call phrase (e.g., "Also... Hey, Assistant, play some music.") followed by another assistant command. The two call phrases and two assistant commands may have been provided within a threshold time period (e.g., 1 minute) in a particular environment, indicating the likelihood that the user may issue those assistant commands again at a subsequent time point and within the threshold time period in the same environment. In some implementations, an example of training data generated from this scenario can characterize one or more features of a particular environment as having a positive or negative correlation with the call phrase and / or the assistant command. For example, an example of training data can include a training instance input corresponding to a feature of a particular environment and a training instance output of a "1" or other positive value indicating that the explicit call to the assistant should be bypassed.
[0006] In some implementations, the nature of one or more computing devices associated with an environment in which a user interacts with an automated assistant can be a basis for bypassing invocation phrase detection. For example, instances of training data can be based on scenarios in which a user interacts with their automated assistant while located in their home kitchen. The kitchen can include one or more smart devices such as a refrigerator, an oven, and / or a tablet device that can be controlled via the automated assistant. When a user provides an invocation phrase followed by an assistant command, one or more properties and / or states of the one or more smart devices can be identifiable. These properties and / or operational states can be used as a basis for generating instances of training data from which they are derived. For example, instances of training data can include training instance inputs that reflect those properties and / or operational states, and training instance outputs of a "1" or other positive value indicating that the explicit invocation of the assistant should be bypassed. For example, when a user provides a first invocation phrase and a first assistant command such as "Assistant, preheat the oven to 350 degrees", a tablet device in the kitchen can be operating in a low power mode. An instance of training data generated from this scenario can be based on the tablet device being in a low power mode and the oven being initially off when the user provides an assistant command in the kitchen to preheat the oven. In other words, instances of training data can provide a positive correlation between device states (e.g., tablet device state and oven device state) and assistant commands (e.g., "preheat the oven"). Thereafter, a machine learning model trained using instances of training data can be used to determine whether to bypass the requirement for an invocation phrase (or other explicit input) from the user to invoke the automated assistant. For example, when a similar scenario occurs in the kitchen or another similar environment in which a user can interact with their automated assistant, the automated assistant can be subsequently invoked based on the trained machine learning model. For example, the trained machine learning model can be used to process scenario features to generate a prediction output, and if the prediction output meets a threshold (e.g., a threshold greater than 0.7 or other value, where the prediction output is a probability), the requirement for an explicit input can be bypassed.
[0007] In some implementations, another instance of training data can be generated based on another scenario where the tablet device is playing music and the oven is operating at 350 degrees Fahrenheit. For example, another instance of training data can provide a correlation between one or more features of the environment, the operating states of various devices, non-invoked actions from one or more users, signals from one or more sensors (e.g., proximity sensors), and / or the user not providing a subsequent assistant command within a threshold time period. For example, the user can provide an invocation phrase and an assistant command such as "Assistant, turn off the oven". Subsequently, and within a specific threshold time period, the user can refrain from providing another invocation phrase and another assistant command. Thus, an instance of training data can be generated based on the tablet device playing music, the automatic assistant being instructed to turn off the oven, and the user not issuing a subsequent invocation phrase or subsequent assistant command - at least not within the threshold time period. For example, an instance of training data can include a training instance input corresponding to the features of a specific environment, and a training instance output indicating a "0" or other negative value that indicates that an explicit invocation of the assistant should be bypassed. This instance of training data can be used to train one or more machine learning models to determine whether to bypass detecting invocation phrases from one or more users in certain scenarios and / or environments.
[0008] In some implementations, data from various different devices including various different sensors and / or technologies for identifying features of the environment can be used to generate various instances of training data. As an example, with the prior permission of the user, training data can be generated from visual data that characterizes the user's proximity, posture, gaze, and / or any other visual features in the environment before, during, and after the user provides an assistant command. Such features can indicate, individually or in combination, whether the user is interested in invoking the automatic assistant. Additionally, the automatic assistant can employ a machine learning model trained using such training data when determining whether to invoke the automatic assistant based on certain environmental features (e.g., features exhibited by one or more users, computing devices, and / or any other features of the environment).
[0009] The above description is provided as an overview of some implementations of the present disclosure. Further descriptions of those implementations and other implementations are described in more detail below.
[0010] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform one or more of the methods such as those described above and / or elsewhere herein. Yet other implementations may include a system of one or more computers including one or more processors operable to run stored instructions to perform one or more of the methods such as those described above and / or elsewhere herein.
[0011] It should be recognized that all combinations of the foregoing concepts and additional concepts described in greater detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1A and Figure 1B Views illustrating instances from which training data is generated for training one or more machine learning models to bypass call phrase detection by an automated assistant.
[0013] Figure 2A and Figure 2B Views illustrating generation of training data based on a user providing an oral utterance to an automated assistant and then subsequently leaving the computing device that received the oral utterance.
[0014] Figure 3A 、 Figure 3B and Figure 3C Views illustrating scenarios in which the automated assistant employs a trained machine learning model to determine when to be invoked to receive an assistant command - instead of requiring an oral call phrase before being invoked.
[0015] Figure 4 Views illustrating a system for providing an automated assistant capable of selectively determining whether to be invoked based on context signals instead of requiring an initially detected oral call phrase.
[0016] Figure 5 Views illustrating a method for selectively bypassing call phrase detection by an automated assistant based on context signals.
[0017] Figure 6 is a block diagram of an example computer system. DETAILED DESCRIPTION
[0018] Figure 1A and Figure 1BViews 100 and 120 respectively illustrate examples from which training data is generated for training one or more machine learning models to bypass scenarios that require call phrase detection by an automated assistant. The machine learning models can be used to determine whether to bypass call phrase detection at the automated assistant. For example, call phrases are often used to invoke the automated assistant and can include one or more words or phrases, such as "Okay, Assistant". In response to the user 110 providing a call phrase, the computing device 106 that provides access to the automated assistant can provide an indication that the user 110 has invoked the automated assistant. Thereafter, the user 110 can be provided with an opportunity to speak or otherwise enter one or more assistant commands to the automated assistant to cause the automated assistant to perform one or more operations.
[0019] In some implementations, and instead of the automated assistant and / or the computing device 106 requiring the user to provide a call phrase, the computing device 106 and / or the automated assistant can process context data representing one or more characteristics of the environment 112 in which the user 110 is located. The context data can be generated based on data from a variety of different sources and can be processed using one or more trained machine learning models. In some implementations, the context data can be generated independently of whether the user has provided a call phrase and / or an assistant command to the automated assistant. In other words, regardless of whether the user has provided a call phrase within a particular environment, the context data can be processed with the prior permission from the user to determine whether the user is required to provide an explicit spoken call phrase before responding to an input from the user. When the context data is processed and indicates a scenario in which the user 110 can otherwise provide a call phrase, the automated assistant can be invoked and wait for further commands from the user without requiring the user to explicitly say the call phrase. This can conserve computing resources that might otherwise be continuously consumed to determine whether the user is providing a call phrase.
[0020] In some implementations, examples of training data for the machine learning models can be based on the interactions between one or more users and one or more automated assistants. For example, a particular user 110 can provide a spoken utterance 102 such as "Assistant, what's the weather like tomorrow?". The user 110 can provide the spoken utterance 102 when located in an environment 112 with a computing device 106. In some implementations, and with the prior permission from the user 110, the computing device 106 can determine one or more characteristics of the environment, such as but not limited to the posture of the user 110, the proximity of the user 110 to the computing device 106 and / or another computing device, the amount of noise in the environment 112 relative to the spoken utterance 102, the presence of one or more other people in the environment 112, the absence of a particular user within the environment 112, the facial expression of the user 110, the trajectory of the user 110, and / or any other characteristic of the environment 112.
[0021] Figure 1B The illustrated user 110 moves from a first user location 124 to a second user location 126 and subsequently provides another spoken utterance 122 in view 120. Figure 1B The scenario provided in can be relative in time to Figure 1A the scenario illustrated in. The feature of this scenario where the user provides two spoken utterances at two different locations can be characterized by the context data generated by computing device 106. Additionally, the context data can characterize the times at which user 110 provides a first invocation phrase, a first assistant command, a second invocation phrase, and a second assistant command (e.g., “… turn on the thermostat.”). In some implementations, the context data can lack any data characterizing invocation phrases from one or more users and / or assistant commands from one or more users. An example of training data can be processed from the context data to generate a positive correlation between consecutive assistant commands provided by user 110 around the time the user moves from the first user location 124 to the second user location 126. Additionally, or alternatively, when user 110 is in environment 112 and moves closer to computing device 106 after providing a first assistant command, the context data can characterize the first assistant command (e.g., “What's the weather like tomorrow?”) as having a positive correlation with the second assistant command (e.g., “… turn on the thermostat.”).
[0022] In some implementations, the trained machine learning model can be a neural network model, and examples of the training data can include input data characterizing Figure 1A and Figure 1B one or more features of the environment and / or scenario characterized in. Examples of the training data can also include output data characterizing user input and / or gestures made by user 110 within the environment and / or within the scenario characterized by Figure 1A and Figure 1B In this way, when the automated assistant employs the trained machine learning model, the trained machine learning model can be used to process context data representing individual scenarios in order to determine whether to bypass the need for an invocation phrase before activating the automated assistant, or alternatively, the need for an invocation phrase before activating the automated assistant. In some implementations, the context data can be based on a separate environment corresponding to the same geographical location as environment 112, except that the separate environment has different features and / or environmental conditions (e.g., user 110 is in a different location, the device is showing a different state, etc.).
[0023] Figure 2A and Figure 2BThe figures illustrate views 200 and 220 of generating training data based on user 210 providing spoken utterance 202 to an automated assistant and then subsequently leaving the computing device 206 that received the spoken utterance 202. The characteristics of this scenario can be characterized by context data and processed to generate training data that can be used to train one or more machine learning models to determine whether to bypass call phrase detection for the automated assistant. In other words, the scenario characterized by Figure 2A and Figure 2B provides an instance where there is a negative correlation between the characteristics of environment 212 and the user 210 providing successive assistant commands. The training data generated based on this scenario can be used to train a machine learning model that can allow the automated assistant to more easily determine when to continue detecting call phrases from the user rather than bypassing call phrase detection.
[0024] Figure 2A Figure 200 illustrates user 210 providing spoken utterance 202 such as "Assistant, set the house alarm to 'Stay'". Computing device 206 can receive the spoken utterance 202 and generate audio data that can be processed to determine whether user 210 provided a call phrase and / or an assistant command. Additionally, computing device 206 and / or any other computing device located in environment 212 can be used to generate context data characterizing the characteristics of environment 212 in which user 210 provided the spoken utterance 202. For example, the context data can characterize the time of day the user provided the spoken utterance 202, the utterance of user 210 with prior permission from the user, the proximity of user 210 to computing device 206, the gaze of user 210 relative to computing device 206, and / or any other objects within environment 212, and / or any other characteristics of the environment.
[0025] After providing the spoken utterance 202, user 210 can relocate from a first location 224 to a second location 226. As Figure 2A shown, user 210 can provide the spoken utterance 202 while in the first location 224 and then move to the second location 226 where user 210 can choose not to provide further input within a threshold time period (as indicated by state 222). Computing device 206 can determine that user 210 does not provide further input within the threshold time period and generate training data characterizing this scenario. In particular, the training data can be based on characterizing the Figure 2A and Figure 2BThe context data that captures the characteristics of the environment during this time period. The context data can characterize the spoken utterance 202 from the user 210, the first location 224 of the user 210, the second location 226 of the user 210, the timestamps corresponding to each location, the gaze of the user 210 at each location, and / or any other characteristics of the scenario where the user 210 provides the spoken utterance 202 and then does not provide further input (or at least does not provide further input until, for example, returning to the environment 212) in the same environment 212.
[0026] The training data can include training inputs related to the training output. The training input can be, for example, a signal vector based on the context data, and the training output can indicate that the user 210 does not provide further input in the scenario characterized by the context data (for example, the training output can be "0" or other negative values). Based on this training data and the training data Figure 1A and Figure 1B The machine learning model trained in association with can be used to determine the probability that the user will provide a call phrase in certain situations. The automatic assistant and / or the computing device can thus use the trained machine learning model to respond to the context signal without necessarily requiring a call phrase from the user.
[0027] Figure 3A 、 Figure 3B and Figure 3C Illustrate scenarios where the automatic assistant employs the trained machine learning model to determine when to be invoked and detect assistant commands instead of requiring a spoken call phrase before being invoked. The automatic assistant and / or the computing device 306 can use one or more trained machine learning models to process data based on the input from one or more sensors connected to and / or otherwise communicating with the computing device 306. Based on the processing of the data, the computing device 306 and / or the automatic assistant can make a determination as to whether to be invoked—without necessarily requiring the user 310 to explicitly say a call phrase.
[0028] For example, Figure 3AFIG. 300 shows a view where user 310 is sitting in environment 312 and listening to output 302 from computing device 306, which can be an assistant-enabled device. User 310 can listen to output 302 while sitting on their couch and facing the camera of computing device 306. Computing device 306 can process the context data characterizing the features of environment 312 before, during, and / or after user 310 is listening to output 302 from computing device 306, with prior permission from user 310. For example, the context data can characterize the location of user 310 and / or changes in the gaze of user 310. User 310 can be in a first position 330 when user 310 provides a spoken utterance to request playback of "ambient natural sounds," and then reposition themselves to a second position 332 after providing the spoken utterance. Based on processing this context data, the automatic assistant can determine the probability that user 310 will request the automatic assistant to perform an additional action. When the probability indicates that user 310 is more likely to provide an additional request than not to provide an additional request, the automatic assistant can bypass the need for user 310 to provide a follow-up invocation phrase (e.g., "Hey, Assistant...") before submitting a subsequent request. In other words, the context data can replace the invocation phrase, at least with prior permission from user 310.
[0029] Figure 3B FIG. 320 shows a view where user 310 has moved to a second position 332 and is using one or more trained machine learning models to process the context data. Based on the processing of the context data characterizing this scenario, the automatic assistant can be invoked to detect and / or receive one or more subsequent assistant commands without the need for a spoken invocation phrase from user 310. In some implementations, the automatic assistant can cause the display panel 324 of computing device 306 to render an interface 328 that is predicted to be useful for any subsequent assistant commands from user 310 in the current context. Additionally or alternatively, the automatic assistant can cause the display panel 324 to provide an indication (e.g., a graphical symbol) that the automatic assistant is operating in a mode where an invocation phrase is not temporarily required before responding to an assistant command from user 310.
[0030] In some implementations, when the automatic assistant is invoked and is waiting for another assistant command from user 310, interface 328 can be rendered to include control element 326 for controlling the thermostat and response output 322 from the automatic assistant. The response output 322 can include natural language content generated based on processing of context data using a trained machine learning model. For example, the natural language content of the response output 322 can characterize an inquiry associated with the predicted assistant command. For example, the predicted assistant command can be a user request to change the settings of the thermostat, and the predicted assistant command can be identified based on processing of context data using a trained machine learning model. Accordingly, the response output 322 can include natural language content such as "What temperature should I set the thermostat to?".
[0031] In some implementations, probabilities can be assigned to one or more actions based on processing of context data using a trained machine learning model. The action having the highest assigned probability relative to the other assigned probabilities for other actions can be identified as the action that user 310 will most likely request. Accordingly, the natural language content of the response output 322 can be generated based on predicting the highest probability action to be requested.
[0032] Regardless of the response output 322, user 310 can provide another spoken utterance 334 to control the automatic assistant without initially providing a spoken invocation phrase. For example, another spoken utterance 334 can be "Set the thermostat to seventy-two degrees", as Figure 3B shown in view 320. In some implementations, user 310 can provide a different spoken utterance that does not correspond to the predicted assistant command and does not include an invocation phrase. For example, due to processing of context data indicating the intention of user 310 to control the automatic assistant, user 310 can provide an assistant command such as "Turn off the lights in the kitchen". Although user 310 does not provide a spoken invocation phrase, the automatic assistant can respond to the spoken utterance 344 and / or the different spoken utterance. Instead, the context data characterizing Figure 3A and Figure 3B the scenario can replace the spoken invocation phrase and thus be used to indicate the intention of user to instruct the automatic assistant to perform one or more actions.
[0033] In response to receiving the other spoken utterance 334, the automatic assistant can perform one or more actions based on the other spoken utterance 334. For example, as Figure 3CAs shown in view 340, the automatic assistant can change the settings of the thermostat from the initial settings to the settings of 72 degrees as specified by user 310. In some implementations, confirmation of other spoken utterances 334 can be provided via a specific modality selected based on the processing of context data using a trained machine learning model. For example, probabilities can be assigned to one or more modalities and / or devices based on the processing of context data using a trained machine learning model. The modality and / or device with the highest probability relative to other modalities and / or devices can be selected to confirm other spoken utterances 334 from user 310. For example, the automatic assistant can cause interface 328 to render the updated settings of the thermostat and the natural language content of other spoken utterances 334 as additional graphical content 342 of interface 328. In this way, the interaction between user 310 and the automatic assistant can be pipelined to bypass the need for call phrases and / or explicit input from the user directly to the automatic assistant to indicate the intention to provide an assistant command.
[0034] In some implementations, when the automatic assistant is operating in a mode for bypassing the need for call phrases, the automatic assistant can rely on speech-to-text processing and / or natural language understanding processing. Such processing can be relied upon to determine, with prior permission from the user, whether the audio detected at one or more microphones embodies natural language content for the automatic assistant. For example, based on the automatic assistant entering a mode for bypassing the need for call phrases, the computing device with the assistant enabled can process audio data embodying a spoken utterance from the user such as "take out the trash". Using speech-to-text, the phrase "take out the trash" can be recognized and further processed to determine whether the phrase is actionable by the automatic assistant.
[0035] When the phrase is determined to be actionable, the automatic assistant can initiate the execution of one or more actions based on phrase recognition. However, when the phrase is determined to be non-actionable by the automatic assistant, the computing device can exit the mode and require a call phrase before responding to an assistant command from that particular user - however, depending on the context data, the automatic assistant can respond to one or more other users when the context data indicates a scenario where one or more other users are predicted to provide a call phrase to the automatic assistant. Alternatively, when the phrase is determined to be non-actionable by the automatic assistant, the computing device can continue to operate in that mode with prior permission from the user until one or more users are determined to have provided an assistant command that is actionable by the automatic assistant.
[0036] In some implementations, once the automatic assistant is operating in this mode, in order to no longer require a call phrase before responding to an assistant command, the automatic assistant can rely on other data to determine whether to remain in this mode or transition out of this mode. For example, and with the prior permission of the user, one or more sensors of one or more computing devices (e.g., proximity sensors and / or other image-based cameras) can be used to generate data characterizing the features of the environment. When one or more sensors provide data that is processed by a trained machine learning model and indicates that the user is not interested in calling the automatic assistant (e.g., the user leaves the room, as detected by a passive infrared sensor), the automatic assistant can transition out of this mode. However, when one or more sensors provide data that is processed by one or more trained machine learning models and indicates the user's intention to call the automatic assistant, the automatic assistant can remain in this mode.
[0037] In some implementations, the context data characterizing the features of one or more environments can be processed periodically to determine whether the automatic assistant should enter a mode where a call phrase is no longer required. For example, a computing device providing access to the automatic assistant can process the context data every T seconds and / or minutes. Based on this processing, the computing device can cause the automatic assistant to enter this mode or avoid entering this mode. Additionally or alternatively, the computing device can process sensor data from a first source (e.g., proximity sensor), and based on this processing, determine whether to process additional data to determine whether to enter this mode. For example, when the proximity sensor indicates that the user has entered a specific room, the computing device can employ a trained machine learning model to process additional context data. Based on the processing of the context data, the computing device can determine that the current context is one in which the user will call the automatic assistant and then enter a mode where a call phrase is temporarily no longer required. However, when the proximity sensor does not indicate that the user has entered a specific room, the computing device (e.g., a client device or a server device) can avoid using the machine learning model to further process the context data.
[0038] In some implementations, the computing device can provide an output perceivable by the user to indicate whether the computing device and / or the automatic assistant is operating in this mode. For example, the computing device or another computing device can be connected to an interface (e.g., a graphical interface, an audio interface, a haptic interface) that can provide an output in response to the automatic assistant transitioning into this mode and / or the automatic assistant transitioning out of this mode. In some implementations, the output can be ambient sound (e.g., natural sound), activation of light, vibration from a haptic feedback device, and / or any other type of output that can alert the user.
[0039] Figure 4The figure illustrates a system 400 for providing an automatic assistant that can selectively determine whether to invoke based on a context signal instead of a spoken invocation phrase that is required. The automatic assistant 404 can operate as part of an assistant application provided at one or more computing devices (such as computing device 402 and / or server device). A user can interact with the automatic assistant 404 via an assistant interface 420, which can be a microphone, a camera, a touch screen display, a user interface, and / or any other device that can provide an interface between the user and the application. For example, the user can initialize the automatic assistant 404 by providing oral, text, and / or graphical input to the assistant interface 420 to cause the automatic assistant 404 to perform functions (such as providing data, controlling a peripheral device, accessing an agent, generating input and / or output, etc.). Alternatively, the automatic assistant 404 can be initialized based on the processing of context data 436 using one or more trained machine learning models. The context data 436 can characterize one or more features of the environment in which the automatic assistant 404 is accessible, and / or one or more features of the user predicted to intend to interact with the automatic assistant 404. The computing device 402 can include a display device, which can be a display panel including a touch interface for receiving touch input and / or gestures to allow the user to control the application 434 of the computing device 402 via the touch interface. In some implementations, the computing device 402 can lack a display device, thereby providing an audible user interface output without providing a graphical user interface output. Additionally, the computing device 402 can provide a user interface for receiving spoken natural language input from the user, such as a microphone. In some implementations, the computing device 402 can include a touch interface and can lack a camera, but can optionally include one or more other sensors.
[0040] The computing device 402 and / or other third-party client devices can communicate with the server device via a network such as the Internet. Additionally, the computing device 402 and any other computing device can communicate with each other via a local area network (LAN) such as a Wi-Fi network. The computing device 402 can offload computing tasks to the server device to conserve computing resources at the computing device 402. For example, the server device can host the automatic assistant 404, and / or the computing device 402 can send input received at one or more assistant interfaces 420 to the server device. However, in some implementations, the automatic assistant 404 can be hosted at the computing device 402, and various processes associated with the operation of the automatic assistant can be performed at the computing device 402.
[0041] In various implementations, all or less than all aspects of the automatic assistant 404 can be implemented on the computing device 402. In some of those implementations, aspects of the automatic assistant 404 are implemented via the computing device 402 and can interface with a server device that can implement other aspects of the automatic assistant 404. The server device can optionally serve multiple users and their associated assistant applications via multiple threads. In implementations where all or less than all aspects of the automatic assistant 404 are implemented via the computing device 402, the automatic assistant 404 can be an application separate from (e.g., installed “on top of”) the operating system of the computing device 402—or alternatively can be directly implemented by the operating system of the computing device 402 (e.g., considered an application of the operating system but integrated with the operating system).
[0042] In some implementations, the automatic assistant 404 can include an input processing engine 406 that can employ multiple different modules to process the input and / or output of the computing device 402 and / or the server device. For example, the input processing engine 406 can include a speech processing engine 408 that can process audio data received at the assistant interface 420 to identify text embodied in the audio data. The audio data can be sent from, for example, the computing device 402 to the server device to conserve computing resources at the computing device 402. Additionally or alternatively, the audio data can be processed exclusively at the computing device 402.
[0043] The process for converting audio data to text can include a speech recognition algorithm that can employ a neural network and / or a statistical model for identifying groups of audio data corresponding to words or phrases. The text converted from the audio data can be parsed by a data parsing engine 410 and provided as text data to an automatic assistant 404, which can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, the output data provided by the data parsing engine 410 can be provided to a parameter engine 412 to determine whether the user has provided an input corresponding to a particular intent, action, and / or routine that can be performed by the automatic assistant 404 and / or an application or agent accessible via the automatic assistant 404. For example, assistant data 438 can be stored at the server device and / or the computing device 402 and can include data defining one or more actions that can be performed by the automatic assistant 404 and the parameters necessary to perform those actions. The parameter engine 412 can generate one or more parameters for the intent, action, and / or slot value and provide the one or more parameters to an output generation engine 414. The output generation engine 414 can use the one or more parameters to communicate with an assistant interface 420 for providing an output to the user and / or with one or more applications 434 for providing an output to one or more applications 434.
[0044] In some implementations, the automatic assistant 404 can be an application that can be installed “on top of” the operating system of the computing device 402 and / or that can itself form part (or all) of the operating system of the computing device 402. The automatic assistant application includes and / or can access on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition can be performed using an on-device speech recognition module that uses an end-to-end speech recognition machine learning model stored locally at the computing device 402 to process audio data (detected by a microphone). The on-device speech recognition generates recognition text for the spoken utterance (if any) present in the audio data. Additionally, for example, on-device natural language understanding (NLU) can be performed using an on-device NLU module that processes the recognition text generated using on-device speech recognition and optionally context data to generate NLU data.
[0045] NLU data can include an intent corresponding to an oral utterance and optionally parameters for the intent (e.g., slot values). On-device fulfillment can be performed using an on-device fulfillment module that utilizes the NLU data (from on-device NLU) and optionally other local data to determine actions to take to resolve the intent of the oral utterance (and optionally parameters for the intent). This can include determining a local response and / or a remote response (e.g., an answer) to the oral utterance, an interaction with a locally installed application to be performed based on the oral utterance, a command to be sent to an Internet of Things (IoT) device (directly or via a corresponding remote system) based on the oral utterance, and / or other resolution actions to be performed based on the oral utterance. The on-device fulfillment can then initiate local and / or remote execution / running of the determined actions to resolve the oral utterance.
[0046] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment can be utilized at least selectively. For example, the recognized text can be sent at least selectively to a remote assistant component for remote NLU and / or remote fulfillment. For example, the recognized text can be optionally sent for remote execution in parallel with on-device execution or in response to a failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device running can be prioritized at least because of the latency reduction they provide in resolving oral utterances (since there is no need for client-server round trips to resolve oral utterances). Additionally, on-device functionality can be the only functionality available in the absence of network connectivity or with limited network connectivity.
[0047] In some implementations, the computing device 402 can include one or more applications 434 that can be provided by a third-party entity different from the entity that provides the computing device 402 and / or the assistant 404. The assistant 404 and / or the application state engine 416 of the computing device 402 can access the application data 430 to determine one or more actions that can be performed by the one or more applications 434, as well as the state of each of the one or more applications 434 and / or the state of the corresponding device associated with the computing device 402. The assistant 404 and / or the device state engine 418 of the computing device 402 can access the device data 432 to determine one or more actions that can be performed by the computing device 402 and / or one or more devices associated with the computing device 402. Additionally, the application data 430 and / or any other data (e.g., the device data 432) can be accessed by the assistant 404 to generate the context data 436, which can characterize the context in which a particular application 434 and / or device is running and / or the context in which a particular user is accessing the computing device 402, accessing the application 434, and / or any other device or module.
[0048] When one or more applications 434 are running at computing device 402, device data 432 can characterize the current operating state of each application 434 running at computing device 402. Additionally, application data 430 can characterize one or more features of the running applications 434, such as the content of one or more graphical user interfaces being rendered under the direction of one or more applications 434. Alternatively or additionally, application data 430 can characterize an action mode that can be updated by the respective application and / or by the automatic assistant 404 based on the current operating state of the respective application. Alternatively or additionally, one or more action modes for one or more applications 434 can remain static, but can be accessed by the application state engine 416 to determine appropriate actions to initialize via the automatic assistant 404.
[0049] Computing device 402 can further include an assistant invocation engine 422 that can use one or more trained machine learning models to process application data 430, device data 432, context data 436, and / or any other data accessible to computing device 402. The assistant invocation engine 422 can process this data to determine whether to wait for the user to explicitly say a call phrase to invoke the automatic assistant 404, or to assume that the data indicates the intent of the user to invoke the automatic assistant — instead of requiring the user to explicitly say a call phrase. For example, one or more trained machine learning models can be trained using instances of training data based on scenarios where the user is in an environment in which multiple devices and / or applications are exhibiting various operating states. Instances of training data can be generated to capture training characterizing scenarios in which the user invokes the automatic assistant and other scenarios in which the user does not invoke the automatic assistant. When training one or more trained machine learning models according to these instances of training data, the assistant invocation engine 422 can cause the automatic assistant 404 to detect or bypass detection of a spoken call phrase from the user based on features of the scenario and / or environment. Additionally or alternatively, the assistant invocation engine 422 can cause the automatic assistant 404 to detect or bypass detection of one or more assistant commands from the user based on features of the scenario and / or environment.
[0050] In some implementations, the auto-assistant 404 can optionally include a training data engine 424 for generating training data based on interactions between the auto-assistant 404 and the user with the prior permission from the user. The training data can characterize instances where the auto-assistant 404 may have been initialized without being explicitly invoked via a spoken invocation phrase and where the user either provided an assistant command or did not provide an assistant command within a threshold time period thereafter. In some implementations, the training data can be shared with a remote server device that also receives data from various different computing devices associated with other users with the prior permission from the user. In this way, one or more trained machine learning models can be further trained so that each corresponding auto-assistant can employ the further trained machine learning model to better assist the user while also conserving computing resources.
[0051] Figure 5 FIG. illustrates a method 500 for selectively bypassing invocation phrase detection by an auto-assistant based on context signals. The method 500 can be performed by one or more computing devices, applications, and / or any other device or module capable of interacting with the auto-assistant. The method 500 can include an operation 502 of processing context data associated with the environment in which the user and the computing device are located. For example, the environment can be the user's home and the context data can be generated based on signals from one or more sensors, applications, and / or devices communicatively coupled to the computing device. The sensors can be, but are not limited to, proximity sensors, infrared sensors, cameras, LIDAR devices, microphones, motion sensors, weight sensors, force sensors, accelerometers, high voltage sensors, temperature sensors, and / or any other device or module capable of responding to direct or indirect interactions with the user.
[0052] The method 500 can further include an operation 502 of determining whether the context data indicates an intent of the user to invoke the auto-assistant. In some implementations, the context data can characterize one of the operating states of one or more of our corresponding computing devices and / or applications. For example, the context data can indicate that a first computing device is operating a first application and a second computing device different from the first computing device is operating a second application different from the first application. For example, the first computing device can be a standalone speaker that is playing music via the first application, and the second computing device can be a thermostat that is operating according to a low energy schedule and / or mode. Thus, the context data can characterize these operating states of these devices.
[0053] When the context data indicates that the user intends to invoke the automatic assistant, method 500 can transition from operation 504 to operation 506. However, when the context data does not indicate that the user intends to invoke the automatic assistant, method 500 can proceed from operation 504 to optional operation 512 and / or operation 502. Operation 506 can include causing the automatic assistant to detect one or more assistant commands based on processing the context data without the user providing a spoken invocation phrase. For example, based on the context data indicating that the user intends to invoke the automatic assistant, a computing device providing access to the automatic assistant can bypass buffering and / or processing audio data to facilitate determining whether the user has provided a spoken invocation phrase. Such an operation can be performed at one or more subsystems of the computing device while one or more other subsystems of the computing device operate in a low-power mode. When the user is predicted to intend to invoke the automatic assistant, the computing device can transition out of the low-power mode to activate one or more other subsystems for processing audio data embodying one or more assistant commands from the user.
[0054] Method 500 can proceed from operation 506 to operation 508 of determining whether the user has provided an assistant command to the automatic assistant interface of the computing device. The automatic assistant interface of the computing device can include one or more microphones, one or more cameras, and / or any other device or module capable of receiving input from the user. When the automatic assistant determines that the user has provided one or more assistant commands to the automatic assistant interface, method 500 can proceed from operation 508 to operation 510. However, when the automatic assistant does not determine that the user has provided an assistant command within a threshold time period, method 500 can proceed from operation 508 to optional operation 512 and / or operation 502. Operation 510 can include causing the automatic assistant to perform one or more actions based on the one or more assistant commands provided by the user. For example, the one or more actions can be actions performed by the automatic assistant, the computing device, a separate computing device, one or more applications, and / or any other device or module capable of interacting with the automatic assistant.
[0055] Method 500 can optionally proceed from operation 510 to optional operation 512. Operation 512 can include generating an instance of training data based on the context data and whether the user has provided an assistant command. This training data can be shared with the prior permission from the user to train one or more machine learning models to improve the probability determination regarding whether the user is about to invoke the auto-assistant. In some implementations, one or more trained machine learning models can be trained at a remote server device that communicates with various different auto-assistants interacting with various different users. The training data generated based on the interactions between different users and different auto-assistants can be used to periodically train one or more machine learning models. The trained machine learning models can then be downloaded by a computing device that provides access to the auto-assistants to improve the functionality of those auto-assistants, thereby improving the efficiency of the computing device and / or any other associated applications and devices.
[0056] For example, by improving the ability to determine whether the user is about to invoke the auto-assistant, the computing device can conserve computing resources that would otherwise be spent processing and buffering audio data to determine whether a call phrase has been detected. Additionally, certain computing devices can reduce power waste by less frequently activating the processor when transitioning between detecting a call phrase and detecting an assistant command. For example, instead of running the subsystem after the user interacts with the computing device, the computing device subsystem can suppress the subsystem based on the determination that the user is likely not going to interact with the auto-assistant further in the short term.
[0057] Figure 6 FIG. 600 is a block diagram of an example computer system 610. Computer system 610 generally includes at least one processor 614 that communicates with a number of peripheral devices via a bus subsystem 612. These peripheral devices can include a storage subsystem 624 (including, for example, a memory 625 and a file storage subsystem 626), a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with computer system 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.
[0058] The user interface input device 622 can include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into the display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways for inputting information into computer system 610 or onto a communication network.
[0059] The user interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, for example, via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways for outputting information from the computer system 610 to a user or to another machine or computer system.
[0060] The storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 can include logic for performing selected aspects of the method 500 and / or implementing one or more of the system 400, the computing device 106, the computing device 206, the computing device 306, and / or any other application, device, apparatus, and / or module discussed herein.
[0061] These software modules are typically run by the processor 614, either alone or in combination with other processors. The memory 625 used in the storage subsystem 624 can include many memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of certain implementations can be stored in the storage subsystem 624 by the file storage subsystem 626 or in other machines accessible by the processor 614.
[0062] The bus subsystem 612 provides a mechanism for enabling the various components and subsystems of the computer system 610 to communicate with each other as expected. Although the bus subsystem 612 is schematically shown as a single bus, alternative implementations of the bus subsystem can use multiple buses.
[0063] The computer system 610 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 6 the description of the computer system 610 depicted is only intended as a specific example for the purpose of illustrating some implementations. Many other configurations of the computer system 610 may have more or fewer components than the computer system depicted in Figure 6 .
[0064] In situations where the systems described in this document collect personal information about a user (or what is often referred to herein as a "participant") or can make use of personal information, users may be provided with an opportunity to control whether programs or features collect user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or control whether and / or how content more relevant to the user is received from a content server. Additionally, certain data may be processed in one or more ways before it is stored or used such that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information cannot be determined for that user, or the user's geographic location may be generalized (such as to the city, zip code, or state level) when location information is obtained so that the user's specific geographic location cannot be determined. Accordingly, users can control how information is collected and / or used about them.
[0065] Although several implementations have been described and illustrated in this document, various other means and / or structures may be utilized for performing the functions described herein and / or obtaining the results and / or advantages described herein, and each such variation and / or modification is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on the particular application or applications for which the teachings are used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific implementations described herein. Accordingly, it is to be understood that the foregoing implementations are presented by way of example only, and that implementations may be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. Implementations of the present disclosure relate to each and every individual feature, system, article of manufacture, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles of manufacture, materials, kits, and / or methods is included within the scope of the present disclosure so long as such features, systems, articles of manufacture, materials, kits, and / or methods are not mutually inconsistent.
[0066] In some implementations, a method implemented by one or more processors is described as including operations such as processing context data associated with a user and an environment in which a computing device is located, where the computing device provides access to an assistant that responds to natural language input from the user, where the processing of the context data is performed independent of whether the user has provided an invocation phrase, and where the processing of the context data is performed using a trained machine learning model trained with instances of training data based on previous interactions between one or more users and one or more assistants. The method can further include an operation of causing the assistant to detect one or more assistant commands being provided by the user based on processing the context data, where instead of the assistant requiring the user to provide an invocation phrase to the assistant, the assistant detects one or more assistant commands being provided by the user, and where the assistant detects the one or more assistant commands independent of whether the user has provided an invocation phrase to the assistant. The method can further include an operation of determining that the user has provided an assistant command to an assistant interface of the computing device based on causing the assistant to detect the one or more assistant commands, where the user has provided the assistant command without explicitly providing an invocation phrase. The method can further include an operation of causing the assistant to perform one or more actions based on the assistant command in response to determining that the user has provided the assistant command.
[0067] In some implementations, at least one instance of the training data is further based on data representing one or more previous states of one or more corresponding computing devices present in the environment. In some implementations, at least one instance of the training data is further based on other data indicating that the user has provided a particular assistant command while one or more corresponding computing devices are exhibiting one or more previous states. In some implementations, the context data represents one or more current states of one or more corresponding computing devices present in the environment. In some implementations, causing the assistant to detect one or more assistant commands being provided by the user includes: causing the computing device to bypass processing audio data to determine whether an invocation phrase has been provided by the user. In some implementations, the method can further include an operation of causing one or more computing devices in the environment to render an output to the user that includes natural language content identifying an inquiry from the assistant based on processing the context data.
[0068] In some implementations, identifying the natural language content of the query is based on the expected assistant commands selected by the automated assistant. In some implementations, the method can further include the operation of determining one or more expected assistant commands based on processing context data, where the one or more expected assistant commands include the expected assistant commands, and where at least one instance of the training data is based on an interaction in which the automated assistant also responds to the expected assistant commands. In some implementations, the expected assistant commands correspond to one or more specific actions that, when run by the automated assistant, cause the automated assistant to control one or more other computing devices associated with the user. In some implementations, the context data lacks data characterizing any invocation phrases provided by the user.
[0069] In some implementations, causing the automated assistant to detect one or more assistant commands being provided by the user includes: performing speech-to-text processing on captured audio data generated using one or more microphones connected to the computing device, where the speech-to-text processing is inactive when the automated assistant no longer detects one or more assistant commands. In some implementations, causing the automated assistant to detect one or more assistant commands being provided by the user includes: determining whether the captured audio data generated using one or more microphones connected to the computing device embodies natural language content identifying one or more actions that can be performed by the automated assistant, where determining whether the natural language content identifies one or more actions is no longer performed when the automated assistant no longer detects one or more assistant commands. In some implementations, the method can further include the operation of causing the computing device to render an output indicating that the computing device is operating to detect one or more assistant commands from the user based on processing context data. In some implementations, at least one instance of the training data is based on an interaction in which the automated assistant responds to an input from the user or another user.
[0070] In other implementations, a method implemented by one or more processors is described as including operations such as determining at a computing device that a user has provided a call phrase and an assistant command to an automated assistant interface of the computing device, where the computing device provides access to an automated assistant in response to natural language input from the user. The method can further include the operation of causing the automated assistant to perform one or more actions based on the assistant command in response to determining that the user has provided the call phrase and the assistant command. The method can further include the operation of processing context data associated with the context in which the user has provided the call phrase and the assistant command, where the context data is processed using a trained machine learning model trained with instances of training data based on one or more previous interactions between one or more users and one or more automated assistants, and where at least one instance of the training data is based on an interaction in which a particular automated assistant responded to multiple call phrases spoken by a particular user in another context within a threshold time period. The method can further include the following operation: after determining that the user has provided the call phrase and the assistant command: instead of the computing device asking the user to provide a subsequent call phrase, causing the automated assistant to detect one or more subsequent assistant commands being provided by the user based on the processed context data in order to respond to the one or more subsequent assistant commands. The method can further include the following operation: determining that the user has provided an additional assistant command, and in response to determining that the user has provided the additional assistant command, causing the automated assistant to perform one or more additional actions based on the additional assistant command and without the user providing a subsequent call phrase.
[0071] In some implementations, at least one instance of the training data is further based on data representing one or more states of one or more corresponding computing devices present in another context. In some implementations, at least one instance of the training data is further based on other data indicating that one or more users have provided a particular assistant command while one or more other computing devices are exhibiting one or more states. In some implementations, the context data represents one or more current states of one or more corresponding computing devices present in the context. In some implementations, causing the automated assistant to determine whether the user is providing one or more subsequent assistant commands includes: causing the computing device to bypass processing audio data to determine whether a call phrase has been provided by the user.
[0072] In some implementations, the method can further include the operation of causing one or more respective computing devices in the environment to render an output to the user that includes natural language content identifying an inquiry from the automated assistant, based on processing context data and an assistant command. In some implementations, identifying the natural language content of the inquiry corresponds to an expected assistant command. In some implementations, the method can further include the operation of determining one or more expected assistant commands based on processing context data, where the one or more expected assistant commands include the expected assistant command, and where at least one instance of the training data is based on an interaction in which a particular automated assistant also responds to the expected assistant command. In some implementations, one or more additional actions, when run by the automated assistant, cause the automated assistant to control one or more other computing devices located in the environment.
[0073] In other additional implementations, a method implemented by one or more processors is set forth as including operations such as processing context data associated with an environment in which a user is present with a computing device providing access to an automated assistant, where the context data is processed using a trained machine learning model trained with instances of training data based on one or more previous interactions between one or more users and one or more automated assistants, and where at least one instance of the training data is based on an interaction in which a particular automated assistant responds to one or more assistant commands provided by a particular user in the environment or another environment within a threshold time period. The method can further include the operation of causing a computing device in the environment or another computing device to render an output representing an inquiry from the automated assistant, based on processing the context data. The method can further include the operation of determining that the user has provided an assistant command to the automated assistant interface of the computing device or another computing device after the output has been rendered at the computing device or another computing device, where the user provides the assistant command without initially providing a call phrase. The method can further include the operation of causing the automated assistant to perform one or more actions based on the assistant command in response to determining that the user has provided the assistant command. In some implementations, the context data characterizes: one or more applications running at the computing device or another computing device, and one or more features of the user's physical location relative to the computing device or another computing device.
[0074] In other implementations, a method implemented by one or more processors is described as including operations such as determining that a user has provided a call phrase and an assistant command to an automated assistant interface of a computing device, where the computing device provides access to an automated assistant in response to natural language input from the user. The method can further include the operation of causing the automated assistant to perform one or more actions based on the assistant command in response to determining that the user has provided the call phrase and the assistant command. The method can further include the operation of generating first training data based on the assistant command being provided by the user, the first training data providing a correlation between the user providing the assistant command and first context data characterizing the context in which the user provided the assistant command. The method can further include the operation of determining that the user has not provided an additional assistant command when the user is present in a separate context including one or more different features from the environment, before or after the automated assistant performs one or more actions, where the separate context includes the computing device providing access to the automated assistant or another computing device. The method can further include the operation of generating second training data based on determining that the user has not provided an additional assistant command in the separate context, the second training data providing an additional correlation between the user not providing an additional assistant command and second context data characterizing the separate context. The method can further include the operation of causing one or more machine learning models to be trained using the first training data and the second training data.
[0075] In some implementations, the first context data is one or more states of one or more computing devices present in the environment. In some implementations, the first context data indicates that one or more users are located in the environment when the user has provided the assistant command. In some implementations, the second context data characterizes one or more other states of one or more other computing devices present in the separate context. In some implementations, the second context data indicates that one or more users have provided a particular assistant command while one or more other computing devices are exhibiting one or more other states. In some implementations, the environment and the separate context correspond to a common geographical location.
Claims
1. A method implemented by one or more processors, the method comprising: Processing context data associated with a user and an environment in which a computing device is located, wherein the computing device provides access to an assistant that responds to natural language input from the user, wherein the processing of the context data is performed independently of whether the user has provided an invocation phrase, wherein the processing of the context data is performed using a trained machine learning model trained with instances of training data that leverage previous interactions between one or more users and one or more assistants, and wherein the context data includes data characterizing a current state of the computing device and additional current states of additional computing devices different from the computing device, wherein the additional computing devices are located in the environment in which the user and the computing device are located; In response to processing the context data, including processing the data characterizing the current state of the computing device and the additional current states of the additional computing devices, determining to bypass a requirement for an explicit invocation of the assistant at the computing device, wherein determining to bypass the requirement for an explicit invocation of the assistant at the computing device is at least partially based on previous instances in which the user provided previous assistance instructions to the assistant while the computing device was operating in the current state and the additional computing devices were operating in the additional current states; Based on determining to bypass the requirement for an explicit invocation of the assistant, causing the assistant to detect one or more assistant commands being provided by the user, wherein instead of the assistant requiring the user to provide an invocation phrase to the assistant, the assistant detects the one or more assistant commands being provided by the user, and wherein the assistant detects the one or more assistant commands independently of whether the user has provided the invocation phrase to the assistant; Based on causing the assistant to detect the one or more assistant commands, determining that the user has provided an assistant command to an assistant interface of the computing device, wherein the user provided the assistant command without explicitly providing the invocation phrase; and In response to determining that the user has provided the assistant command, causing the assistant to perform one or more actions based on the assistant command.
2. The method according to claim 1, wherein, At least one instance of the instances of the training data is further based on data characterizing one or more previous states of one or more corresponding computing devices present in the environment.
3. The method according to claim 2, wherein The at least one instance of the training data is further based on other data indicating that the user provided a specific assistant command while the one or more corresponding computing devices were exhibiting the one or more previous states.
4. The method according to claim 1, further comprising: Based on processing the context data, causing one or more computing devices in the environment to render an output to the user that includes natural language content identifying an inquiry from the assistant.
5. The method according to claim 4, wherein, Identifying the natural language content of the query is based on an expected assistant command selected by the automated assistant.
6. The method according to claim 5, further comprising: Determining one or more expected assistant commands based on processing the context data, wherein the one or more expected assistant commands include the expected assistant command, and wherein at least one instance of the training data is based on an interaction in a previous interaction in which the automated assistant also responds to the expected assistant command.
7. The method according to claim 1, wherein, The context data further includes additional data indicating a further current state of a further computing device that is different from the computing device and different from the additional computing device.
8. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-7.
9. A computer-readable storage medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 7.
10. A system for causing an automated assistant to perform one or more actions, the system comprising one or more processors for performing the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Speech recognition method and apparatus
US20180173494A1