Dynamic and / or context-specific hotwords for invoking an automation assistant
By enabling dynamic hotwords based on context-specific conditions, the automated assistant can efficiently manage resource usage and enhance user convenience, addressing the limitations of traditional restricted hot word listening states.
Patent Information
- Application Number
- JP2022164025
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2038-08-21
AI Technical Summary
Existing automated assistants operate in a restricted hot word listening state to conserve power and resources, but this can be cumbersome in scenarios where users need to invoke actions frequently, leading to potential resource wastage and user inconvenience.
Implementing dynamic hotwords that can be selectively enabled based on context-specific conditions, such as proximity to the device or user recognition, allowing the automated assistant to temporarily expand or change its hotword vocabulary.
This approach reduces the need for users to invoke the automated assistant before performing actions, conserves resources by minimizing unnecessary speech recognition processing, and enhances user convenience by allowing context-specific command execution.
Smart Images

Figure 0007700087000001 
Figure 0007700087000002 
Figure 0007700087000003
Abstract
Description
Background Art
[0001] A human can be involved in a human-computer dialogue with an interactive software application called an "automation assistant" (also referred to as a "chatbot", "conversational personal assistant", "intelligent personal assistant", "personal voice assistant", "conversational agent", "virtual assistant", etc.) in this specification. For example, a human (sometimes called a "user" when interacting with an automation assistant) can use free-form natural language input that may include free-form natural language input that has been converted to text and then processed and / or typed free-form natural language input to provide commands, queries, and / or requests (collectively referred to as "queries" in this specification).
[0002] Often, before an automated assistant can interpret and respond to a user's request, the user's request must first be "invoked" using a pre-defined verbal call phrase, often referred to as, for example, a "hot word" or "wake word". Thus, many automated assistants operate in what is herein referred to as a "restricted hot word listening state" or "default listening state", where the automated assistant is constantly "listening" for audio data sampled by a microphone for a restricted (or finite or "default") set of hot words. Any utterance captured within audio data outside of the default set of hot words is ignored. When the automated assistant is invoked using one or more of the default set of hot words, the automated assistant may operate in what is herein referred to as a "speech recognition state", in which, for at least some time interval after the invocation, the automated assistant performs a speech-to-text ("STT") process on the audio data sampled by the microphone to generate a text input, which is then semantically processed to determine (and to fulfill) the user's intent.
[0003] Operating an automated assistant in a default listening state provides various advantages. Limiting the number of "listened for" hotwords enables conservation of power and / or computing resources. For example, an on-device machine learning model can be trained to generate an output indicating that one or more hotwords have been detected. Implementing such a model may require only minimal computing resources and / or power, which can be particularly beneficial for assistant devices with resource constraints. Storing such a trained model locally on a client also provides privacy-related advantages. For example, most users do not want STT processing to be performed on everything they say within earshot of the computing device operating the automated assistant. Additionally, an on-device model also prevents data indicating user utterances not intended to be processed by the automated assistant from often being provided to a semantic processor that operates at least partially in the cloud.
[0004] Along with these advantages, operating an automated assistant in a limited hot-word listening state also presents various challenges. To avoid inadvertent invocation of the automated assistant, hot words are typically selected to be words or phrases that are not uttered very often in everyday conversation (e.g., "long-tail" words or phrases). However, there exist various scenarios where it may be cumbersome to require the user to utter a long-tail hot word before invoking the automated assistant to perform some action. Some automated assistants may provide an option for "continuous listening" after the user utters a command so that the user does not need to use the hot word to "wake up" the automated assistant again before executing subsequent commands. However, transitioning the automated assistant to a continuous listening mode means that the automated assistant may perform much more SST processing on much more utterances, potentially wasting power and / or computing resources. Additionally, as noted above, most users prefer that only utterances directed to the automated assistant be STT processed. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0005] This specification describes techniques for enabling the use of "dynamic" hotwords for an automated assistant. In various situations, an automated assistant configured in accordance with selected aspects of the present disclosure can often more intelligently listen for specific context-specific hotwords embodied in what is referred to herein as an "expanded" set of hotwords. In various implementations, the automated assistant can listen for these context-specific hotwords in addition to, or instead of, the default hotwords used to invoke the automated assistant. In other words, in various implementations, an automated assistant configured in accordance with selected aspects of the present disclosure can, in certain situations, at least temporarily expand or change its hotword vocabulary.
[0006] In various implementations, the dynamic hotwords can be selectively enabled under various different circumstances, such as the user being in sufficient proximity to the assistant device, or the user being recognized based on the user's appearance, voice, and / or other identifying characteristics. These other identifying characteristics can include, for example, RFID badges, visual markers, uniforms, a piece of flair, the size of the user (which can indicate that the user is an adult), physical disabilities that can affect the user's ability to communicate verbally and / or via conventional user inputs (such as a mouse, touch screen, etc.).
[0007] As an example, in some implementations, for instance, within a predetermined proximity of a computing device that can be used to interact with an automated assistant, if a user is detected by a proximity sensor, one or more additional hotwords can be activated to enable the nearby user to more easily invoke the automated assistant. In some implementations, the closer the user gets to the assistant device, the more dynamic hotwords can be activated. For example, a user detected within 3 meters of the assistant device may activate a first set of dynamic hotwords. If the user is detected closer, for example, within 1 meter, additional or alternative hotwords can be activated.
[0008] As another example, in some implementations, one or more dynamic hotwords associated with the user (e.g., custom), such as hotwords manually selected or entered by the user, can be downloaded and / or "activated". The user can then utter one or more of these custom hotwords to invoke the automated assistant without uttering one or more of the default set of hotwords generally used by people to invoke the automated assistant.
[0009] In some implementations, the dynamic hotword can be activated based on some combination of proximity and recognition. This can be particularly beneficial in noisy and / or crowded environments. For example, any number of other people in such an environment may also be trying to interact with the automated assistant using default hotwords that are likely to be relatively standardized. To avoid inadvertent activation of the automated assistant by others, when the user is sufficiently close to the assistant device (e.g., holding the user's smartphone, or wearing the user's smartwatch, e.g., it is detected that the snap or fastener is engaged, which can trigger activation of an extended set of hotwords), and recognized, the automated assistant can stop listening for standard or default hotwords and instead listen for custom hotwords tailored to the specific user. Alternatively, the automated assistant can continue to listen for the default hotwords but increase the confidence threshold required for calls based on those default hotwords while activating or lowering the threshold associated with the custom hotwords.
[0010] As another example, in some implementations, a group of people (e.g., employees, by gender, age range, sharing visual characteristics, etc.) can be associated with and have activated dynamic hotwords that are detected and / or recognized by one or more of those people. For example, a group of people can share visual characteristics. In some implementations, when one or more of these visual characteristics are detected by one or more hardware sensors of the assistant device, one or more dynamic hotwords that may otherwise not be available to non-group members can be activated.
[0011] Hotwords can be detected by, or instead of, an automated assistant in a variety of ways. In some implementations, a machine learning model, such as a neural network, can be trained to detect an ordered or unordered sequence of one or more hotwords within an audio data stream. In some such implementations, a separate machine learning model can be trained for each applicable hotword (or "hot phrase" including multiple hotwords).
[0012] In some implementations, a machine learning model trained to detect these dynamic hotwords can be downloaded as needed. For example, assume that a particular user is detected based on, for example, face recognition processing performed on one or more images of the user, or speech recognition processing performed on audio data generated from the user's speech. If not yet available on the device, these models can be downloaded from the cloud based on, for example, the user's online profile.
[0013] In some implementations, to improve the user experience and reduce latency, response actions that should be performed by the automated assistant upon detection of context-specific hotwords can be pre-cached on the device where the user interacts with the automated assistant. Then, as soon as a context-specific hotword is detected, the automated assistant can immediately take action. This is in contrast to cases where the automated assistant may first need to perform round-trip network communication with one or more computing devices (e.g., the cloud) to fulfill the user's request.
[0014] The techniques described in this specification yield a variety of technical advantages. At least temporarily expanding or modifying the vocabulary available for invoking an automated assistant in a particular context can reduce the need to first invoke the automated assistant before having the automated assistant perform an action related to some context, such as stopping a timer, pausing music, etc. Some existing assistant devices make it easy to pause media playback or stop a timer by allowing a user to simply tap on an active portion of the device's surface (e.g., a capacitive touchpad or display) without the user having to first invoke the automated assistant. However, users with physical disabilities and / or other busy users (e.g., cooking, driving, etc.) may not be able to easily touch the device. Thus, the techniques described in this specification enable those users to more easily and quickly have the automated assistant perform some response action, such as stopping a timer, without first invoking the automated assistant.
[0015] In addition, as described herein, in some implementations, the automated assistant may actively download content that responds to context-specific hotwords. For example, assume that a user has a tendency to repeat certain requests such as "what's the weather forecast today?", "what's on my schedule?", etc. Information responding to these requests can be pre-downloaded and cached in memory, and commands (or portions thereof) can be used to activate specific dynamic hotwords. As a result, when one or more hotwords are spoken, the automated assistant can provide response information more quickly than if the automated assistant had to first make contact with one or more remote resources via one or more networks to obtain the response information. This can also be beneficial if the assistant device is in a vehicle that may enter and exit zones where the data network is available. For example, by pre-downloading and caching content that responds to specific context-specific hotwords while the vehicle is within the data coverage zone, that data can then be available if the user requests that data while moving outside of the data coverage zone.
[0016] As yet another example, the techniques described herein may enable a user to trigger response actions without requiring a comprehensive speech-to-text (``STT'') process. For example, when a particular context invocation model is activated and context-specific hotwords are detected, response actions may be triggered based on the output of these models without the user's utterance having to be converted to text using STT processing. This may conserve computing resources on the client device and / or avoid round-trip communication with a cloud infrastructure to perform STT and / or semantic processing of the user's utterance, which conserves network resources. Also, avoiding round-trip communication may improve latency and avoid sending at least some data to the cloud infrastructure, which may be advantageous and / or desirable from the perspective of user privacy.
[0017] In some implementations, a method is provided that is executed by one or more processors. The method includes: executing an automated assistant in a default listening state, where the automated assistant is at least partially executed on one or more computing devices operated by a user; monitoring, by one or more microphones, audio data captured for one or more of a default set of one or more hotwords while in the default listening state, where detection of one or more of the default set of hotwords triggers a transition of the automated assistant from the default listening state to an audio recognition state; detecting one or more sensor signals generated by one or more hardware sensors integral with one or more of the computing devices; analyzing the one or more sensor signals to determine an attribute of the user; transitioning, based on the analysis, the automated assistant from the default listening state to an extended listening state; and monitoring, by one or more of the microphones, audio data captured for one or more of an extended set of one or more hotwords while in the extended listening state, where detection of one or more of the extended set of hotwords triggers the automated assistant to perform a response action without requiring detection of one or more of the default set of hotwords, and where one or more of the extended set of hotwords are not within the default set.
[0018] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0019] In various implementation forms, one or more hardware sensors may include proximity sensors, and user attributes may include that the user is detected by the proximity sensor, or that the user is detected by the proximity sensor within a predetermined distance of one or more of the computing devices.
[0020] In various implementation forms, one or more hardware sensors may include a camera. In various implementation forms, user attributes may include that the user is detected by the camera, or that the user is detected by the camera within a predetermined distance of one or more of the computing devices. In various implementation forms, the analysis may include face recognition processing. In various implementation forms, user attributes may include the user's identification information. In various implementation forms, user attributes may include the user's membership in a group.
[0021] In various implementation forms, one or more hardware sensors may include one or more of microphones. In various implementation forms, one or more attributes may include that the user is auditorily detected based on audio data captured by one or more of the microphones. In various implementation forms, the analysis may include speech recognition processing. In various implementation forms, user attributes may include the user's identification information. In various implementation forms, user attributes may include the user's membership in a group.
[0022] In various implementation forms, the response action may include transitioning the automated assistant to a speech recognition state. In various implementation forms, the response action may include the automated assistant executing a task requested by the user using one or more of an extended set of hotwords.
[0023] In addition, some implementations include one or more processors of one or more computing devices, the one or more processors being operable to execute instructions stored in associated memory, the instructions being configured to cause execution of any of the methods described above. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to perform any of the methods described above.
[0024] It should be understood that all combinations of the foregoing concepts and additional concepts described in greater detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein.
Brief Description of the Drawings
[0025]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
DETAILED DESCRIPTION OF THE INVENTION
[0026] Turning now to FIG. 1, an exemplary environment in which the techniques disclosed herein may be implemented is shown. The exemplary environment includes one or more client computing devices 106. Each client device 106 may execute a respective instance of an automation assistant client 108, which may sometimes be referred to herein as the “client portion” of the automation assistant. One or more cloud-based automation assistant components 119, which may sometimes be collectively referred to herein as the “server portion” of the automation assistant, may be implemented on one or more computing systems (collectively referred to as “cloud” computing systems) communicatively coupled to the client devices 106 via one or more local and / or wide area networks (e.g., the Internet), as shown generally at 114.
[0027] In various implementations, an instance of the automation assistant client 108 can form, from the user's perspective, something that appears as a logical instance of an automation assistant 120 with which the user can engage in a human-computer interaction, by virtue of its interaction with one or more cloud-based automation assistant components 119. One instance of such an automation assistant 120 is shown by the dashed line in FIG. 1. Thus, it should be understood that each user interacting with the automation assistant client 108 running on the client device 106 can effectively interact with a logical instance of the automation assistant 120 that is the user's own. For the sake of brevity and simplicity, the term "automation assistant" as used herein to refer to what "serves" a particular user refers to the combination of the automation assistant client 108 running on the client device 106 operated by the user and one or more cloud-based automation assistant components 119 (which may be shared among multiple automation assistant clients 108). In some implementations, it should also be understood that the automation assistant 120 can respond to requests from any user, regardless of whether the user is actually "served" by that particular instance of the automation assistant 120.
[0028] One or more client devices 106 can include, for example, a desktop computing device, a laptop computing device, a tablet computing device, a cellular phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker (which may optionally include a visual sensor), a smart home appliance such as a smart TV (or a standard TV equipped with a networked dongle having automated assistant capabilities), and / or a wearable device of the user that includes a computing device (e.g., a user's wristwatch having a computing device, a user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices can be provided. Some client devices 106, such as a stand-alone interactive speaker (or "smart speaker"), can take the form of an assistant device that is primarily designed to facilitate interaction between the user and the automated assistant 120. Some such assistant devices can take the form of a stand-alone interactive speaker with a display that may or may not be a touchscreen device.
[0029] In some implementations, client device 106 may include one or more vision sensors 107 having one or more fields of view, but this is not essential. The vision sensors 107 can take various forms such as digital cameras, passive infrared ( "PIR") sensors, stereo cameras, RGB cameras, etc. The one or more vision sensors 107 can be used, for example, by an image capture module 111 to capture an image frame (still image or video) of the environment in which the client device 106 is deployed. These image frames can then be analyzed, for example, by a visual cue module 1121 to detect user-provided visual cues contained within the image frame. These visual cues can include, but are not limited to, hand gestures, fixation on a particular reference point, facial expressions, pre-defined movements by the user, etc. These detected visual cues can be used for various purposes such as invoking the automated assistant 120 and / or causing the automated assistant 120 to perform various actions.
[0030] Additionally or alternatively, in some implementations, client device 106 may include one or more proximity sensors 105. Proximity sensors can take various forms such as passive infrared ( "PIR"), radio frequency identification ( "RFID"), components that receive signals emitted from another nearby electronic component (e.g., a Bluetooth signal from a nearby user's client device, a high-frequency or low-frequency sound emitted from a device), etc. Additionally or alternatively, the vision sensors 107 and / or the microphone 109 can also be used as proximity sensors, for example, by visually and / or audibly detecting the proximity of a user.
[0031] As described in more detail herein, the automation assistant 120 is involved in a human-computer dialogue session with one or more users via the user interface input and output devices of one or more client devices 106. In some implementations, the automation assistant 120 can be involved in a human-computer dialogue session with a user in response to user interface input provided by the user via one or more user interface input devices of one of the client devices 106. In some of those embodiments, the user interface input is explicitly directed to the automation assistant 120. For example, the user may verbally provide (e.g., type, speak) a predetermined call phrase such as "OK, Assistant" or "Hey, Assistant" to initiate the automation assistant 120 to actively listen or monitor typed text. Additionally or alternatively, in some implementations, the automation assistant 120 can be invoked based on one or more detected visual cues, alone or in combination with a verbal call phrase.
[0032] In some implementations, even when the user interface input is not explicitly directed to the automation assistant 120, the automation assistant 120 can participate in a human-computer dialogue session in response to the user interface input. For example, the automation assistant 120 can examine the content of the user interface input and participate in the dialogue session in response to the presence of specific terms in the user interface input and / or based on other cues. In many implementations, the automation assistant 120 utilizes speech recognition to convert the user's utterance into text and, in response thereto, can respond to the text, for example, by providing search results, general information, and / or performing one or more response actions (e.g., playing media, launching a game, ordering food, etc.). In some implementations, the automation assistant 120 can additionally or alternatively respond to the utterance without converting the utterance into text. For example, the automation assistant 120 can convert the voice input into an embedding, an entity representation (indicating an entity present in the voice input), and / or other "non-text" representations and can operate on such non-text representations. Thus, the implementations described herein that operate based on text converted from voice input can additionally and / or alternatively operate directly on the voice input and / or on other non-text representations of the voice input.
[0033] Each of the client computing device 106 and the computing device operating the cloud-based automated assistant component 119 may include one or more memories for storing data and software applications, one or more processors for accessing the data and executing the applications, and other components that facilitate communication via a network. Operations performed by the client computing device 106 and / or by the automated assistant 120 may be distributed across multiple computer systems. The automated assistant 120 may be implemented, for example, as a computer program executed on one or more computers at one or more locations coupled to each other via a network.
[0034] As described above, in various implementations, the client computing device 106 may operate the automated assistant client 108 or the "client portion" of the automated assistant 120. In various implementations, the automated assistant client 108 may include a voice capture module 110, the aforementioned image capture module 111, a visual cue module 1121, and / or a call module 113. In other implementations, one or more aspects of the voice capture module 110, the image capture module 111, the visual cue module 112, and / or the call module 113 may be implemented separately from the automated assistant client 108, for example, by one or more cloud-based automated assistant components 119. For example, in FIG. 1, there is also a cloud-based visual cue module 1122 that can detect visual cues in image data.
[0035] In various implementation forms, the voice capture module 110, which can be implemented using any combination of hardware and software, can interface with hardware such as the microphone 109 or other pressure sensors to capture the audio recording of the user's speech. Various types of processing can be performed on this audio recording for various purposes. In some implementation forms, the image capture module 111, which can be implemented using any combination of hardware and software, can be configured to interface with the camera 107 to capture one or more image frames (e.g., digital photos) corresponding to the field of view of the visual sensor 107.
[0036] In various implementation forms, the visual cue module 1121 (and / or the cloud-based visual cue module 1122) can be implemented using any combination of hardware and software and can be configured to analyze one or more image frames provided by the image capture module 111 to detect one or more visual cues captured within one image frame and / or across multiple image frames. The visual cue module 1121 can use various techniques to detect visual cues. For example, the visual cue module 1122 can use one or more artificial intelligence (or machine learning) models trained to generate an output indicating the detected user-provided visual cues within the image frame.
[0037] As described above, the voice capture module 110 can be configured to capture a user's voice, for example, via the microphone 109. Additionally or alternatively, in some implementations, the voice capture module 110 can be further configured to convert the captured audio into text and / or other representations or embeddings, for example, using a voice text conversion ("STT") processing technique. Additionally or alternatively, the voice capture module 110 can be configured to convert text into computer synthesized voice, for example, using one or more voice synthesizers. However, in some cases, since the client device 106 may be relatively constrained with respect to computing resources (e.g., processor cycles, memory, battery, etc.), the voice capture module 110 local to the client device 106 can be configured to convert a finite number of different utterance phrases, particularly phrases that invoke the automated assistant 120, into text (or other forms such as lower dimensional embeddings). Other voice inputs can be sent to a cloud-based automated assistant component 119 that can include a cloud-based text to speech ("TTS") module 116 and / or a cloud-based STT module 117.
[0038] In various implementation forms, the calling module 113 may be configured to determine whether to call the automation assistant 120 based on, for example, the output provided by the voice capture module 110 and / or the visual cue module 1121 (which may be combined with the image capture module 111 in a single module in some implementation forms). For example, the calling module 113 may determine whether the user's utterance is qualified as a calling phrase to start a human-computer dialogue session with the automation assistant 120. In some implementation forms, the calling module 113 may analyze data indicating the user's utterance, such as an audio recording or a vector of features (e.g., embedding) extracted from the audio recording, alone or in combination with one or more visual cues detected by the visual cue module 1121. In some implementation forms, the threshold used by the calling module 113 to determine whether to call the automation assistant 120 in response to an audio utterance may be lowered when a specific visual cue is also detected. As a result, even if the user provides an audio utterance that is somewhat acoustically similar but different from the appropriate calling phrase "OK assistant", that utterance may still be accepted as an appropriate call if it is detected in combination with a visual cue (e.g., a hand gesture by the speaker, the speaker directly gazing at the visual sensor 107, etc.).
[0039] In some implementations, for example, one or more on-device invocation models stored in the on-device model database 114 can be used by the invocation module 113 to determine whether utterances and / or visual cues are eligible as invocations. Such on-device invocation models can be trained to detect variations in invocation phrases / gestures. For example, in some implementations, an on-device invocation model (e.g., one or more neural networks) can be trained using training examples that each include an audio recording of an utterance from a user (or an extracted feature vector), as well as data indicating one or more image frames captured simultaneously with the utterance and / or detected visual cues.
[0040] In FIG. 1, the on-device model database 114 can store one or more on-device invocation models 1141-114 N In some implementations, the default on-device invocation model 1141 can be trained to detect one or more default invocation phrases or hotwords such as those described above (e.g., "OK Assistant", "Hey, Assistant", etc.) in an audio recording or other data representing it. In some such implementations, these models can always be available and used to transition the automated assistant 120 to a general listening state, in which any audio recording captured by the audio capture module 110 (for at least some period following the invocation) can be processed using other components of the automated assistant 120 described below (e.g., on the client device 106 or by one or more cloud-based automated assistant components 119).
[0041] In addition, in some implementations, the on-device model database 114 can store, at least temporarily, one or more additional "context call models" 1142-114 N These context call models 1142-114 N can be used by, and / or be available to, a calling module 113 (e.g., activated) in a particular context. The context call models 1142-114 N can be trained, for example, to detect one or more context-specific hotwords in an audio recording or other data indicative thereof. In some implementations, the context call models 1142-114 N can be selectively downloaded as needed from, for example, a dynamic hotword engine 128 that forms part of a cloud-based automated assistant component 119, as described in more detail below.
[0042] In various implementations, when the calling module 113 uses the context call models 1142-114 N to detect various dynamic hotwords, it can transition the automated assistant 120 to the aforementioned general listening state. Additionally or alternatively, the calling module 113 can transition the automated assistant 120 to a context-specific state, in which one or more context-specific response actions are performed regardless of whether the automated assistant 120 is transitioned to the general listening state. Often, the audio data that triggered the transition of the automated assistant 120 to the context-specific state may not be sent to the cloud. Instead, one or more context-specific responses can be fully executed on the client device 106, which can reduce both the response time and the amount of information sent to the cloud, which can be beneficial from a privacy perspective.
[0043] The cloud-based TTS module 116 can be configured to utilize the virtually unlimited resources of the cloud to convert text data (e.g., natural language responses formulated by the automation assistant 120) into computer-generated voice output. In some implementations, the TTS module 116 can provide the computer-generated voice output to the client device 106 for direct output, for example, using one or more speakers. In other implementations, the text data (e.g., natural language responses) generated by the automation assistant 120 can be provided to the voice capture module 110, and the voice capture module 110 can then convert the text data into computer-generated voice output for local output.
[0044] The cloud-based STT module 117 can be configured to utilize the virtually unlimited resources of the cloud to convert audio data captured by the voice capture module 110 into text, and the text can then be provided to the intent matcher 135. In some implementations, the cloud-based STT module 117 can convert the audio recording of the voice into one or more phonemes and then convert the one or more phonemes into text. Additionally or alternatively, in some implementations, the STT module 117 can use a state decoding graph. In some implementations, the STT module 117 can generate multiple candidate text interpretations of the user's utterance. In some implementations, the STT module 117 can weight or bias a particular candidate text interpretation higher than others depending on whether there are visually detected cues detected simultaneously.
[0045] The automated assistant 120 (and in particular, the cloud-based automated assistant component 119) may include an intent integrator 135, the aforementioned TTS module 116, the aforementioned STT module 117, and other components that will be described in more detail below. In some implementations, one or more of the modules and / or components of the automated assistant 120 may be omitted, combined, and / or implemented within a component separate from the automated assistant 120. In some implementations, to protect privacy, one or more of the components of the automated assistant 120, such as the natural language processor 122, the TTS module 116, the STT module 117, etc., may be implemented at least partially on the client device 106 (e.g., excluding the cloud).
[0046] In some implementations, the automated assistant 120 generates response content in response to various inputs generated by a user of one of the client devices 106 during a human-computer dialogue session with the automated assistant 120. The automated assistant 120 may provide the response content for presentation to the user as part of the dialogue session (e.g., via one or more networks if separated from the user's client device). For example, the automated assistant 120 may generate response content in response to free-form natural language input provided via the client device 106. As used herein, free-form input is input that is formulated by the user and not restricted to a group of options presented for the user's selection.
[0047] As used herein, an "interaction session" may include a logically self - contained exchange of one or more messages between a user and an automated assistant 120 (and, optionally, other human participants). The automated assistant 120 can distinguish between multiple interaction sessions with a user based on various cues such as the passage of time between sessions, changes in the user context between sessions (e.g., location, before / during / after a scheduled meeting, etc.), detection of one or more intervening interactions between the user and the client device other than the interaction between the user and the automated assistant (e.g., the user leaves and later returns to a stand - alone voice - activated product), locking / sleeping of the client device between sessions, and changes in the client device used to interface with one or more instances of the automated assistant 120.
[0048] The intent integrator 135 may be configured to determine the user's intent based on input provided by the user (e.g., voice utterances, visual cues, etc.) and / or based on other signals such as sensor signals, online signals (e.g., data obtained from a web service). In some implementations, the intent integrator 135 may include a natural language processor 122 and the aforementioned cloud - based visual cue module 1122. In various implementations, the cloud - based visual cue module 1122 may operate similarly to the visual cue module 1121, except that the cloud - based visual cue module 1122 may have more resources available for its use. In particular, the cloud - based visual cue module 1122 may be able to detect visual cues that can be used by the intent integrator 135, either alone or in combination with other signals, to determine the user's intent.
[0049] The natural language processor 122 may be configured to process natural language input generated by a user via the client device 106 and generate an annotated output (e.g., in text form) for use by one or more other components of the automated assistant 120. For example, the natural language processor 122 may process natural language free-form input generated by a user via one or more user interface input devices of the client device 106. The generated annotated output includes one or more annotations of the natural language input and one or more (e.g., all) of the terms of the natural language input.
[0050] In some implementations, the natural language processor 122 is configured to identify and annotate various types of grammatical information within the natural language input. For example, the natural language processor 122 may include a morphological module that separates individual words into morphemes and / or annotates the morphemes, e.g., with their classes. The natural language processor 122 may also include a part-of-speech tagger configured to annotate terms with their grammatical roles. For example, the part-of-speech tagger may tag each term with its part of speech, such as "noun", "verb", "adjective", "pronoun", etc. Also, for example, in some implementations, the natural language processor 122 may additionally and / or alternatively include a dependency parser (not shown) configured to determine syntactic relationships between terms within the natural language input. For example, the dependency parser may determine which terms modify other terms, the subject, the verb, etc. of a sentence (e.g., a parse tree) and may annotate such dependencies.
[0051] In some implementations, the natural language processor 122 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references in one or more segments, such as references to people (e.g., including literary characters, celebrities, public figures, etc.), organizations, places (real and fictional), and the like. In some implementations, information about the entities may be stored in one or more databases, such as in a knowledge graph (not shown). In some implementations, the knowledge graph may include nodes that represent known entities (and, in some cases, entity attributes), as well as edges that connect the nodes and represent relationships between the entities. For example, a "banana" node may be connected (e.g., as a child) to a "fruit" node, which may in turn be connected (e.g., as a child) to a "produce" node and / or a "food" node. As another example, a restaurant called "Hypothetical Cafe" may be represented by a node that also includes attributes such as its address, the type of food served, hours, contact information, and the like. In some implementations, the "Virtual Cafe" node may be connected by edges (e.g., representing child-parent relationships) to one or more other nodes, such as "Restaurant" nodes, "Business" nodes, nodes representing the city and / or state in which the restaurant is located, etc.
[0052] The entity tagger of the natural language processor 122 may annotate references to entities at a high level of granularity (e.g., allowing for identification of all references to an entity class, such as a person) and / or at a lower level of granularity (e.g., allowing for identification of all references to a particular entity, such as a particular person). The entity tagger may rely on the content of the natural language input to resolve particular entities and / or may optionally communicate with a knowledge graph or other entity database to resolve particular entities.
[0053] In some implementations, the natural language processor 122 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, the coreference resolver may be utilized to resolve the term "there" in the natural language input "I liked Hypothetical Cafe last time we ate there" to "Hypothetical Cafe".
[0054] In some implementations, one or more components of the natural language processor 122 may depend on annotations from one or more other components of the natural language processor 122. For example, in some implementations, the named entity tagger may depend on annotations from the coreference resolver and / or the dependency parser when annotating all mentions of a particular entity. Also, for example, in some implementations, the coreference resolver may depend on annotations from the dependency parser when clustering references to the same entity. In some implementations, when processing a particular natural language input, one or more components of the natural language processor 122 may use related previous inputs and / or other related data external to the particular natural language input to determine one or more annotations.
[0055] The intent integrator 135 can use various techniques to determine the user's intent based on, for example, the output from a natural language processor 122 (which may include annotations and terms of natural language input) and / or based on the output from a visual cue module (e.g., 1121 and / or 1122). In some implementations, the intent integrator 135 can access one or more databases (not shown) that include, for example, multiple mappings between grammar, visual cues, and response actions (or more generally, intents). Often, these grammars can be selected and / or learned over time and can represent the most common intents of the user. For example, "play <artist>" <artist>)」One grammar is on the client device 106 operated by the user <artist>( <artist>) can be mapped to an intention to call a response action to play music. Another grammar, "[weather|forecast] today", can potentially match user queries such as "What's the weather today?" and "What's the forecast for today?".
[0056] In addition to, or instead of, grammar, in some implementations, the intent matcher 135 can use one or more trained machine learning models, either alone or in combination with one or more grammars and / or visual cues. These trained machine learning models can also be stored in one or more databases, for example, by embedding data indicating a user utterance and / or any detected user-provided visual cue into a space with reduced dimensions, and then determining which other embeddings (and thus intents) are closest, for example, using techniques such as Euclidean distance, cosine similarity, etc., to be trained to identify intents.
[0057] "Play <artist>)」As can be seen in the exemplary grammar of "", some grammars have slots (e.g., <artist>) that can be filled with slot values (or "parameters") <artist>)) has. The slot value can be determined in various ways. Often, the user actively provides the slot value. For example, "Please order a <topping>Regarding the grammar of "pizza)", the user may have spoken the phrase "order me a sausage pizza", in which case the slot <topping> ( <topping>) is automatically satisfied. Additionally or alternatively, when the user invokes a grammar that includes slots to be filled with slot values without the user actively providing the slot values, the automation assistant 120 may request those slot values from the user (e.g., "what type of crust do you want on your pizza?"). In some implementations, the slots may be filled with slot values based on visual cues detected by the visual cue modules 1121 - 1122. For example, the user can utter something like "Order me this many cat bowls" while holding up three fingers to the visual sensor 107 of the client device 106. Or, the user can utter something like "Find me more movies like this" while holding a DVD case of a particular movie.
[0058] In some implementations, the automation assistant 120 may function as an intermediary between the user and one or more third - party computing services 130 (or "third - party agents" or "agents"). These third - party computing services 130 can be independent software processes that receive inputs and provide response outputs. Some third - party computing services can take the form of third - party applications that may or may not operate on a computing system separate from the one operating the cloud - based automation assistant component 119. One type of user intent that can be identified by the intent integrator 135 is to interact with third - party computing services 130. For example, the automation assistant 120 can provide access to an application programming interface ("API") to a service for controlling smart devices. The user can call the automation assistant 120 and provide a command such as "I'd like to turn the heating on". The intent integrator 135 can trigger the automation assistant 120 to interact with the third - party service and thereby map this command to the grammar for turning the user's heating on. The third - party service 130 can provide the automation assistant 120 with a minimal list of slots that need to be filled to fulfill (or "resolve") the command to turn the heating on. In this example, the slots can include the temperature to which the heating should be set and the duration for which the heating should be on. The automation assistant 120 can generate a natural - language output asking for parameters for the slots and provide it to the user (via the client device 106).
[0059] The fulfillment module 124 may receive the predicted / estimated intent output by the intent integrator 135, as well as the related slot values (regardless of whether actively provided by the user or requested from the user), and may be configured to fulfill (or "resolve") the intent. In various implementations, the fulfillment (or "resolution") of the user's intent may cause the fulfillment module 124, for example, to generate / obtain various fulfillment information (also referred to as "response" information or "resolution" information). As described below, the fulfillment information may be provided to a natural language generator (referred to as "NLG" in some figures) 126 in some implementations, which may generate a natural language output based on the fulfillment information.
[0060] Since the intent can be satisfied (or "resolved") in various ways, the fulfillment (or "resolution") information can take various forms. Assume that the user requests pure information such as "Where were the outdoor shots of 'The Shining' filmed?". The user's intent may be determined, for example, by the intent integrator 135 as a search query. The intent and content of the search query may be provided to the fulfillment module 124, and the fulfillment module 124 may communicate with one or more search modules 150 configured to search a corpus of documents and / or other data sources (such as a knowledge graph, etc.) for response information, as shown in FIG. 1. The fulfillment module 124 may provide data indicating the search query (such as the text of the query, a reduced-dimensional embedding, etc.) to the search module 150. The search module 150 may provide response information such as GPS coordinates or other more explicit information such as "Timberline Lodge, Mt. Hood, Oregon". This response information may form part of the fulfillment information generated by the fulfillment module 124.
[0061] Additionally or alternatively, the fulfillment module 124 may be configured to receive, from, for example, the intent integrator 135, the user's intent provided by the user or determined using other means and any slot values (e.g., the user's GPS coordinates, user settings, etc.), and trigger a response action. The response action may include, for example, ordering a good / service, starting a timer, setting a reminder, initiating a phone call, playing media, sending a message, and the like. In some such implementations, the fulfillment information may include slot values related to the fulfillment, a confirmation response (which may, in some cases, be selected from a predetermined set of responses), and the like.
[0062] The natural language generator 126 may be configured to generate and / or select a natural language output (e.g., words / phrases designed to mimic human speech) based on data obtained from various sources. In some implementations, the natural language generator 126 may be configured to receive, as input, fulfillment information related to the fulfillment of an intent and generate a natural language output based on the fulfillment information. Additionally or alternatively, the natural language generator 126 may receive information (e.g., required slots) from other sources such as third-party applications that may be used to construct a natural language output for the user.
[0063] Figure 2 schematically shows an exemplary state machine that can be implemented by an automated assistant (e.g., 120) and / or an assistant device (e.g., 106) configured in accordance with selected aspects of the present disclosure in various implementation forms. The upper left is the "default inactive state" where the automated assistant 120 can exist when not involved by the user. In the default inactive state, one or more microphones of one or more client devices (106) are activated, and the audio data captured by it can be analyzed using the techniques described herein. The automated assistant 120 can transition to the "general listening state" in response to the detection of one or more default call words (such as "OK, Assistant" or "Hey, Assistant", "DIW" in Figure 2, also referred to as "hot words" herein), for example, by a call module 113 and / or a visual cue module 112 based on a default call model 1141. Utterances other than the default hot words (such as ambient conversations, etc.) may be ignored and not processed.
[0064] In the general listening state, the automated assistant 120 can capture the audio data spoken after the default call word and transition to the "general processing" state. In the general processing state, the automated assistant 120 can process the data indicating the voice input as described above with respect to Figure 1, including STT processing, natural language processing, intent matching, fulfillment, etc. When the processing is completed, the automated assistant 120 can return to the default inactive state. If no audio input is received after the detection of the default call word, a timeout (labeled "TO" in Figure 2) can return the automated assistant 120 from the general listening state to the default inactive state so that subsequent utterances not intended for processing by the automated assistant are not captured or processed.
[0065] As described above, the techniques described herein facilitate context-specific hotwords that can be activated and detected to transition the automated assistant 120 to various different states such as a general listening state, or other context-specific states in which the automated assistant 120 performs various actions. In some implementations, in a particular context, the vocabulary of invocation words that can be spoken to transition the automated assistant 120 from a default inactive state to a general listening state can be extended, at least temporarily (e.g., for a limited time until the context is no longer applicable).
[0066] For example, in FIG. 2, a first context-specific signal CS1 can transition the automated assistant 120 from a default inactive state to a first context-specific listening state "CSLS1". In CSLS1, the automated assistant 120 can listen for both a default invocation word ("DIW") and a first context-specific hotword ("C1 hotword"). If either is detected, the automated assistant 120 can transition to the general listening state as described above. Thus, in the first context-specific listening state, the vocabulary of hotwords that can transition the automated assistant 120 to the general listening state is extended to include both the default invocation word and the first context-specific hotword. Also, in some implementations, if sufficient time elapses while the automated assistant 120 is in the first context-specific listening state without detection of an activated hotword, a timeout ("TO") can return the automated assistant 120 to the default inactive state.
[0067] Additionally or alternatively, in some implementations, in a particular context, the automation assistant 120 may transition to either a general listening state that uses, for example, an extended vocabulary of hotwords, or a context-specific state in which one or more context-specific actions can be performed. For example, in FIG. 2, the automation assistant 120 may transition from a default inactive state to a second context-specific listening state "CSLS2" in response to a second context signal ("CS2"). In this second context-specific listening state, the automation assistant 120 may, for example, detect one or more default call words and / or, in some cases, one or more context-specific hotwords ("C2 hotwords A ") that effectively expand the vocabulary available to transition the automation assistant 120 to a general listening state, thereby transitioning to the general listening state.
[0068] Additionally or alternatively, the automation assistant 120 may transition, for example, in response to one or more additional second context-specific hotwords ("C2 hotwords B ") from the second context-specific state ("CSLS2") to one or more states in which one or more context-specific response actions ("second context-specific response actions") are performed. Exemplary response actions are described below. In some implementations, a particular second context-specific hotword may be mapped to a particular context-specific response action, but this is not required. Although not shown in FIG. 2 for clarity, after the execution of one or more of these second context-specific response actions, the automation assistant 120 may return to the default inactive state.
[0069] In some implementations, in a particular context, the automated assistant 120 may no longer listen for the default hotword. Instead, the automated assistant 120 may listen only for context-specific hotwords and execute response actions. For example, in FIG. 2, the automated assistant 120 may transition from a default inactive state to a context-specific listening state "CSLSM" in response to the Mth (where M is a positive integer) context signal ( "CS M "). In this situation, the automated assistant 120 may listen for the Mth context-specific hotword ( "CM hotword"). In response to detecting one or more Mth context-specific hotwords, the automated assistant 120 may execute one or more Mth response actions ( "Mth context-specific response actions").
[0070] In various implementations, the automated assistant 120 may activate context-specific hotwords in various ways. For example, referring to both FIGS. 1 and 2, in some implementations, upon transitioning to a particular context, the automated assistant 120 may, for example, detect from the dynamic hotword engine 128 one or more context-specific machine learning models or classifiers (e.g., neural networks, hidden Markov models, etc.) that have been pre-trained to detect the hotwords to be activated in a particular context (e.g., 1142, 1143,..., 114 N ) can be downloaded. For example, in a particular context, assume that the vocabulary for transitioning the automated assistant 120 from its default inactive state to a general listening state is extended to include the word "howdy". In various implementations, the automated assistant 120 can obtain, for example, from a database 129 available to the dynamic hotword engine 128, a classifier trained to generate an output indicating whether the word "howdy" has been detected. In various implementations, this classifier can be binary (e.g., output "1" if the hotword is detected and "0" otherwise), or can generate a probability. If the probability meets some confidence threshold, the hotword can be detected.
[0071] Additionally or alternatively, in some implementations, one or more on-device models 114 can take the form of a dynamic hotword classifier and / or a machine learning model (e.g., a neural network, a hidden Markov model, etc.) that is adjustable on the fly to generate one output for one or more predetermined phonemes and another output for other phonemes. Assume that the hotword "howdy" should be activated. In various implementations, the dynamic hotword classifier can be adjusted, for example, by changing one or more parameters and / or by providing a specific input either with or embedded in the audio data, to "listen" for the phonemes "how" and "dee". When those phonemes are detected in the audio input, the dynamic hotword classifier can generate an output that triggers the automated assistant 120 to perform response actions such as transitioning to a general listening state, performing some context-specific response action, etc. Other phonemes can generate an output that is ignored or not considered. Additionally or alternatively, the output can be generated only by the dynamic hotword classifier in response to the activated phonemes, and other phonemes may not generate any output at all.
[0072] Figures 3A and 3B show an example of how a human-computer dialogue session between user 101 and an instance of an automated assistant (not shown in FIGS. 3A-3B) can occur via the microphone and speaker of client computing device 306 (shown as a stand-alone interactive speaker, but this is not meant to be limiting), according to the implementations described herein. One or more aspects of the automated assistant 120 may be implemented on the computing device 306, and / or on one or more computing devices that are network communicating with the computing device 306. The client device 306 includes a proximity sensor 305, which, as described above, can take the form of, for example, a camera, a microphone, a PIR sensor, a wireless communication component, and the like.
[0073] In FIG. 3A, the client device 306 may determine, based on a signal generated by the proximity sensor 305, that the user 101 is located at a distance D1 from the client device 306. The distance D1 may be greater than a predetermined distance (e.g., a threshold). As a result, the automated assistant 120 may remain in a default listening state that can only be invoked using one or more of a default set of hotwords (at least audibly). In FIG. 3A, for example, the user 101 provides a natural language input of "Hey assistant, set a timer for five minutes". "Hey Assistant" may include a default hotword that together forms a valid "call phrase" that can be used to invoke the automated assistant 120. Accordingly, the automated assistant 120 replies with "OK. Timer starting...now" and starts a five-minute timer.
[0074] In contrast, in FIG. 3B, the client device 306 can detect, for example, based on a signal generated by the proximity sensor 305, that the user 101 is currently at a distance D2 from the client device. The distance D2 may be less than a predetermined distance (or threshold). As a result, the automated assistant 120 can transition from the default listening state of FIG. 3A to an extended listening state. In other words, the determination that the user 101 is at a distance D2 can constitute a context signal (e.g., "CS1" in FIG. 2) that causes the automated assistant 120 to transition to a context-specific state, in which additional or alternative hotwords are activated and available for use to transition the automated assistant 120 from the default listening state to the extended listening state.
[0075] While in the extended listening state, the user 101 can call the automated assistant 120 using one or more hotwords from an extended set of hotwords in addition to, or instead of, the default set of hotwords. In FIG. 3B, for example, the user 101 does not need to start the user's utterance with "Hey Assistant". Instead, the user simply says "set a timer for five minutes". Words such as "set" or phrases such as "set a timer for" may include hotwords that are part of an extended set of hotwords currently available for calling the automated assistant 120. As a result, the automated assistant 120 replies with "OK. Timer starting...now" and starts a five-minute timer.
[0076] In this example, the automation assistant 120 transitions from a default listening state to one context-specific state in which the automation assistant 120 can be invoked using an additional / alternative hotword. However, the automation assistant 120 can transition to additional or alternative states in response to such events. For example, in some implementations, when the user 101 is detected to be sufficiently close to the assistant device (e.g., D2), the automation assistant 120 can transition to another context-specific state. For example, the automation assistant 120 can transition to a context-specific state in which comprehensive speech recognition is not available (at least if the automation assistant 120 is not first invoked using the default hotword), but other context-specific commands are available.
[0077] For example, assume that the user 101 is preparing dinner in the kitchen according to an online recipe accessed by the automation assistant 120 that operates at least partially on the client device 306. Further assume that the client device 306 is located within the cooking area such that the user 101 executes the steps of the recipe within D2 of the client device 306. In some implementations, the fact that the user 101 is detected at D2 of the client device 306 can cause the automation assistant 120 to transition to a context-specific state in which at least some commands are available without first invoking the automation assistant 120 using the default hotword. For example, the user 101 can say "next" to proceed to the next step of the recipe, or say "back" or "repeat last step" to ask about previous steps such as "start a timer", "stop a timer", etc.
[0078] In some implementations, other actions may also be available in this context. For example, in some implementations, a plurality of assistant devices deployed across a household (e.g., forming an adjusted ecosystem of client devices) may each be associated with a particular room. The association with a particular room can activate different dynamic hotwords when someone gets close enough to that device. For example, when user 101 first sets up client device 306 in the kitchen, user 101 can specify client device 306 as being located within the kitchen, for example, by selecting the kitchen on a menu. This can activate the kitchen-specific state machine of automated assistant 120 that activates certain kitchen-centered hotwords when the user is in close proximity to the client device. Thus, when user 101 is within D2 of client device 306 in the kitchen, certain kitchen-related commands such as "set a timer", "go to the next step" to operate smart kitchen appliances, etc., can become active without user 101 having to first invoke automated assistant 120 using the default hotword.
[0079] In this proximity example, user 101 must be detected within D2 of the client device, but this does not mean it is limiting. In various implementations, it may be sufficient for user 101 to simply be detected by proximity sensor 305. For example, a PIR sensor can detect the presence of a user only when the user is relatively close.
[0080] In addition to, or instead of, proximity, in some implementations, the user's identification information and / or membership in a group can be determined and used to activate one or more dynamic hotwords. FIGS. 4A-4B illustrate one such example. Again, the user (101A in FIG. 4A, 101B in FIG. 4B) interacts with an automated assistant 120 that operates at least partially on the client device 406. The client device 406 includes a sensor 470 that generates signals that can be used to determine the user's identification information. Thus, the sensor 470 can be one or more cameras configured to generate visual data that can be analyzed using an object recognition process that can identify visual features such as face recognition processing, or uniform visual markers (e.g., a badge or a QR code (registered trademark) on a shirt) to identify the user (or that the user is a member of a group). Additionally or alternatively, the sensor 470 can be one or more microphones configured to perform speech recognition (sometimes called "speaker recognition") on one or more utterances made by the user 101. As yet another option, the sensor 470 can be a wireless receiver (e.g., Bluetooth, Wi-Fi, Zigbee, Z-wave, ultrasonic, RFID, etc.) that receives wireless signals from a device (not shown) carried by the user and analyzes the wireless signals to determine the user's identification information.
[0081] In FIG. 4A, user 101A is not identifiable. Thus, user 101A is required to use one or more default hotwords to invoke the automated assistant 120. In FIG. 4A, user 101A speaks, "Hey assistant, play 'We Wish You a Merry Christmas'". In response to "Hey assistant", the automated assistant 120 is invoked. The remainder of the utterance is processed using the pipeline described above to cause the automated assistant 120 to start generating the song on the client device 406.
[0082] In contrast, in FIG. 4B, user 101B is identified, for example, based on a signal generated by sensor 470. The act of identifying user 101B may constitute a context signal (e.g., "CS1" in FIG. 2) that transitions the automated assistant to a context-specific state, in which additional or alternative hotwords are activated and available to transition the automated assistant 120 from a default listening state to an extended listening state. For example, as shown in FIG. 4B, user 101B may speak something like "Play 'We Wish You a Merry Christmas'" without first invoking the automated assistant 120. Nevertheless, the automated assistant 120 starts playing the song. Other hotwords that may be active in such a context include, but are not limited to, "fast forward", "skip ahead <number seconds>", "stop", "pause", "rewind <time increment>)」, including "turn up / down the volume".
[0083] Additionally or alternatively, it may not be necessary to detect specific identification information of user 101B. In some implementations, it may be sufficient to recognize some visual attribute of the user to activate a specific dynamic hotword. For example, in FIG. 4B, user 101B is a doctor wearing clothing typically worn by doctors. This clothing (or a badge with an indicator, RFID, etc.) can be detected and used to determine that user 101B is a member of a group (e.g., medical personnel) for which a specific hotword should be activated.
[0084] FIG. 5 shows a situation where proximity and / or identity can be used. In FIG. 5, user 501 (shown hatched in white) holding client device 506 is in an area crowded with a plurality of other users (shown hatched in gray). Assume that the other people also have client devices (not shown) used to interact with automated assistant 120. If another nearby person utters the default hotword to call their automated assistant, it may inadvertently call automated assistant 120 on client device 506. Similarly, if user 501 utters the default hotword to call automated assistant 120 on client device 506, it may inadvertently call the automated assistant on other devices carried by other people.
[0085] Accordingly, in some implementations, the automation assistant 120 can detect the identification information of the user 501 and / or that the user 501 is within a predetermined proximity, and activate a custom set of dynamic hotwords necessary to invoke the automation assistant 120. In FIG. 5, for example, when the user 501 is within (and optionally identified within) a first proximity 515A of the client device 506, a first custom set of hotwords, which are unique to the user 501 (e.g., manually selected by the user 501), can be activated. In this way, the user 501 can utter these custom hotwords without fear of accidentally invoking the automation assistant on someone else's device. In some implementations, different proximity ranges can activate different dynamic hotwords. For example, in FIG. 5, if the user 501 is outside of proximity range 515A but is detected within another proximity range 515B, additional / alternative hotwords can be activated. The same applies to the third proximity range 515C, or any number of proximity ranges. In some implementations, if the user 501 is outside of proximity range 515C, only the default hotwords may be activated, or no hotwords may be activated.
[0086] Alternatively, in some implementations, upon recognizing the identification information of the user 501, the automation assistant 120 can implement speaker recognition to respond only to the voice of the user 501 and ignore other voices.
[0087] The implementations described in this specification have focused on causing the automated assistant 120 to perform various actions (e.g., searching for information, controlling media playback, stopping a timer, etc.) in response to context-specific hotwords, but this is not meant to be limiting. The techniques described in this specification can be extended to other use cases. For example, the techniques described in this specification may be applicable when a user desires to, e.g., enter into a form field on a search web page. In some implementations, if a search bar or other similar text input element exists within a web page, one or more additional context-specific hotwords may be activated. For example, when a user navigates an assistant-enabled device to a web page having a search bar, the hotword "search for" may be activated such that, e.g., the user can simply say "<desired topic> search for" without first having to invoke the automated assistant 120, and the user's utterance following "search for" is automatically transcribed into the search bar.
[0088] In various implementations, a transition to a particular context of a computing device may activate one or more context-specific gestures in addition to, or instead of, one or more context-specific hotwords. For example, assume that a user is detected in a particular proximity of an assistant device. In some implementations, one or more context-specific gestures may be activated. For example, detection of those gestures by the invocation module 113 may trigger a transition to a general listening state of the automated assistant 120 and / or may cause the automated assistant 120 to perform some context-specific response action.
[0089] FIG. 6 is a flowchart illustrating an exemplary method 600 according to the implementation forms disclosed in this specification. For convenience, the operations of the flowchart are described with reference to the system that executes the operations. This system may include various components of various computer systems, such as one or more components of the automated assistant 120. Further, although the operations of method 600 are shown in a particular order, this does not mean it is limiting. One or more operations may be reordered, omitted, or added.
[0090] In block 602, the system may operate the automated assistant 120 at least partially on a computing device (e.g., client devices 106, 306, 406, 506). For example, as described above, often the automated assistant 120 may be implemented partially on the client device 106 and partially on the cloud (e.g., cloud-based automated assistant component 119). In block 604, the system may monitor audio data captured by a microphone (e.g., 109) for one or more default hotwords. For example, the audio data (or other data indicating the audio data, such as an embedding) may be applied as input across one or more currently active invocation models 114 to generate an output. The output may indicate the detection of one or more of the default hotwords (block 606). In block 608, the system may transition the automated assistant 120 from a restricted hotword listening state (e.g., the default inactive state in FIG. 2) to an audio recognition state (e.g., the general listening state in FIG. 2).
[0091] In some implementations, in parallel with (or serially to) the operations of blocks 604-608, the system may monitor the state of the client device at block 610. For example, the system may monitor one or more context signals such as the detected presence of a user, the detected identification information of the user, the detection of user membership in a group, the detection of visual attributes, etc.
[0092] At block 612, if the system detects a context signal, at block 614, the system may transition the computing device to a given state. For example, the system may detect context signals such as the user being in sufficient proximity to the client device, the identification information of the user, etc. After the transition of block 614, at block 616, the system may monitor the audio data captured by the microphone for one or more context-specific hotwords in addition to, or instead of, one or more default hotwords monitored at block 604.
[0093] As described above, in some contexts, some hotwords can transition the automated assistant 120 to a general listening state, and other hotwords can cause the automated assistant 120 to perform context-specific response actions (e.g., stopping a timer, pausing media playback, etc.). Thus, at block 618, if the system detects a first one or more context hotwords (e.g., hotwords intended to cause the automated assistant 120 to perform context-specific tasks), at block 620, the system can perform one or more context-specific response actions or cause the automated assistant 120 to perform one or more context-specific response actions. On the other hand, if the first one or more contexts are not detected at block 618 but a second one or more context hotwords (e.g., hotwords generally intended to simply invoke the automated assistant 120) are detected at block 622, method 600 can return to block 606 where the automated assistant 120 is in a general listening state.
[0094] In some implementations, one or more timeouts can be used to ensure that the automated assistant 120 returns to a stable or default state when context-specific actions are not required. For example, if the first or second context-specific hotwords are not detected at blocks 618 and 622 respectively, at block 624, a determination can be made as to whether the timeout has expired (e.g., 10 seconds, 30 seconds, 1 minute, etc.). If the answer at block 624 is yes, method 600 can return to block 604 where the automated assistant 120 is transitioned to a default inactive state. However, if the answer at block 624 is no, in some implementations, method 600 can return to block 616, at which point the system can monitor for context-specific hotwords.
[0095] FIG. 7 shows an exemplary method 700 similar to method 600 for implementing selected aspects of the present disclosure according to various embodiments. For convenience, the operations of the flowchart are described with reference to a system that executes the operations. This system can include various components of various computer systems. Further, the operations of method 700 are shown in a particular order, but this is not meant to be limiting. One or more operations can be rearranged, omitted, or added.
[0096] In block 702, the system may execute an automated assistant in a default listening state. While in the default listening state, in block 704, the system may monitor audio data captured by one or more microphones for one or more of a default set of one or more hotwords. Detection of one or more of the default set of hotwords can trigger a transition from the default listening state to an audio recognition state (e.g., block 608 of FIG. 6).
[0097] In block 706, the system may detect one or more sensor signals generated by one or more hardware sensors integral with one or more of a computing device, such as a microphone, camera, PIR sensor, wireless components, etc. In block 708, the system may analyze one or more sensor signals to determine an attribute of the user. This attribute of the user can be, for example, a presence attribute regarding the user himself or herself, or the user's physical proximity to one or more assistant devices. Non-limiting examples of attributes that can be determined in block 708 include, but are not limited to, the user being within a predetermined proximity of the assistant device, the user's identification information, the user's membership in a group, the user being detected by a proximity sensor (in a situation where distance is not considered), visual attributes of the user such as a uniform, badge, small adornment, etc.
[0098] In block 710, the system can transition the automated assistant from the default listening state to an extended listening state based on the analysis. While in the extended listening state, in block 712, the system can monitor audio data captured by one or more microphones for one or more of an extended set of one or more hotwords. Detection of one or more of the extended set of hotwords can trigger the automated assistant to perform a response action without requiring detection of one or more of the default set of hotwords. One or more of the extended set of hotwords may not be within the default set.
[0099] In situations where the particular implementations discussed in this specification may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, as well as the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether the information is collected, whether the personal information is stored, whether the personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed in this specification collect, store, and / or use personal information about a user only when they have received explicit permission from the relevant user to do so.
[0100] For example, a user is provided with control over whether a program or function collects user information about that particular user or other users related to the program or function. Each user from whom personal information is collected is presented with one or more options that enable control of information collection related to that user, in order to provide permission or approval regarding whether the information is collected and which portions of the information should be collected. For example, a user may be provided with one or more such control options via a communication network. Additionally, certain data may be processed in one or more ways before being stored or used such that information that can identify an individual is removed. As an example, a user's identification information may be processed such that it cannot be used to determine information that can identify an individual. As another example, a user's geographical location may be generalized to a larger area such that the user's specific location cannot be identified.
[0101] FIG. 8 is a block diagram of an exemplary computing device 810 that may optionally be utilized to execute one or more aspects of the techniques described herein. In some implementations, one or more of a client computing device, a user control resource engine 134, and / or other components may comprise one or more components of the exemplary computing device 810.
[0102] Computing device 810 typically includes at least one processor 814 that communicates with several peripheral devices via a bus subsystem 812. These peripheral devices can include, for example, a storage subsystem 824 that includes a memory subsystem 825 and a file storage subsystem 826, a user interface output device 820, a user interface input device 822, and a network interface subsystem 816. The input and output devices enable user interaction with the computing device 810. The network interface subsystem 816 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.
[0103] The user interface input device 822 can include a keyboard, a mouse, a trackball, a pointing device such as a touchpad or a graphics tablet, a scanner, a touch screen incorporated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 810 or a communication network.
[0104] The user interface output device 820 can include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 810 to a user or another machine or computing device.
[0105] The memory subsystem 824 stores programming structures and data structures that provide some or all of the functionality of some of the modules described herein. For example, the memory subsystem 824 can include logic for executing selected aspects of the methods of FIGS. 6-7 and for implementing the various components shown in FIG. 1.
[0106] These software modules are generally executed by the processor 814 alone or in combination with other processors. The memory 825 used in the memory subsystem 824 can include several memories including a main random access memory (RAM) 830 for storing instructions and data during program execution and a read-only memory (ROM) 832 in which fixed instructions are stored. The file storage subsystem 826 can provide permanent storage for program files and data files and can include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular implementation can be stored by the file storage subsystem 826 within the memory subsystem 824 or within another machine accessible by the processor 814.
[0107] The bus subsystem 812 provides a mechanism for enabling various components and subsystems of the computing device 810 to communicate with each other as intended. Although the bus subsystem 812 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0108] The computing device 810 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the constantly changing nature of computers and networks, the description of the computing device 810 shown in FIG. 8 is intended only as a specific example for the purpose of explaining some implementations. Many other configurations of the computing device 810 with more or fewer components than the computing device shown in FIG. 8 are possible.
[0109] Although several implementations are described and illustrated herein, various other means and / or structures may be utilized for performing the functions and / or for obtaining one or more of the results and / or advantages described herein, and each such change and / or modification is to be regarded as being within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be illustrative, and the actual parameters, dimensions, materials, and / or configurations will depend upon the particular application for which the teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the foregoing implementations are presented by way of example only, and it is to be understood that within the scope of the appended claims and their equivalents, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each of the individual features, systems, articles, materials, kits, and / or methods described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure provided such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
Description of Reference Numerals
[0110] 101 User 101A User 101B User 105 Proximity Sensor 106 Client Computing Device, Client Device 107 Visual Sensor, Camera 108 Automation Assistant Client 109 Microphone 110 Voice Capture Module 111 Image Capture Module 112 Visual Cue Module 1121 Visual Cue Module 1122 Visual Cue Module 113 Call Module 114 Local and / or Wide Area Network, On-Device Model Database, Call Model, On-Device Model 1141 On-Device Call Model, Default On-Device Call Model, Default Call Model 1142~114 N On-Device Call Model, Context Call Model, Context-Specific Machine Learning Model or Classifier 116 Cloud-Based Text-to-Speech (「TTS」) Module, Cloud-Based TTS Module, TTS Module 117 Cloud-Based STT Module, STT Module 119 Cloud-Based Automation Assistant Component 120 Automation Assistant 122 Natural Language Processor 124 Execution Module 126 Natural Language Generator 128 Dynamic Hotword Engine 129 Database 130 Third-Party Computing Service, Third-Party Service 134 User Control Resource Engine 135 Intent Integrator 150 Search Module 305 Proximity Sensor 306 Client Computing Device, Computing Device, Client Device 406 Client Device 470 Sensor 501 User 506 Client Device 515A Proximity, Proximity Range 515B Proximity Range 515C Proximity Range 810 Computing Device 812 Bus Subsystem 814 Processor 816 Network Interface Subsystem 820 User Interface Output Device 822 User Interface Input Device 824 Memory Subsystem 825 Memory Subsystem, Memory 826 File Memory Subsystem 830 Main Random Access Memory (RAM) 832 Read Only Memory (ROM)< / time> < / topping> < / topping> < / artist> < / artist> < / artist> < / artist>
Claims
1. A method performed by one or more processors, comprising: executing an automated assistant in a default listening state, wherein the automated assistant is executed at least partially on one or more computing devices operated by a user; monitoring, using a default on-device invocation model, audio data captured by one or more microphones for one or more of a default set of one or more hotwords while in the default listening state; wherein detection of one or more of the hotwords in the default set triggers a transition of the automated assistant from the default listening state to an audio recognition state in which some or additional audio data is processed using a speech-to-text (“STT”) process; detecting one or more sensor signals generated by one or more hardware sensors integral with one or more of the computing devices; analyzing the one or more sensor signals to determine an attribute of the user; activating a contextual on-device invocation model trained to detect an extended set of one or more hotwords specific to the user based on the analysis; monitoring, for one or more of the hotwords in the extended set, the audio data captured by one or more of the microphones, where one or more of the hotwords in the extended set are based on the contextual on-device invocation model and not within the default set; wherein detection of one or more of the hotwords in the extended set triggers the automated assistant to transition the automated assistant from the default listening state to the audio recognition state without requiring detection of one or more of the hotwords in the default set.
2. A method performed by one or more processors, comprising: executing an automated assistant in a default listening state, steps in which the automated assistant is at least partially executed on one or more computing devices operated by a user; while in the default listening state, monitoring audio data captured by one or more microphones for one or more of a default set of one or more hotwords, wherein detection of one or more of the hotwords in the default set triggers a transition of the automated assistant from the default listening state to an audio recognition state; detecting one or more sensor signals generated by one or more hardware sensors integrated with one or more of the computing devices; analyzing the one or more sensor signals to determine an attribute of the user; based on the analysis, transitioning the automated assistant from the default listening state to an extended listening state, and while in the extended listening state, monitoring the audio data captured by one or more of the microphones for one or more of an extended set of one or more hotwords specific to the user, wherein detection of one or more of the hotwords in the extended set triggers the automated assistant to perform a response action without performing a speech-to-text process, the response action being performed without detection of one or more of the hotwords in the default set, and one or more of the hotwords in the extended set not being within the default set; A method comprising the steps of.
3. The one or more hardware sensors include proximity sensors, The attribute of the user is, the user being detected by the proximity sensor, or the user being detected by the proximity sensor within a predetermined distance of one or more of the computing devices The method according to claim 1 or 2, comprising the steps of.
4. The method according to any one of claims 1 to 3, wherein the one or more hardware sensors include a camera.
5. The attribute of the user is, the user is detected by the camera; or the user is detected by the camera within a predetermined distance of one or more of the computing devices; 5. The method of claim 4, comprising:
6. The method of claim 5 , wherein the analysis includes a facial recognition process.
7. The method of claim 6 , wherein the attributes of the user include an identity of the user.
8. The method of claim 4 , wherein the attributes of the user include the user's membership in a group.
9. The method of claim 1 , wherein the one or more hardware sensors include one or more of the microphones.
10. 10. The method of claim 9, wherein the one or more attributes include the user being audibly detected based on the audio data captured by one or more of the microphones.
11. The method of claim 9 , wherein the analysis includes a voice recognition process.
12. The method of claim 11 , wherein the attributes of the user include an identity of the user.
13. The method of claim 9 , wherein the attributes of the user include the user's membership in a group.
14. the response action includes transitioning the automated assistant to the speech recognition state; and / or 14. The method of claim 2 or any one of claims 3 to 13 that cites claim 2, wherein the response action includes the automated assistant performing a task requested by the user using one or more of the hotwords in the expansion set.
15. 15. A system comprising one or more processors and a memory operatively coupled to the one or more processors, the memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to perform a method according to any one of claims 1 to 14.
16. At least one non-transitory computer-readable recording medium storing instructions, wherein execution of the instructions by one or more processors causes the one or more processors to perform the method according to any one of claims 1 to 14. [
17. ] A method performed by one or more processors, comprising: Executing an automated assistant in a first listening state, the automated assistant being at least partially executed on one or more computing devices operated by a user; While in the first listening state, Monitoring audio data captured by one or more microphones for one or more of a default set of one or more hotwords, detection of one or more of the hotwords in the default set triggering a transition of the automated assistant from the first listening state to an audio recognition state, in which the audio data or a portion of additional audio data is processed using a speech-to-text (STT) process; Detecting one or more sensor signals generated by one or more hardware sensors integral with one or more of the computing devices; Analyzing the one or more sensor signals to determine an attribute of the user; Based on the analysis, transitioning the automated assistant from the first listening state to a second listening state; While in the second listening state, Stopping monitoring of the audio data captured by one or more of the microphones for one or more of the hotwords in the default set; Instead of one or more of the hotwords of the default set, for one or more of a custom set of one or more hotwords specific to the user, monitoring the audio data captured by one or more of the microphones, wherein detection of one or more of the hotwords of the custom set triggers the automated assistant to perform a response action without performing speech-to-text processing, steps and including, Method. **Claim 18** The one or more hardware sensors include a proximity sensor, and the attributes of the user are the user being detected by the proximity sensor, or the user being detected by the proximity sensor within a predetermined distance of one or more of the computing devices, including, The method according to claim 17. **Claim 19** The one or more hardware sensors include a camera, The attributes of the user are the user being detected by the camera, or the user being detected by the camera within a predetermined distance of one or more of the computing devices, including, The method according to claim 17. **Claim 20** The analysis includes face recognition processing, The method according to claim 19. **Claim 21** The attributes of the user include the identification information of the user or the membership of the user in a group, The method according to claim 20. **Claim 22** The response action includes transitioning the automated assistant from the second listening state to the speech recognition state, The method according to claim 17. **Claim 23** The one or more hardware sensors include one or more of the microphones, The method according to any one of claims 17. **Claim 24** The one or more attributes include the user being auditorily detected based on the audio data captured by one or more of the microphones, The method according to claim 23. **Claim 25** The analysis includes speech recognition processing, The method according to claim 23. **Claim 26** The attributes of the user include the identification information of the user or the membership of the user in a group, The method according to claim 25. **Claim 27** A method performed by one or more processors, comprising the step of executing an automated assistant in a first listening state, wherein the automated assistant is at least partially executed on one or more computing devices operated by a user; while in the first listening state, monitoring audio data captured by one or more microphones for one or more of a default set of one or more hotwords, wherein detection of one or more of the hotwords in the default set triggers a transition of the automated assistant from the first listening state to an audio recognition state, and in the audio recognition state, a portion of the audio data or additional audio data is processed using a speech-to-text (STT) process; detecting one or more sensor signals generated by one or more hardware sensors integral with one or more of the computing devices; analyzing the one or more sensor signals to determine an attribute of the user; transitioning the automated assistant from the first listening state to a second listening state based on the analysis; while in the second listening state, raising a confidence threshold required to invoke the automated assistant using one or more of the hotwords in the default set; monitoring the audio data captured by one or more of the microphones for one or more of a custom set of one or more hotwords, wherein detection of one or more of the hotwords in the custom set triggers the automated assistant to perform a response action. A method. **Claim 28** The method according to claim 27, further comprising, in the second listening state, activating a confidence threshold required to invoke the automated assistant using one or more of the hotwords in the custom set. The method according to claim 27. **Claim 29** In the second listening state, further comprising the step of lowering a confidence threshold required for calling the automation assistant by using one or more of the hotwords of the custom set. The method according to claim 27.
30. The one or more hardware sensors include a proximity sensor, and the attribute of the user is the user being detected by the proximity sensor, or the user being detected by the proximity sensor within a predetermined distance of one or more of the computing devices. The method according to claim 27.
31. The one or more hardware sensors include a camera, and the attribute of the user is the user being detected by the camera, or the user being detected by the camera within a predetermined distance of one or more of the computing devices. The method according to claim 27.
32. The analysis includes a face recognition process. The method according to claim 31.
33. The attribute of the user includes identification information of the user or the user's membership in a group. The method according to claim 32.
34. The response action includes transitioning the automation assistant from the second listening state to the speech recognition state. The method according to claim 27.
Citation Information
Patent Citations
Voice recognizer, control method therefor, and program
JP2009128723A
Information processing device, information processing method and program
JP2017144521A
Audio recognizing device
WO2007066433A1
Contextual hotwords
WO2018125292A1