Detecting and preventing commands in media that can trigger another automated assistant

By detecting and blocking commands in the media that may trigger another automated assistant, and by using machine learning models to identify device status and audio clips, the problem of unintentional activation of automated assistants in multi-device environments is solved, improving resource utilization efficiency and user experience.

CN115668362BActive Publication Date: 2025-11-21GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180035649.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-01
Filing Date
2021-11-30
Publication Date
2025-11-21
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

In environments where multiple automated assistant devices coexist, automated assistants may be unintentionally activated, leading to negative user experiences, wasted resources, and unnecessary actions.

Method used

By detecting and blocking commands in the media that might trigger another automated assistant, machine learning models are used to identify device status and audio clips, predict whether hot words will be triggered, and thus prevent unnecessary activation.

Benefits of technology

It reduces accidental activation of automated assistants, avoids resource waste and negative user experience, and improves the resource utilization efficiency of devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668362B_ABST
    Figure CN115668362B_ABST
Patent Text Reader

Abstract

Described herein are techniques for detecting and preventing media that can trigger another automated assistant. One method includes determining, for each of a plurality of automated assistant devices that each execute at least one automated assistant in an environment, an active state capability of the automated assistant device; initiating, by an automated assistant, playback of digital media; in response to initiating playback, processing the digital media based on the active state capability of at least one of the plurality of automated assistant devices to identify an audio segment in the digital media that, when played back, is predicted to trigger activation of at least one automated assistant executing on at least one of the plurality of automated assistant devices in the environment; and in response to identifying the audio segment in the digital media, modifying the digital media to prevent the activation of the at least one automated assistant.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Humans can engage in human-computer dialogue with interactive software applications, referred to herein as “automated assistants” (also known as “digital agents,” “interactive personal assistants,” “intelligent personal assistants,” “assistant applications,” “conversational agents,” etc.). For example, humans (who may be referred to as “users” when interacting with an automated assistant) can provide commands and / or requests to the automated assistant using spoken natural language input (i.e., speech), by providing text (e.g., typing) natural language input, and / or by touch and / or free physical movements of speech (e.g., hand gestures, eye gaze, facial movements, etc.), which in some cases can be converted to text and then processed. The automated assistant responds to requests by providing responsive user interface outputs (e.g., auditory and / or visual user interface outputs), controlling one or more intelligent devices, and / or controlling one or more functions of the device implementing the automated assistant (e.g., controlling other functions of the device).

[0002] As mentioned above, many automated assistants are configured to interact via spoken words. To protect user privacy and / or conserve resources, automated assistants avoid executing one or more automated assistant functions based on all spoken words appearing in audio data detected via the microphone of the client device implementing (at least partially) the automated assistant. Instead, some processing based on spoken words occurs only in response to the determination that certain conditions have occurred.

[0003] For example, many client devices that include automated assistants and / or interface with automated assistants include hot word detection models. When the microphone of such a client device is not deactivated, the client device can use the hot word detection model to continuously process audio data detected via the microphone to generate predictive output indicating whether one or more hot words (including multi-word phrases) have appeared, such as “Hey Assistant,” “OK Assistant,” and / or “Assistant.” When the predictive output indicates the appearance of a hot word, any audio data following within a threshold time period (optionally determined to include voice activity) can be processed by one or more on-device and / or remote automated assistant components, such as speech recognition components, voice activity detection components, etc. Furthermore, the recognized text (from the speech recognition component) can be processed using a natural language understanding engine and / or actions can be performed based on the output of the natural language understanding engine. For example, actions can include generating and providing responses and / or controlling one or more applications and / or smart devices. Other hot words (such as "No", "Stop", "Cancel", "Volume Up", "Volume Down", "Next Track", "Previous Track", etc.) can be mapped to various commands, and the mapped command can be processed by the client device when the predicted output indicates that one of these hot words has appeared. However, when the predicted output does not indicate a hot word, the corresponding audio data is discarded without further processing, thus saving resources and protecting user privacy.

[0004] Some environments may include multiple client devices, each including one or more automation assistants. These automation assistants may be different (e.g., automation assistants from multiple providers) and / or multiple instances of the same automation assistant. In some cases, multiple client devices may be placed close enough that they each receive audio data via one or more microphones, capturing the user's spoken words, and each may be able to respond to queries included in the spoken words.

[0005] The aforementioned and / or other machine learning models (e.g., the additional machine learning models described below) perform well in many cases, and their predictive outputs determine whether an automated assistant function is activated. However, in certain situations, such as when multiple automated assistant devices (client devices), each including one or more automated assistants, are located in an environment, audio clips of digital media (e.g., podcasts, movies, TV shows, etc.) played by an automated assistant included on a first automated assistant device in the environment can be processed by an automated assistant included on a second automated assistant device in the environment (or by multiple other automated assistants included on multiple other automated assistant devices in the environment). In cases where the audio data processed by the second automated assistant (or by multiple other automated assistants) includes audio clips of digital media played by the first automated assistant, the hot word detection model of the second automated assistant (or multiple other automated assistants) can detect one or more hot words or words with similar pronunciations from the audio clips when one or more hot words or words with similar pronunciations appear in the audio clips. This can be particularly likely to occur when the number of client devices including automated assistants or the density of client devices including automated assistants in an environment (e.g., a home) increases, or when the number of active hot words (including warm words) increases. Then, a second automated assistant (or multiple other automated assistants) can process any audio from the media content and respond to requests within the media content that follow one or more hot words or words with similar pronunciations within a threshold time period. Alternatively, the automated assistant can perform one or more actions corresponding to the detected hot words (e.g., increasing or decreasing the audio volume).

[0006] Activation of a second (or multiple) automated assistant based on one or more hot words or similar-sounding words appearing in audio clips of digital media played by a first automated assistant may be unintentional on the part of the user and could result in a negative user experience (e.g., the second automated assistant may respond and / or perform actions that the user does not want to perform, even when not requested by the user). Unintentional activation of automated assistants may waste network and / or computing resources and potentially force humans to request the automated assistant to retract unwanted actions (e.g., a user may issue a "lower volume" command to counteract an unexpected "increase volume" command from the media content that triggers the automated assistant to increase the volume level). Summary of the Invention

[0007] Some implementations disclosed herein relate to improving the performance of machine learning models by detecting and blocking commands in media that could trigger another automated assistant. As described in more detail herein, such machine learning models may, for example, include hot word detection models and / or other machine learning models. Various implementations include automated assistants that detect other client devices performing other automated assistants and proactively take action to prevent other automated assistants running on other client devices from erroneously triggering digital media played by the automated assistant. In some implementations, in response to detecting an audio segment of media content that might trigger a nearby client device performing an automated assistant, the system automatically blocks the triggering of the audio segment.

[0008] In some implementations, the system determines which devices are currently located around the media playback device and their current state and active capabilities. For example, a nearby device might be listening for the hot phrase "Hey Assistant 1," while another device might be listening for "OK ​​Assistant 2." Alternatively, one of the devices might be engaged in an active conversation with the user and therefore accept any voice input.

[0009] In some implementations, the system discovers other devices in the environment and maintains awareness of the current state of each other device in the environment via an application programming interface (API) (e.g., if other devices include other instances of the same automation assistant, e.g., from the same provider) and / or based on analysis of previously overheard interactions and taking into account the types or capabilities of nearby devices (e.g., if other devices include different automation assistants, e.g., from different providers).

[0010] In some implementations, in response to detecting an audio segment of media content that may trigger a nearby client device executing an automation assistant, the system prevents unnecessary automation assistant triggering while the media is playing, thereby automatically blocking triggering from the audio segment. In some implementations, by automatically blocking triggering, the system can avoid negative user experience and unnecessary bandwidth and power consumption by reducing unintended activation of the automation assistant.

[0011] In various implementations, a method implemented by one or more processors may include: determining the active capabilities of each of a plurality of automation assistant devices in an environment, each executing at least one automation assistant, via a client device; initiating playback of digital media via an automation assistant executing at least partially on the client device; in response to initiating playback, processing the digital media based on the active capabilities of at least one of the plurality of automation assistant devices to identify an audio segment in the digital media, which, during playback, is expected to trigger activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment; and in response to identifying the audio segment in the digital media, modifying the digital media to prevent activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment.

[0012] In some embodiments, the method may further include: using a wireless communication protocol, a client device to discover each of a plurality of automation assistant devices in an environment. In some embodiments, the method may further include: using calibrated sound or wireless / radio signal strength analysis, a client device to determine the proximity of each of the plurality of automation assistant devices to the client device. In some embodiments, the client device processing digital media to identify audio segments is further in response to determining that the volume level associated with the playback of the digital media meets a threshold determined based on the proximity of each of the plurality of automation assistant devices to the client device.

[0013] In some implementations, the ability to determine the activity state of an automated assistant device may include one or more of the following: determining whether the automated assistant device is performing hot word detection; determining whether the automated assistant device is performing open automatic speech recognition; and determining whether the automated assistant device is listening to specific speech.

[0014] In some implementations, the client device processing digital media to identify audio segments is further based on the output of a speaker recognition model and the ability to determine the active state of one or more automated assistant devices among a plurality of automated assistant devices, including performing hot word detection or performing open-ended automatic speech recognition. In some implementations, the client device processing digital media to identify audio segments includes: processing the digital media using a hot word detection model to detect potential triggers in the digital media. In some implementations, the client device modifying the digital media includes: inserting a digital watermark into the audio track of the digital media.

[0015] In some additional or alternative embodiments, the computer program product may include one or more computer-readable storage media having program instructions commonly stored on the one or more computer-readable storage media. The program instructions may be executable to: determine the active capability of each of a plurality of automation assistant devices in an environment, each executing at least one automation assistant, via a client device; initiate playback of digital media via an automation assistant executing at least partially on the client device; in response to initiating playback, process the digital media via the client device based on the active capability of at least one of the plurality of automation assistant devices to identify an audio segment in the digital media, which, during playback, is expected to trigger activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment; and in response to identifying the audio segment in the digital media, communicate via the client device with at least one of the plurality of automation assistant devices in the environment to prevent activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment during playback of the audio segment.

[0016] In some implementations, communicating with at least one of a plurality of automated assistant devices in the environment to prevent activation of at least one automated assistant executing on at least one of the plurality of automated assistant devices in the environment during playback of an audio segment includes: providing at least a portion of the audio segment or a fingerprint based on the audio segment to at least one of the plurality of automated assistant devices; or providing instructions to at least one of the plurality of automated assistant devices to stop listening at a time corresponding to the audio segment.

[0017] In some additional or alternative embodiments, the system may include a processor, a computer-readable storage device, one or more computer-readable storage media, and program instructions collectively stored on the one or more computer-readable storage media. The program instructions may be executable to: determine the active capability of each of a plurality of automation assistant devices, each executing at least one automation assistant in the environment, via a client device; initiate playback of digital media via an automation assistant executing at least partially on the client device; in response to initiating playback, process the digital media via the client device based on the active capability of at least one of the plurality of automation assistant devices to identify an audio segment in the digital media, which, during playback, is expected to trigger activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment; and in response to identifying the audio segment in the digital media, modify the digital media via the client device to prevent activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment.

[0018] By utilizing one or more of the techniques described herein, automated assistants can improve performance by detecting and blocking commands in the media to avoid the accidental activation of another automated assistant that could waste network and / or computing resources and potentially force a human to request the automated assistant to rescind an unwanted action.

[0019] The above description is provided as an overview of some embodiments of this disclosure. Further descriptions of those and other embodiments are given below in more detail.

[0020] Various implementations may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), and / or tensor processing unit (TPU)) to perform methods, such as one or more methods described herein. Other implementations may include an automation assistant client device (e.g., a client device including at least an automation assistant interface for interacting with cloud-based automation assistant components), the automation assistant client device including a processor operable to execute the stored instructions to perform methods, such as one or more methods performed herein. Other implementations may include a system of one or more servers including one or more processors operable to execute the stored instructions to perform methods, such as one or more methods described herein. Attached Figure Description

[0021] Figure 1A and Figure 1B Example process flows illustrating various aspects of this disclosure are described according to various embodiments.

[0022] Figure 2 A block diagram depicts an example environment, which includes components from... Figure 1A and Figure 1B The various components, and the implementation methods disclosed herein can be implemented in this example environment.

[0023] Figure 3 A flowchart illustrating example methods for detecting and blocking commands in the media that could trigger another automation assistant, according to various implementations, is provided.

[0024] Figure 4 A flowchart illustrating example methods for detecting and blocking commands in the media that could trigger another automation assistant, according to various implementations, is provided.

[0025] Figure 5 An example architecture for a computing device is described. Detailed Implementation

[0026] Figure 1A and Figure 1B An example process flow illustrating various aspects of this disclosure is depicted. Client device 110, such as... Figure 1A As shown, and including Figure 1A The component contained within the frame of client device 110 is represented by [reference]. Machine learning engine 122A may receive audio data 101 corresponding to spoken utterance detected via one or more microphones of client device 110 and / or other sensor data 102 corresponding to free physical movements of the utterance detected via one or more non-microphone sensor components of client device 110 (e.g., hand gestures and / or movements, body gestures and / or body movements, eye gaze, facial movements, mouth movements, etc.). One or more non-microphone sensors may include a camera or other visual sensor, proximity sensor, pressure sensor, accelerometer, magnetometer, and / or other sensors. Machine learning engine 122A processes audio data 101 and / or other sensor data 102 using machine learning model 152A to generate a predictive output 103. As described herein, machine learning engine 122A may be a hot word detection engine 122B or an alternative engine, such as a speech activity detector (VAD) engine, an endpoint detector engine, an automatic speech recognition (ASR) engine, and / or other engines.

[0027] In some implementations, when the machine learning engine 122A generates a prediction output 103, the prediction output 103 may be locally stored on a client device in on-device storage 111 and optionally associated with corresponding audio data 101 and / or other sensor data 102. In some versions of these implementations, the prediction output may be retrieved by the gradient engine 126 for later use in generating gradient 106, such as when one or more conditions described herein are met. For example, the on-device storage 111 may include read-only memory (ROM) and / or random access memory (RAM). In other implementations, the prediction output 103 may be provided to the gradient engine 126 in real time.

[0028] Based on whether the predicted output 103 meets the threshold in box 182, the client device 110 decides whether to initiate the currently dormant automated assistant function (e.g., Figure 2 The automation assistant 110 may activate an currently dormant automation assistant function (295) and / or disable an currently active automation assistant function using the assistant activation engine 124. The automation assistant function may include: speech recognition for generating recognized text, natural language understanding (NLU) for generating natural language understanding (NLU) output, generating a response based on the recognized text and / or NLU output, transmitting audio data to a remote server, transmitting recognized text to a remote server, and / or directly triggering one or more actions (e.g., common tasks, such as changing device volume) in response to audio data 101. For example, suppose the predicted output 103 is a probability (e.g., 0.80 or 0.90), and the threshold in box 182 is a threshold probability (e.g., 0.85). If the client device 110 determines that the predicted output 103 (e.g., 0.90) satisfies the threshold in box 182 (e.g., 0.85), the assistant activation engine 124 may activate a currently dormant automation assistant function.

[0029] In some implementations, and as such Figure 1B As shown, the machine learning engine 122A can be a hot word detection engine 122B. It is evident that various automated assistant functions, such as the on-device voice recognizer 142, the on-device NLU engine 144, and / or the on-device execution engine 146, are currently dormant (i.e., as shown by the dashed lines). Further, assuming that the predicted output 103 generated using the hot word detection model 152B and based on the audio data 101 satisfies the threshold in box 182, and the voice activity detector 128 detects user voice directed at the client device 110.

[0030] In some versions of these implementations, the assistant activation engine 124 activates the on-device voice recognizer 142, the on-device NLU engine 144, and / or the on-device fulfillment engine 146 as currently dormant automated assistant functions. For example, the on-device voice recognizer 142 can process audio data 101 of spoken words, including the hotword "OK Assistant" and additional commands and / or phrases following the hotword "OK Assistant," using the on-device voice recognition model 142A to generate recognized text 143A; the on-device NLU engine 144 can process the recognized text 143A using the on-device NLU model 144A to generate NLU data 145A; the on-device fulfillment engine 146 can process the NLU data 145A using the on-device fulfillment model 146A to generate fulfillment data 147A; and the client device 110 can use the fulfillment data 147A in the execution 150 of one or more actions in response to the audio data 101.

[0031] In other versions of these implementations, the assistant activation engine 124 activates only the on-device fulfillment engine 146, without activating the on-device voice recognizer 142 and the on-device NLU engine 144, to handle various commands such as "No," "Stop," "Cancel," "Volume Up," "Volume Down," "Next Track," "Previous Track," and / or other commands that do not require the on-device voice recognizer 142 and the on-device NLU engine 144 to process. For example, the on-device fulfillment engine 146 processes audio data 101 using the on-device fulfillment model 146A to generate fulfillment data 147A, and the client device 110 may use fulfillment data 147A in the execution 150 of one or more actions in response to the audio data 101. Furthermore, in versions of these implementations, the assistant activation engine 124 may initially activate currently dormant automation functions to verify that the audio data 101 includes the hotword "OK Assistant" by initially activating only the on-device voice recognizer 142, thereby confirming that the decision made in box 182 is correct (e.g., the audio data 101 actually includes the hotword "OK Assistant"). And / or the assistant activation engine 124 may transmit the audio data 101 to one or more servers (e.g., remote server 160) to verify that the decision made in box 182 is correct (e.g., the audio data 101 actually includes the hotword "OK Assistant").

[0032] Back Figure 1AIf client device 110 determines that the predicted output 103 (e.g., 0.80) does not meet the threshold (e.g., 0.85) in box 182, then assistant activation engine 124 may avoid initiating currently dormant automated assistant functions and / or disable any currently active automated assistant functions. Further, if client device 110 determines that the predicted output 103 (e.g., 0.80) does not meet the threshold (e.g., 0.85) in box 182, then client device 110 may determine whether further user interface input has been received in box 184. For example, further user interface input may include additional spoken words including hot words, additional spoken words acting as proxies for hot words with free physical movement, actuation of an explicit automated assistant invocation button (e.g., a hardware or software button), sensing of a "squeeze" by client device 110 (e.g., when client device 110 is squeezed with at least a threshold amount of force to invoke the automated assistant), and / or other explicit automated assistant invocations. If client device 110 determines that no further user interface input has been received in box 184, then client device 110 may terminate in box 190.

[0033] However, if client device 110 determines that further user interface input has been received in box 184, the system can then determine whether the further user interface input received in box 184 includes the correction in box 186, which contradicts the decision made in box 182. If client device 110 determines that the further user interface input received in box 184 does not include the correction in box 186, then client device 110 can stop identifying the correction and end in box 190. However, if client device 110 determines that the further user interface input received in box 184 includes the correction in box 186, which contradicts the initial decision made in box 182, then client device 110 can determine the place name real-time output 105.

[0034] In some implementations, gradient engine 126 may generate gradient 106 based on the predicted output 103 to the ground reality output 105. For example, gradient engine 126 may generate gradient 106 based on comparing the predicted output 103 with the ground reality output 105. In some versions of these implementations, client device 110 locally stores the predicted output 103 and the corresponding ground reality output 105 in on-device storage 111, and gradient engine 126 retrieves the predicted output 103 and the corresponding ground reality output 105 to generate gradient 106 when one or more conditions are met. For example, one or more conditions may include: the client device is charging, the client device has at least a threshold charging state, the temperature of the client device (based on one or more on-device temperature sensors) is less than a threshold, and / or the client device is not being held by a user. In other versions of these implementations, client device 110 provides the predicted output 103 and the ground reality output 105 to gradient engine 126 in real time, and gradient engine 126 generates gradient 106 in real time.

[0035] Furthermore, gradient engine 126 can provide the generated gradient 106 to on-device machine learning training engine 132A. Upon receiving gradient 106, on-device machine learning training engine 132A uses gradient 106 to update on-device machine learning model 152A. For example, on-device machine learning training engine 132A can utilize backpropagation and / or other techniques to update on-device machine learning model 152A. Notably, in some embodiments, on-device machine learning training engine 132A can utilize batch processing techniques to update on-device machine learning model 152A based on gradients and additional gradients determined locally at client device 110 based on additional corrections.

[0036] Furthermore, client device 110 can transmit the generated gradient 106 to remote system 160. When remote system 160 receives gradient 106, its remote training engine 162 updates the global weights of global hot word model 152A1 using gradient 106 and additional gradient 107 from additional client device 170. The additional gradients 107 from additional client device 170 can each be generated based on the same or similar techniques described above for gradient 106 (but according to locally identified failed hot word attempts, which are specific to these client devices).

[0037] In response to the satisfaction of one or more conditions, the updated distribution engine 164 may provide updated global weights and / or the updated global hot word model itself to client device 110 and / or other devices, as shown in 108. For example, one or more conditions may include the duration and / or quantity of training thresholds since the last provision of updated weights and / or the updated speech recognition model. For example, one or more conditions may additionally or alternatively include a measurement improvement in the updated speech recognition model and / or the elapsed duration of thresholds since the last provision of updated weights and / or the updated speech recognition model. When updated weights are provided to client device 110, client device 110 may replace the weights of on-device machine learning model 152A with the updated weights. When updated global hot word models are provided to client device 110, client device 110 may replace on-device machine learning model 152A with the updated global hot word models. In other embodiments, client device 110 may download a more suitable hot word model (or multiple models) from a server based on the type of command the user wishes to utter, and replace on-device machine learning model 152A with the downloaded hot word model.

[0038] In some implementations, based on the geographic region and / or other attributes of the client device 110 and / or the user of the client device 110, the on-device machine learning model 152A is transmitted (e.g., via a remote system 160 or other component) for storage and use on the client device 110. For example, the on-device machine learning model 152A may be one of N available machine learning models for a given language, but may be trained based on modifications specific to a particular geographic region, device type, context (e.g., music playback), etc., and provided to the client device 110 based on the client device 110 primarily located in that particular geographic region.

[0039] Turn now Figure 2 The client device 110 provided an explanation of the implementation method, in which... Figure 1A and Figure 1B The machine learning engine is included as part of the automation assistant client 240 on various devices (or communicates with the automation assistant client 240). Also applicable to... Figure 1A and Figure 1B The corresponding machine learning models for the machine learning engine interfaces on various devices are described. For simplicity, Figure 2 No right Figure 1A and Figure 1B The other components will be explained. Figure 2 The illustration shows an example of how the automation assistant client 240 utilizes... Figure 1A and Figure 1BMachine learning engines and their corresponding machine learning models are used on various devices to perform various actions.

[0040] Figure 2 The client device 110 is illustrated as having one or more microphones 211, one or more speakers 212, one or more cameras and / or other vision components 213, and a display 214 (e.g., a touch-sensitive display). The client device 110 may further include pressure sensors, proximity sensors, accelerometers, magnetometers, and / or other sensors for generating sensor data other than audio data captured by the one or more microphones 211. The client device 100 at least selectively executes the automation assistant client 240. Figure 2 In the example, the automation assistant client 240 includes an on-device hot word detection engine 122B, an on-device voice recognizer 142, an on-device natural language understanding (NLU) engine 144, and an on-device execution engine 146. The automation assistant client 240 further includes a voice capture engine 242 and a visual capture engine 244. The automation assistant client 140 may include additional and / or alternative engines, such as a voice activity detector (VAD) engine, an endpoint detector engine, and / or other engines.

[0041] One or more cloud-based automation assistant components 280 may optionally be implemented on one or more computing systems (collectively referred to as “cloud” computing systems), which are communicatively coupled to the client device 110 via one or more local area networks and / or wide area networks (e.g., the Internet), generally indicated by 290. For example, the cloud-based automation assistant component 280 may be implemented via a high-performance server cluster.

[0042] In various implementations, an instance of the automation assistant client 240 can form an instance that appears from the user's perspective as a logical instance of the automation assistant 295 through its interaction with one or more cloud-based automation assistant components 280, which the user can use to engage in human-computer interaction (e.g., verbal interaction, gesture-based interaction, and / or touch-based interaction).

[0043] For example, client device 110 may be: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a user's vehicle computing device (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker, a smart home appliance—such as a smart TV (or a standard TV equipped with a networked electronic dog with automated assistant capabilities)—and / or a user's wearable device, which includes a computing device (the user's watch with a computing device, the user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.

[0044] One or more vision components 213 can take various forms, such as a thematic camera, a stereo camera, a LiDAR component (or other laser-based component), a radar component, etc. One or more vision components 213 can be used—for example, by the vision capture engine 242—to capture visual frames (e.g., image frames, laser-based visual frames) of the environment in which the client device 110 is deployed. In some embodiments, such visual frames can be used to determine whether a user is present in the vicinity of the client device 110 and / or the distance of the user (e.g., the user's face) relative to the client device 110. For example, this determination can be used to determine whether activation is enabled. Figure 2 The machine learning engine and / or other engines shown on the various devices.

[0045] Voice capture engine 242 can be configured to capture user speech and / or other audio data captured by microphone 211. Further, client device 110 may include pressure sensors, proximity sensors, accelerometers, magnetometers, and / or other sensors for generating sensor data other than the audio data captured by microphone 211. As described herein, such audio data and other sensor data can be utilized by hot word detection engine 122B and / or other engines to determine whether to initiate one or more currently dormant automated assistant functions, avoid initiating one or more currently dormant automated assistant functions, and / or disable one or more currently active automated assistant functions. Automated assistant functions may include on-device voice recognizer 142, on-device NLU engine 144, on-device execution engine 146, and additional and / or alternative engines. For example, on-device voice recognizer 142 may utilize on-device voice recognition model 142A to process audio data capturing spoken utterances to generate recognized text 143A corresponding to the spoken utterances. On-device NLU engine 144 optionally utilizes on-device NLU model 144A to perform on-device natural language understanding on recognized text 143A to generate NLU data 145A. For example, NLU data 145A may include an intent corresponding to a spoken utterance and optionally include parameters of the intent (e.g., slot values). Further, on-device fulfillment engine 146 generates fulfillment data 147A based on NLU data 145A, optionally utilizing on-device fulfillment model 146A. This fulfillment data 147A may define local and / or remote responses to spoken utterances (e.g., answers), interactions performed with locally installed applications based on spoken utterances, commands transmitted to Internet of Things (IoT) devices (directly or via corresponding remote systems) based on spoken utterances, and / or other parsed actions performed based on spoken utterances. The fulfillment data 147A is then provided to enable the local and / or remote execution of defined actions to parse the spoken utterance. For example, execution may include presenting local and / or remote responses (e.g., visually and / or audibly (optionally utilizing a local text-to-speech module)), interacting with locally installed applications, transmitting commands to IoT devices, and / or other actions.

[0046] Display 214 may be used to display recognized text 143A and / or further recognized text 143B from the voice recognizer 122 on the device and / or one or more results from execution 150. Display 214 may further be one of the user interface output components through which a visual portion of the response from the automation assistant client 240 is presented.

[0047] In some implementations, the cloud-based automation assistant component 280 may include a remote ASR engine 281 performing speech recognition, a remote NLU engine 282 performing natural language understanding, and / or a remote fulfillment engine 283 generating fulfillment. A remote execution module may also be optionally included, which performs remote execution based on locally or remotely determined fulfillment data. Additional and / or alternative remote engines may be included. As described herein, in various implementations, at least on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be preferred because they provide reduced latency and / or network usage when parsing spoken utterances (since no client-server round trip is required to parse spoken utterances). However, at least one or more cloud-based automation assistant components may be selectively utilized. For example, such components may be utilized in parallel with on-device components, and the output from such components may be utilized when local components fail. For example, on-device fulfillment engine 146 may fail in certain situations (e.g., due to the relatively limited resources of client device 110), and remote fulfillment engine 283 may utilize the more robust resources of the cloud in such cases to generate fulfillment data. The remote execution engine 283 can operate in parallel with the on-device execution engine 146, and can utilize its results when the on-device execution fails, or can be invoked in response to determining that the on-device execution engine 146 has failed.

[0048] In various implementations, an NLU engine (on-device and / or remote) can generate NLU data that includes one or more annotations of the identified text and one or more (e.g., all) terms of the natural language input. In some implementations, the NLU engine is configured to identify and annotate various types of syntactic information in the natural language input. For example, the NLU engine may include a morphological module that can segment individual words into morphemes and / or annotate morphemes, for example, by their classes. The NLU engine may also include a part-of-speech tagger configured to annotate terms with their syntactic functions. Similarly, for example, in some implementations, the NLU engine may additionally and / or alternatively include a dependency parser (not shown) configured to determine syntactic relations between terms in the natural language input.

[0049] In some implementations, the NLU engine may additionally and / or alternatively include entity annotators configured to annotate entity references in one or more fragments, such as references to people (e.g., including literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), etc. In some implementations, the NLU engine may additionally and / or alternatively include coreference resolvers (not shown) configured to group or "cluster" references to the same entity based on one or more contextual cues. In some implementations, one or more components of the NLU engine may depend on annotations from one or more other components of the NLU engine.

[0050] The NLU engine may also include an intent matcher configured to determine the intent of a user engaging in interaction with the automation assistant 295. The intent matcher can use various techniques to determine the user's intent. In some implementations, the intent matcher may access one or more local and / or remote data structures, such as those including multiple mappings between syntax and response intents. For example, the syntax included in the mappings may be selected and / or learned over time and may represent the user's general intent. For example, a syntax, "play..." <artist>"(Play <Artist>)" can be mapped to an intent that invokes a response action, which results in <artist>The music of (<artist>) is played on client device 110. Another syntax, "[weather|forecast]today", can be mapped to user queries such as "what's the weather today" and "what's the forecast for today?". In addition to or instead of the syntax, in some implementations, the intent matcher can employ one or more trained machine learning models individually, or in combination with one or more syntaxes. These trained machine learning models can be trained to identify intents, for example, by embedding identified text from spoken utterances into a dimensionality-reduced space, and then determining which other embeddings are closest (and thus, which intents are closest), for example, using techniques such as Euclidean distance, cosine similarity, etc. As in the above "play" <artist>As seen in the example syntax, some syntaxes have slots that can be filled with slot values ​​(or "parameters") (e.g., <artist>Slot values ​​can be determined in various ways. Typically, the user will actively provide the slot value. For example, for the syntax "Order me a..." <topping>"pizza (order me a pizza with toppings)" might be used, and the user might say the phrase "order me a sausage pizza," in which case the slot "toppings" is automatically filled. Other slot values ​​can be inferred based on factors such as user location, currently displayed content, user preferences, and / or other prompts.

[0051] The fulfillment engine (local and / or remote) can be configured to receive predicted / estimated intents and any associated slot values ​​output by the NLU engine, and fulfill (or "parse") the intents. In various implementations, fulfilling (or "parses") a user's intent may result in the generation / acquisition of various fulfillment information (also known as fulfillment data), for example, through the fulfillment engine. This can include determining local and / or remote responses to spoken utterances (e.g., answers), interactions performed with locally installed applications based on spoken utterances, commands transmitted to Internet of Things (IoT) devices (directly or via corresponding remote systems) based on spoken utterances, and / or other parsed actions performed based on spoken utterances. On-device fulfillment can then initiate / execute the determined actions locally and / or remotely to parse the spoken utterances.

[0052] Figure 3 A flowchart is provided illustrating an example method 300 for detecting and blocking commands in the media that could trigger another automation assistant. For convenience, the operation of method 300 is described with reference to a system that performs the operation. Such a system for method 300 includes one or more processors and / or other components of a client device. Furthermore, although the operations of method 300 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0053] In block 310, the system uses a wireless communication protocol via a client device to discover each of a plurality of automation assistant devices in an environment, each executing at least one automation assistant. In some implementations, the automation assistant running on the client device may use an API based on a wireless communication protocol (e.g., WiFi or Bluetooth) to discover each of the plurality of automation assistant devices in an environment (e.g., a user's home). Block 310 may be repeated on a continuous basis (e.g., at predetermined time intervals) so that the client device can discover the addition or removal of automation assistant devices in the environment.

[0054] In box 320, the system uses a client device to determine the proximity of each of a plurality of automated assistant devices (found in box 310) in the environment using calibrated sound or wireless signal strength analysis. Box 320 may be repeated on a continuous basis (e.g., at predetermined time intervals) to identify changes in proximity (e.g., due to relocation of automated assistant devices in the environment).

[0055] In box 330, the system determines the active capabilities of each of the multiple automated assistant devices in the environment (discovered in box 310) via a client device. In some implementations, determining the active capabilities of an automated assistant device includes one or more of the following: determining whether the automated assistant device is performing hot word detection (or key phrase detection), determining whether the automated assistant device is performing open automatic speech recognition, and determining whether the automated assistant device is listening for specific speech.

[0056] Referring again to box 330, in some implementations, the client device may determine active capabilities using an API that indicates the current state of the automation assistant device. In other implementations, the client device may determine active capabilities via static knowledge of the automation assistant device's capabilities (e.g., automation assistant device type 1 might always be listening for "HeyDevice 1"). In still other implementations, the client device may rely on environmental perception to determine active capabilities. For example, the client device may process and detect interactions with other nearby automation assistants and track the state of other automation assistant devices. Box 330 may be repeated on a continuous basis (e.g., at predetermined time intervals) to identify changes in active capabilities (e.g., due to changes in the state of automation assistant devices). For some devices, certain active capabilities may remain constant over time, for example, if a hot word model is always active on a particular automation assistant device.

[0057] In some implementations, by discovering multiple automated assistant devices in the environment in box 310, discovering the proximity of each of the multiple automated assistant devices in the environment in box 320, and determining the active capabilities of each of the multiple automated assistant devices in the environment in box 330, the client device maintains a set of information about nearby automated assistant devices and a set of information about corresponding capabilities known to be active on these nearby automated assistant devices. There may be some overlap between capabilities. For example, multiple automated assistant devices may be listening for specific hot words or specific key phrases (e.g., "play music").

[0058] In box 340, the system initiates playback of digital media via an automated assistant that executes at least partially on the client device. In some implementations, playback may be initiated via a user's verbal utterance, including a request to initiate playback (e.g., "Hey Computer, play the latest episode of my favorite podcast"). In other implementations, playback may be initiated by the user delivering content to a specific target device.

[0059] In box 350, the system determines whether the volume level associated with the playback of digital media (initiated in box 340) meets a threshold determined based on the proximity of each of the multiple automation assistant devices to the client device (determined in box 320). Taking into account the proximity of the automation assistant devices in the environment (determined in box 320), the threshold may be a sufficiently high volume level so that the automation assistant device will be triggered by any hot words or commands appearing in the audio track of the digital media, the playback of which was initiated in box 340. If, during the iteration in box 350, the system determines that the volume level associated with the playback of digital media does not meet the threshold, the system proceeds to box 360, and the process ends. Conversely, if, during the iteration in box 350, the system determines that the volume level associated with the playback of digital media meets the threshold, the system proceeds to box 370.

[0060] In box 370, the system determines whether the active capabilities of one or more of the multiple automation assistant devices (determined in box 330) include performing hot word detection or performing open-ended automatic speech recognition. If, during the iteration of box 370, the system determines that the active capabilities of one or more of the multiple automation assistant devices do not include performing hot word detection or performing open-ended automatic speech recognition, the system proceeds to box 360, and the process ends. On the other hand, if, during the iteration of box 370, the system determines that the active capabilities of one or more of the multiple automation assistant devices include performing hot word detection or performing open-ended automatic speech recognition, the system proceeds to box 380.

[0061] In other implementations, in block 370, the client device may also run a speaker recognition model to detect whether speech in the content is expected to trigger activation of at least one automated assistant. If, during the iteration at block 370, the system determines that the active capabilities of one or more automated assistant devices do not include performing hot word detection or performing open-ended automatic speech recognition, and that speech in the content is expected to trigger activation of at least one automated assistant, the system proceeds to block 360, and the process ends. Alternatively, if, during the iteration at block 370, the system determines that the active capabilities of one or more automated assistant devices include performing hot word detection or performing open-ended automatic speech recognition, and that speech in the content is expected to trigger activation of at least one automated assistant, the system proceeds to block 380.

[0062] In box 380, in response to the system initiating playback of digital media (in box 340), determining that the volume level meets a threshold (in box 350) and determining the active capabilities of one or more of the multiple automation assistant devices, including performing hot word detection or performing open automatic speech recognition (in box 370), the client device processes the digital media based on the active capabilities of at least one of the multiple automation assistant devices (determined in box 330) to identify audio segments in the digital media that, during playback, are expected to trigger activation of at least one automation assistant executing on at least one of the multiple automation assistant devices in the environment. In some embodiments, the client device may continuously analyze the digital media stream (e.g., using a lookahead buffer) to identify potential audio segments in the digital media that are expected to trigger automation assistant functions in other nearby devices. In some embodiments, multiple audio segments may be identified in box 380 as audio segments expected to trigger activation of at least one automation assistant executing on at least one of the multiple automation assistant devices in the environment.

[0063] Referring again to box 380, in some embodiments, the client device processes digital media using one or more hot word detection models to detect potential triggers within the digital media. For example, a first hot word detection model may be associated not only with the client device and a first automated assistant device with a first set of detected hot words found in box 310, but also with a second hot word detection model that is a proxy model associated with a second automated assistant device with a second set of detected hot words found in box 310, the second automated assistant device having capabilities different from those of the client device. In some embodiments, the hot word detection model may include one or more machine learning models that generate predictive outputs indicating the probability of one or more hot words appearing in audio segments of the digital media. For example, the one or more machine learning models may be on-device hot word detection models and / or other machine learning models. Each of the machine learning models may be a deep neural network or any other type of model and may be trained to identify one or more hot words. Further, for example, the generated output may be a probability and / or other likelihood measure.

[0064] Referring again to box 380, in some implementations, in response to a predicted output from one or more hot word detection models satisfying a threshold indicating that one or more hot words appear in an audio segment in digital media, a client device can identify an audio segment as an audio segment in digital media that, during playback, is expected to trigger activation of at least one automated assistant (i.e., an automated assistant associated with the one or more hot word detection models that generate the predicted output satisfying the threshold). In other implementations, the client device can also run a speaker identification model to detect whether speech in the content is expected to trigger activation of at least one automated assistant. In the example, assume the predicted output is a probability, and the probability must be greater than 0.85 to satisfy the threshold indicating that one or more hot words appear in an audio segment in digital media, and the predicted probability is 0.88. Based on the predicted probability of 0.88 satisfying the threshold of 0.85, and further based on the output of the speaker identification model, the system identifies an audio segment in digital media as an audio segment in digital media that, during playback, is expected to trigger activation of at least one automated assistant. In some implementations, in the case of words or phrases that are auditorily similar to hot words, the hot word detection model can generate a predicted output that satisfies the threshold. In this context, audio segments in digital media that include auditory similar words or phrases can be identified as audio segments in digital media, and during playback, such audio segments are expected to trigger the activation of at least one automated assistant executing on at least one of a plurality of automated assistant devices in the environment.

[0065] Referring again to box 380, in other embodiments, the client device uses a voice recognizer to detect potential triggers in the digital media. In other embodiments, the client device utilizes captions or subtitles included in the digital media to detect potential triggers in the digital media.

[0066] In block 390, in response to identifying an audio segment in the digital media (in block 380), the client device modifies the digital media to prevent activation of at least one automation assistant executing on at least one of a plurality of automation assistant devices in the environment. In some embodiments, in response to detecting a potential triggering by at least one of a plurality of automation assistant devices in the environment, the client device performs active blocking to prevent other devices from erroneously treating the audio segment as automation assistant input.

[0067] Referring again to box 390, in some embodiments, a client device modifies the digital media by inserting one or more digital watermarks into an audio track (e.g., at a position corresponding to or adjacent to an audio segment within the audio track). The digital watermark can be an audio watermark imperceptible to humans, recognizable by an automation assistant running on the automation assistant device, and upon recognition, can prevent the automation assistant from accidentally activating its functions due to playback of digital media including audio segments on the client device. In some embodiments, the client device can pass information about the used audio watermark to the automation assistant running on the automation assistant device, so that the automation assistant on the automation assistant device recognizes the audio watermark as a command to prevent activation of the automation assistant.

[0068] Referring again to box 390, in other embodiments, the client device may modify audio segments before playback to ensure that the audio segments do not trigger automated assistants running on one or more automated assistant devices in the environment. For example, the client device may perform adversarial modifications on the audio segments. In other embodiments, the client device may perform other transformations on the audio segments based on the type of trigger to be avoided. For example, if an automated assistant running on one or more automated assistant devices in the environment is expected to be triggered because the speech voice is too similar to the speech voice to be received (e.g., as determined in box 330), the client device may apply a pitch change to the audio segments to avoid triggering the automated assistants running on one or more automated assistant devices in the environment.

[0069] In some implementations, the system may provide a feedback mechanism to the user (e.g., via a client device) that allows the user to identify unintended activation of the automation assistant. For example, when an automation assistant running on the client device is triggered during playback of digital media, the automation assistant on the client device may ask the user, "Was this a mis-trigger?"

[0070] Figure 4 A flowchart is provided illustrating an example method 400 for detecting and blocking commands in the media that could trigger another automation assistant. For convenience, the operation of method 400 is described with reference to a system that performs the operation. Such a system for method 400 includes one or more processors and / or other components of a client device. Furthermore, although the operations of method 400 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0071] In block 410, the system uses a wireless communication protocol via a client device to discover each of a plurality of automation assistant devices in an environment, each executing at least one automation assistant. In some implementations, the automation assistant running on the client device may use an API based on a wireless communication protocol (e.g., WiFi or Bluetooth) to discover each of the plurality of automation assistant devices in an environment (e.g., a user's home). Block 410 may be repeated on a continuous basis (e.g., at predetermined time intervals) so that the client device can discover the addition or removal of automation assistant devices in the environment.

[0072] In box 420, the system uses a client device to perform calibrated sound or wireless signal strength analysis to determine the proximity of each of a plurality of automated assistant devices in the environment (found in box 410). Box 420 may be repeated on a continuous basis (e.g., at predetermined time intervals) to identify changes in proximity (e.g., due to relocation of automated assistant devices in the environment).

[0073] In box 430, the system determines the active capabilities of each of the multiple automation assistant devices in the environment (discovered in box 410) via a client device. In some implementations, determining the active capabilities of an automation assistant device includes one or more of the following: determining whether the automation assistant device is performing hot word detection (or key phrase detection), determining whether the automation assistant device is performing open automatic speech recognition, and determining whether the automation assistant device is listening for specific speech.

[0074] Referring again to box 430, in some implementations, the client device may use an API to determine active capabilities, which indicates the current state of the automation assistant device. In other implementations, the client device may determine active capabilities via static knowledge of the capabilities of a particular automation assistant device (e.g., automation assistant device type 1 may always be listening for "Hey Device 1"). In other implementations, the client device may rely on environmental perception to determine active capabilities. For example, the client device may process and detect interactions with other nearby automation assistants and track the state of other automation assistant devices. Box 430 may be repeated on a continuous basis (e.g., at predetermined time intervals) to identify changes in active capabilities (e.g., due to changes in the state of automation assistant devices). For some devices, certain active capabilities may remain constant over time, for example, if a hot word model is always active on a particular automation assistant device.

[0075] In some implementations, by discovering multiple automated assistant devices in the environment in box 410, discovering the proximity of each of the multiple automated assistant devices in the environment in box 420, and determining the active capabilities of each of the multiple automated assistant devices in the environment in box 430, the client device maintains a set of information about nearby automated assistant devices and a set of information about corresponding capabilities known to be active on these nearby automated assistant devices. There may be some overlap between capabilities. For example, multiple automated assistant devices may be listening for specific hot words or specific key phrases (e.g., "play music").

[0076] In box 440, the system initiates playback of digital media via an automated assistant that executes at least partially on the client device. In some implementations, playback may be initiated via a user's verbal utterance, including a request to initiate playback (e.g., "Hey Computer, play the latest episode of my favorite podcast"). In other implementations, playback may be initiated by the user delivering content to a specific target device.

[0077] In box 450, the system determines whether the volume level associated with the playback of digital media (initiated in box 440) meets a threshold determined based on the proximity of each of the multiple automation assistant devices to the client device (determined in box 420). Taking into account the proximity of the automation assistant devices in the environment (determined in box 420), the threshold can be a sufficiently high volume level so that the automation assistant device will be triggered by any hot words or commands appearing in the audio track of the digital media, the playback of which was initiated in box 440. If, during the iteration in box 450, the system determines that the volume level associated with the playback of digital media does not meet the threshold, the system proceeds to box 460, and the process ends. Conversely, if, during the iteration in box 450, the system determines that the volume level associated with the playback of digital media meets the threshold, the system proceeds to box 470.

[0078] In box 470, the system determines whether the active capabilities of one or more of the multiple automation assistant devices (determined in box 430) include performing hot word detection or performing open-ended automatic speech recognition. If, during the iteration of box 470, the system determines that the active capabilities of one or more of the multiple automation assistant devices do not include performing hot word detection or performing open-ended automatic speech recognition, the system proceeds to box 460, and the process ends. On the other hand, if, during the iteration of box 470, the system determines that the active capabilities of one or more of the multiple automation assistant devices include performing hot word detection or performing open-ended automatic speech recognition, the system proceeds to box 480.

[0079] In box 480, in response to the system initiating playback of digital media (in box 440), determining that the volume level meets a threshold (in box 450) and determining the active capabilities of one or more of the multiple automation assistant devices, including performing hot word detection or performing open automatic speech recognition (in box 470), the client device processes the digital media based on the active capabilities of at least one of the multiple automation assistant devices (determined in box 430) to identify audio segments in the digital media that, during playback, are expected to trigger activation of at least one automation assistant executing on at least one of the multiple automation assistant devices in the environment. In some embodiments, the client device may continuously analyze the digital media stream (e.g., using a lookahead buffer) to identify potential audio segments in the digital media that are expected to trigger automation assistant functions in other nearby devices. In some embodiments, multiple audio segments may be identified in box 480 as audio segments expected to trigger activation of at least one automation assistant executing on at least one of the multiple automation assistant devices in the environment.

[0080] Referring again to box 480, in some embodiments, the client device processes digital media using one or more hot word detection models to detect potential triggers within the digital media. For example, a first hot word detection model may be associated not only with the client device and a first automated assistant device with a first set of detected hot words found in box 410, but also with a second hot word detection model that is a proxy model associated with a second automated assistant device with a second set of detected hot words found in box 410, the second automated assistant device having capabilities different from those of the client device. In some embodiments, the hot word detection model may include one or more machine learning models that generate predictive outputs indicating the probability of one or more hot words appearing in audio segments of the digital media. For example, the one or more machine learning models may be on-device hot word detection models and / or other machine learning models. Each of the machine learning models may be a deep neural network or any other type of model and may be trained to identify one or more hot words. Further, for example, the generated output may be a probability and / or other likelihood measure.

[0081] Referring again to box 480, in some implementations, in response to a predicted output from one or more hot word detection models satisfying a threshold indicating that one or more hot words appear in an audio segment in digital media, a client device can identify the audio segment as an audio segment in digital media, which, upon playback, is expected to trigger the activation of at least one automated assistant (i.e., an automated assistant associated with the one or more hot word detection models that generate the predicted output that satisfies the threshold). In the example, assume the predicted output is a probability, and the probability must be greater than 0.85 to satisfy the threshold indicating that one or more hot words appear in an audio segment in digital media, and the predicted probability is 0.88. Based on the predicted probability of 0.88 satisfying the threshold of 0.85, the system identifies the audio segment in digital media as an audio segment in digital media, which, upon playback, is expected to trigger the activation of at least one automated assistant. In some implementations, in the case of words or phrases that are auditorily similar to hot words, the hot word detection model can generate a predicted output that satisfies the threshold. In this context, audio segments in digital media that include auditory similar words or phrases can be identified as audio segments in digital media, and during playback, such audio segments are expected to trigger the activation of at least one automated assistant executing on at least one of a plurality of automated assistant devices in the environment.

[0082] In block 490, in response to identifying an audio segment in digital media (in block 480), the client device communicates with at least one of a plurality of automation assistant devices in the environment to prevent activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment during playback of the audio segment. In some embodiments, in response to detecting a potential triggering by at least one of the plurality of automation assistant devices in the environment, the client device performs active blocking to prevent other devices from erroneously treating the audio segment as automation assistant input.

[0083] Referring again to box 490, in some embodiments, communicating with at least one of a plurality of automation assistant devices in the environment to prevent activation of at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment during playback of an audio segment includes: providing at least a portion of the audio segment or a fingerprint based on the audio segment to at least one of the plurality of automation assistant devices, or providing instructions to at least one of the plurality of automation assistant devices to stop listening at a time corresponding to the audio segment.

[0084] In some implementations, if a fingerprint is provided and an automated assistant running on an automated assistant device is triggered, the automated assistant can compare the fingerprint with a fingerprint captured on the audio, and if they match, prevent triggering. If an audio clip is provided, it can be used as a reference for an acoustic echo cancellation algorithm to prevent activation of at least one automated assistant running on at least one of a plurality of automated assistant devices in the environment during playback of the audio clip. In this case, the genuine user voice can still be allowed to pass through.

[0085] In some implementations, the system may provide a feedback mechanism to the user (e.g., via a client device) that allows the user to identify unintended activation of the automation assistant. For example, when an automation assistant running on the client device is triggered during playback of digital media on the client device, the automation assistant running on the client device may ask the user, "Was this a mis-trigger?"

[0086] Figure 5 This is a block diagram of an example computing device 510, which may optionally be used to perform one or more aspects of the techniques described herein. In some implementations, one or more of a client device, a cloud-based automation assistant component, and / or other components may include one or more components of the example computing device 510.

[0087] Computing device 510 typically includes at least one processor 514 that communicates with a plurality of peripheral devices via a bus subsystem 512. These peripheral devices may include a storage subsystem 524—for example, including a memory subsystem 525 and a file storage subsystem 526—a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices allow users to interact with computing device 510. The network interface subsystem 516 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0088] User interface input device 522 may include a keyboard, pointing devices—such as a mouse, trackball, touchpad or graphics tablet, scanner, touchscreen integrated into a display, audio input devices—such as a voice recognition system, microphone, and / or other types of input devices. Generally, the term "input device" is used to include all possible types of devices and methods for inputting information into computing device 510 or into a communication network.

[0089] User interface output device 520 may include a display subsystem, printer, fax machine, or non-visual display, such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and methods for outputting information from computing device 510 to a user or another machine or computing device.

[0090] Storage subsystem 524 stores the programming and data structures that provide some or all of the functionality of the modules described herein. For example, storage subsystem 524 may include logic for performing selected aspects of the methods described herein and implementing... Figure 1A and Figure 1B The various components described in the document.

[0091] These software modules are typically executed by processor 514 alone or in combination with other processors. The memory subsystem 525 included in storage subsystem 524 may include multiple memories, including main random access memory (RAM) 530 for storing instructions and data during program execution and read-only memory (ROM) 532 storing fixed instructions. File storage subsystem 526 can provide permanent storage for program and data files and may include hard disk drives, floppy disk drives and associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored by file storage subsystem 526 in storage subsystem 524 or in other machines accessible by processor 514.

[0092] The bus subsystem 512 provides a mechanism for allowing the various components and subsystems of the computing device 510 to communicate with each other as intended. Although the bus subsystem 512 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0093] The computing device 510 can be of various types, including workstations, servers, computing clusters, blade servers, server groups, or any other data processing system or computing device. Due to the diverse nature of computers and networks, this is for the purpose of illustrating some implementation methods. Figure 5 The description of computing device 510 is merely a specific example. Many other configurations of computing device 510 may have... Figure 5 The computing device depicted in the text has more or fewer components.

[0094] In cases where the system described herein collects or otherwise monitors personal information about a user or may use personal information and / or monitored information, the user may be given the opportunity to control whether a program or feature collects user information (e.g., information about the user's social networks, social actions or activities, occupation, user preferences, or the user's current location), or to control whether and / or how content is received from a content server that may be more relevant to the user. Similarly, certain data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be processed to make it impossible to determine personally identifiable information for the user, or, where geolocation information is available, the user's geolocation may be generalized (e.g., for city, ZIP code, or state), making the user's specific geolocation undeterminable. Therefore, the user can control how information about themselves and / or information used is collected.

[0095] While several implementations have been described and illustrated herein, various other ways and / or structures may be utilized for performing functions and / or obtaining results and / or one or more advantages described herein, and each of such variations and / or modifications is considered to be within the scope of the implementations described herein. More generally, all parameters, sizes, materials, and configurations described herein are intended to be exemplary, and actual parameters, sizes, materials, and / or configurations will depend on one or more specific applications using one or more teachings. Those skilled in the art will recognize or be able to identify many equivalents of the specific implementations described herein using only conventional experimentation. Therefore, it should be understood that the foregoing implementations are presented by way of example only, and that implementations other than those specifically described and claimed may be practiced within the scope of the appended claims and their equivalents. The embodiments of this disclosure relate to the various individual features, systems, articles, materials, kits, and / or methods described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of this disclosure if these features, systems, articles, materials, kits, and / or methods are not contradictory.< / topping> < / artist> < / artist> < / artist> < / artist>

Claims

1. A method implemented by one or more processors, the method comprising: For each of a plurality of automated assistant devices in an environment that each executes at least one automated assistant, the active state capability of the automated assistant device is determined by a client device, wherein determining the active state capability of the automated assistant device includes determining whether the automated assistant device is currently performing open automated speech recognition, and wherein currently performing open automated speech recognition includes the automated assistant device being in a state in which it is currently using a speech recognition model to recognize speech rather than simply using a hot word detection model to detect one or more hot words; Playback of digital media is initiated by an automated assistant that executes at least in part on the client device; In response to initiating playback, the client device processes the digital media to identify audio segments in the digital media based on the activity state capability of at least one of the plurality of automation assistant devices currently performing open automatic speech recognition in the environment, wherein the audio segment is identified based on the expectation that the audio segment will trigger the activation of at least one automation assistant performing on at least one of the plurality of automation assistant devices during playback, and based on the state in which the at least one of the plurality of automation assistant devices is currently using the speech recognition model to identify speech rather than simply using the hot word detection model to detect the one or more hot words; and In response to identifying the audio segment in the digital media, the client device modifies the digital media to prevent the activation of the at least one automation assistant executing on at least one of the plurality of automation assistant devices in the environment.

2. The method according to claim 1, further comprising: The client device uses a wireless communication protocol to discover each of the plurality of automation assistant devices in the environment.

3. The method according to claim 1, further comprising: The proximity of each of the plurality of automation assistant devices to the client device is determined by the client device using calibrated sound or wireless signal strength analysis.

4. The method according to claim 3, wherein, The client device processes the digital media to identify the audio segment in response to determining that the volume level associated with the playback of the digital media satisfies a threshold determined based on the proximity of each of the plurality of automation assistant devices to the client device.

5. The method according to claim 1, wherein, Determining the active state capability of the automated assistant device further includes determining whether the automated assistant device is currently listening to a specific voice.

6. The method according to claim 5, wherein, The client device processes the digital media to identify the audio segment as an output based on a speaker identification model and, in response to determining the active state capability of one or more of the plurality of automated assistant devices, includes performing open-ended automatic speech recognition.

7. The method according to any one of claims 1-6, wherein, The client device modifies the digital media by inserting a digital watermark into the audio tracks of the digital media.

8. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 7.

9. A computer-readable storage medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 7.

10. A system comprising a processor, a computer-readable storage medium, one or more computer-readable storage media, and program instructions co-stored on the one or more computer-readable storage media, the program instructions being executable to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice Control of a Media Playback System

    US20170242653A1

  • Wake-Word Detection Suppression

    US20190043492A1

  • Networked devices, systems, & methods for intelligently deactivating wake-word engines

    US20200090646A1