Multi-factor audio watermarking
By embedding multi-factor audio watermarking technology into audio content and combining it with voice signal detection of media content, the problem of unintentional activation of automated assistants by media content is solved, achieving resource savings and improved user experience.
Patent Information
- Application Number
- CN202180035299.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-07
- Filing Date
- 2021-11-05
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-11-05
AI Technical Summary
When processing audio data, automated assistants are easily activated unintentionally by trending words in media content such as TV commercials and TV programs, resulting in wasted resources and a negative user experience. Existing machine learning models have difficulty effectively distinguishing between user intent and media content, leading to unintentional activation and wasted resources.
By employing multi-factor audio watermarking technology, combined with additional factors such as speech signals, media content is detected and query processing in its audio data is suppressed. By embedding audio signals that are imperceptible to humans into the audio content, hot word detection and speech recognition models are adjusted to reduce unintentional activation.
It effectively reduces unintentional activation of automated assistants, saves network and computing resources, improves user experience, and reduces misuse and waste of resources of automated assistants.
Smart Images

Figure CN115605951B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to multi-factor audio watermarking. BACKGROUND
[0002] Humans can engage in human-to-computer dialog with interactive software applications referred to herein as "automated assistants" (also referred to as "digital agents," "interactive personal assistants," "intelligent personal assistants," "assistant applications," "conversational agents," etc.). For example, humans (which when they interact with automated assistants can be referred to as "users") can provide commands and / or requests to an automated assistant by using spoken natural language input (i.e., utterances) that can be converted, in some cases, to text and then processed, by providing textual (e.g., typed) natural language input, and / or through touch and / or speech-free physical movements (e.g., gestures, eye gaze, facial expressions, etc.). Automated assistants respond to requests by providing responsive user interface output (e.g., aural and / or visual user interface output), controlling one or more smart devices, and / or controlling one or more functions of a device that implements the automated assistant (e.g., controlling other applications of the device).
[0003] As noted above, many automated assistants are configured to be interacted with via spoken utterances. To protect user privacy and / or conserve resources, automated assistants refrain from performing one or more automated assistant functions based on all spoken utterances that are present in audio data detected via a microphone of a client device that (at least partially) implements the automated assistant. Rather, certain processing of spoken utterances occurs only in response to a determination that certain conditions are present.
[0004] For example, many client devices that include and / or interact with automated assistants include a hotword detection model. When a microphone of such a client device is not deactivated, the client device can continuously process audio data detected via the microphone using the hotword detection model to generate a predicted output that indicates whether one or more hotwords (including multiword phrases), such as "Hey Assistant," "OK Assistant," and / or "Assistant," are present. When the predicted output indicates that a hotword is present, any audio data that follows within a threshold amount of time (and optionally is determined to include speech activity) can be processed by one or more on-device and / or remote automated assistant components, such as speech recognition components, speech activity detection components, and / or the like. In addition, recognized text (from the speech recognition components) can be processed using natural language understanding engines and / or actions can be performed based on natural language understanding engine outputs. Actions can include, for example, generating and providing responses and / or controlling one or more applications and / or smart devices. Other hotwords (e.g., "No," "Stop," "Cancel," "Volume Up," "Volume Down," "Next Track," "Previous Track," and / or the like) can be mapped to various commands, and when the predicted output indicates that one of these hotwords is present, the mapped command can be processed by the client device. However, when the predicted output indicates that a hotword is not present, the corresponding audio data will be discarded without any further processing, thereby conserving resources and user privacy.
[0005] The predicted output of the above-described and / or other machine learning models (e.g., additional machine learning models described below) that specify whether an automated assistant functionality is activated performs well in many cases. However, in certain cases, the audio data processed by the automated assistant can include audio from a television advertisement, a television program, a movie, and / or other media content in addition to utterances spoken by the user. In cases where the audio data processed by the automated assistant includes audio from media content, when one or more hotwords are present in the audio from the media content, a hotword detection model of the automated assistant can detect the one or more hotwords from the media content. The automated assistant can then process any audio from the media content that follows the one or more hotwords within a threshold amount of time and respond to a request in the media content. Alternatively, the automated assistant can perform one or more actions (e.g., increase or decrease a volume) that correspond to the detected hotword.
[0006] Activation of an automated assistant based on the presence of one or more hotwords in audio from a television advertisement, a television program, a movie, and / or other media content can be unintentional by the user and can result in a negative user experience (e.g., the automated assistant can respond and / or perform an operation that the user does not wish to perform even though the user did not make a request). The occurrence of unintentional activation of an automated assistant can waste network and / or computing resources and potentially force a human to make a request to the automated assistant to undo or cancel an undesired action (e.g., a user can issue a "lower volume" command to counteract an unintentional "raise volume" command present in media content that triggered the automated assistant to raise the volume). SUMMARY
[0007] Some implementations disclosed herein aim to improve the performance of machine learning models with multi-factor audio watermarking. As described in more detail herein, such machine learning models can include, for example, hotword detection models and / or other machine learning models. Various implementations combine audio watermarking with additional factors (e.g., speech-based signals) to detect media content. In response to detecting media content, a system can suppress processing of queries included in audio data of the media content, thereby reducing the occurrence of unintentional activation of an automated assistant. In other implementations, in response to detecting media content, a system can adjust hotword detection and / or automatic speech recognition.
[0008] In some implementations, an audio watermark can be a human-imperceptible audio signal that is embedded into a piece of audio content. In some implementations, by combining detection of the audio watermark with additional factors, misuse of the watermark (e.g., to disable an automated assistant throughout playback of a piece of content) can be reduced. Moreover, by combining detection of the audio watermark with additional factors, media content can be reliably detected even when an audio fingerprint of a known piece of media content is unavailable on a device or a cloud database is impractical to maintain (e.g., because the media content is too new or because the amount of media content is too large to be included in such a database).
[0009] In some implementations, an audio watermarking system performs an offline preprocessing step to modify the soundtrack of some pre-recorded content (media) in a way that is imperceptible to a human listener. A corresponding process is used on an automated assistant-enabled device to detect the presence of a particular watermark or one of a known set of watermarks and suppress processing of a query or action that would otherwise be triggered by a hotword that accompanies the watermark in the soundtrack. In some implementations, the watermark can be positioned before, concurrently with, or after a hotword (and / or a query associated with a hotword) in the soundtrack of the media content.
[0010] In some implementations, the system reduces the risk of misuse of the audio watermark by malicious entities (e.g., disabling the automated assistant throughout playback of a piece of content, rather than only during playback of a hotword that is present in the audio track of the content) by combining audio watermark detection with one or more additional factors. In some implementations, the system can automatically learn the additional factors on the fly.
[0011] In other implementations, the system can use the watermark to trigger adaptation of behavior of any voice systems running in parallel. In an example, in response to detecting the watermark, the system can trigger a temporary reduction in the threshold for hotwords if the user appears likely to want to issue a query while a particular media content is playing.
[0012] In some implementations, the system combines detection of the watermark with satisfaction of one or more additional factors (e.g., matching a speech transcription to an expected set of ambiguous speech transcriptions, matching a speaker identification vector to an expected set of speaker identification vectors, and / or other signals). In response to both detection of the watermark and satisfaction of the one or more additional factors, the system can suppress processing of a query or action that would otherwise be triggered by a hotword in the watermark’s accompanying audio track. In some implementations, by combining detection of the watermark with satisfaction of the one or more additional factors, the system can reduce the risk of misuse of the watermark and limit suppression of query processing to cases in which the watermark is detected in conjunction with a particular query or a similarly pronounced variant of the query. In some implementations, using a set of ambiguous speech transcriptions to match against a speech transcription allows for errors and inaccuracies in automated speech recognition processing.
[0013] In some implementations, in a watermark registration step, the system can insert an audio watermark into an audio track of media content. The audio watermark can be a human- imperceptible audio signal that is embedded into an audio track of a piece of audio content. The audio watermark can be inserted into a particular temporal position in the audio track (e.g., before, concurrently with, or after a hotword and / or a query associated with the hotword).
[0014] Additionally, during the watermark registration step, the system can extract one or more features from the audio track of the media content for use in detection. In some implementations, the extracted features can include the particular watermark that is embedded, which can be selected from a range of watermarks available for use. In some implementations, the extracted features can include a unique identifier that the system used to encode the watermark.
[0015] In some implementations, features extracted during the watermark registration step can include speech transcripts at points in the audio track where watermarks are inserted, as well as potentially alternative erroneous speech transcripts that can be limited by an edit distance (or speech algorithm) and can be derived from the top N hypotheses of the speech recognition engine. For example, in some implementations, additional copies of sample audio with varying "noise" / spectral enhancements can be used to generate erroneous transcripts that simulate situations that can be encountered when detected in real-world environments (e.g., various environmental noises). In another example, additional automatic speech recognition systems with different models and / or intentionally incorrect settings (e.g., incorrect localization, such as British English instead of American English) can be used to generate additional erroneous transcripts that simulate behavior that can be seen on client devices.
[0016] In some implementations, in addition to or instead of the above-mentioned features, other features can be extracted during the watermark registration step, including hotword activation thresholds and timing relative to watermarks, associated speaker identification d-vectors, voice activity detection (VAD) binary masks, endpointer cutoff times, and / or language identification (Langld) confusion matrices. In some implementations, in addition to or instead of the above-mentioned features, time can be used as an additional feature that can be used when the expected future broadcast information of the media content is known. For example, timestamps or windows and / or locations, metropolitan areas, and / or broadcasters from which information can be inferred can be used.
[0017] In some implementations, some of all of the above-mentioned features can be output from the watermarking process, pushed to client devices executing automated assistants that perform watermark detection, and stored locally on those client devices as factors used in the watermark detection process. In other implementations, some of all of the above-mentioned features can be output from the watermarking process, pushed to one or more cloud computing nodes, and stored by those cloud computing nodes as factors used in the watermark detection process (e.g., by one or more cloud-based automated assistant components). In some implementations, the multifactor watermark includes one or more specific watermark identifiers that the client device is looking for as well as one or more other stored factors, which typically include expected speech transcripts, although this can be omitted.
[0018] In some implementations, in the multi-factor audio watermark detection step performed on the client device, audio watermark detection can be performed on a continuous stream of audio or only in response to detecting a particular voice event such as a hotword. Once a known audio watermark is detected, the system can load the associated additional factors (which can be stored locally on the client device or on a cloud computing node, as described above) and extract features from the audio to compare to those additional factors. For example, in some implementations, the inference of the additional factors can include running an automatic speech recognition engine to generate speech transcription features to compare to those stored as second factors. These can be compared, for example, using an edit distance or some other measure of speech or text similarity. Multiple hypotheses can be considered. In other implementations, the inference of the additional factors can include running a speaker ID model to verify the speaker of the query from a stored speaker ID vector, for example, by performing an embedding distance comparison to ensure that the two are sufficiently similar. In other implementations, the inference of the additional factors can include determining whether the current time and / or date is within a window in which the watermark is known to be active.
[0019] In some implementations, the audio watermark detected by the client device can be verified by matching at least one other feature. In response to verifying the watermark, the system can suppress processing of the query included in the audio data. If the detected watermark cannot be verified by matching at least one other feature, the system can refrain from suppressing processing of the query included in the audio data and can process the query in a typical manner. In such cases, the system can use techniques based on joint analysis to discover both unregistered use of the watermark and misuse of the watermark. In some implementations, the above-described feature extraction process can be performed for unknown watermarks (e.g., watermarks that cannot be verified) at the time of watermark detection.
[0020] In some implementations, if the system determines that there is a watermark that cannot be verified by matching at least one other feature, the system can extract one or more of the above-described features (e.g., speech transcription) and aggregate these features using techniques such as joint analysis. By performing this aggregation across multiple users and / or devices, the system can learn, in run, the characteristics of content with watermarks, which can then be stored and used to verify watermarks in future instances of detection. In other implementations, the system can detect misuse of the watermark system (e.g., based on a variety of speech transcriptions or other associated features).
[0021] In other implementations, the client device, in response to detecting and verifying the audio watermark, can adjust speech processing after detecting the audio watermark. For example, the watermark can be placed in a piece of content and used to temporarily lower a hotword detection threshold to make the automated assistant more responsive to queries that will be expected with the content. More generally, the watermark can be used to adjust any speech parameter, e.g., bias, language selection, etc. For example, when the media content is in a certain language, upon detection of the audio watermark, the system can bias the language selection to that language because the system has a strong prior that the follow-up query will be in the same language. In this setup, the system can perform one or more of the following: a machine transcription feature can generate a bias instruction for directing automated speech recognition to a particular transcription; a hotword threshold feature can generate an instruction for increasing or decreasing a threshold; and / or a Langld confusion matrix can generate an instruction that a language-specific ASR system should be used during factor generation and during detection.
[0022] In various implementations, a method implemented by one or more processors can include receiving, via one or more microphones of a client device, audio data that captures a spoken utterance; processing the audio data using one or more machine learning models to generate a predicted output that indicates a probability that one or more hotwords are present in the audio data; determining that the predicted output satisfies a threshold that indicates that the one or more hotwords are present in the audio data; in response to determining that the predicted output satisfies the threshold, processing the audio data using automated speech recognition to generate a speech transcription feature; detecting a watermark embedded in the audio data; and in response to detecting the watermark: determining that the speech transcription feature corresponds to one of a plurality of stored speech transcription features or an intermediate embedding; and in response to determining that the speech transcription feature corresponds to one of the plurality of stored speech transcription features or the intermediate embedding, inhibiting processing of a query included in the audio data.
[0023] In some implementations, detecting the watermark is in response to determining that the predicted output satisfies the threshold. In some implementations, the watermark is a human- imperceptible audio watermark. In some implementations, and the plurality of stored speech transcription features or the intermediate embedding are stored on the client device.
[0024] In some implementations, determining that the speech transcription feature corresponds to one of the plurality of stored speech transcription features includes determining that an edit distance between the speech transcription feature and the one of the plurality of stored speech transcription features satisfies a threshold edit distance. In some implementations, determining that the speech transcription feature corresponds to one of the plurality of stored speech transcription features includes determining that an embedding-based distance between the speech transcription feature and the one of the plurality of stored speech transcription features satisfies a threshold embedding-based distance.
[0025] In some implementations, the method can further include, in response to detecting the watermark: determining, using speaker recognition of the audio data, a speaker vector corresponding to the query included in the audio data; and determining that the speaker vector corresponds to one of the plurality of stored speaker vectors. In some implementations, the suppressing the processing of the query included in the audio data is further in response to determining that the speaker vector corresponds to one of the plurality of stored speaker vectors.
[0026] In some implementations, the method can further include determining whether a current time or date is within an active window of the watermark, and further in response to determining that the current time or date is within the active window of the watermark, suppressing the processing of the query included in the audio data. In some implementations, the plurality of stored speech transcription features or intermediate embeddings includes an erroneous transcription.
[0027] In some additional or alternative implementations, a method implemented by one or more processors can include receiving, via one or more microphones of a client device, first audio data capturing a first spoken utterance; processing the first audio data using automatic speech recognition to generate a speech transcription feature or intermediate embedding; detecting a watermark embedded in the first audio data; and in response to detecting the watermark: determining that the speech transcription feature or intermediate embedding corresponds to one of a plurality of stored speech transcription features or intermediate embeddings; and in response to determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings, modifying a threshold indicating a presence of the one or more hotwords in the audio data.
[0028] In some implementations, the method can further include: receiving, via the one or more microphones of the client device, second audio data capturing a second spoken utterance; processing the second audio data using one or more machine learning models to generate a predicted output, the predicted output indicating a probability of a presence of the one or more hotwords in the second audio data; determining that the predicted output satisfies a modified threshold indicating a presence of the one or more hotwords in the second audio data; and in response to determining that the predicted output satisfies the modified threshold, processing a query included in the second audio data.
[0029] In some implementations, the method can further include, in response to detecting the watermark: determining, using speaker recognition of the audio data, a speaker vector corresponding to the query included in the audio data; and determining that the speaker vector corresponds to one of the plurality of stored speaker vectors. In some implementations, the modifying the threshold indicating a presence of the one or more hotwords in the audio data is further in response to determining that the speaker vector corresponds to one of the plurality of stored speaker vectors.
[0030] In some additional or alternative implementations, a system can include a processor, a computer-readable memory, one or more computer-readable storage media, and program instructions collectively stored on the one or more computer-readable storage media. The program instructions are executable to: receive, via one or more microphones of a client device, audio data that captures a spoken utterance; process the audio data using one or more machine learning models to generate a predicted output that indicates a probability that one or more hotwords are present in the audio data; determine that the predicted output satisfies a threshold that indicates that the one or more hotwords are present in the audio data; in response to determining that the predicted output satisfies the threshold, process the audio data using automatic speech recognition to generate a speech transcription feature or intermediate embedding; detect a watermark embedded in the audio data; and in response to detecting the watermark: determine that the speech transcription feature or intermediate embedding corresponds to one of a plurality of stored speech transcription features or intermediate embeddings; and in response to determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings, suppress processing of a query included in the audio data.
[0031] Through use of one or more techniques described herein, occurrences of unintentional activation of an automated assistant that can waste network and / or computing resources and potentially force a human to make a request to the automated assistant to undo an undesirable action of the automated assistant can be reduced. This results in improved performance by allowing the automated assistant to suppress processing of a query included in an audio track of media content.
[0032] The above description is provided as an overview of some implementations of the present disclosure. Further description of those implementations and other implementations are described in greater detail below.
[0033] Various implementations can include non-transitory computer-readable storage media storing instructions executable by one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), and / or tensor processing units (TPUs)) to perform a method, such as one or more of the methods described herein. Other implementations can include automated assistant client devices (e.g., client devices that include an automated assistant interface at least for interfacing with cloud-based automated assistant components), that include a processor operable to execute stored instructions to perform a method, such as one or more of the methods described herein. Yet other implementations can include a system of one or more servers that include one or more processors operable to execute stored instructions to perform a method, such as one or more of the methods described herein. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1A and Figure 1BExample processing flows according to various implementations are described, which illustrate various aspects of this disclosure.
[0035] Figure 2 A block diagram depicts an example environment, which includes elements from... Figure 1A and 1B Various components, and implementations disclosed herein can be implemented therein.
[0036] Figure 3 A flowchart illustrating example methods for detecting media content and, in response, suppressing the processing of queries included in the audio data of the media content, according to various embodiments, is provided.
[0037] Figure 4 A flowchart illustrating example methods for detecting media content and adjusting hot word detection and / or automatic speech recognition in response, according to various implementations, is provided.
[0038] Figure 5 An example architecture for a computing device is described. Detailed Implementation
[0039] Figure 1A and 1B An example processing flow illustrating various aspects of this disclosure is depicted. Client device 110 is shown in... Figure 1A In, and included in representing client device 110 Figure 1A The components covered within the frame. Machine learning engine 122A may receive: audio data 101 corresponding to spoken utterances detected via one or more microphones of client device 110; and / or other sensor data 102 corresponding to non-vocal physical motions (e.g., gestures and / or movements, body postures and / or body movements, eye gaze, facial movements, mouth movements, etc.) detected via one or more non-microphone sensor components of client device 110. The one or more non-microphone sensors may include cameras or other visual sensors, proximity sensors, pressure sensors, accelerometers, magnetometers, and / or other sensors. Machine learning engine 122A processes audio data 101 and / or other sensor data 102 using machine learning model 152A to generate a predictive output 103. As described herein, machine learning engine 122A may be a hot word detection engine 122B or an alternative engine, such as a voice activity detector (VAD) engine, an endpoint detector engine, an automatic speech recognition (ASR) engine, and / or other engines.
[0040] In some implementations, when the machine learning engine 122A generates a predicted output 103, the predicted output 103 may be locally stored on a client device in on-device storage 111 and optionally associated with corresponding audio data 101 and / or other sensor data 102. In some versions of those implementations, the predicted output may be retrieved by the gradient engine 126 for use in generating gradient 106 at a later time, such as when one or more conditions described herein are met. On-device storage 111 may include, for example, read-only memory (ROM) and / or random access memory (RAM). In other implementations, the predicted output 103 may be provided to the gradient engine 126 in real time.
[0041] Client device 110 can determine whether to initiate the currently dormant automated assistant function based on whether the predicted output 103 in box 182 meets a threshold (e.g., Figure 2 The automation assistant 110 may initiate a currently dormant automation assistant function (e.g., 295) and / or disable the currently active automation assistant function using the assistant activation engine 124. The automation assistant function may include: speech recognition for generating recognized text, NLU for generating natural language understanding (NLU) output, generating a response based on the recognized text and / or NLU output, transmitting audio data to a remote server, transmitting recognized text to a remote server, and / or directly triggering one or more actions in response to audio data 101 (e.g., common tasks such as changing device volume). For example, suppose the predicted output 103 is a probability (e.g., 0.80 or 0.90), and the threshold at box 182 is a threshold probability (e.g., 0.85). If the client device 110 determines at box 182 that the predicted output 103 (e.g., 0.90) meets the threshold (e.g., 0.85), the assistant activation engine 124 may initiate a currently dormant automation assistant function.
[0042] In some implementations, and as such Figure 1B As shown, the machine learning engine 122A can be a hot word detection engine 122B. Note that various automated assistant functions, such as the on-device speech recognizer 142, the on-device NLU engine 144, and / or the on-device fulfillment engine 146, are currently dormant (i.e., as shown by the dashed lines). Furthermore, assume that the predicted output 103 generated using the hot word detection model 152B and based on the audio data 101 meets the threshold at box 182, and the voice activity detector 128 detects user speech directed at the client device 110.
[0043] In some versions of these implementations, the assistant activation engine 124 activates the on-device speech recognizer 142, the on-device NLU engine 144, and / or the on-device fulfillment engine 146 as currently dormant automated assistant functionality. For example, the on-device speech recognizer 142 can use the on-device speech recognition model 142A to process the audio data 101 of the spoken utterance, including the hotword "OK Assistant" and the additional commands and / or phrases that follow the hotword "OK Assistant," to generate recognized text 143A, the on-device NLU engine 144 can use the on-device NLU model 144A to process the recognized text 143A to generate NLU data 145A, the on-device fulfillment engine 146 can use the on-device fulfillment model 146A to process the NLU data 145A to generate fulfillment data 147A, and the client device 110 can use the fulfillment data 147A in performance 150 of one or more actions in response to the audio data 101.
[0044] In other versions of these implementations, the assistant activation engine 124 activates only the on-device fulfillment engine 146 without activating the on-device speech recognizer 142 and the on-device NLU engine 144 to process various commands, such as "No," "Stop," "Cancel," "Volume Up," "Volume Down," "Next Track," "Previous Track," and / or other commands that can be processed without the on-device speech recognizer 142 and the on-device NLU engine 144. For example, the on-device fulfillment engine 146 uses the on-device fulfillment model 146A to process the audio data 101 to generate fulfillment data 147A, and the client device 110 can use the fulfillment data 147A in performance 150 of one or more actions in response to the audio data 101. Moreover, in versions of these implementations, the assistant activation engine 124 can initially activate the currently dormant automated functionality to verify that the decision made at block 182 is correct (e.g., the audio data 101 actually does include the hotword "OK Assistant") by initially activating only the on-device speech recognizer 142 to determine that the audio data 101 includes the hotword "OK Assistant," and / or the assistant activation engine 124 can transmit the audio data 101 to one or more servers (e.g., the remote servers 160) to verify that the decision made at block 182 is correct (e.g., the audio data 101 actually does include the hotword "OK Assistant").
[0045] Returning to Figure 1AIf the client device 110 determines at block 182 that the predicted output 103 (e.g., 0.80) fails to satisfy the threshold (e.g., 0.85), the assistant activation engine 124 can refrain from initiating the currently dormant automated assistant function and / or shutting down any currently active automated assistant functions. In addition, if the client device 110 determines at block 182 that the predicted output 103 (e.g., 0.80) fails to satisfy the threshold (e.g., 0.85), the client device 110 can determine at block 184 whether additional user interface input is received. For example, the additional user interface input can be an additional spoken utterance that includes a hotword, an additional non-utterance physical motion that serves as a proxy for a hotword, an initiation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed“squeeze” of the client device 110 (e.g., invoking the automated assistant when the client device 110 is squeezed with at least a threshold amount of force), and / or other explicit automated assistant invocation. If the client device 110 determines at block 184 that no additional user interface input is received, the client device 110 can end at block 190.
[0046] However, if the client device 110 determines at block 184 that additional user interface input is received, the system can determine whether the additional user interface input received at block 184 includes a correction at block 186 that contradicts the decision made at block 182. If the client device 110 determines that the additional user interface input received at block 184 does not include a correction at block 186, the client device 110 can cease recognizing corrections and end at block 190. However, if the client device 110 determines that the additional user interface input received at block 184 includes a correction at block 186 that contradicts the initial decision made at block 182, the client device 110 can determine the true value output 105.
[0047] In some implementations, the gradient engine 126 can generate the gradient 106 based on the predicted output 103 to the ground truth output 105. For example, the gradient engine 126 can generate the gradient 106 based on comparing the predicted output 103 to the ground truth output 105. In some versions of those implementations, the client device 110 stores the predicted output 103 and the corresponding ground truth output 105 locally in the on-device storage 111, and the gradient engine 126 retrieves the predicted output 103 and the corresponding ground truth output 105 to generate the gradient 106 when one or more conditions are met. The one or more conditions can include, for example, that the client device is charging, that the client device has at least a threshold state of charge, that a temperature of the client device (based on one or more on-device temperature sensors) is less than a threshold, and / or that the client device is not being held by a user. In other versions of those implementations, the client device 110 provides the predicted output 103 and the ground truth output 105 to the gradient engine 126 in real-time, and the gradient engine 126 generates the gradient 106 in real-time.
[0048] Further, the gradient engine 126 can provide the generated gradient 106 to the on-device machine learning training engine 132A. The on-device machine learning training engine 132A uses the gradient 106 to update the on-device machine learning model 152A as it receives the gradient 106. For example, the on-device machine learning training engine 132A can utilize backpropagation and / or other techniques to update the on-device machine learning model 152A. Note that, in some implementations, the on-device machine learning training engine 132A can utilize batch techniques to update the on-device machine learning model 152A based on the gradient 106 and additional gradients determined locally at the client device 110 based on additional corrections.
[0049] Further, the client device 110 can communicate the generated gradient 106 to the remote system 160. When the remote system 160 receives the gradient 106, the remote training engine 162 of the remote system 160 uses the gradient 106 and additional gradients 107 from additional client devices 170 to update the global weights of the global hotword model 152A1. The additional gradients 107 from the additional client devices 170 can each be generated based on the same or similar techniques as described above with respect to the gradient 106 (but based on locally-identified failed hotword attempts that are specific to those client devices).
[0050] The update distribution engine 164 can provide the updated global weights and / or the updated global hotword model itself to the client device 110 and / or other client devices in response to one or more conditions being met, as indicated at 108. The one or more conditions can include, for example, a threshold duration and / or amount of training since the updated weights and / or the updated speech recognition model were last provided. The one or more conditions can additionally or alternatively include, for example, a measured improvement in the updated speech recognition model and / or a threshold duration has elapsed since the updated weights and / or the updated speech recognition model were last provided. When the updated weights are provided to the client device 110, the client device 110 can replace the weights of the on-device machine learning model 152A with the updated weights. When the updated global hotword model is provided to the client device 110, the client device 110 can replace the on-device machine learning model 152A with the updated global hotword model. In other implementations, the client device 110 can download a more suitable hotword model(s) from the server based on the type of command the user is expected to speak and replace the on-device machine learning model 152A with the downloaded hotword model.
[0051] In some implementations, the on-device machine learning model 152A is transmitted (e.g., by the remote system 160 or other component) for storage and use at the client device 110 based on a geographic region and / or other attributes of the client device 110 and / or a user of the client device 110. For example, the on-device machine learning model 152A can be one of N available machine learning models for a given language, but can be trained based on corrections specific to a particular geographic region, device type, context (e.g., music playing), etc., and provided to the client device 110 based on the client device 110 being primarily located in the particular geographic region.
[0052] Turning now to Figure 2 , the client device 110 is illustrated in implementations in which various on-device machine learning engines of Figure 1A and 1B are included as part of (or in communication with) the automated assistant client 240. The respective machine learning models are also illustrated as interfacing with the various on-device machine learning engines of Figure 1A and 1B . For simplicity, Figure 2 other components from Figure 1A and 1B are not shown in Figure 2 FIG. 13 illustrates one example of how the various on-device machine learning engines of Figure 1A and 1B and their respective machine learning models can be used by the automated assistant client 240 to perform various actions.
[0053] Figure 2 Client device 110 in the example of FIG. 1 is illustrated as having one or more microphones 211, one or more speakers 212, one or more cameras and / or other vision components 213, and a display 214 (e.g., a touch-sensitive display). Client device 110 can further include a pressure sensor, a proximity sensor, an accelerometer, a magnetometer, and / or other sensors for generating other sensor data in addition to audio data captured by the one or more microphones 211. Client device 110 at least selectively executes an automated assistant client 240. Automated assistant client 240 includes, in the example of FIG. 1, an on-device hotword detection engine 122B, an on-device speech recognizer 142, an on-device natural language understanding (NLU) engine 144, and an on-device fulfillment engine 146. Automated assistant client 240 further includes a speech capture engine 242 and a vision capture engine 244. Automated assistant client 140 can include additional and / or alternative engines, such as a voice activity detector (VAD) engine, an end-point detector engine, and / or other engines. Figure 2
[0054] One or more cloud-based automated assistant components 280 can optionally be implemented on one or more computing systems (collectively, a "cloud" computing system) that are communicatively coupled to client device 110 via one or more local and / or wide area networks (e.g., the Internet), generally indicated at 290. Cloud-based automated assistant components 280 can be implemented, for example, via a high-performance server cluster.
[0055] In various implementations, instances of automated assistant client 240, through their interaction with one or more cloud-based automated assistant components 280, can form what appears to be, from the perspective of a user, a logical instance of automated assistant 295 with which the user can engage in human-to-computer interactions (e.g., spoken language interactions, gesture-based interactions, and / or touch-based interactions).
[0056] Client device 110 can be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance such as a smart television (or a standard television equipped with a networked dongle having automated assistant functionality), and / or a wearable of a user that includes a computing device (e.g., a watch of a user that has a computing device, glasses of a user that have a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices can be provided.
[0057] The one or more vision components 213 can take various forms, such as a single camera, a stereo camera, a LIDAR component (or other laser-based component), a radar component, etc. The one or more vision components 213 can be used, e.g., by the vision capture engine 242, to capture vision frames (e.g., image frames, laser-based vision frames) of the environment in which the client device 110 is deployed. In some implementations, such vision frame(s) can be used to determine whether a user is present in the vicinity of the client device 110 and / or a certain distance of the user (e.g., the user’s face) relative to the client device 110. Such determinations can be used, e.g., to determine whether to activate the client device 110 and / or to activate the on-device machine learning engine 220. Figure 2 Various on-device machine learning engines and / or other engines depicted in FIG. 2.
[0058] The speech capture engine 242 can be configured to capture user speech and / or other audio data captured via the microphone 211. In addition, the client device 110 can include a pressure sensor, a proximity sensor, an accelerometer, a magnetometer, and / or other sensors for generating other sensor data in addition to audio data captured via the microphone 211. As described herein, the hotword detection engine 122B and / or other engines can utilize such audio data and other sensor data to determine whether to initiate one or more currently dormant automated assistant functions, refrain from initiating one or more currently dormant automated assistant functions, and / or close one or more currently active automated assistant functions. The automated assistant functions can include the on-device speech recognizer 142, the on-device NLU engine 144, the on-device fulfillment engine 146, and additional and / or alternative engines. For example, the on-device speech recognizer 142 can utilize an on-device speech recognition model 142A to process audio data capturing a spoken utterance to generate recognized text 143A corresponding to the spoken utterance. The on-device NLU engine 144 optionally performs on-device natural language understanding on the recognized text 143A utilizing an on-device NLU model 144A to generate NLU data 145A. The NLU data 145A can include, for example, an intent corresponding to the spoken utterance and optionally parameters (e.g., slot values) of the intent. In addition, the on-device fulfillment engine 146 optionally utilizes an on-device fulfillment model 146A to generate fulfillment data 147A based on the NLU data 145A. The fulfillment data 147A can define a local and / or remote response (e.g., an answer) to the spoken utterance, an interaction to be performed with a locally installed application based on the spoken utterance, a command to be transmitted to an Internet of Things (IoT) device based on the spoken utterance (directly or via a corresponding remote system), and / or other resolved actions to be performed based on the spoken utterance. The fulfillment data 147A is then provided for local and / or remote execution / implementation of the determined actions to resolve the spoken utterance. Execution can include, for example, presenting the local and / or remote response (e.g., visually and / or aurally presenting (optionally using a local text-to-speech module)), interacting with the locally installed application, transmitting the command to the IoT device, and / or other actions.
[0059] The display 214 can be used to display recognized text 143A and / or further recognized text 143B from the on-device speech recognizer 122, and / or one or more results from the execution 150. The display 214 can further be one of the user interface output components through which a visual portion of a response from the automated assistant client 240 is presented.
[0060] In some implementations, the cloud-based automated assistant components 280 can include a remote ASR engine 281 that performs speech recognition, a remote NLU engine 282 that performs natural language understanding, and / or a remote fulfillment engine 283 that generates fulfillment. A remote execution module that performs remote execution based on locally or remotely determined fulfillment data can also optionally be included. Additional and / or alternative remote engines can be included. As described herein, in various implementations, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution are prioritized at least due to the latency and / or network usage reductions they provide in resolving spoken utterances (due to not requiring client-server round trips to resolve spoken utterances). However, one or more cloud-based automated assistant components 280 can be used at least selectively. For example, such components can be used in parallel with on-device components, and output from such components in the event of a failure of the local components. For example, the on-device fulfillment engine 146 can fail in certain situations (e.g., due to the relatively limited resources of the client device 110), and the remote fulfillment engine 283 can leverage the more robust resources of the cloud to generate fulfillment data in such situations. The remote fulfillment engine 283 can operate in parallel with the on-device fulfillment engine 146, and its results can be used in the event of a failure of the on-device fulfillment, or can be invoked in response to a determination of a failure of the on-device fulfillment engine 146.
[0061] In various implementations, the NLU engine (on-device and / or remote) can generate NLU data that includes one or more annotations of the recognized text and one or more (e.g., all) terms of the natural language input. In some implementations, the NLU engine is configured to identify and annotate various types of grammatical information in the natural language input. For example, the NLU engine can include a morpheme module that can break individual words into morphemes and / or annotate the morphemes, e.g., with their part of speech. The NLU engine can also include a part-of-speech tagger that is configured to annotate terms with their grammatical roles. In addition, for example, in some implementations, the NLU engine can additionally and / or alternatively include a dependency parser that is configured to determine syntactic relationships between terms in the natural language input.
[0062] In some implementations, the NLU engine can additionally and / or alternatively include an entity tagger configured to annotate entity references in the one or more fragments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and imaginary), and the like. In some implementations, the NLU engine can additionally and / or alternatively include a coreference resolver (not depicted) configured to group or "cluster" references to the same entity based on one or more contextual cues. In some implementations, one or more components of the NLU engine can rely on annotations from one or more other components of the NLU engine.
[0063] The NLU engine can also include an intent matcher configured to determine an intent of a user engaged in an interaction with the automated assistant 295. The intent matcher can use various techniques to determine the intent of the user. In some implementations, the intent matcher can have access to one or more local and / or remote data structures that include, for example, a plurality of mappings between grammars and intents of responses. For example, grammars included in the mappings can be selected and / or learned over time and can represent common intents of users. For example, one grammar "play <artist>(“Play <artist>”) can be mapped to an intent that invokes a responsive action that causes <artist>The music of the artist "The Beatles" is played on the client device 110. Another syntax "[weather | forecast] today" can be matched to user queries such as "what's the weather today" and "what's the forecast for today?" In addition to or instead of syntax, in some implementations the intent matcher can employ one or more trained machine learning models, alone or in combination with one or more syntaxes. These trained machine learning models can be trained to recognize an intent, for example, by embedding recognized text from a spoken utterance into a reduced dimension space, and then determining which other embeddings (and thus, intents) are closest, e.g., using techniques such as Euclidean distance, cosine similarity, etc. From the above "play <artist>As can be seen in the example syntax, some syntaxes have slots (e.g., <artist>Slot values can be determined in various ways. Typically, the user will actively provide the slot value. For example, for the syntax "Order mea..." <topping>pizza (order me a <great> pizza)," the user can say the phrase "order me a sausage pizza" in which case the slot <topping>(<Excellent>) is automatically populated. Other slot values can be inferred based on, for example, user location, currently presented content, user preferences, and / or other cues.
[0064] A fulfillment engine (local and / or remote) can be configured to receive the predicted / estimated intent output by the NLU engine along with any associated slot values and fulfill (or "resolve") the intent. In various implementations, fulfillment (or "resolution") of the user intent can result in, for example, various fulfillment information (also referred to as fulfillment data) generated / obtained by the fulfillment engine. This can include determining a local and / or remote response (e.g., an answer) to the spoken utterance, an interaction with a locally installed application to be performed based on the spoken utterance, a command to be transmitted (directly or via a corresponding remote system) to an Internet of Things (IoT) device based on the spoken utterance, and / or other resolution actions to be performed based on the spoken utterance. On-device fulfillment can then initiate local and / or remote performance / execution of the determined actions to resolve the spoken utterance.
[0065] Figure 3 A flow diagram depicting an example method 300 of detecting media content and, in response, suppressing processing of a query included in audio data of the media content is illustrated. For convenience, the operations of method 300 are described with reference to a system that performs the operations. This system of method 300 includes one or more processors and / or other components of a client device. Moreover, while operations of method 300 are illustrated in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.
[0066] At block 310, the system receives, via one or more microphones of the client device, audio data that captures a spoken utterance.
[0067] At block 320, the system processes the audio data received at block 310 using one or more machine learning models to generate a predicted output that indicates a probability that the one or more hotwords are present in the audio data. The one or more machine learning models can be, for example, an on-device hotword detection model and / or other machine learning models. Each machine learning model can be a deep neural network or any other type of model and can be trained to recognize the one or more hotwords. Moreover, the generated output can be, for example, a probability and / or other likelihood measure.
[0068] At block 330, the system determines whether the predicted output generated at block 320 satisfies a threshold that indicates that the one or more hotwords are present in the audio data. If, at an iteration of block 330, the system determines that the predicted output generated at block 320 does not satisfy the threshold, the system proceeds to block 340 and the flow ends. On the other hand, if, at an iteration of block 330, the system determines that the predicted output generated at block 320 satisfies the threshold, the system proceeds to block 350.
[0069] Still referring to block 330, in examples, assume that the predicted output generated at block 320 is a probability and that the probability must be greater than 0.85 to satisfy the threshold at block 330, and that the predicted probability is 0.88. Based on the predicted probability 0.88 satisfying the threshold 0.85, the system proceeds to block 350.
[0070] At block 350, in response to determining at block 330 that the predicted output satisfies the threshold, the system processes the audio data received at block 310 using automatic speech recognition to generate a speech transcription feature or an intermediate embedding. In some implementations, the system generates a speech transcription feature, which is a transcription of the speech captured in the audio data. In other implementations, the system generates an intermediate embedding based on the audio data, which is an acoustic-based signal derived from the ASR engine.
[0071] Still referring to block 350, in other implementations, in response to determining at block 330 that the predicted output satisfies the threshold, the system further processes the audio data using speaker recognition techniques to determine a speaker vector corresponding to the query included in the audio data.
[0072] At block 360, in response to determining at block 330 that the predicted output satisfies the threshold, the system determines whether a watermark is embedded in the audio data received at block 310. In some implementations, the watermark is a human-imperceptible audio watermark. If the system does not detect a watermark embedded in the audio data in the iteration of block 360, the system proceeds to block 370 and initiates one or more currently dormant automated assistant functions. In some implementations, the one or more automated assistant functions include: speech recognition to generate recognized text, natural language understanding (NLU) to generate NLU output, generating a response based on the recognized text (e.g., from block 350) and / or the NLU output, transmitting the audio data to a remote server, transmitting the recognized text to a remote server, and / or directly triggering one or more actions (e.g., common tasks such as changing a device volume) in response to the first audio data. If, on the other hand, the system detects a watermark embedded in the audio data in the iteration of block 360, the system proceeds to block 380.
[0073] At block 380, the system determines whether the speech transcription feature or intermediate embedding generated at block 350 corresponds to one of a plurality of stored speech transcription features or intermediate embeddings. In some implementations, the plurality of stored speech transcription features or intermediate embeddings are stored on the client device and include a set of ambiguous speech transcriptions (e.g., erroneous transcriptions) for matching against errors and inaccuracies in the automated speech recognition process. In some implementations, determining whether the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an edit distance between the speech transcription feature and one of the plurality of stored speech transcription features satisfies a threshold edit distance. In other implementations, determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an embedding-based distance satisfies a threshold embedding-based distance. The embedding-based distance can be determined based on individual terms included in the query in the audio data or based on the entirety of the query included in the audio data.
[0074] If the system determines at the iteration of block 380 that the speech transcription feature or intermediate embedding does not correspond to one of the plurality of stored speech transcription features or intermediate embeddings, the watermark is not verified, and the system proceeds to block 370 and initiates one or more currently dormant automated assistant functions. In some implementations, the one or more automated assistant functions include speech recognition to generate recognized text or intermediate acoustic features (e.g., embeddings from an ASR engine), natural language understanding (NLU) to generate NLU output, generating a response based on the recognized text and / or NLU output, transmitting the audio data to a remote server, transmitting the recognized text to a remote server, and / or directly triggering one or more actions (e.g., common tasks such as changing a device volume) in response to the first audio data. If, on the other hand, the system determines at the iteration of block 380 that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings, the watermark is verified, and the system proceeds to block 390.
[0075] Still referring to block 380, in other implementations, the system determines whether the speaker vector generated at block 370 corresponds to one of a plurality of stored speaker vectors. If the system determines at the iteration of block 380 that the speaker vector does not correspond to one of the plurality of stored speaker vectors, the system proceeds to block 370 and initiates one or more currently dormant automated assistant functions. If, on the other hand, the system determines at the iteration of block 380 that the speaker vector corresponds to one of the plurality of stored speaker vectors, the watermark is verified, and the system proceeds to block 390.
[0076] In some implementations, the system can use one or more of the speech transcription features, the speaker vector, and any other stored factors (e.g., a date and / or time window in which the watermark is known to be active) to verify the watermark, as described above. If all factors successfully match (e.g., the speech transcription corresponds to a stored speech transcription, the speaker vector corresponds to a stored speaker vector, the current time falls within a stored time window in which the watermark is known to be active, etc.), the watermark is verified, and the system proceeds to block 390. On the other hand, if one or more factors do not successfully match, the watermark is not verified, and the system proceeds to block 370.
[0077] At block 390, in response to determining that the speech transcription features correspond to one of a plurality of stored speech transcription features or otherwise verifying the watermark at block 380 (e.g., by determining that the speaker vector corresponds to one of a plurality of stored speaker vectors or by determining that the current date and / or time falls within a stored date and / or time window in which the watermark is known to be active), the system suppresses processing of the query included in the audio data received at block 310.
[0078] Figure 4 A flow diagram depicting an example method 400 of detecting media content and adjusting hotword detection and / or automatic speech recognition in response is illustrated. For convenience, the operations of method 400 are described with reference to a system that performs the operations. This system of method 400 includes one or more processors and / or other components of a client device. Moreover, while operations of method 400 are illustrated in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted or added.
[0079] At block 405, the system receives, via one or more microphones of the client device, first audio data that captures a first spoken utterance.
[0080] At block 410, the system processes the first audio data received at block 405 using automatic speech recognition to generate a speech transcription feature or an intermediate embedding. In some implementations, the system generates a speech transcription feature that is a speech transcription captured in the first audio data. In other implementations, the system generates an intermediate embedding based on the first audio data, the intermediate embedding being an acoustic-based signal derived from an ASR engine.
[0081] Still with reference to block 410, in other implementations, the system further processes the audio data using speaker recognition techniques to determine a speaker vector corresponding to the query included in the audio data.
[0082] At block 415, the system determines whether the watermark is embedded in the first audio data received at block 405. In some implementations, the watermark is an inaudible audio watermark that is imperceptible to humans. If the system does not detect a watermark embedded in the first audio data at the iteration of block 415, the system proceeds to block 420 and the process ends. On the other hand, if the system detects a watermark embedded in the first audio data at the iteration of block 415, the system proceeds to block 425.
[0083] At block 425, the system determines whether the speech transcription feature or intermediate embedding generated at block 410 corresponds to one of a plurality of stored speech transcription features or intermediate embeddings. In some implementations, the plurality of stored speech transcription features or intermediate embeddings are stored on the client device and include a set of ambiguous speech transcriptions (e.g., erroneous transcriptions) for matching against errors and inaccuracies in the automated speech recognition process. In some implementations, determining whether the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an edit distance between the speech transcription feature and one of the plurality of stored speech transcription features satisfies a threshold edit distance. In other implementations, determining whether the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an embedding-based distance satisfies a threshold embedding-based distance. The embedding-based distance can be determined based on individual terms included in the query in the audio data or based on the entirety of the query included in the audio data.
[0084] If the system determines that the speech transcription feature or intermediate embedding does not correspond to one of the plurality of stored speech transcription features or intermediate embeddings at the iteration of block 425, the watermark is not verified, and the system proceeds to block 420 and the process ends. On the other hand, if the system determines that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings at the iteration of block 425, the watermark is verified, and the system proceeds to block 430.
[0085] Still referring to block 425, in other implementations, the system determines whether the speaker vector generated at block 410 corresponds to one of a plurality of stored speaker vectors. If the system determines that the speaker vector does not correspond to one of the plurality of stored speaker vectors at the iteration of block 425, the system proceeds to block 420 and the process ends. On the other hand, if the system determines that the speaker vector corresponds to one of the plurality of stored speaker vectors at the iteration of block 425, the watermark is verified, and the system proceeds to block 430.
[0086] In some implementations, the system can use one or more of the speech transcription features, the intermediate embedding, the speaker vector, and any other stored factors (e.g., a date and / or time window in which the watermark is known to be active) to verify the watermark, as described above. If all factors successfully match (e.g., the speech transcription corresponds to a stored speech transcription, the speaker vector corresponds to a stored speaker vector, the current date and / or time falls within a stored date and / or time window in which the watermark is known to be active, etc.), the watermark is verified, and the system proceeds to block 430. On the other hand, if one or more factors do not successfully match, the watermark is not verified, and the system proceeds to block 420, and the process ends.
[0087] In other implementations, the system can verify the watermark and proceed to block 430 if any one factor successfully matches or if a predetermined portion of the factors successfully match. In other implementations, the system can verify the watermark and proceed to block 430 based on a weighted combination of confidence scores generated based on the stored factors.
[0088] At block 430, in response to determining that the speech transcription feature or the intermediate embedding corresponds to one of a plurality of stored speech transcription features or intermediate embeddings or otherwise verifying the watermark (e.g., by determining that the speaker vector corresponds to one of a plurality of stored speaker vectors or by determining that the current date and / or time falls within a stored date and / or time window in which the watermark is known to be active), the system modifies a threshold indicating the presence of the one or more hotwords in the audio data.
[0089] At block 435, the system receives, via the one or more microphones of the client device, second audio data capturing a second spoken utterance.
[0090] At block 440, the system processes the second audio data received at block 435 using one or more machine learning models to generate an output predicting a probability of the presence of the one or more hotwords in the second audio data. The one or more machine learning models can be, for example, the on-device hotword detection model and / or other machine learning models. Each machine learning model can be a deep neural network or any other type of model and can be trained to recognize the one or more hotwords. Further, the generated output can be, for example, a probability and / or other likelihood measure.
[0091] At block 445, the system determines whether the predicted output generated at block 440 satisfies the modified threshold (from block 430) that indicates the presence of the one or more hotwords in the second audio data. If, at an iteration of block 445, the system determines that the predicted output generated at block 440 does not satisfy the modified threshold, the system proceeds to block 420 and the flow ends. On the other hand, if, at an iteration of block 445, the system determines that the predicted output generated at block 440 satisfies the modified threshold, the system proceeds to block 450.
[0092] Still referring to block 445, in an example, assume that the predicted output generated at block 440 is a probability and that the probability must be greater than 0.75 to satisfy the modified threshold at block 445, and that the predicted probability is 0.78. Based on the predicted probability of 0.78 satisfying the threshold of 0.75, the system proceeds to block 450.
[0093] At block 450, in response to determining that the predicted output satisfies the modified threshold, the system processes the query included in the second audio data. For example, as a result of processing the query, the system can initiate one or more currently dormant automated assistant functions. In some implementations, the one or more automated assistant functions include: speech recognition to generate recognized text, natural language understanding (NLU) to generate NLU output, generating a response based on the recognized text and / or the NLU output, transmitting the audio data to a remote server, transmitting the recognized text to a remote server, and / or directly triggering one or more actions in response to the first audio data (e.g., common tasks such as changing a device volume).
[0094] Figure 5 FIG. 5 is a block diagram of an example computing device 510 that can optionally be used to implement one or more aspects of the technology described herein. In some implementations, one or more of the client devices, cloud-based automated assistant components, and / or other components can include one or more components of the example computing device 510.
[0095] The computing device 510 typically includes at least one processor 514 which communicates with a number of peripheral devices via bus subsystem 512. These peripheral devices can include a storage subsystem 524, including, for example, a memory subsystem 525 and a file storage subsystem 526, user interface input devices 522, user interface output devices 520, and a network interface subsystem 516. The input and output devices allow user interaction with the computing device 510. Network interface subsystem 516 provides an interface to an external network (e.g., the Internet) and is coupled to corresponding interface devices in other computing devices.
[0096] User interface input devices 522 can include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information to computing device 510 or to communication network.
[0097] User interface input devices 522 can include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information to computing device 510 or to communication network.
[0098] Storage subsystem 524 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 524 can include the Figure 1A and 1B logic of the various components depicted in FIG. 5.
[0099] These software modules are generally executed by processor 514 alone or in combination with other processors. Memory subsystem 525 included in storage subsystem 524 can include a number of memories including a main random access memory (RAM) 530 for storage of instructions and data during program execution and a read only memory (ROM) 532 in which fixed instructions are stored. A file storage subsystem 526 can provide persistent (nonvolatile) storage for program and data files, and can include a hard disk drive, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical drive, or removable media cartridges. The modules implementing the functionality of certain implementations can be stored by file storage subsystem 526 in storage subsystem 524, or in another machine accessible by the processor(s) 514.
[0100] Bus subsystem 512 provides a mechanism for letting the various components and subsystems of computing device 510 communicate with each other as intended. Although bus subsystem 512 is shown schematically as a single bus, alternative implementations of the bus subsystem can use multiple busses.
[0101] Computing device 510 can be various types of devices including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 5 The description of the computing device 510 depicted is merely intended to be illustrative of some implementations for purposes of discussion. Many other configurations of the computing device 510 are possible having more or fewer components than those depicted. Figure 5 The computing device depicted can have more or fewer components than those depicted.
[0102] In situations in which the systems described herein collect or otherwise monitor personal information about users, or can make use of personal and / or monitored information, the users can be provided with an opportunity to control whether programs or features collect user information (e.g., information about a user's social network, social actions or activities, profession, a user's preferences, or a user's current geographic location), or to control whether and / or how to receive content from the content server that can be more relevant to the user. In addition, certain data can be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user's identity can be treated so that no personally identifiable information can be determined for the user, or a user's geographic location can be generalized where location information is obtained (such as to a city, postal code, or state level), so that a particular geographic location of a user cannot be determined. Thus, the user can have control over how information is collected about the user and / or used.
[0103] While several embodiments have been described and illustrated herein, a wide variety of other embodiments devised to perform the same functionality are possible which embody the same principles of the embodiments described herein, and should be considered within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which a teaching is used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that within the scope of the appended claims and equivalents thereto, embodiments can be practiced otherwise than as specifically described and claimed. Embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.< / topping> < / topping> < / artist> < / artist> < / artist> < / artist>
Claims
1. A method for audio processing, the method comprising: receiving, via one or more microphones of a client device, audio data capturing a spoken utterance; processing the audio data using one or more machine learning models to generate a prediction output indicating a probability that one or more hotwords are present in the audio data; determining that the prediction output satisfies a threshold indicating that the one or more hotwords are present in the audio data; in response to determining that the prediction output satisfies the threshold, processing the audio data using automatic speech recognition to generate a speech transcription feature or intermediate embedding; detecting a watermark embedded in the audio data; and in response to detecting the watermark: determining that the speech transcription feature or intermediate embedding corresponds to one of a plurality of stored speech transcription features or intermediate embeddings; and based on detecting the watermark and determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings, inhibiting processing of a query included in the audio data. The detecting the watermark is in response to determining that the prediction output satisfies the threshold. The watermark is a human-imperceptible audio watermark.
2. The method of claim 1, wherein, The plurality of stored speech transcription features or intermediate embeddings are stored on the client device.
3. The method of claim 1, wherein, Determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an edit distance between the speech transcription feature and one of the plurality of stored speech transcription features satisfies a threshold edit distance.
4. The method of claim 1, wherein, Determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an embedding-based distance satisfies a threshold embedding-based distance.
5. The method of claim 1, wherein, 7. The method of claim 1, further comprising, in response to detecting the watermark:
6. The method of claim 1, wherein, determining, using speaker recognition on the audio data, a speaker vector corresponding to the query included in the audio data; and determining that the speaker vector corresponds to one of a plurality of stored speaker vectors, the inhibiting processing of the query included in the audio data is further in response to determining that the speaker vector corresponds to one of the plurality of stored speaker vectors. determining whether a current time or date is within an active window of the watermark, and wherein wherein the inhibiting processing of the query included in the audio data is further in response to determining that the current time or date is within the active window of the watermark.
8. The method of claim 1, further comprising: The plurality of stored speech transcription features or intermediate embeddings include erroneous transcriptions.
10. A method for audio processing, the method comprising:
9. The method of any one of claims 1 to 8, wherein, receiving, via one or more microphones of a client device, first audio data capturing a first spoken utterance; processing the first audio data using automatic speech recognition to generate a speech transcription feature or intermediate embedding; detecting a watermark embedded in the first audio data; and in response to detecting the watermark: determining that the speech transcription feature or intermediate embedding corresponds to one of a plurality of stored speech transcription features or intermediate embeddings; and based on detecting the watermark and determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings, inhibiting processing of a query included in the audio data. modify a threshold indicating presence of the one or more hot words in audio data based on detecting the watermark and determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings.
11. The method of claim 10, further comprising: receiving, via the one or more microphones of the client device, second audio data capturing a second spoken utterance; processing the second audio data using one or more machine learning models to generate a prediction output indicating a probability of presence of one or more hot words in the second audio data; determining that the prediction output satisfies the modified threshold indicating presence of the one or more hot words in the second audio data; and processing a query included in the second audio data in response to determining that the prediction output satisfies the modified threshold.
12. The method of claim 10, wherein, the watermark is a human-imperceptible audio watermark.
13. The method of claim 10, wherein, the plurality of stored speech transcription features or intermediate embeddings are stored on the client device.
14. The method of claim 10, wherein, determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an edit distance between the speech transcription feature and one of the plurality of stored speech transcription features satisfies a threshold edit distance.
15. The method of claim 10, wherein, determining that the speech transcription feature or intermediate embedding corresponds to one of the plurality of stored speech transcription features or intermediate embeddings includes determining that an embedding-based distance satisfies a threshold embedding-based distance.
16. The method of claim 10, further comprising, in response to detecting the watermark: determining, using speaker recognition on the audio data, a speaker vector corresponding to a query included in the audio data; and determining that the speaker vector corresponds to one of a plurality of stored speaker vectors, wherein modifying the threshold indicating presence of the one or more hot words in audio data is further in response to determining that the speaker vector corresponds to one of the plurality of stored speaker vectors.
17. The method of claim 10, further comprising: determining whether a current time or date is within an active window of the watermark, and wherein modifying the threshold indicating presence of the one or more hot words in audio data is further in response to determining that the current time or date is within the active window of the watermark.
18. The method of any one of claims 10 to 17, wherein, the plurality of stored speech transcription features include erroneous transcriptions.
19. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to carry out the method of any one of claims 1 to 18.
20. A computer-readable storage medium comprising instructions which, when executed by one or more processors, cause the one or more processors to carry out the method of any one of claims 1 to 18.
21. A system comprising a processor, computer-readable memory, one or more computer-readable storage media, and program instructions collectively stored on the one or more computer-readable storage media and executable to carry out the method of any one of claims 1 to 18.
Citation Information
Patent Citations
System and method for measuring confusion among words in an adaptive speech recognition system
US20060064177A1
Speaker identification
US20150127342A1
Recorded media hotword trigger suppression
US20180130469A1
Recorded media hotword trigger suppression
US20180350356A1
Hotword Suppression
US20190362719A1