Hotword suppression
The audio watermarking-based approach using a convolutional neural network effectively distinguishes between live and recorded speech to prevent false hotword triggers, reducing energy consumption and enhancing privacy in voice-responsive systems.
Patent Information
- Application Number
- JP2023200953
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-21
- Filing Date
- 2023-11-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-05-22
AI Technical Summary
Existing voice-responsive systems face challenges in distinguishing between live and recorded audio inputs, leading to unintentional activation of virtual assistants due to false hotword triggers, which can result in energy consumption, equipment wear, and privacy concerns.
An audio watermarking-based approach using a convolutional neural network to detect watermarks in audio inputs, allowing the system to differentiate between live and recorded speech and suppress false hotword triggers.
Reduces unintentional activation of virtual assistants, conserving battery power and processing capacity, and preventing unnecessary network traffic while enhancing privacy and equipment reliability.
Smart Images

Figure 0007711152000001 
Figure 0007711152000002 
Figure 0007711152000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Patent Application No. 16 / 418,415, filed May 21, 2019, which claims the benefit of U.S. Patent Application No. 62 / 674,973, filed May 22, 2018, the contents of both applications being incorporated by reference herein.
[0002] This disclosure generally relates to automated speech processing.
Background Art
[0003] The reality of voice - enabled homes or other environments, i.e., environments where a user can simply speak a query or command out loud and a computer - based system receives the query, answers, and / or executes the command, is up to us. Voice - enabled environments (e.g., homes, workplaces, schools, etc.) can be implemented using a network of connected microphone devices distributed throughout various rooms or areas of the environment. Through such a network of microphones, a user has the ability to verbally query the system essentially from anywhere in the environment without the need for a computer or other device to be in front of or even near them. For example, while cooking in the kitchen, a user may ask the system "How many milliliters are in 3 cups?" and in response, receive an answer from the system, for example, in the form of a synthetic voice output. Alternatively, a user may ask the system questions such as "When does the nearest gas station close?" or "Should I wear a coat today?" when getting ready to leave the house.
[0004] Furthermore, a user may query the system and / or issue commands regarding the user's personal information. For example, a user may ask the system "When will I meet John?" or command the system to "Remind me to call John when I get home."
Summary of the Invention
Means for Solving the Problem
[0005] For a voice-responsive system, the method of interaction between the user and the system is designed to mainly use voice input, even if it is not exclusive. Therefore, the system may pick up all utterances made in the surrounding environment, including those not directed at the system, and must have some way to distinguish when an arbitrary given utterance is directed at the system, as opposed to, for example, being directed at an individual in the environment. One way to accomplish this is to use "hotwords" that are reserved as a given word or words spoken to draw the system's attention by agreement among the users in the environment. In one exemplary environment, the hotword used to draw the system's attention is the phrase "OK Computer". Thus, each time the phrase "OK Computer" is spoken, it is picked up by the microphone, transmitted to the system, and the system may either implement voice recognition techniques or use audio features and neural networks to determine whether the hotword has been spoken. If it has been spoken, it then awaits subsequent commands or queries. Thus, an utterance directed at the system takes the general form [HOTWORD][QUERY], where "HOTWORD" in this example is "OK Computer" and "QUERY" can be any question, command, statement, or other request that can be voice recognized, analyzed, and acted upon by the system, either alone or in conjunction with a server via a network.
[0006] The present disclosure discusses an audio watermarking-based approach for distinguishing recorded audio, such as broadcast audio or text-to-speech audio, from live audio. This distinction enables the detection of false hotword triggers in an input containing recorded audio and suppresses the false hotword triggers. However, live audio inputs from users are not watermarked, and hotwords in the audio inputs that are determined to be unwatermarked need not be suppressed. The watermark detection mechanism can use a convolutional neural network-based detector that is robust to noisy and reverberant environments and is designed to meet the goals of small footprint, both memory and computation, and low latency. The scalability advantage of this approach is prominent in preventing simultaneous hotword triggers on millions of devices during a television event with a large audience.
[0007] Hot-word based triggering can be a mechanism for activating a virtual assistant. Distinguishing hot words in live speech from recorded speech, such as in an advertisement, can be problematic as false hot word triggers can lead to unintentional activation of the virtual assistant. Moreover, where a user has installed a virtual assistant on multiple devices, it is possible for the audio output from one virtual assistant to include a hot word that unintentionally triggers another virtual assistant. Unintentional activation of a virtual assistant is generally undesirable. For example, where a virtual assistant is used to control home automation devices, unintentional activation of the virtual assistant can lead to, for example, lighting, heating or air conditioning equipment being unintentionally turned on, which can result in unnecessary energy consumption and be inconvenient for the user. Also, when a device is turned on, it may send messages to other devices (e.g., to retrieve information from other devices, to signal its status to other devices, to communicate with a search engine to perform a search, etc.), and thus turning on a device unintentionally can lead to unnecessary network traffic and / or unnecessary use of processing capacity, as well as unnecessary power consumption. Moreover, unintentional activation of equipment such as lighting, heating or air conditioning equipment can cause unnecessary wear on the equipment and potentially reduce its reliability. Further, as the range of virtual assistant controlled equipment and devices increases, the potential for unintentional activation of the virtual assistant to become potentially dangerous also increases. Also, unintentional activation of a virtual assistant can raise privacy concerns.
[0008] According to certain inventive aspects of the subject matter described in this application, a method for suppressing hotwords includes actions by a computing device to receive audio data corresponding to the playback of speech, and by the computing device to: provide the audio data as input to a model trained using watermarked audio data samples each including an audio watermark sample and non-watermarked audio data samples each not including an audio watermark sample, the model being configured to determine whether a given audio data sample includes an audio watermark; receive from the model trained using watermarked audio data samples each including an audio watermark sample and non-watermarked audio data samples each not including an audio watermark sample, data indicating whether the audio data includes an audio watermark, the model being configured to determine whether a given audio data sample includes an audio watermark; and based on the data indicating whether the audio data includes an audio watermark, determine whether to continue or stop processing the audio data.
[0009] These and other implementations may each optionally include one or more of the following features. The action of receiving data indicating whether audio data includes an audio watermark includes receiving data indicating that the audio data includes an audio watermark. The action of determining whether to continue or stop processing the audio data includes determining to stop processing the audio data based on receiving data indicating that the audio data includes an audio watermark. The action further includes causing the computing device to stop processing the audio data based on determining to stop processing the audio data. The action of receiving data indicating whether audio data includes an audio watermark includes receiving data indicating that the audio data does not include an audio watermark. The action of determining whether to continue or stop processing the audio data includes determining to continue processing the audio data based on receiving data indicating that the audio data does not include an audio watermark.
[0010] The action further includes, based on a determination that the audio data processing is to continue, the computing device continuing to process the audio data. The action of processing the audio data includes generating a transcription of the utterance by performing speech recognition on the audio data. The action of processing the audio data includes determining whether the audio data includes an utterance of a particular, predefined hotword. The action further includes the computing device determining that the audio data includes an utterance of a particular, predefined hotword before providing the audio data as input to a model trained using (i) watermarked audio data samples each including an audio watermark sample and (ii) non-watermarked audio data samples each not including an audio watermark sample, the action being configured to determine whether a given audio data sample includes an audio watermark. The action further includes the computing device determining that the audio data includes an utterance of a particular, predefined hotword. The action of providing the audio data as input to a model trained using (i) watermarked audio data samples each including an audio watermark sample and (ii) non-watermarked audio data samples each not including an audio watermark sample, the action being configured to determine whether a given audio data sample includes an audio watermark, responds to a determination that the audio data includes an utterance of a particular, predefined hotword.
[0011] An action includes receiving, by a computing device, watermarked audio data samples each including an audio watermark, watermarkless audio data samples each not including an audio watermark, and data indicating whether each watermarked and watermarkless audio sample includes an audio watermark; and training, by the computing device and using machine learning, a model using the watermarked audio data samples each including an audio watermark, the watermarkless audio data samples each not including an audio watermark, and the data indicating whether each watermarked and watermarkless audio sample includes an audio watermark. At least a portion of the watermarked audio data samples each include an audio watermark at a plurality of periodic locations. An audio watermark in one of the watermarked audio data samples is different from an audio watermark in another of the watermarked audio data samples. The action further includes determining, by the computing device, a first time of receipt of audio data corresponding to reproduction of an utterance; receiving, by the computing device, a second time at which an additional computing device provided audio data corresponding to reproduction of the utterance for output, and data indicating whether the audio data included a watermark; determining, by the computing device, that the first time matches the second time; and updating, by the computing device, the model using the data indicating whether the audio data included a watermark based on determining that the first time matches the second time.
[0012] Other implementations of this aspect include corresponding systems, devices, and computer programs recorded on computer storage devices, each configured to perform the operations of the method. Other implementations of this aspect include computer-readable media storing software comprising instructions executable by one or more computers, which instructions, when so executed, cause one or more computers to perform operations including any of the methods described herein.
[0013] Certain implementations of the subject matter described herein may be implemented to realize one or more of the following advantages. A computing device may respond to hotwords included in live audio but not to hotwords included in recorded media. This reduces or prevents unintentional activation of the device, thereby saving battery power and processing capacity of the computing device. Network bandwidth may also be conserved as fewer computing devices perform search queries when receiving hotwords with audio watermarks.
[0014] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0016] In the drawings, like reference numerals represent corresponding parts throughout.
[0017] FIG. 1 shows an exemplary system 100 for suppressing a hotword trigger when a “hotword” is detected in a recorded medium. Briefly and as described in more detail below, computing device 104 outputs an utterance 108 that includes an audio watermark 116 and an utterance of a predefined hotword 110. Computing device 102 detects the utterance 108 and determines, using an audio watermark identification model 158, that the utterance 108 includes an audio watermark 134. Based on the utterance 108 including the audio watermark 134, computing device 102 does not respond to the predefined hotword 110.
[0018] More specifically, computing device 104 is playing a Nugget World commercial. During the commercial, an actor in the commercial says utterance 108, namely, "Ok Computer, what's in a nugget?" Utterance 108 includes hotword 110 "Ok Computer" and query 112, the other words "what's in a nugget?" Computing device 104 outputs utterance 108 through a speaker. Any computing device near the microphone can detect utterance 108.
[0019] The audio of utterance 108 includes voice portion 114 and audio watermark 116. The creator of the commercial may add audio watermark 116 so that a computing device that detects utterance 108 does not respond to hotword 110. In some implementations, audio watermark 116 may include audio frequencies higher or lower than the human audible range. For example, audio watermark 116 may include frequencies greater than 20 kHz or less than 20 Hz. In some implementations, audio watermark 116 may include audio that is within the human audible range but is not detectable by humans because the sound is similar to noise. For example, audio watermark 116 may include a frequency pattern between 8 and 10 kHz. The intensity of different frequency bands may be imperceptible to humans but may be detectable by a computing device. As shown by frequency domain representation 115, utterance 108 includes audio watermark 116 in a frequency range higher than audible portion 114.
[0020] In some implementations, computing device 104 may use audio watermarking 120 to add a watermark to audio data 118. The audio data 118 may be a recorded utterance 108 such as "Ok computer, what's in the naget?" The audio watermarking 120 can add a watermark into the audio data 118 at periodic intervals. For example, the audio watermarking 120 may add a watermark every 200 milliseconds. In some implementations, computing device 104 can identify portions of the audio data 118 that include hotword 110, for example, by performing speech recognition. The audio watermarking 120 can add periodic watermarks across the audio of the hotword 110 before and / or after the hotword 110. For example, the audio watermarking 120 can add three (or any other number) watermarks across the audio of "ok computer" at periodic intervals.
[0021] Techniques for adding watermark 116 will be discussed in detail later with respect to FIGS. 3-7. Generally, each watermark 116 is different for each audio data sample. Audio watermarker 120 adds an audio watermark to the audio of utterance 108 every 200 or 300 milliseconds, and can add different or the same audio watermark every 200 or 300 milliseconds to the audio of the utterance "Ok Computer, order a cheese pizza". Audio watermarker 120 may generate a watermark for each audio sample such that the watermark minimizes distortion of the audio sample. This can be important because audio watermarker 120 can add a watermark within a frequency range that can be detected by a human. Computing device 104 may store the watermarked audio samples in watermarked speech 112 for later output by computing device 104.
[0022] In some implementations, each time computing device 104 outputs watermarked audio, computing device 104 may store data indicating the output audio in playback log 124. Playback log 124 may include data identifying any combination of the output audio 108, the date and time the audio 108 was output, computing device 104, the location of computing device 104, the transcription of audio 108, and the audio 108 without the watermark.
[0023] Computing device 102 detects utterance 108 through a microphone. Computing device 102 may be any type of device capable of receiving audio. For example, computing device 102 may be a desktop computer, laptop computer, tablet computer, wearable computer, cellular phone, smartphone, music player, e-book reader, navigation system, smart speaker and home assistant, wireless (e.g., Bluetooth) headset, hearing aid, smartwatch, smart glasses, activity tracker, or any other suitable computing device. As shown in FIG. 1, computing device 102 is a smartphone. Computing device 104 may be any device capable of outputting audio, such as, for example, a television, radio, music player, desktop computer, laptop computer, tablet computer, wearable computer, cellular phone, or smartphone. As shown in FIG. 1, computing device 104 is a television.
[0024] The microphone of computing device 102 may be part of audio subsystem 150. Audio subsystem 150 may include buffers, filters, analog-to-digital converters, each designed to first process the audio received through the microphone. The buffer may store the current audio received through the microphone and processed by audio subsystem 150. For example, the buffer stores the audio data for the previous five seconds.
[0025] Computing device 102 includes an audio watermark identifier 152. The audio watermark identifier 152 is configured to process audio received through a microphone and / or stored in a buffer, and identify an audio watermark included in the audio. The audio watermark identifier 152 may be configured to provide the processed audio as an input to an audio watermark identification model 158. The audio watermark identification model 158 may be configured to receive audio data and output data indicating whether the audio data includes a watermark. For example, the audio watermark identifier 152 may continuously provide audio processed through the audio subsystem 150 to the audio watermark identification model 158. As the audio watermark identifier 152 provides more audio, the accuracy of the audio watermark identification model 158 may increase. For example, after 300 milliseconds, the audio watermark identification model 158 may receive audio containing one watermark. After 500 milliseconds, the audio watermark identification model 158 may receive audio containing two watermarks. In an embodiment where all watermarks in any one audio sample are identical to each other, the audio watermark identification model 158 may improve accuracy by processing more audio.
[0026] In some implementations, the audio watermark identifier 152 may be configured to remove any detected watermark from the audio received from the audio subsystem 150. After removing the watermark, the audio watermark identifier 152 may provide the watermark-free audio to the hotword detector 154 and / or the speech recognizer 162. In some implementations, the audio watermark identifier 152 may be configured to pass the audio received from the audio subsystem 150 to the hotword detector 154 and / or the speech recognizer 162 without removing the watermark.
[0027] The hotword detector 154 is configured to identify hotwords in audio received through a microphone and / or stored in a buffer. In some implementations, the hotword detector 154 may be active at any time when the computing device 102 is powered on. The hotword detector 154 can continuously analyze the audio data stored in the buffer. The hotword detector 154 calculates a hotword reliability score that reflects the likelihood that the current audio data in the buffer contains a hotword. To calculate the hotword reliability score, the hotword detector 154 may extract audio features such as filter bank energy or mel-frequency cepstral coefficients from the audio data. The hotword detector 154 may use a classification window to process these audio features, for example, by using a support vector machine or a neural network. In some implementations, the hotword detector 154 does not perform speech recognition to determine the hotword reliability score (for example, by comparing the audio features extracted from the received audio with the corresponding audio features for one or more of the hotwords, without using the extracted audio features to perform speech recognition on the audio data). If the hotword reliability score satisfies a hotword reliability score threshold, the hotword detector 154 determines that the audio contains a hotword. For example, if the hotword reliability score is 0.8 and the hotword reliability score threshold is 0.7, the hotword detector 154 determines that the audio corresponding to utterance 108 contains the hotword 110. In some cases, the hotword may be referred to as an activation word or a watchword.
[0028] The speech recognizer 162 may perform any type of process that generates a transcription based on the incoming audio. For example, the speech recognizer 162 may use an acoustic model to identify phonemes in the audio data in the buffer. The speech recognizer 162 may use a language model to determine the transcription corresponding to the phonemes. As another example, the speech recognizer 162 may use a single model that processes the audio data in the buffer and outputs a transcription.
[0029] In the case where the audio watermark identification model 158 determines that the audio contains a watermark, the audio watermark identifier 152 may deactivate the speech recognizer 162 and / or the hotword 154. By deactivating the speech recognizer 162 and / or the hotword 154, the audio watermark identifier 152 can prevent further processing of the audio that may trigger the computing device 102 to respond to the hotword 110 and / or the query 112. As shown in FIG. 1, the audio watermark identifier 152 sets the hotword 154 to the inactive state 156 and the speech recognizer 162 to the inactive state 160.
[0030] In some implementations, the default state of the hotword detector 154 may be the active state, and the default state of the speech recognizer 162 may be the active state. In this case, the inactive state 156 and the inactive state 160 may expire after a predetermined amount of time. For example, after 5 seconds (or another predetermined amount of time), the states of both the hotword detector 154 and the speech recognizer 162 may return to the active state. The 5-second period may be updated each time the audio watermark identifier 152 detects an audio watermark. For example, if the audio 115 of the utterance 108 includes a watermark throughout the duration of the audio, the hotword detector 154 and the speech recognizer 162 may be set to the inactive state 156 and the inactive state 160, and may remain in that state for an additional 5 seconds after the computing device 104 outputs the utterance 108. As another example, if the audio 115 of the utterance 108 includes a watermark throughout the utterance of the hotword 110, the hotword detector 154 and the speech recognizer 162 may be set to the inactive state 156 and the inactive state 160, and may remain in that state for an additional 5 seconds after the computing device 104 outputs the hotword 110 that overlaps with the output of the query 112.
[0031] In some implementations, the audio watermark identifier 152 may store data indicating the date and time at which the audio watermark identifier 152 identified a watermark in the identification log 164. For example, the audio watermark identifier 152 may identify a watermark in the audio of the utterance 108 at 3:15 PM on June 10, 2019. The identification log 164 may store data identifying any combination of the time and date of receipt of the watermark, the transcription of the utterance including the watermark 134, the computing device 102, the watermark 134, the location of the computing device 102 when the watermark was detected, the base audio 132, the combination of the audio and the watermark, and any audio detected before or after a certain time period of the utterance 108 or the watermark 134.
[0032] In some implementations, the audio watermark identifier 152 may store in the identification log 164 data indicating the date and time when the hot word recognizer 154 recognized a hot word while the audio watermark identifier 152 did not identify a watermark. For example, at 7:15 PM on June 20, 2019, the audio watermark identifier 152 may not identify a watermark in the uttered audio, and the hot word recognizer 154 may identify a hot word in the uttered audio. The identification log 164 may store data identifying the time and date of receipt of audio without a watermark and hot words, the transcription of the utterance, the computing device 102, the location of the computing device, and any combination of audio detected before or after a certain time period from the utterance or hot word.
[0033] In some implementations, the hot word recognizer 154 can process the audio received from the audio subsystem 150 before, after, or simultaneously with the audio watermark identifier 152. For example, the audio watermark identifier 152 may determine that the audio of the utterance 108 includes a watermark, and at the same time, the hot word recognizer 154 may determine that the audio of the utterance 108 includes a hot word. In this case, the audio watermark identifier 152 may set the state of the speech recognizer 162 to the inactive state 160. The audio watermark identifier 152 may not be able to update the state 156 of the hot word recognizer 154.
[0034] In some implementations, before the audio watermark identifier 152 uses the audio watermark identification model 158, the computing device 106 generates watermark identification data 130 and provides the watermark identification data 130 to the computing device 102. The computing device 106 uses the watermark-free voice samples 136, the audio watermarker 138, and a trainer 144 that uses machine learning to generate the audio watermark identification model 148.
[0035] The watermark-free voice samples 136 can include various voice samples collected under various conditions. The watermark-free voice samples 136 can include audio samples of different users, such as speaking different words, speaking the same words, speaking words with different types of background noise, speaking words in different languages, speaking words with different accents, speaking words recorded by different devices, etc. In some implementations, each of the watermark-free voice samples 136 includes the utterance of a hotword. In some implementations, only some of the watermark-free voice samples 136 include the utterance of a hotword.
[0036] The audio watermarker 138 can generate different watermarks for each watermark-free audio sample. The audio watermarker 138 can generate one or more watermarked audio samples 140 for each watermark-free audio sample. Using the same watermark-free audio sample, the audio watermarker 138 can generate a watermarked audio sample containing a watermark every 200 milliseconds and another watermarked audio sample containing a watermark every 300 milliseconds. The audio watermarker 138 can also generate, if any, a watermarked audio sample containing a watermark that only overlaps with the hotword. The audio watermarker 138 can also generate a watermarked audio sample that overlaps with the hotword and contains a watermark preceding the hotword. In this case, the audio watermarker 138 can create four different watermarked audio samples using the same watermark-free audio sample. The audio watermarker 138 can also create more or fewer than four. In some cases, the audio watermarker 138 may operate similarly to the audio watermarker 120.
[0037] The trainer 144 generates an audio watermark identification model 148 using machine learning and training data including the watermark-free audio samples 136 and the watermarked audio samples 140. Since the watermark-free audio samples 136 and the watermarked audio samples 140 are labeled as either including or not including a watermark, the trainer 144 may use training data including the watermark-free audio samples 136, and a label indicating that each sample does not include a watermark, as well as the watermarked audio samples 140, and a label indicating that each sample includes a watermark. The trainer 144 uses machine learning to generate an audio watermark identification model 148 that can receive an audio sample and output whether the audio sample includes a watermark.
[0038] The computing device 106 can provide the model 128 to the computing device 102 for use when accessing the audio watermark identification model 148 and processing the received audio data. The computing device 102 can store the model 128 in the audio watermark identification model 158.
[0039] Computing device 106 may update the audio watermark identification model 148 based on the playback log 142 and the identification log 146. The playback log 142 may include data such as playback data 126 received from computing device 104 and stored in playback log 124. The playback log 142 may include playback data from a plurality of computing devices that output watermarked audio. The identification log 146 may include data such as identification data 130 received from computing device 102 and stored in identification log 164. The identification log 146 may include additional identification data from a plurality of computing devices configured to identify an audio watermark and prevent execution of any commands or queries included in the watermarked audio.
[0040] The trainer 144 may compare the playback log 142 and the identification log 146 to identify matching entries indicating that a computing device output watermarked audio and another computing device identified the watermark in the watermarked audio. The trainer 144 may also identify watermark identification errors in the identification log 146 and the playback log 142. A first type of watermark identification error may occur when the identification log 146 indicates that a computing device identified a watermark, but the playback log 142 does not indicate output of watermarked audio. A second type of watermark identification error may occur when the playback log 142 indicates output of watermarked audio, but the identification log 146 indicates that a computing device in the vicinity of the watermarked audio did not identify the watermark.
[0041] The trainer 144 may update the error and use the corresponding audio data as additional training data for updating the audio watermark identification model 148. The trainer 144 may use the audio to update the audio watermark identification model 148 if the computing device properly identifies the watermark. The trainer 144 may use both the audio output by the computing device and the audio detected by the computing device as training data. The trainer 144 may use machine learning and the audio data stored in the playback log 142 and the identification log 146 to update the audio watermark identification model 148. The trainer 144 may use the watermarking labels given in the playback log 142 and the identification log 146 and the corrected labels from the error identification technique described above as part of the machine learning training process.
[0042] In some implementations, computing device 102 and some other computing devices may be configured to send audio 115 to a server for processing by a server-based hotworder and / or a server-based speech recognizer running on the server. Audio watermark identifier 152 may indicate that audio 115 does not contain an audio watermark. Based on that determination, computing device 102 may send the audio to the server for further processing by a server-based hotworder and / or a server-based speech recognizer. The audio watermark identifiers of some other computing devices may also indicate that audio 115 does not contain an audio watermark. Based on those determinations, each of the other computing devices may send its respective audio to the server for further processing by a server-based hotworder and / or a server-based speech recognizer. The server may determine whether the audio from each computing device contains a hotword and / or generate a transcription of the audio and return the results to each computing device.
[0043] In some implementations, the server may receive data indicating a watermark reliability score for each of the watermark determinations. The server may determine that the audio received by computing device 102 and other computing devices is from the same source based on the locations of computing device 102 and the other computing devices, the characteristics of the received audio, the fact that each audio portion was received at a similar time, and any other similar indicators. In some cases, each of the watermark reliability scores may be within a specific range that includes a watermark reliability score threshold at one end of the range and another reliability score, such as less than 5 percent, which is a percent difference from the watermark reliability score threshold. For example, the range may be a watermark reliability score threshold of 0.80 to 0.76. In other cases, the other end of the range may be a fixed distance from the watermark reliability score threshold, such as 0.05. For example, the range may be a watermark reliability score threshold of 0.80 to 0.75.
[0044] If the server determines that each of the watermark reliability scores is within a range that is close to but does not meet the watermark reliability score threshold, the server may determine that the watermark reliability score threshold should be adjusted. In this case, the server may adjust the watermark reliability score threshold to the lower end of the range. In some implementations, the server may update the watermarked audio sample 140 by including the audio received from each computing device in the watermarked audio sample 140. The trainer 144 may update the audio watermark identification model 148 using machine learning and the updated watermarked audio sample 140.
[0045] FIG. 1 shows three different computing devices that implement the different functions described above, but any combination of one or more computing devices can implement any combination of functions. For example, instead of separate computing device 106 training audio watermark identification model 148, computing device 102 could train audio watermark identification model 148.
[0046] FIG. 2 shows an exemplary process 200 for suppressing a hotword trigger when a hotword is detected in a recorded medium. Generally, process 200 processes received audio to determine whether the audio contains an audio watermark. If the audio contains an audio watermark, process 200 may suppress further processing of the audio. If the audio does not contain an audio watermark, process 200 processes the audio and continues to execute any queries or commands included in the audio. Process 200 is described as being implemented by a computer system comprising one or more computers, such as computing devices 102, 104, and / or 106 shown in FIG. 1.
[0047] The system receives (210) audio data corresponding to the playback of speech. For example, a television may be playing a commercial, and an actor in the commercial may say, "Ok computer, turn on the light." The system includes a microphone that detects the audio of the commercial, including the actor's speech.
[0048] The system is configured to (i) determine whether a given audio data sample includes an audio watermark, and (ii) provide the audio data as an input to a model trained using watermarked audio data samples each including an audio watermark sample and unwatermarked audio data samples each not including an audio watermark sample (220). In some implementations, the system may determine that the audio data includes a hotword. Based on detecting the hotword, the system provides the audio data as an input to the model. For example, the system may determine that the audio data includes "ok computer". Based on detecting "ok computer", the system provides the audio data to the model. The system may provide the portion of the audio data that included the hotword and the audio received after the hotword. In some cases, the system may provide the portion of the audio from before the hotword.
[0049] In some implementations, the system may analyze the audio data to determine whether the audio data includes a hotword. The analysis may occur before or after providing the audio data as an input to the model. In some implementations, the system can train the model using machine learning and watermarked audio data samples each including an audio watermark, unwatermarked audio data samples each not including an audio watermark, and data indicating whether each watermarked and unwatermarked audio sample includes an audio watermark. The system can train the model to output data indicating whether the audio input to the model includes a watermark or does not include a watermark.
[0050] In some implementations, audio signals with different watermarks may include different watermarks (all watermarks in any one audio sample may be the same as each other, but the watermark in one audio signal may be different from the watermark in another audio signal). The system may generate different watermarks for each audio signal to minimize distortion in the audio signal. In some implementations, the system may place the watermarks at periodic intervals in the audio signal. For example, the system may place the watermarks every 200 milliseconds. In some implementations, the system may place the watermarks over an audio including a hotword and / or a time period before the hotword.
[0051] The system is configured to (i) determine whether a given audio data sample includes an audio watermark, and (ii) receive data indicating whether the audio data includes an audio watermark from a model trained using watermarked audio data samples that include an audio watermark and watermark-free audio data samples that do not include an audio watermark (230). The system may receive an indication that the audio data includes a watermark or an indication that the audio data does not include a watermark.
[0052] The system continues or stops processing the audio data (240) based on data indicating whether the audio data contains an audio watermark. In some implementations, the system may stop processing the audio data if the audio data contains an audio watermark. In some implementations, the system may continue processing the audio data if the audio data does not contain an audio watermark. In some implementations, processing the audio data may include performing speech recognition on the audio data and / or determining whether the audio data contains a hotword. In some implementations, the processing may include executing a query or command included in the audio data.
[0053] In some implementations, the system logs the time and date at which the system received the audio data. The system may compare this time and date with the time and date received from the computing device that outputs the audio data. If the system determines that the date and time of receipt of the audio data match the date and time at which the audio data was output, the system may update the model using the audio data as additional training data. When determining whether the audio data contains a watermark, the system may identify whether the model was correct and ensure that the audio data contains the correct watermark label when added to the training data.
[0054] More specifically, software agents capable of performing tasks for users are generally referred to as "virtual assistants". Virtual assistants may be activated, for example, by voice input from a user, i.e., they may be programmed to recognize one or more trigger words that, when spoken by the user, activate the virtual assistant and cause it to perform tasks associated with the spoken trigger words. Such trigger words are often referred to as "hot words". A virtual assistant may be provided, for example, on a user's computer, mobile phone, or other user device. Alternatively, the virtual assistant may be integrated into another device, such as a so-called "smart speaker" (a type of wireless speaker with an integrated virtual assistant that provides interactive actions and hands-free activation with the help of one or more hot words).
[0055] With the widespread adoption of smart speakers, additional problems arise. During events with a large audience, such as a sports event that attracts over 100 million viewers, ads with hot words may lead to the simultaneous triggering of virtual assistants. The presence of a large number of viewers may result in a significant increase in simultaneous queries to the speech recognition server, which can lead to a denial of service (DOS).
[0056] Two possible mechanisms for filtering out fake hot words are: (1) audio fingerprinting, where a fingerprint from the query audio is compared to a database of fingerprints from known audio, such as ads, to filter out false triggers; and (2) audio watermarking, where the audio is watermarked by the source and the queries recorded by the virtual assistant are examined for watermarks for filtering.
[0057] The present disclosure describes the design of a low-latency, small-footprint watermark detector using convolutional neural networks. This watermark detector is trained to be robust against noisy and reverberant environments that may be frequent in the target scenario.
[0058] Audio watermarking can be used in copyright protection and second-screen applications. In copyright protection, watermark detection generally does not need to be sensitive to latency when the entire audio signal can be responded to for detection. In the case of second-screen applications, the latency introduced by high-latency watermark detection may be acceptable. Different from these two scenarios, watermark detection in virtual assistants is very sensitive to latency.
[0059] In known applications involving watermark detection, the embedded message that constitutes the watermark is usually not known in advance, and the watermark detector has to decrypt the message sequence before it can determine whether the message sequence contains a watermark. If it does, the watermark has to be determined. However, in some applications described herein, the watermark detector may be detecting a watermark pattern that is exactly known to the decoder / watermark detector. That is, the originator or provider of the recorded audio content may watermark it with a watermark and make the details of the watermark available, for example, to the provider of the virtual assistant and / or the provider of the device that includes the virtual assistant. Similarly, the provider of the virtual assistant may arrange for the audio output from the virtual assistant to be watermarked and make the details of the watermark available. As a result, when a watermark is detected in the received message, it can be determined that the received message is not a live audio input from the user, and any activation of the virtual assistant resulting from a hot word in the received message can be suppressed without having to wait until the entire message is received and processed. This reduces latency.
[0060] Some implementations for hot word suppression use audio fingerprinting techniques. This technique requires a database of known audio fingerprint data. Maintaining this database on a device is non-trivial, so the device-side deployment of such a solution is not feasible. However, a major advantage of the audio fingerprinting technique is that it does not require modification to the audio distribution process. Thus, this technique can also address adversarial scenarios where the audio originator is not a collaborator.
[0061] The present disclosure describes a watermark-based hotword suppression mechanism. The hotword suppression mechanism can be used for device deployment that brings design constraints on memory and computational footprint. Further, there are constraints on latency to avoid affecting the user experience.
[0062] Watermark-based approaches may require modification of the audio publishing process to add watermarks. Thus, these approaches may only be used to detect audio published by collaborators. However, these approaches do not need to maintain a fingerprint database. This feature enables several advantages.
[0063] The first advantage can be the feasibility of device deployment. This can be an advantage during high-audience events where several virtual assistants may be triggered simultaneously. Server-based solutions for detecting these false triggers can lead to service denial due to the scale of simultaneous triggers. The second advantage can be the detection of unknown audio published by collaborators, such as text-to-speech (TTS) synthesizer output where the publisher may be cooperative but the audio is not known in advance. The third advantage can be scalability. Entities such as audio / video publishers on online platforms can watermark their audio to avoid triggering virtual assistants. In some implementations, these platforms host millions of hours of content that cannot actually be handled using audio fingerprinting-based approaches.
[0064] In some implementations, the watermark-based approaches described herein can be combined with audio fingerprinting-based approaches that may have the ability to deal with adversarial agents.
[0065] The following description relates to a watermark embedder and a watermark detector.
[0066] The watermark embedder can be based on spectrum-spreading-based watermarking in the FFT domain. The watermark embedder may use a psychoacoustic model to estimate the minimum masking threshold (MMT) used to form the amplitude of the watermark signal.
[0067] To summarize this technique, regions of the host signal that are to receive the watermark addition are selected based on a minimum energy criterion. Discrete Fourier Transform (DFT) coefficients are estimated for every host signal frame (25 ms window - 12.5 ms hop) in these regions. These DFT coefficients are used to estimate the minimum masking threshold (MMT) using a psychoacoustic model. The MMT is used to form the magnitude spectrum for a frame of the watermark signal. Figure 3 presents the estimated MMT along with the host signal energy and the absolute threshold of hearing. The phase of the host signal may be used for the watermark signal, and the sign of the DFT coefficients is determined from the message payload. The message bit payload can be spread over chunks of frames using multiple scrambling. In some implementations, the system may be detecting whether an inquiry is watermarked and may not need to transmit any payload. Thus, the system may randomly select a sign matrix over chunks of frames (e.g., 16 frames or 200 ms) and repeat this sign matrix over the watermarking region. This repetition of the sign matrix can be utilized to post-process the watermark detector output and improve the detection performance. By adding an overlap of individual watermark frames, a watermark signal can be generated. Subplots (a) and (b) of Figure 2 represent the magnitude spectra of the host signal and the watermark signal, and subplot (c) represents the sign matrix. The vertical lines represent the boundaries between two replicas of the matrix.
[0068] The watermark signal can be added to the host signal after scaling the host signal in the time domain by a certain factor (e.g., α ∈ [0, 1]) in order to further ensure the inaudibility of the watermark. In some implementations, α is determined iteratively using an objective evaluation metric such as Perceptual Evaluation of Audio Quality (PEAQ). In some implementations, the system can use traditional scaling factors (e.g., α ∈ {0.1, 0.2, 0.3, 0.4, 0.5}) and evaluate the detection performance for each of these scaling factors.
[0069] In some implementations, the design requirements for the watermark detector can be device - deployable, imposing significant constraints on both the memory footprint of the model and the complexity of its computation. The following description is about a convolutional neural network - based model architecture for keyword detection on a device. In some implementations, the system can use a temporal convolutional neural network.
[0070] In some implementations, the neural network is trained to estimate the cross - correlation of an embedded watermark symbol matrix (Figure 4, sub - plot (c)), which can be a replica of the same 200 - ms pattern as one instance of the 200 - ms pattern. Sub - plot (d) of Figure 4 shows the cross - correlation. The cross - correlation can encode information about the start of each symbol matrix block and can be non - zero over the entire duration of the watermark signal in the host signal.
[0071] The system can train a neural network using a multitask loss function. The primary task may be the estimation of the ground truth cross-correlation, and the auxiliary task may be the estimation of the energy perturbation pattern and / or the watermark scale spectrum. The mean squared error between the ground truth and the network output can be calculated. After scaling the auxiliary loss by a normalization constant, some or all of the loss can be interpolated. In some implementations, the performance can be improved by bounding each network output to exactly cover the dynamic range of the corresponding ground truth.
[0072] In some implementations, the system may post-process the network output. In some implementations, the watermark may not have a payload message and a single code matrix is replicated throughout the watermarking area. This can result in a periodic cross-correlation pattern (Figure 4, subplot (d)). This aspect can be utilized to eliminate spurious peaks in the network output. In some implementations, and to improve performance, the system may use a matching filter created by replicating the cross-correlation pattern (see Figure 6) across a bandpass filter that separates the frequencies of interest. Figure 7 compares the network output generated for the signal without a watermark before and after matched filtering. In some implementations, spurious peaks that do not have periodicity can be significantly suppressed. The ground truth 705 may be about 0.0 (e.g., between -0.01 and 0.01) and may track the x-axis more carefully than the network output 710 and the matched-filtered network output 720. The network output 710 may vary with respect to the x-axis more than the ground truth 705 and the matched-filtered network output 720. The matched-filtered network output 720 may track the x-axis more carefully than the network output 710 and may not need to track the x-axis as carefully as the ground truth 705. The matched-filtered network output 720 may be smoother than the network output 710. The matched-filtered network output 720 may stay within a narrower range than the network output 710. For example, the matched-filtered network output 720 may stay between -0.15 and 0.15. The network output 710 may stay between -0.30 and 0.60.
[0073] Once trained, a neural network can be used in a method for determining whether a given audio data sample contains an audio watermark by applying a model embodying the neural network to the audio data sample. The method may include determining a confidence score reflecting the likelihood that the audio data contains an audio watermark, comparing the confidence score reflecting the likelihood that the audio data contains an audio watermark with a confidence score threshold, and determining whether to perform additional processing on the audio data based on comparing the confidence score reflecting the likelihood that the audio data contains an audio watermark with the confidence score threshold.
[0074] In one embodiment, the method includes determining that a reliability score satisfies a reliability score threshold based on comparing the reliability score, which reflects the likelihood that audio data includes an audio watermark, to the reliability score threshold, and determining whether to perform additional processing on the audio data includes determining to refrain from performing additional processing on the audio data. In one embodiment, the method includes determining that a reliability score does not satisfy the reliability score threshold based on comparing the reliability score, which reflects the likelihood that speech includes an audio watermark, to the reliability score threshold, and determining whether to perform additional processing on the audio data includes determining to perform additional processing on the audio data. In one embodiment, the method includes receiving, from a user, data for confirming performance of additional processing on the audio data, and updating a model based on receiving the data for confirming performance of additional processing on the audio data. In one embodiment, performing additional processing on the audio data includes performing an action based on a transcription of the audio data or determining whether the audio data includes a particular, predefined hotword. In one embodiment, the method includes, prior to applying a model trained using (i) audio data samples with watermarks that include an audio watermark and (ii) audio data samples without watermarks that do not include an audio watermark, determining that the audio data includes a particular, predefined hotword, configured to determine whether a given audio data sample includes an audio watermark.In one embodiment, the method includes determining that the audio data includes a particular, predefined hotword, and in response to determining that the audio data includes a particular, predefined hotword, applying to the audio data a model configured to (i) determine whether a given audio data sample includes an audio watermark and (ii) trained with watermarked audio data samples that include an audio watermark and unwatermarked audio data samples that do not include an audio watermark. In one embodiment, the method includes receiving watermarked audio data samples that include an audio watermark and unwatermarked audio data samples that do not include an audio watermark, and training a model that uses machine learning with the watermarked audio data samples that include an audio watermark and the unwatermarked audio data samples that do not include an audio watermark. In one embodiment, at least a portion of the watermarked audio data samples includes audio watermarks at multiple periodic locations.
[0075] FIG. 8 shows examples of a computing device 800 and a mobile computing device 850 that can be used to implement the techniques described herein. Computing device 800 is intended to represent various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. Mobile computing device 850 is intended to represent various forms of mobile devices, such as a personal digital assistant, cellular phone, smartphone, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are intended to be exemplary only, and are not intended to be limiting.
[0076] The computing device 800 includes a processor 802, a memory 804, a storage device 806, a high-speed interface 808 that connects to the memory 804 and a plurality of high-speed expansion ports 810, and a low-speed interface 812 that connects to a low-speed expansion port 814 and the storage device 806. Each of the processor 802, the memory 804, the storage device 806, the high-speed interface 808, the high-speed expansion ports 810, and the low-speed interface 812 are interconnected using various buses and may be mounted on a common motherboard or in other manners as required. The processor 802 can process instructions for execution within the computing device 800, including instructions stored in the memory 804 or on the storage device 806 for displaying graphical information about a GUI on an external input / output device such as a display 816 coupled to the high-speed interface 808. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and types of memory, as required. Also, multiple computing devices may be connected, and each device may provide a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
[0077] The memory 804 stores information within the computing device 800. In some implementations, the memory 804 is one or more volatile memory units. In some implementations, the memory 804 is one or more non-volatile memory units. The memory 804 may also be another form of computer-readable medium, such as a magnetic or optical disk.
[0078] The memory device 806 can provide large-capacity storage to the computing device 800. In some implementations, the memory device 806 is a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other devices in other configurations, or may include them. Instructions can be stored in an information carrier. When executed by one or more processing devices (e.g., the processor 802), the instructions implement one or more methods as described above. The instructions can also be stored by one or more memory devices (e.g., the memory 804, the memory device 806, or the memory on the processor 802) such as a computer or machine-readable medium.
[0079] The high-speed interface 808 manages bandwidth-consuming operations for the computing device 800, and the low-speed interface 812 manages more bandwidth-low-consuming operations. Such an allocation of functions is merely an example. In some implementations, the high-speed interface 808 is coupled to the memory 804, to the display 816 (e.g., through a graphics processor or accelerator), and to the high-speed expansion port 810 that can receive various expansion cards. In this implementation, the low-speed interface 812 is coupled to the memory device 806 and the low-speed expansion port 814. The low-speed expansion port 814 can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), but can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, etc., or a network device such as a switch or a router, e.g., through a network adapter.
[0080] Computing device 800 may be implemented in several different forms, as shown in the figures. For example, it may be implemented as a standard server 820, or multiple times as a group of such servers. Additionally, it may be implemented in a personal computer such as a laptop computer 822. It may also be implemented as part of a rack server system 824. Alternatively, components from computing device 800 may be combined with other components in a mobile device such as mobile computing device 850. Each of such devices may include one or more of computing device 800 and mobile computing device 850, and the overall system may be made up of multiple computing devices that communicate with each other.
[0081] Mobile computing device 850 includes, among other components, a processor 852, a memory 864, input / output devices such as a display 854, a communication interface 866, and a transceiver 868. A storage device such as a microdrive or other device may be provided in mobile computing device 850 to provide additional storage. Each of processor 852, memory 864, display 854, communication interface 866, and transceiver 868 are interconnected using various buses, and some of the components may be mounted on a common motherboard or in other ways as needed.
[0082] Processor 852 can execute instructions within mobile computing device 850, including instructions stored in memory 864. Processor 852 may be implemented as a chipset of chips including separate and multiple analog and digital processors. Processor 852 may enable the coordination of other components of mobile computing device 850, such as, for example, a user interface, applications run by mobile computing device 850, and the control of wireless communication by mobile computing device 850.
[0083] The processor 852 can communicate with the user through a control interface 858 and a display interface 856 coupled to a display 854. The display 854 may be, for example, a TFT (thin film transistor liquid crystal display) display or an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 856 may comprise appropriate circuitry for driving the display 854 to present graphical and other information to the user. The control interface 858 can receive commands from the user and convert them for submission to the processor 852. Further, an external interface 862 can provide communication with the processor 852 to enable short-range communication of the mobile computing device 850 with other devices. The external interface 862 can provide, for example, wired communication in some implementations or wireless communication in other implementations, and multiple interfaces may be used.
[0084] Memory 864 stores information within mobile computing device 850. Memory 864 can be implemented as one or more of one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. An expansion memory 874 may be provided and connected to mobile computing device 850 through an expansion interface 872, and interface 872 may include, for example, a SIMM (Single In-line Memory Module) card interface. Expansion memory 874 can provide additional storage space for mobile computing device 850, or can store applications or other information for mobile computing device 850. In particular, expansion memory 874 can include instructions for practicing or supplementing the processes described above, and can also include secure information. Thus, for example, expansion memory 874 may be provided as a security module for mobile computing device 850 and can be programmed with instructions that permit secure use of mobile computing device 850. Further, secure applications may be provided via the SIMM card, along with additional information, such as placing identification information on the SIMM card so that it cannot be hacked.
[0085] The memory can include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as discussed below. In some implementations, the instructions are stored in an information carrier such that when the instructions are executed by one or more processing devices (e.g., processor 852), one or more of the methods as described above are implemented. The instructions can also be stored by one or more storage devices such as one or more computer or machine-readable media (e.g., memory 864, expansion memory 874, or memory on processor 852). In some implementations, the instructions can be received in a propagated signal, for example, via transceiver 868 or external interface 862.
[0086] Mobile computing device 850 can communicate wirelessly through a communication interface 866 that may include a digital signal processing circuitry if necessary. The communication interface 866 may provide communication under various modes or protocols, such as, among others, GSM voice calls (wide area mobile communication system), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS message communication (Multimedia Messaging Service), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA (registered trademark) (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service). Such communication may occur, for example, through a transceiver 868 that uses radio frequencies. Further, short - range communication may occur using, for example, Bluetooth, WiFi, or other such transceivers. Additionally, a GPS (Global Positioning System) receiver module 870 may provide additional navigation and location - related wireless data to the mobile computing device 850, and this data may be used by applications running on the mobile computing device 850 as needed.
[0087] Mobile computing device 850 can also communicate audibly using an audio codec 860, which can receive speech information from a user and convert it into usable digital information. The audio codec 860 can similarly generate audible sound for the user, such as through a speaker in the handset of the mobile computing device 850. Such sound may include sound from a voice call, may include recorded sound (such as voice messages, music files, etc.), and may also include sound generated by an application running on the mobile computing device 850.
[0088] Mobile computing device 850 may be implemented in several different forms, as shown in the figures. For example, it may be implemented as a cellular phone 880. It may also be implemented as part of a smartphone 882, a personal digital assistant, or other similar mobile device.
[0089] The various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include an implementation in one or more computer programs executable and / or translatable on a programmable system including at least one programmable processor, the programmable processor being coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device, which may be of special or general purpose.
[0090] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. The terms machine-readable medium and computer-readable medium as used herein refer to any computer program product, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0091] To enable interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball) for the user to provide input to the computer. Other types of devices can also be used to enable interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form, including acoustic, speech, or tactile input.
[0092] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or includes middleware components (e.g., an application server), or includes front-end components (e.g., a client computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0093] The computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises from computer programs that run on respective computers and have a client-server relationship with each other.
[0094] Several implementation forms have been described in detail above, but other modifications are possible. For example, the logical flow described in this application does not require a specific order or sequence shown to achieve the desired result. Further, other actions may be provided, or actions may be eliminated from the described flow, and other components may be added to or removed from the described system. Therefore, other implementation forms are within the scope of the following claims. Also, the features described in one aspect or implementation form may be applied in any other aspect or implementation form.
Description of Reference Numerals
[0095] 100 System 102 Computing Device 104 Computing Device 106 Computing Device 120 Audio Watermarker 124 Playback Log 138 Audio Watermarker 142 Playback Log 144 Trainer 146 Identification Log 148 Audio Watermark Identification Model 150 Audio Subsystem 152 Audio Watermark Identifier 154 Hot Water 158 Audio Watermark Identification Model 162 Speech Recognizer 164 Identification Log 800 Computing Device 802 Processor 804 Memory 806 Storage Device 808 High-Speed Interface 810 High-Speed Expansion Port 812 Low-Speed Interface 814 Low-Speed Expansion Port 816 Display 820 Standard Server 822 Laptop Computer 824 Rack Server System 850 Mobile Computing Device 852 Processor 854 Display 856 Display Interface 858 Control Interface 860 Audio Codec 862 External Interface 864 Memory 866 Communication Interface 868 Transceiver 870 GPS (Global Positioning System) Receiver Module 872 Expansion Interface 874 Expansion Memory 880 Cellular Phone 882 Smartphone
Claims
1. A step of obtaining a plurality of watermark-free voice samples by data processing hardware, wherein each of the plurality of watermark-free voice samples does not include an audio watermark sample; A step of generating one or more corresponding watermarked voice samples from each watermark-free voice sample of the plurality of watermark-free voice samples by the data processing hardware, wherein each of the one or more corresponding watermarked voice samples includes at least one audio watermark; A step of training a model by the data processing hardware to determine whether a given audio data sample includes an audio watermark using the plurality of watermark-free voice samples and the corresponding watermarked voice samples; A step of transmitting the trained model to a user computing device by the data processing hardware after training the model, wherein the user computing device receives audio data corresponding to the playback of the utterance, uses the trained model to obtain data indicating whether the audio data includes the audio watermark, and determines whether to continue or stop processing the received audio data based on the data indicating whether the audio data includes the audio watermark configured as; A method including.
2. The method according to claim 1, wherein each of the plurality of watermark-free voice samples includes the utterance of a hot word.
3. The method according to claim 1, wherein some of the plurality of watermark-free voice samples include the utterance of a hot word.
4. The method according to claim 1, wherein at least one of the plurality of watermark-free voice samples is recorded by a device different from the device used to record at least one of the other plurality of watermark-free voice samples.
5. The method according to claim 1, wherein at least one of the plurality of watermark-free voice samples includes spoken terms different from at least one of the other plurality of watermark-free voice samples.
6. The method according to claim 1, wherein at least some of the plurality of watermark-free voice samples include background noise.
7. The step of generating the one or more corresponding watermarked voice samples each including at least one audio watermark includes generating at least one corresponding watermarked voice sample including at least one audio watermark that overlaps only with the hot word from at least one of the plurality of watermark-free voice samples including the utterance of the hot word. The method according to claim 1.
8. The step of generating the one or more corresponding watermarked voice samples each including at least one audio watermark includes generating at least one corresponding watermarked voice sample including an audio watermark sequence that overlaps with the hot word and precedes the hot word from at least one of the plurality of watermark-free voice samples including the utterance of the hot word. The method according to claim 1.
9. The step of generating the one or more corresponding watermarked voice samples includes, from at least one of the plurality of watermark-free voice samples, generating a first corresponding watermarked voice sample including a first sequence of equally spaced audio watermarks, and generating a second corresponding watermarked voice sample including a second sequence of equally spaced audio watermarks and wherein a period between the audio watermarks in the second sequence of equally spaced audio watermarks is different from a period between the audio watermarks in the first sequence of equally spaced audio watermarks. The method according to claim 1.
10. The method according to claim 1, wherein at least one of the audio watermarks in one of the watermarked voice samples is different from at least one of the audio watermarks in another one of the watermarked voice samples.
11. Before training the model, labeling, by the data processing hardware, each of the plurality of voice samples without watermark used to train the model as not including an audio watermark; labeling, by the data processing hardware, each of the corresponding watermarked voice samples used to train the model as including an audio watermark The method according to claim 1, further comprising.
12. The step of training the model includes using machine learning to train the model with the plurality of voice samples without watermark and the corresponding watermarked voice samples so as to determine whether the given audio data sample includes the audio watermark. The method according to claim 1.
13. The user computing device uses the trained model to obtain the data indicating whether the audio data includes the audio watermark in response to a determination that the audio data includes the pronunciation of a specific predefined hotword. The method according to claim 1, configured as such.
14. A system, data processing hardware; memory hardware communicating with the data processing hardware The memory hardware records instructions that cause the data processing hardware to execute a plurality of operations when executed on the data processing hardware, and the plurality of operations include: an operation of obtaining a plurality of voice samples without watermark, each of the plurality of voice samples without watermark not including an audio watermark sample; An operation of generating one or more corresponding watermarked voice samples from each of the plurality of watermark-free voice samples, wherein each of the one or more corresponding watermarked voice samples includes at least one audio watermark, and An operation of training a model to determine whether a given audio data sample includes an audio watermark using the plurality of watermark-free voice samples and the corresponding watermarked voice samples, and After training the model, an operation of transmitting the trained model to a user computing device, wherein the user computing device Receives audio data corresponding to the playback of speech, Uses the trained model to obtain data indicating whether the audio data includes the audio watermark, Based on the data indicating whether the audio data includes the audio watermark, determines whether to continue or stop processing the received audio data An operation configured as A system including
15. The system according to claim 14, wherein each of the plurality of watermark-free voice samples includes the utterance of a hotword.
16. The system according to claim 14, wherein some of the plurality of watermark-free voice samples include the utterance of a hotword.
17. The system according to claim 14, wherein at least one of the plurality of watermark-free voice samples is recorded by a device different from the device used to record at least one of the other watermark-free voice samples of the plurality of watermark-free voice samples.
18. The system according to claim 14, wherein at least one of the plurality of watermark-free voice samples includes spoken terms different from at least one of the other watermark-free voice samples of the plurality of watermark-free voice samples.
19. The system according to claim 14, wherein at least some of the plurality of watermark-free voice samples include background noise.
20. The operation of generating the one or more corresponding watermarked audio samples each including at least one audio watermark includes an operation of generating at least one corresponding watermarked audio sample including at least one audio watermark that overlaps only with the hotword from at least one of the plurality of non-watermarked audio samples including the utterance of the hotword, the system according to claim 14.
21. The step of generating the one or more corresponding watermarked audio samples each including at least one audio watermark includes a step of generating at least one corresponding watermarked audio sample including a sequence of audio watermarks that overlaps with the hotword and precedes the hotword from at least one of the plurality of non-watermarked audio samples including the utterance of the hotword, the system according to claim 14.
22. The operation of generating the one or more corresponding watermarked audio samples includes, from at least one of the plurality of non-watermarked audio samples, an operation of generating a first corresponding watermarked audio sample including a first sequence of equally spaced audio watermarks, and an operation of generating a second corresponding watermarked audio sample including a second sequence of equally spaced audio watermarks and wherein a period between the audio watermarks in the second sequence of equally spaced audio watermarks is different from a period between the audio watermarks in the first sequence of equally spaced audio watermarks, the system according to claim 14.
23. At least one audio watermark in one of the watermarked audio samples is different from at least one audio watermark in another one of the watermarked audio samples, the system according to claim 14.
24. The plurality of operations are performed before training the model, The operation of labeling each of the plurality of watermark - free voice samples used for training the model by the data - processing hardware as not including an audio watermark, and The operation of labeling each of the corresponding watermark - attached voice samples used for training the model by the data - processing hardware as including an audio watermark The system according to claim 14, further comprising.
25. The operation of training the model includes the operation of using machine learning to train the model with the plurality of watermark - free voice samples and the corresponding watermark - attached voice samples so as to determine whether the given audio data sample includes the audio watermark. The system according to claim 14.
26. The user computing device Is configured to use the trained model to obtain the data indicating whether the audio data includes the audio watermark in response to a determination that the audio data includes the utterance of a specific predefined hot - word. The system according to claim 14.
Citation Information
Patent Citations
Method for embedding watermark information, its device, watermark information embedding program, and computer readable recording medium having the program recorded thereon
JP2003263182A
Multi-User Personalization at a Voice Interface Device
US20180096690A1
Recorded media hotword trigger suppression
US20180130469A1
Speech synthesizer, electronic watermark information detection device, speech synthesis method, electronic watermark information detection method, speech synthesis program, and electronic watermark information detection program
WO2014112110A1