Recorded media hotword trigger suppression

By embedding audio watermarks in audio and analyzing them on a computing device, the problem of false triggering in speech recognition systems is solved, ensuring that hot words are responded to only under specific conditions. This achieves energy saving and accurate response.

CN116597836BActive Publication Date: 2026-04-07GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-03-13
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing speech recognition systems are prone to mis-triggering hot words when receiving non-directional speech, leading to unnecessary consumption of computing resources and battery power. This is especially true in multi-device environments, where it is difficult to distinguish between directional speech input and background noise.

Method used

By embedding audio watermarks in audio, computing devices analyze the audio watermarks to determine whether to respond to hot words, and perform speech recognition only when specific conditions are met, including device location, user settings, and matching of audio watermarks, reducing unnecessary speech recognition processes.

Benefits of technology

It effectively reduces the consumption of battery power and processing power of computing devices, ensures that hot words are responded to only under specific conditions, improves the system's response accuracy and user experience, and is especially convenient for users with hearing impairments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597836B_ABST
    Figure CN116597836B_ABST
Patent Text Reader

Abstract

Methods, systems, and user devices for suppressing hotword triggering when a hotword in a recorded media is detected are disclosed, including computer programs encoded on computer storage media. In an aspect, a method includes receiving, at data processing hardware, audio data corresponding to playback of a media content item, the audio data including an audio watermark and an utterance of a command preceded by a hotword; determining, by the data processing hardware, that the received audio data includes the hotword; processing, by the data processing hardware, the audio data to: identify the audio watermark included in the audio data; and determine a corresponding bitstream of the audio watermark; and based on the determined corresponding bitstream of the audio watermark, determining, by the data processing hardware, to bypass performing the command preceded by the hotword without accessing an audio watermark database to identify a matching audio watermark.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with the application number 201880008785.3 and the title "Recorded Media Hotword Trigger Suppression" filed on 13 March 2018. TECHNICAL FIELD

[0002] This specification generally relates to automatic speech recognition. BACKGROUND

[0003] The reality of voice-enabled homes or other environments - that is, homes or other environments in which a user need only speak a query or command out loud and a computer-based system will process and respond to the query and / or cause the command to be executed - is upon us. A voice-enabled environment (e.g., a home, a workplace, a school, etc.) can be implemented using a network of connected microphone devices distributed throughout various rooms or areas of the environment. With such a network of microphones, a user has the ability to orally query the system from virtually anywhere in the environment without needing to have a computer or other device in front of him / her or even in close proximity. For example, while cooking in the kitchen, a user can ask the system "What is three cups in milliliters?" and receive a response from the system in response, e.g., in the form of a synthetic voice output. Alternatively, a user can ask the system such as "What time does the gas station closest to me close?" or, while preparing to leave the house, "Should I wear a coat today?"

[0004] In addition, a user can ask queries of and / or issue commands to the system that involve the user's personal information. For example, a user can ask the system "When do I meet John?" or command the system "Remind me to call John when I get home." SUMMARY

[0005] For voice-enabled systems, user interaction is designed primarily (if not specifically) via voice input. Therefore, a system that may pick up all utterances from the surrounding environment, including those not directed at the system, must somehow discern when any given utterance is directed at the system and not, for example, at an individual present in the environment. One way to achieve this is by using hotwords, which are reserved as pre-defined words through an agreement among users in the environment and spoken to attract the system's attention. In the example environment, the hotword used to attract the system's attention is the phrase "OK computer." Therefore, each time the phrase "OK computer" is spoken, it is picked up by the microphone, transmitted to the system, which performs speech recognition technology or uses audio features and neural networks to determine whether the hotword has been spoken, and if so, waits for a subsequent command or query. Therefore, the discourse directed to the system takes the general form of [HOTWORD][QUERY], where the "hot word" in this example is "OK computer" and the "query" can be any question, command, statement, or other request that can be performed by the system alone or via a network and server to recognize, parse, and act upon.

[0006] In addition to detecting hot words, computing devices can be configured to detect audio watermarks embedded in the audio of hot words. Audio watermarks can be high-frequency watermarks inaudible to humans, or watermarks that sound like background noise or static. Computing devices can be configured to respond differently to hot words based on the audio watermark. For example, a portion of media content might be created where an actor says, “OK computer, give me directions to the train station.” To prevent any computing device within range of the television playing the media content from providing directions to the train station, the creator of the media content could include an audio watermark overlapping with the hot word. The computing device detecting the audio of the media content can extract the audio watermark and compare it to an audio watermark database. The audio watermark database can include rules on when the computing device should perform speech recognition on audio following the hot word and that particular audio watermark, as well as additional rules for other audio watermarks. It should be understood that at least some of the methods and systems described herein enable computing devices to selectively determine when to respond to spoken utterances output through speakers near the computing device.

[0007] Rules for audio watermarking of media content can include conditions that a computing device should meet before further processing the audio from the media content. Example rules could be: the computing device should respond if it is not currently providing directions, is moving at a speed not exceeding 10 miles per hour, is a smart speaker, and is located at the user's residence. Another example of a rule could be: the computing device should only respond if it is located in a place associated with the owner or creator of the media content and the watermark indicates that the utterance has high priority. If the computing device meets these rules, it can perform speech recognition on the portion following the hot word. If the transcription is "Give me directions to the train station," the computing device can provide directions to the train station, either explicitly or audibly. If the computing device does not meet the rules, it does not perform speech recognition on the portion of the audio following the hot word and does not respond to the audio with any further action.

[0008] In some implementations, audio watermarking can encode data, which eliminates the need for the computing device to compare the audio watermark with an audio watermark database. The encoded data may include rules, identifiers, actions, or any other similar data indicating when the computing device should perform speech recognition. In some implementations, the computing device may use the encoded data in conjunction with an audio watermark database to determine whether to perform speech recognition on audio following a hot word.

[0009] According to the innovative aspects of the subject matter described in this application, a method for suppressing hot word triggering when hot words are detected in recorded media includes the following actions: receiving audio corresponding to the playback of a media content item by a computing device; determining by the computing device that the audio includes utterances of predefined hot words and determining that the audio includes an audio watermark; analyzing the audio watermark by the computing device; and, based on the analysis of the audio watermark, determining by the computing device whether to perform speech recognition on the audio portion following the predefined hot words.

[0010] These and other implementations may optionally include one or more of the following features. The action of analyzing the audio watermark includes comparing the audio watermark with one or more audio watermarks. The action of determining whether to perform speech recognition on the audio portion following a predefined hot word is also based on comparing the audio watermark with one or more audio watermarks. The audio watermark is the inaudible portion of the audio corresponding to the playback of the media content item. These actions also include: identifying the audio source corresponding to the playback of the media content item based on the analysis of the audio watermark. The action of determining whether to perform speech recognition on the audio portion following a predefined hot word is also based on the audio source corresponding to the playback of the media content item. These actions also include: identifying the audio source corresponding to the playback of the media content item based on the analysis of the audio watermark; and updating the log file to indicate the audio source corresponding to the playback of the media content item.

[0011] Audio watermarking is included in the audio portion of the utterance including predefined hot words. These actions also include determining the type of the additional computing device. The action of determining whether to perform speech recognition on the audio portion following the predefined hot words is also based on the type of the additional computing device. The action of the computing device determining whether to perform speech recognition on the audio portion following the predefined hot words includes determining whether to perform speech recognition on the audio portion following the predefined hot words. These actions also include generating an audio transcription following the predefined hot words by an automatic speech recognizer; and performing actions corresponding to the audio transcription following the predefined hot words. The action of the computing device determining whether to perform speech recognition on the audio portion following the predefined hot words includes determining not to perform speech recognition on the audio portion following the predefined hot words. These actions also include suppressing actions corresponding to the playback of audio items related to media content.

[0012] These actions also include determining the location of the additional computing device. The action of determining whether to perform speech recognition on the audio portion following the predefined hot words is also based on the location of the additional computing device. These actions also include determining the user settings of the additional computing device. The action of determining whether to perform speech recognition on the audio portion following the predefined hot words is also based on the user settings of the additional computing device. The actions of the computing device determining that the audio includes utterances with predefined hot words and determining that the audio includes an audio watermark include determining that the audio includes utterances with predefined hot words; and based on determining that the audio includes utterances with predefined hot words, determining that the audio includes an audio watermark. The actions of the computing device determining that the audio includes utterances with predefined hot words and determining that the audio includes an audio watermark include determining that the audio includes utterances with predefined hot words; and after determining that the audio includes utterances with predefined hot words, determining that the audio includes an audio watermark.

[0013] The actions of analyzing the audio watermark include extracting the data encoded in the audio watermark. Determining whether to perform speech recognition on the audio portion following the predefined hot words is also based on the data encoded in the audio watermark. These actions also include: identifying the type of media content corresponding to the playback of the audio content item based on the analysis of the audio watermark; and updating the log file to indicate the type of media content corresponding to the playback of the media content item. These actions also include: identifying the type of media content corresponding to the playback of the audio content item based on the analysis of the audio watermark. Determining whether to perform speech recognition on the audio portion following the predefined hot words is also based on the type of media content corresponding to the playback of the media content item. These actions also include: determining, by a computing device, whether to perform natural language processing on the audio portion following the predefined hot words based on the analysis of the audio watermark.

[0014] Other embodiments of this aspect include corresponding systems, apparatuses, and computer programs recorded on computer storage devices, each configured to perform the operations of these methods.

[0015] Specific embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. The computing device can respond to only hot words containing a specific audio watermark, thereby conserving battery power and processing power. When hot words with an audio watermark are received, fewer computing devices can perform search queries to maintain network bandwidth. Additionally, the audio watermark can be used to enable a user's computing device to convey information (e.g., a response to a verbal inquiry or some kind of warning) to the user even if the user may not be able to hear it (e.g., if it is information output via a speaker located near the user). Such users can include those with hearing impairments or those listening to other audio via personal speakers (e.g., headphones) connected to their audio devices. For example, a specific audio watermark can be interpreted by the computing device as indicating high priority, in which case the computing device can respond to queries received via the main audio.

[0016] A method according to an embodiment of the present disclosure includes: receiving at data processing hardware audio data corresponding to playback of a media content item, the audio data including an audio watermark and utterances of a command preceded by a hot word; determining by the data processing hardware that the received audio data includes a hot word; processing the audio data by the data processing hardware to: identify the audio watermark included in the audio data; and determine a corresponding bitstream of the audio watermark; and based on the determined corresponding bitstream of the audio watermark, determining by the data processing hardware, without accessing an audio watermark database to identify a matching audio watermark, to bypass execution of the command preceded by the hot word.

[0017] A system according to an embodiment of the present disclosure includes: data processing hardware; and memory hardware communicating with the data processing hardware and storing instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform operations, including: receiving audio data corresponding to playback of a media content item, the audio data including an audio watermark and utterances of a command preceded by a hot word; determining that the received audio data includes a hot word; processing the audio data to: identify the audio watermark included in the audio data; and determine a corresponding bitstream of the audio watermark; and based on the determined corresponding bitstream of the audio watermark, determining to bypass the execution of the command preceded by the hot word without accessing an audio watermark database to identify a matching audio watermark.

[0018] A computer-implemented method according to an embodiment of the present disclosure, when executed on data processing hardware of a user device, causes the data processing hardware to perform operations, including: when the user device is in a sleep mode, receiving audio data captured by the microphone of the user device, the audio data corresponding to playback of media content items output from an audio source near the user device; when the user device remains in sleep mode, processing the audio data to determine a corresponding bitstream of an audio watermark encoded in the audio data; and based on the determined corresponding bitstream of the audio watermark encoded in the audio data, determining to bypass performing speech recognition on the audio data without accessing an audio watermark database to identify a matching audio watermark.

[0019] A user equipment according to an embodiment of the present disclosure includes: data processing hardware; and memory hardware communicating with the data processing hardware and storing instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform operations, including: when the user equipment is in a sleep mode, receiving audio data captured by the user equipment's microphone, the audio data corresponding to playback of media content items output from an audio source near the user equipment; when the user equipment remains in sleep mode, processing the audio data to determine a corresponding bitstream of an audio watermark encoded in the audio data; and based on the determined corresponding bitstream of the audio watermark encoded in the audio data, determining to bypass performing speech recognition on the audio data without accessing an audio watermark database to identify a matching audio watermark.

[0020] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, drawings, and claims. Attached Figure Description

[0021] Figure 1 An example system is shown for suppressing hot word triggering when hot words are detected in recorded media.

[0022] Figure 2 This is a flowchart of an example process for suppressing hot word triggering when hot words are detected in recorded media.

[0023] Figure 3 These are examples of computing devices and mobile computing devices.

[0024] The same reference numerals and names in different figures denote the same elements. Detailed Implementation

[0025] Figure 1An example system 100 for suppressing hot word triggering when hot words are detected in recorded media is shown. Briefly, as described in more detail below, computing devices 102 and 104 receive audio 108 output from an audio source 106 (e.g., a television set). Audio 108 includes utterances of predefined hot words and an audio watermark. Both computing devices 102 and 104 process audio 108 and determine that audio 108 includes predefined hot words. Computing devices 102 and 104 identify the audio watermark. Based on the audio watermark and the context or characteristics of computing devices 102 and 104, each of computing devices 102 and 104 can perform speech recognition on the audio.

[0026] exist Figure 1 In the example shown, audio source 106 is playing media content related to Nugget World. During the media content, an actor in the media content says the utterance 108, “OK computer, what’s in the chicken nuggets?”. The utterance 108 includes the hot word 110 “Ok computer” and the query 112 “What’s in the chicken nuggets?”. Audio source 106 outputs the utterance 108 through a speaker. Any nearby computing device with a microphone can detect the utterance 108.

[0027] The audio of speech 108 includes an audible portion 114 and an audio watermark 116. Creators of media content can add the audio watermark 116 to ensure that a specific computing device responds correctly to speech 108. In some embodiments, the audio watermark 116 may include audio frequencies above or below the range of human hearing. For example, the audio watermark 116 may include frequencies greater than 20 kHz or less than 20 Hz. In some embodiments, the audio watermark 116 may include audio within the range of human hearing but undetectable to humans because its sound resembles noise. For example, the audio watermark 116 may include frequency patterns between 8 and 10 kHz. The intensity of different frequency bands may be imperceptible to humans but can be detected by computing devices. As shown in the frequency domain representation 118 of audio 108, the audio watermark 116 includes frequencies in a higher frequency range than the audible portion 114.

[0028] Computing devices 102 and 104 can be any type of device capable of receiving audio via a microphone. For example, computing devices 102 and 104 can be desktop computers, laptop computers, tablet computers, wearable computers, cellular phones, smartphones, music players, e-book readers, navigation systems, smart speakers and home assistants, wireless (e.g., Bluetooth) headphones, hearing aids, smartwatches, smart glasses, activity trackers, or any other suitable computing device. Figure 1As shown, computing device 102 is a smartphone, and computing device 104 is a desktop computer. Audio source 106 can be any audio source, such as, for example, a television, radio, music player, desktop computer, laptop computer, tablet computer, wearable computer, cellular phone, or smartphone. Figure 1 As shown, audio source 106 is a television set.

[0029] Computing devices 102 and 104 each receive audio via a microphone. Regarding computing device 102, the microphone may be part of an audio subsystem 120. Audio subsystem 120 may include buffers, filters, and analog-to-digital converters, each designed to initially process the audio received via the microphone. The buffer may store the current audio received via the microphone and processed by audio subsystem 120. For example, the buffer stores the first five seconds of audio data. Similarly, the microphone of computing device 104 may be part of an audio subsystem 122. Audio subsystem 122 may include buffers, filters, and analog-to-digital converters, each designed to initially process the audio received via the microphone. The buffer may store the current audio received via the microphone and processed by audio subsystem 122. For example, the buffer stores the first three seconds of audio data.

[0030] Computing devices 102 and 104 each include hotword generators 124 and 126, respectively. Hotword generators 124 and 126 are each configured to identify hotwords in audio received via a microphone and / or stored in a buffer. In some embodiments, hotword generators 124 and 126 can be active at any time while computing devices 102 and 104 are powered on. Hotword generator 124 continuously analyzes the audio data stored in the buffer. Hotword generator 124 calculates a hotword confidence score reflecting the probability that the current audio data in the buffer contains a hotword. To calculate the hotword confidence score, hotword generator 124 can extract audio features from the audio data, such as filterbank energy or mel-frequency cepstral coefficients. Hotword generator 124 can process these audio features using a classification window, such as by using a support vector machine or a neural network. In some embodiments, hotword generator 124 does not perform speech recognition to determine the hotword confidence score. If the hot word confidence score meets the hot word confidence score threshold, then hot word generator 124 determines that the audio includes the hot word. For example, if the hot word confidence score is 0.8 and the hot word confidence score threshold is 0.7, then hot word generator 124 determines that the audio corresponding to utterance 108 includes hot word 110. Hot word generator 126 is functionally similar to hot word generator 124.

[0031] Computing devices 102 and 104 each include audio watermark identifiers 128 and 130, respectively. Audio watermark identifiers 128 and 130 are each configured to process audio received via a microphone and / or stored in a buffer, and to identify audio watermarks included in the audio. Audio watermark identifiers 128 and 130 may each be configured to detect spread-spectrum and psychoacoustic shaping watermarks. These types of watermarks may reside in frequency bands overlapping with the corresponding audio frequency band. Humans may perceive these types of watermarks as noise. Audio watermark identifiers 128 and 130 may also each be configured to detect high-frequency watermarks. These types of watermarks may reside in frequency bands above the corresponding audio frequency band. The frequency band of high-frequency watermarks may be above the threshold of human hearing. Audio watermark identifiers 128 and 130 may also each be configured to detect low-frequency watermarks. These types of watermarks may reside in frequency bands below the corresponding audio frequency band. The frequency band of low-frequency watermarks may be below the threshold of human hearing. In some implementations, audio watermark identifiers 128 and 130 process audio in response to hot words detected by corresponding hot word identifiers 124 and 126.

[0032] Audio watermark identifiers 128 and 130 can each be configured to separate the audio watermark and the main audio. The main audio can be the audio portion to which the audio watermark is added. For example, the main audio may include an audible portion 114 that includes audio corresponding to “OK computer, what’s in the chicken nuggets?” but without the watermark 116. Audio watermark identifier 128 separates the audio 118 received through the microphone of computing device 102 into the main audio 132 and the audio watermark 134. Similarly, audio watermark identifier 130 separates the audio 118 received through the microphone of computing device 104 into the main audio 136 and the audio watermark 138. In some embodiments, the audio watermark and the main audio may overlap in the temporal domain.

[0033] In some implementations, audio watermark identifiers 128 and 130 can process audio watermarks 134 and 138 respectively to identify the corresponding bitstream of the audio watermark. For example, audio watermark identifier 128 can process audio watermark 134 and determine that audio watermark 134 corresponds to bitstream 0101101110101. Audio watermark identifier 130 can perform similar processing on audio watermark 138.

[0034] Audio watermark comparators 140 and 144 each compare the corresponding audio watermarks 134 and 138 with audio watermarks 142 and 146, respectively. For example, audio watermark comparator 140 can compare the frequency pattern or bitstream of watermark 134 with audio watermark 142. Audio watermark comparator 140 can determine that audio watermark 134 matches the audio watermark of Chicken Nuggets World. Audio watermark comparator 144 can make a similar determination.

[0035] Audio watermarks 142 and 146 may contain various entities embedded in the audio of media content containing hot words or other distributed or broadcast audio. Chicken Nuggets World may include a watermark in audio 108 to ensure that only specific devices respond to the hot words, perform speech recognition on the audio, and execute query 112. Chicken Nuggets World may provide audio watermark 116 to be included in audio watermarks 142 and 146, as well as instructions on when a device should respond to a hot word with audio watermark 116. For example, Chicken Nuggets World may include instructions in audio watermarks 142 and 146 for any device located in a Chicken Nuggets World restaurant with a Chicken Nuggets World user identifier to respond to a hot word with audio watermark 116. In some embodiments, audio watermarks 142 and 146 are stored on computing devices 102 and 104 and are updated periodically, for example, once daily. In some embodiments, audio watermarks 142 or 146, audio watermark identifiers 128 and 130, and / or audio watermark comparators 140 and 144 may reside on a remote server. In this scenario, computing devices 102 or 104 can communicate with a remote server via a network.

[0036] The computing device 102 extracts the audio watermark 134 and matches it with the Chicken Nuggets World watermark. Based on the instructions for the Chicken Nuggets World watermark in the audio watermark 142, the computing device 102 can perform speech recognition on the main audio 132 and execute any queries or commands included in the corresponding transcription. The instructions may include a set of rules that the computing device 102 must follow to determine whether to perform speech recognition.

[0037] Computing device 102 includes a location detector 156. Location detector 156 can generate geographic location data reflecting the location of the computing device. Location detector 156 can use any geolocation technology, such as Global Positioning System (GPS), triangulation, and / or any other similar positioning technology. In some embodiments, location detector 156 can access maps or location data indicating the locations of various points of interest. Location detector 156 can further identify the points of interest where the computing device is located. For example, location detector 156 can determine that computing device 102 is located in Chicken Nuggets World.

[0038] Computing device 102 includes a device identifier 158. Device identifier 158 includes a device identity 160 that identifies the device type of computing device 102. Device identity 160 can be a desktop computer, laptop computer, tablet computer, wearable computer, cellular phone, smartphone, music player, e-book reader, navigation system, smart speaker, home assistant, or any other suitable computing device. For example, device identity 160 of computing device 102 is a telephone.

[0039] Computing device 102 includes a user identifier 162. User identifier 162 includes a user identity 164 that identifies a user of computing device 102. User identity 164 can be an email address, phone number, or any other similar unique user identifier. For example, user identity 164 of computing device 102 is user@example.com. User identifier 162 can be entered by user 154.

[0040] Computing device 102 includes user settings 152. User settings 152 may be provided by user 154 and may include additional rules regarding how computing device 102 should respond to hot words. For example, user settings 152 may include a rule that computing device 102 does not respond to any hot words that include an audio watermark unless computing device 102 receives the hot word containing the audio watermark while at home. As another example, user settings 152 may include a rule that computing device 102 does not respond to hot words that include an audio watermark corresponding to a specific entity (e.g., the owner or creator of media content) (such as Chicken World). In some implementations, user 154 may consent to allowing computing device 102 to respond to hot words with watermarks of specific entities.

[0041] exist Figure 1 In the example shown, speech recognizer 166 remains inactive, as indicated by speech recognizer state 168. Computing device 102 sets speech recognizer state 168 to inactive based on instructions stored in the audio watermark corresponding to audio watermark 134, applied to device location, user settings 152, device identity 160, and user identity 164. For example, the instruction corresponding to audio watermark 134 could be to set speech recognizer state 168 to active if user identity 164 is a Chicken Nuggets World identifier and the device is located at a Chicken Nuggets World restaurant. For computing device 102, user identity 164 is not a Chicken Nuggets World identifier. Therefore, speech recognizer state 168 is inactive.

[0042] In some implementations, the user interface generator 148 of the computing device 102 may provide data for a graphical interface to the display of the computing device. The graphical interface may indicate the processes or actions of the computing device 102 before, during, or after the execution of a process or action by the computing device. For example, the user interface generator 148 may display an interface indicating that the computing device 102 is processing received audio, that the computing device 102 is identifying an audio watermark 134, a speech recognizer state 168, and / or any attributes or rules of the identified audio watermark 134.

[0043] In some implementations, the user interface generator 148 can generate an interface indicating that the speech recognizer state 168 is inactive. This interface may also include user-selectable options to override the speech recognizer state 168. For example, user 154 can choose to set the speech recognizer state 168 to active. After hearing query 112 "What's in the chicken nuggets?", user 154 might become curious and request computing device 102 to process query 112 and provide output.

[0044] In some implementations, computing device 102 may include an audio watermark log 170. The audio watermark log 170 may include data indicating the number of times each audio watermark has been received by computing device 102. For example, whenever computing device 102 receives and identifies audio watermark 134, computing device 102 may store data in the audio watermark log 170 indicating the reception of audio watermark 134. The data may include timestamps, device location, any relevant user settings, user identifiers, and any other similar information. In some implementations, computing device 102 may provide the data from the audio watermark log 170 to an aggregated audio watermark log on a server, which combines audio watermark logs from different computing devices that received the audio watermarks. The aggregated audio watermark log may include user identity, device identifiers, and data stored in the audio watermark log 170 for the receiving computing device. In some implementations, the aggregated audio watermark log and the data in the audio watermark log 170 may be synchronized. In this case, the audio watermark log 170 may include additional log data from different devices, as well as data identifying different devices, different users, location information, timestamp data, and other relevant information.

[0045] In some implementations, instructions for a specific audio watermark may include instructions related to data stored in the audio watermark log 170. The instructions may relate to a specific number of times a hot word marked with a specific audio watermark should activate the speech recognizer. For example, the instructions may indicate that audio watermark 116 should activate speech recognizer 166 only once within a 24-hour period.

[0046] In some implementations, the creator of media content on audio device 106 can access the aggregated audio watermark log to identify details associated with each activation of the speech recognizer for hot word 110 and its corresponding audio watermark 116. In some implementations, a user can instruct the computing device not to upload the audio watermark log to the aggregated audio watermark log through user settings on the device.

[0047] Computing device 104 processes audio watermark 138 in a similar manner to computing device 102, which processes audio watermark 134. Specifically, computing device 104 extracts audio watermark 138 and matches it with the Chicken Nuggets World watermark. Based on the instructions for the Chicken Nuggets World watermark in audio watermark 146, computing device 102 can perform speech recognition on the main audio 136 and execute any queries or commands included in the corresponding transcription. The instructions may include a set of rules that computing device 104 must follow to determine whether to perform speech recognition.

[0048] The computing device 104 includes a location detector 176. The location detector 176 can generate geographic location data reflecting the location of the computing device. The location detector 176 can use any geolocation technology, such as GPS, triangulation, and / or any other similar positioning technology. In some embodiments, the location detector 176 can access maps or location data indicating the locations of various points of interest. The location detector 176 can further identify the points of interest where the computing device 104 is located. For example, the location detector 176 can determine that the computing device 104 is located in Chicken Nuggets World.

[0049] Computing device 104 includes a device identifier 178. Device identifier 178 includes a device identity 180 that identifies the device type of computing device 104. Device identity 180 can be a desktop computer, laptop computer, tablet computer, wearable computer, cellular phone, smartphone, music player, e-book reader, navigation system, smart speaker, home assistant, or any other suitable computing device. For example, device identity 180 of computing device 104 is a desktop computer.

[0050] Computing device 104 includes a user identifier 182. User identifier 182 includes a user identity 184 that identifies the user of computing device 104. User identity 184 can be an email address, phone number, or any other similar unique user identifier. For example, user identity 184 for computing device 108 is store@nuggetworld.com. User identifier 182 can be entered by the user.

[0051] Computing device 104 includes user settings 186. User settings 186 may be provided by a user and may include additional rules regarding how computing device 104 should respond to hot words. For example, user settings 186 may include a rule that computing device 104 does not respond to any hot words that include an audio watermark unless computing device 104 is located at the Chicken Nuggets World restaurant. As another example, user settings 186 may include a rule that computing device 104 does not respond to any hot words other than those marked with an audio watermark from Chicken Nuggets World. As yet another example, user settings 186 may instruct computing device 104 not to respond to any hot words with any type of audio watermark outside of Chicken Nuggets World's opening hours.

[0052] exist Figure 1 In the example shown, speech recognizer 172 is active, as indicated by speech recognizer state 174. Computing device 104 sets speech recognizer state 174 to active based on instructions stored in the audio watermark corresponding to audio watermark 138, applied to device location, user settings 186, device identity 180, and user identity 184. For example, the instruction corresponding to audio watermark 134 could be to set speech recognizer state 174 to active if user identity 184 is a Chicken Nuggets World identifier and the device is located at the Chicken Nuggets World restaurant. For computing device 104, user identity 184 is a Chicken Nuggets World identifier, and the location is at Chicken Nuggets World. Therefore, speech recognizer state 174 is active.

[0053] Speech recognizer 172 performs speech recognition on main audio 136. Speech recognizer 172 generates a transcription of "What's in the chicken nuggets?". If the transcription corresponds to a query, computing device 104 can provide the transcription to a search engine. If the transcription corresponds to a command, computing device can execute the command. Figure 1In the example, computing device 104 provides a transcription of main audio 136 to a search engine. The search engine returns results, and computing device 104 can output the results via a speaker, which can be, for example, the speaker of the computing device, or a personal speaker connected to the computing device, such as a headphone, earphone, or earbud. Outputting results via a personal speaker can be useful, for example, if the information is output as part of the main audio, then it is useful to provide the information to the user when the user will not be able to hear it. For example, in the Chicken Nuggets example, computing device 104 could output audio 190 “Chicken Nuggets contain chicken meat.” In some implementations, user interface generator 150 can display search results on the display of computing device 104. This is particularly useful for providing information to users who may not be able to hear the information, such as those with hearing impairments, if the information is output as part of the main audio or via a speaker associated with the computing device.

[0054] In some implementations, the user interface generator 150 may provide additional interfaces. The graphical interface may indicate the processes or actions of the computing device 104 before, during, or after the execution of those processes or actions. For example, the user interface generator 150 may display an interface indicating that the computing device 104 is processing received audio, that the computing device 104 is identifying an audio watermark 138, a speech recognizer state 174, and / or any attributes or rules of the identified audio watermark 138.

[0055] In some implementations, the user interface generator 150 may generate an interface indicating that the speech recognizer state 174 is active. This interface may also include user-selectable options to override the speech recognizer state 174. For example, the user may select an option to set the speech recognizer state 174 to suppress any transcription-related actions. In some implementations, the user interface generator 150 may generate an interface that updates the user settings 186 based on recently received overrides and the current attributes of the computing device 104. The user interface generator 148 may also provide a similar interface upon receiving an override command.

[0056] In some implementations, computing device 104 may include an audio watermark log 188. Audio watermark log 188 may store data similar to audio watermark log 170 based on audio watermarks received by computing device 104. Audio watermark log 188 may interact with aggregated audio watermark logs in a manner similar to audio watermark log 170.

[0057] In some implementations, computing devices 102 and 104 may perform speech recognition on the main audio 132 and 136, respectively, without relying on rules stored in audio watermarks 142 and 146. Audio watermarks 142 and 146 may include rules related to actions performed on the main audio based in part on transcription.

[0058] Figure 2 An example process 200 is shown for suppressing hot word triggering when hot words are detected in recorded media. Typically, process 200 performs speech recognition based on an audio pair corresponding to the media content, including hot words and a watermark. Process 200 will be described as being performed by a computer comprising one or more computers (e.g., Figure 1 The computer system of the computing device 102 or 104 shown is used to execute the operation.

[0059] The system receives audio (210) corresponding to the playback of a media content item. In some embodiments, the audio may be received via the system's microphone. The audio may correspond to the audio of media content played on a television or radio.

[0060] The system determines that the audio includes utterances of predefined hot words and an audio watermark (220). In some embodiments, the audio watermark is an inaudible portion of the audio. For example, the audio watermark may be located in a frequency band above or below the range of human hearing. In some embodiments, the audio watermark is audible but sounds like noise. In some embodiments, the audio watermark overlaps with the audio of the predefined hot words. In some embodiments, the system determines that the audio includes the predefined hot words. In response to this determination, the system processes the audio to determine whether the audio includes an audio watermark.

[0061] The system compares the audio watermark with one or more audio watermarks (230). In some embodiments, the system may compare the audio watermark with an audio watermark database. The database may be stored on the system or on different computing devices. The system may compare digital or analog representations of the audio watermarks in the time and / or frequency domains. The system may identify matching audio watermarks and process the audio according to rules specified in the database for the identified audio watermarks. In some embodiments, the system may identify the source or owner of the audio watermark. For example, the source or owner might be Entity Chicken Nuggets World. The system may update a log file to indicate that the system has received hot words with the Chicken Nuggets World audio watermark.

[0062] The system determines whether to perform speech recognition (240) on the audio portion following a predefined hot word by comparing the audio watermark with one or more audio watermarks. Based on rules specified for the identified audio watermark in a database, the source of the audio watermark, and the system context, the system determines whether to perform speech recognition on the audio following the predefined hot word. The system context can be based on any combination of system type, system location, and any user settings. For example, the rule could specify that a mobile phone located at a user's residence should perform speech recognition on audio when it receives a hot word with a specific watermark from the management company of the user's apartment. In some implementations, the system determines whether to perform natural language processing on the audio portion following the predefined hot word based on a comparison of the audio watermark with one or more watermarks or based on analysis of the audio watermark. The system can perform natural language processing in addition to or in lieu of speech recognition.

[0063] Once the system determines to perform speech recognition, it generates an audio transcription following the hot words. The system then executes commands included in the transcription, such as adding an appointment to an apartment building meeting or submitting a query to a search engine. The system can output the search results through the system's speakers, on the system's display, or both.

[0064] If the system determines that it will not perform speech recognition, it can remain in sleep mode, standby mode, or low-power mode. The system can be in sleep mode, standby mode, or low-power mode while processing audio, and it can remain in sleep mode, standby mode, or low-power mode if it does not perform speech recognition on the audio. In some embodiments, user 154 may be using computing device 102 while computing device 102 receives audio 118. For example, user 154 may be listening to music or viewing a photo application. In this case, hot word and audio watermarking processing can occur in the background, and the user's activity can be uninterrupted. In some embodiments, the audio may not include an audio watermark. In this case, the system can perform speech recognition on the audio after the hot words and execute any commands or queries included in the audio.

[0065] In some implementations, the system can determine the type of media content for the audio. The system can compare the audio watermark with audio watermarks included in an audio watermark database. The system can identify matching audio watermarks in the audio watermark database, and the matching audio watermark can identify the type of media content for that particular audio watermark. The system can apply rules to the identified type of media content. For example, the audio watermark database may indicate that the audio watermark is included in sales media, targeted media, commercial media, political media, or any other type of media. In this case, the system can follow general rules for the media type. For example, the rule might be to perform speech recognition for commercial media only when the system is located at a residence. The rule can also be a rule specific to the received audio watermark. In some implementations, the system can also record the type of media content in an audio watermark log.

[0066] In some implementations, the system can analyze audio watermarks. The system can analyze audio watermarks instead of comparing audio watermarks with an audio watermark database, or in combination with comparing audio watermarks with an audio watermark database. Audio watermarks can encode actions, identifiers, rules, or any other similar data. The system can decode audio watermarks and process audio based on the decoded audio watermark. Audio watermarks can be encoded as a header and a payload. The system can identify the header, which can be common to all or almost all audio watermarks, or can identify a specific group of audio watermarks. The payload can follow the header and encode actions, identifiers, rules, or other similar data.

[0067] The system can apply rules encoded in the audio watermark. For example, a rule could be that if the system is a smart speaker located at a position corresponding to a user identifier stored in the system, then the system performs speech recognition on the audio portion following the hot word. In this case, the system may not need to access the audio watermark database. In some implementations, the system can add the rules encoded in the audio watermark to the audio watermark database.

[0068] The system can utilize data encoded in audio watermarks in conjunction with an audio watermark database. For example, data encoded in an audio watermark could indicate that the audio is political media content. The system can access rules corresponding to the audio watermark, specifying that when the system is located in a user's residence, it performs speech recognition on audio containing either political or commercial media content watermarks. In this case, the audio watermark could include a header or other portions of the audio watermark database that the system can use to identify the corresponding audio watermark. The payload can encode the type of media content or other data, such as actions, identifiers, or rules.

[0069] Figure 3Examples of computing devices 300 and 350 that can be used to implement the techniques described herein are shown. Computing device 300 is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Mobile computing device 350 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wireless (e.g., Bluetooth) headsets, hearing aids, smartwatches, smart glasses, activity trackers, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended as limitations.

[0070] The computing device 300 includes a processor 302, a memory 304, a storage device 306, a high-speed interface 308 connected to the memory 304 and multiple high-speed expansion ports 310, and a low-speed interface 312 connected to a low-speed expansion port 314 and the storage device 306. Each of the processor 302, memory 304, storage device 306, high-speed interface 308, high-speed expansion port 310, and low-speed interface 312 is interconnected using various buses and can be mounted on a common motherboard or otherwise suitably mounted. The processor 302 can process instructions for execution within the computing device 300, including instructions stored in the memory 304 or on the storage device 306 for displaying graphical information for a GUI on an external input / output device, such as a display 316 coupled to the high-speed interface 308. In other embodiments, multiple processors and / or multiple buses, along with multiple memories and various types of memory, may be used as appropriate. Furthermore, multiple computing devices can be connected, with each device providing a portion of the necessary operation (e.g., as a server group, blade server cluster, or multiprocessor system).

[0071] Memory 304 stores information within computing device 300. In some embodiments, memory 304 is one or more volatile memory cells. In some embodiments, memory 304 is one or more non-volatile memory cells. Memory 304 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0072] Storage device 306 provides large-scale storage for computing device 300. In some embodiments, storage device 306 may be or contain computer-readable media, such as floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory or other similar solid-state storage devices, or device arrays, including devices in a storage area network or other configuration. Instructions may be stored in an information carrier. When executed by one or more processing devices (e.g., processor 302), the instructions perform one or more methods such as those described above. The instructions may also be stored by one or more storage devices such as computer or machine-readable media (e.g., memory 304, storage device 306, or memory on processor 302).

[0073] High-speed interface 308 manages bandwidth-intensive operations of computing device 300, while low-speed controller 312 manages lower bandwidth-intensive operations. This functional allocation is merely exemplary. In some embodiments, high-speed interface 308 is coupled to memory 304, display 316 (e.g., coupled via a graphics processor or accelerator), and high-speed expansion port 310 that can accept various expansion cards (not shown). In this embodiment, low-speed interface 312 is coupled to storage device 306 and low-speed expansion port 314. Low-speed expansion port 314, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices such as a keyboard, indicating device, scanner, microphone, speaker, or, for example, coupled to a networking device such as a switch or router via a network adapter.

[0074] As shown, computing device 300 can be implemented in a variety of different forms. For example, computing device 300 can be implemented as a standard server 320, or multiple times in a group of such servers. Alternatively, computing device 300 can be implemented in a personal computer such as a laptop computer 322. Computing device 300 can also be implemented as part of a rack-mount server system 324. Alternatively, components from computing device 300 can be combined with other components in a mobile device (not shown), such as mobile computing device 350. Each of such devices can contain one or more of computing device 300 and mobile computing device 350, and the entire system can consist of multiple computing devices communicating with each other.

[0075] Mobile computing device 350 includes processor 352, memory 364, input / output devices such as a touch-enabled display 354, communication interface 366, and transceiver 368, as well as other components. Mobile computing device 350 may also provide storage devices, such as microdrives or other devices, to provide additional storage. Each of processor 352, memory 364, display 354, communication interface 366, and transceiver 368 is interconnected using various buses, and several components may be mounted on a common motherboard or otherwise suitably mounted.

[0076] Processor 352 can execute instructions within mobile computing device 350, including instructions stored in memory 364. Processor 352 can be implemented as a chipset including individual and multiple analog and digital processors. For example, processor 352 can provide coordination of other components of mobile computing device 350, such as control of user interface, applications running on mobile computing device 350, and wireless communications performed by mobile computing device 350.

[0077] Processor 352 can communicate with the user via control interface 358 and display interface 356 coupled to display 354. For example, display 354 can be a TFT (Thin-Film-Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display, or other suitable display technologies. Display interface 356 may include appropriate circuitry for driving display 354 to present graphics and other information to the user. Control interface 358 can receive commands from the user and translate those commands for submission to processor 352. Additionally, an external interface 362 can be provided to communicate with processor 352, enabling mobile computing device 350 to perform near-field communication with other devices. For example, Ethernet interface 363 may provide wired communication in some embodiments, or wireless communication in others, and multiple interfaces may also be used.

[0078] Memory 364 stores information within the mobile computing device 350. Memory 364 may be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. Extended memory 374 may also be provided and connected to the device 350 via an extended interface 372, which may include a SIMM (Single In Line Memory Module) card interface. Extended memory 374 may provide additional storage space for the mobile computing device 350, or it may store applications and other information for the mobile computing device 350. Specifically, extended memory 374 may include instructions to execute or supplement the processes described above, and may also include security information. Therefore, for example, extended memory 374 may be provided as a security module for the mobile computing device 350 and may be programmed with instructions authorizing secure use of the mobile computing device 350. Additionally, security applications along with additional information may be provided via a SIMM card, such as setting identification information on the SIMM card in a non-intrusive manner.

[0079] As discussed below, for example, the memory may include flash memory and / or NVRAM (non-volatile random access memory). In some embodiments, instructions are stored in an information carrier. When executed by one or more processing devices (e.g., processor 352), the instructions perform one or more methods such as those described above. The instructions can also be stored by one or more storage devices such as one or more computer-readable or machine-readable media (e.g., memory 364, extended memory 374, or memory on processor 352). In some embodiments, for example, the instructions can be received in a propagated signal manner via transceiver 368 or external interface 362.

[0080] When necessary, the mobile computing device 350 can communicate wirelessly via a communication interface 366, which may include digital signal processing circuitry. The communication interface 366 can provide communication under various modes or protocols, such as GSM voice calls (Global System for Mobile communication), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), and others. For example, such communication can be initiated using radio frequency via transceiver 368. Additionally, short-range communication can be initiated using transceivers such as Bluetooth, WiFi, or others (not shown). Additionally, the GPS (Global Positioning System) receiver module 370 can provide additional navigation-related and location-related wireless data to the mobile computing device 350, which can be used as appropriate by applications running on the mobile computing device 350.

[0081] The mobile computing device 350 can also communicate audibly using an audio codec 360, which can receive spoken information from a user and convert it into usable digital information. The audio codec 360 can also generate audible sounds for the user, such as through a speaker, for example, in the mobile phone of the mobile computing device 350. Such sounds can include sounds from voice call, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications running on the mobile computing device 350.

[0082] As shown in the figure, the mobile computing device 350 can be implemented in a variety of different forms. For example, the mobile computing device 350 can be implemented as a cellular phone 380. The mobile computing device 350 can also be implemented as part of a smartphone 382, ​​a personal digital assistant, or other similar mobile device.

[0083] Various implementations of the systems and techniques described herein can be implemented using digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system comprising at least one programmable processor, at least one input device, and at least one output device, the programmable processor being dedicated or general-purpose and coupled to receive data and instructions from and send data and instructions to the storage system.

[0084] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level programming languages ​​and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0085] To provide interaction with the user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including sound, speech, or tactile input.

[0086] The systems and technologies described herein can be implemented as computing systems, including: back-end components (e.g., as data servers), or middleware components (e.g., application servers), or front-end components (e.g., client computers having a graphical user interface or web browser through which users can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0087] A computing system can include clients and servers. Clients and servers are typically geographically separated and interact via communication networks. The client-server relationship is established by computer programs running on their respective computers and having a client-server relationship with each other.

[0088] While some implementations have been described in detail above, other modifications are possible. For example, although the client application is described as accessing (multiple) delegates, in other implementations, (multiple) delegates may be accessed by other applications implemented by one or more processors, such as applications running on one or more servers. Furthermore, the logical flow depicted in the figures does not require the specific or sequential order shown to achieve the desired result. Additionally, other actions may be provided, or actions may be eliminated from the described flow, and other components may be added to or removed from the described system. Therefore, other implementations are within the scope of the following claims.

Claims

1. A computer-implemented method, comprising: The audio data corresponding to the playback of the media content item is received at the data processing hardware. The audio data includes an audio watermark and a command phrase preceded by a hot word. The data processing hardware determines that the received audio data includes hot words; Audio data is processed by data processing hardware, in order to: Identify audio watermarks included in audio data; as well as Determine the corresponding bitstream for the audio watermark; as well as Based on the corresponding bitstream of the determined audio watermark, the data processing hardware determines how to bypass the execution of speech recognition for commands preceded by hot words without accessing the audio watermark database to identify matching audio watermarks.

2. The computer-implemented method according to claim 1, wherein, The audio watermark is added to the audio data by the creator of the media content project.

3. The computer-implemented method according to claim 1, wherein, Processing the audio data to identify audio watermarks in the audio data includes detecting watermarks of spread spectrum shaping type.

4. The computer-implemented method according to claim 1, wherein, The received audio data is determined to include hot words, including: Without performing speech recognition, calculate the hot word confidence score indicating the probability that the audio data includes hot words; and Determine if the confidence score of the hot words meets the hot word confidence score threshold.

5. The computer-implemented method according to claim 1, wherein: The data processing hardware resides on the user device; When receiving audio data, determining that the received audio data includes hot words, and processing the audio data, the user device is in one of three modes: sleep mode, standby mode, or low-power mode; and After determining that bypassing the execution of commands preceded by hot words is possible, the user device remains in one of the following modes: sleep mode, standby mode, or low power mode.

6. The computer-implemented method according to claim 1, further comprising: The audio watermark is analyzed by data processing hardware to identify the source of the audio data corresponding to the playback of the media content item. The determination to bypass the execution of the command preceded by the hot word is also based on the source of audio data corresponding to the playback of the media content item.

7. The computer-implemented method of claim 6 further includes updating a log file by data processing hardware to indicate the source of audio data corresponding to the playback of a media content item.

8. The computer-implemented method according to claim 1, wherein, The audio watermark is included in a portion of the audio data that includes hot words.

9. A computer-implemented system, comprising: Data processing hardware; and Memory hardware that communicates with and stores instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform operations, including: Receive audio data corresponding to the playback of a media content item, the audio data including an audio watermark and a command phrase preceded by a hot word; The received audio data was confirmed to include hot words; Process audio data to: Identify audio watermarks included in audio data; and Determine the corresponding bitstream for the audio watermark; and Based on the corresponding bitstream of a given audio watermark, without accessing an audio watermark database to identify matching audio watermarks, it is determined to bypass the execution of speech recognition for commands preceded by hot words.

10. The computer-implemented system according to claim 9, wherein, The audio watermark is added to the audio data by the creator of the media content project.

11. The computer-implemented system according to claim 9, wherein, Processing the audio data to identify audio watermarks in the audio data includes detecting watermarks of spread spectrum shaping type.

12. The computer-implemented system according to claim 9, wherein, The received audio data is determined to include hot words, including: Without performing speech recognition, calculate the hot word confidence score indicating the probability that the audio data includes hot words; and Determine if the confidence score of the hot words meets the hot word confidence score threshold.

13. The computer-implemented system according to claim 9, wherein: The data processing hardware resides on the user device; When receiving audio data, determining that the received audio data includes hot words, and processing the audio data, the user device is in one of three modes: sleep mode, standby mode, or low-power mode; and After determining that bypassing the execution of commands preceded by hot words is possible, the user device remains in one of the following modes: sleep mode, standby mode, or low power mode.

14. The computer-implemented system according to claim 9, wherein, The operation also includes: Analyze audio watermarks to identify the source of audio data corresponding to the playback of media content items. The determination to bypass the execution of the command preceded by the hot word is also based on the source of audio data corresponding to the playback of the media content item.

15. The computer-implemented system according to claim 14, wherein, The operation also includes updating log files to indicate the source of audio data corresponding to the playback of media content items.

16. The computer-implemented system according to claim 9, wherein, The audio watermark is included in a portion of the audio data that includes hot words.

17. A computer-implemented method, when executed on data processing hardware of a user equipment, causing the data processing hardware to perform an operation, comprising: When the user device is in sleep mode, it receives audio data captured by the user device's microphone, which corresponds to the playback of media content items output from an audio source near the user device; While the user device is in sleep mode, the audio data is processed to determine the corresponding bitstream of the audio watermark encoded in the audio data; Before processing the audio data to determine the corresponding bitstream of the audio watermark encoded in the audio data, it is determined that the received audio data includes hot words; and Based on the corresponding bitstream of the audio watermark encoded in the audio data, without accessing the audio watermark database to identify matching audio watermarks, it is determined to bypass performing speech recognition on the portion of the audio data where the hot words come first.

18. The computer-implemented method according to claim 17, wherein, The audio data includes utterances of commands preceded by hot words.

19. The computer-implemented method according to claim 18, wherein, Determining to bypass speech recognition on audio data includes bypassing speech recognition on a portion of the received audio data that corresponds to the utterance of the command.

20. The computer-implemented method according to claim 18, wherein, The audio watermark is encoded in a portion of the audio data, which includes hot words.

21. The computer-implemented method according to claim 17, wherein, The received audio data is determined to include hot words, including: Without performing speech recognition, calculate the hot word confidence score indicating the probability that the audio data includes hot words; and Determine if the confidence score of the hot words meets the hot word confidence score threshold.

22. The computer-implemented method according to claim 17, wherein, When the user device is in sleep mode, processing audio data also includes processing the audio data to identify audio watermarks encoded in the audio data by detecting spread spectrum shaping type watermarks.

23. The computer-implemented method according to claim 17, wherein, After determining that speech recognition on the audio data is bypassed, the user equipment remains in sleep mode.

24. The computer-implemented method according to claim 17, wherein, The operation also includes: Analyze audio watermarks to identify the audio source of audio data corresponding to the playback of media content items. The determination of bypassing speech recognition on the audio data is also based on the audio source of the audio data corresponding to the playback of the media content item.

25. The computer-implemented method according to claim 24, wherein, The operation also includes updating log files to indicate the audio source of the audio data corresponding to the playback of the media content item.

26. A user equipment, comprising: Data processing hardware; and Memory hardware that communicates with and stores instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform operations, including: When the user device is in sleep mode, it receives audio data captured by the user device's microphone, which corresponds to the playback of media content items output from an audio source near the user device; While the user device is in sleep mode, the audio data is processed to determine the corresponding bitstream of the audio watermark encoded in the audio data; Before processing the audio data to determine the corresponding bitstream of the audio watermark encoded in the audio data, it is determined that the received audio data includes hot words; and Based on the corresponding bitstream of the audio watermark encoded in the audio data, without accessing the audio watermark database to identify matching audio watermarks, it is determined to bypass performing speech recognition on the portion of the audio data where the hot words come first.

27. The user equipment according to claim 26, wherein, The audio data includes utterances of commands preceded by hot words.

28. The user equipment according to claim 27, wherein, Determining to bypass speech recognition on audio data includes bypassing speech recognition on a portion of the received audio data that corresponds to the utterance of the command.

29. The user equipment according to claim 27, wherein, The audio watermark is encoded in a portion of the audio data, which includes hot words.

30. The user equipment according to claim 26, wherein, The received audio data is determined to include hot words, including: Without performing speech recognition, calculate the hot word confidence score indicating the probability that the audio data includes hot words; and Determine if the confidence score of the hot words meets the hot word confidence score threshold.

31. The user equipment according to claim 26, wherein, When the user device is in sleep mode, processing audio data also includes processing the audio data to identify audio watermarks encoded in the audio data by detecting spread spectrum shaping type watermarks.

32. The user equipment according to claim 26, wherein, After determining that speech recognition on the audio data is bypassed, the user equipment remains in sleep mode.

33. The user equipment according to claim 26, wherein, The operation also includes: Analyze audio watermarks to identify the audio source of audio data corresponding to the playback of media content items. The determination of bypassing speech recognition on the audio data is also based on the audio source of the audio data corresponding to the playback of the media content item.

34. The user equipment according to claim 33, wherein, The operation also includes updating log files to indicate the audio source of the audio data corresponding to the playback of the media content item.

Citation Information

Patent Citations

  • Additional information embedding method, additional information reading method, and speech recognition system

    JP2003295894A

  • Audible command filtering

    US9548053B1