Dynamic Wake Word for Voice-Enabled Devices
The method of dynamically assigning wake words through natural language queries addresses wake word collisions by constructing new wake word spotters quickly, enhancing user experience and power efficiency in voice-enabled devices.
Patent Information
- Application Number
- JP2020201933
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-05
- Filing Date
- 2020-12-04
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2040-12-04
AI Technical Summary
Existing voice-enabled devices face issues with wake word collisions, leading to inappropriate device activation and degraded user experience due to the inability to dynamically change wake words without a large dataset of audio samples.
A method for dynamically assigning wake words through natural language queries, allowing for the quick construction of a wake word spotter by parsing oral requests into natural language and phonetic segments, and constructing a new wake word spotter using segmentation, sampling, or continuous transformation approaches.
Enables immediate activation of new wake words within seconds, reducing power consumption and preventing device collisions by allowing personalized and customized wake word settings.
Smart Images

Figure 0007715494000001 
Figure 0007715494000002 
Figure 0007715494000003
Abstract
Description
Technical Field
[0001] Field The present technology relates to wake words for voice-responsive devices, and more particularly to dynamically assigning wake words using natural language queries and quickly constructing a wake word spotter for wake words for one or more voice-responsive devices of a single user.
Background Art
[0002] Background An automatic speech recognition (ASR) system that recognizes human speech offers great potential as an easy and natural way to connect to voice-responsive devices, along with a natural language understanding (NLU) function that extracts the meaning of the speech. Such systems are partially enabled by the vast computing and communication resources available on current devices. Advanced voice understanding systems, such as virtual assistants, have been developed that can recognize a wide range of speech and handle complex requests in various languages and dialects.
[0003] Virtual assistants do not respond to verbal requests when they are idle. They wake up or activate and switch their state from idle to active when they receive an activation signal such as a tap, button press, or a verbal wake-up phrase (or wake phrase) called a wake word. The use of wake words is an important hands-free and eye-free operation for voice-responsive devices. In the active state, virtual assistants respond to user requests. They usually return to the idle state after responding to the request. When idle, the voice-responsive device continuously monitors incoming voice to detect the wake word. To reduce power consumption, some devices operate in a low-power mode when the virtual assistant is idle and return to the full-power mode when the virtual assistant is activated.
[0004] A wake word is generally a word or short phrase. A continuously operating module that monitors incoming audio to detect a wake word is called a wake word spotter. Some commercial examples of wake words for voice-enabled devices include "Hey, Siri," "Okay Google," and "Alexa." Voice-enabled devices may be sold with a wake word installed at the factory and a wake word spotter that is ready to detect these predefined wake words.
[0005] A wake word spotter is a voice processing algorithm specifically designed to detect an assigned wake word or set of assigned wake words in a continuous audio stream. This algorithm typically runs continuously at a fixed frame rate and is necessarily very efficient. On low-power mode devices, the spotter can operate continuously without drawing excessive power, thus saving battery life. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0006] There are times when it may be desirable to customize the wake word installed at the factory on one or more voice-enabled devices. For example, in a home or office setting, there may be several devices that use the same wake word installed at the factory as an activation phrase. This can lead to the inappropriate activation of a device or a collision where a common wake word is detected and multiple devices are activated. The activation of multiple devices can lead to various problems depending on the type of request that follows the wake word. For example, a request to play music can result in multiple devices playing the same song (asynchronously) or different songs simultaneously. A request to send a message can result in multiple copies of the message being sent. These and other collision scenarios lead to a degraded user experience.
[0007] A major challenge in providing a dynamic wake word is training a new wake word spotter in a very short period of time. Generally, the wake word spotter installed at the factory is trained using a large dataset of audio samples that include positive examples recorded specifically for one or more given wake words and perhaps some negative examples. Such labeled samples are used to train a classifier algorithm, such as a recurrent neural network, to distinguish a given wake word (or multiple wake words) in an audio stream from utterances that are not wake words. Unfortunately, traditional approaches to collecting audio sample data are not usable for dynamic wake words where the spotter must be constructed quickly without the aid of collecting a large dataset of audio samples for the dynamic wake word.
Means for Solving the Problem
[0008] Summary In certain aspects of the present disclosure, a method for changing a set of one or more wake words of a voice-responsive device is provided. The method may include receiving an oral request from a user and parsing the oral request into a natural language request and an uttered voice segment, where the natural language request instructs the device to accept the uttered voice segment as a new wake word, and the method may further include constructing a new wake word spotter to recognize the new wake word as an activation trigger.
[0009] In another aspect of the present disclosure, a method for changing a set of one or more wake words of a voice-responsive device is provided. The method may include receiving an oral utterance and parsing the utterance into a natural language request and an uttered voice segment, where the natural language request instructs the device to accept the uttered voice segment as a new wake word, and the method may further include using automatic speech recognition to map the new wake word to a new wake word phonetic sequence, and constructing a new wake word spotter to recognize the new wake word phonetic sequence as an activation trigger. The step of constructing the new wake word spotter may be performed by splitting the new wake word phonetic sequence into a sequence of two or more consecutive partial phonetic segments, providing a corresponding partial wake word spotter for each partial phonetic segment, and sequentially assembling the provided partial wake word spotters into the new wake word spotter for the entire new wake word phonetic sequence.
[0010] In yet another aspect of the present disclosure, a method for changing a set of one or more wake words of a voice-responsive device is provided. The method may include receiving an oral request and parsing the oral request into a natural language request and an utterance voice segment, the natural language request instructing the device to accept the utterance voice segment as a new wake word, and the method may further include defining a new wake word spotter to recognize the new wake word as an activation trigger, the step of defining the new wake word spotter being performed by a step of requesting further utterance voice samples of the utterance voice segment, a step of converting the utterance voice segment and the further utterance voice samples into a phoneme sequence, and a step of defining the new wake word spotter based on one or more of the phoneme sequences of the phoneme sequence.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0012] Detailed description An automatic speech recognition (ASR) system that recognizes human speech offers great potential as an easy and natural way to connect with voice-enabled devices, along with a natural language understanding (NLU) function that extracts the meaning of the speech. Such systems are partially enabled by the vast computing and communication resources available on current devices. Advanced speech understanding systems have been developed that can process complex utterances and recognize a wide range of speech in various languages and dialects.
[0013] Here, the present technology will be described with reference to the drawings. In an embodiment, the present technology relates to a system that can analyze a received utterance into a natural language request and an utterance audio segment, and this request instructs the system to use the utterance audio segment as a new wake word. This type of request is called a wake word assignment directive (WAD). In response to such a request, the system can further construct a new wake word spotter to recognize the new wake word, and the construction of the spotter is fast enough that the new wake word can be used immediately after the system's response to the WAD.
[0014] In the context of a voice-enabled system, the terms utterance, query, and request are closely related and can sometimes be used synonymously. A verbal natural language request from a user is conveyed simultaneously as an utterance audio (utterance) and, when properly written, as a word (query). The device can perform various actions in response to common queries.
[0015] A wake word assignment directive, or simple directive or WAD, is a request to a device to change its wake word set by addition or replacement.
[0016] One such action may be a natural language request in the form of a wake word assignment directive to assign a new wake word. Upon recognizing such a wake word assignment directive, the technology may quickly construct a new wake word spotter for this dynamic wake word. The term "quickly" in this context means that the new wake word spotter can be constructed within seconds after receiving the new wake word, as will be explained in more detail below. Without the aid of a large dataset of audio instances of the new wake word, the use of dynamic wake words requires other approaches to quickly construct a new wake word spotter. These approaches include at least the following three and their variations: 1. A wake word segmentation approach, 2. A wake word sampling approach, and 3. A continuous transformation approach.
[0017] Each of these will be described in detail below. Immediately after the construction of the new wake word spotter is completed (e.g., within seconds), the dynamic wake word and its spotter may be stored and made ready to activate the device.
[0018] The parser may further identify any parameters that define variations of the directive as part of a predefined wake word assignment directive template, and these parameters may be characteristics of the dynamic wake word (i.e., how it is recognized) or characteristics of the directive (i.e., how the WAD is satisfied). For example, the user may specify that the new wake word is public, which means that other users may use the same dynamic wake word to wake up the device. Or, the user may specify that the new wake word is private, which means that the new wake word only functions for that user and excludes others. As a further example, the user may specify whether the new wake word replaces the previous wake word or is used in addition to the previous wake word.
[0019] Directive parameters have natural language phrases that convey specific parameter values when found within the directive. If there are no arbitrary parameters, it may have an implementation-dependent implicit default value, or the system may prompt the user for a value. In the present disclosure, any directive parameter is simply referred to as a parameter.
[0020] It is understood that the present invention may be embodied in many different forms and should not be construed as limited to the embodiments described herein. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the invention to those skilled in the art. In fact, the present invention is intended to cover alternatives, modifications, and equivalents of these embodiments that are within the scope and spirit of the invention as defined by the appended claims. Further, the following detailed description of the present invention includes numerous specific details for the purpose of providing a complete understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without these specific details.
[0021] FIG. 1 is a schematic block diagram of an example of a voice-responsive device 100 in which the present technology can be implemented. The device 100 may be an agent having any of various electronic or electromechanical components configured to receive voice requests, or may include such an agent. Examples of agents include mobile phones, digital assistants, tablets and other computing devices, automotive control systems, and the like.
[0022] A more detailed description of an example of the voice-responsive device 100 is provided below with reference to FIG. 9, but generally, the device 100 may include a processor 102 configured to control operations within the device 100 and facilitate communication between various components within the device 100. The processor 102 may include a standardized processor, a specialized processor, a microprocessor, or the like that can execute instructions for controlling the device 100.
[0023] The processor 102 may receive and process inputs from various input devices including one or more microphones 106. The microphone 106 may include a transducer or sensor that can receive sound and convert it into an electrical signal. According to one embodiment, the microphone 106 may be used to receive voice signals, which are processed as requests, i.e., requests and inputs to the device 100 described below.
[0024] As described above, device 100 may operate in a low-power mode when not in use to conserve energy. A power circuit 108 may be provided to control the power level within device 100 under the management of processor 102. In the low-power mode, most of the systems of device 100 are shut down, and only some components are operated by processor 102. One such component is wake word spotter 112, which will be described below. Microphone 106 also operates in the low-power mode to continuously monitor the environment surrounding device 100 for voice input.
[0025] In an embodiment, device 100 may operate at 0.3 to 0.5 watts in the low-power mode and at 5 to 10 watts in the full-power mode. However, it is understood that in further embodiments, device 100 may operate at various power levels in the idle or active states. In one example, when wake word spotter detects the occurrence of any one of one or more current wake words 120 in the input stream, processor 102 may exit the low-power mode and instruct power circuit 108 to power on the device. After a predetermined period (e.g., 2 to 10 seconds) likely elapses after the user request is completed, the device may return to the idle state, and processor 102 may instruct power circuit 108 to return to the low-power mode. In some embodiments, such as when device 100 is plugged into an outlet or has a large battery, power circuit 108 may be omitted and there may be no low-power mode. However, since wake word spotting is still required to activate the device, it is ready to listen for queries.
[0026] In the illustrated embodiment, the wake word spotter 112 is present on the device 100. Executing the wake word spotter locally can be a preferred embodiment. The operation of the word spotter is driven by a data structure 120 that includes the current wake word, wake word spotter, and their associated parameters. Detection of any wake word causes a transition to the active state.
[0027] The wake word spotter for a wake word may be implemented by a classifier that has two results (the wake word matches or does not match). In some embodiments, multiple wake words are used simultaneously. Detecting any one of N wake words in parallel can be realized by an integrated classifier that has N + 1 results, i.e., one result for each wake word and one result for not matching any wake word. In such an embodiment, the wake word with the highest score may be the wake word that activates the device. In other embodiments, detecting multiple wake words in parallel may be realized by running multiple wake word spotters in parallel from the same incoming audio stream. In such an embodiment, the earliest matching wake word becomes the wake word that activates the device. The ability to use parallel spotters in this way is a great advantage for the use of dynamic wake words. Because the addition of a new wake word spotter can be done without considering the previously existing spotters.
[0028] In any of the above spotter embodiments, it must not be forgotten that a private wake word requires positive speaker verification before the device can be activated. The speaker verification engine 114 may operate continuously, sacrificing power to achieve low latency. The power consumption can be low if the speaker verification engine 114 is triggered only when the private wake word matches.
[0029] The voice-responsive device 100 may further include a memory 104 that can store algorithms executable by the processor 102. According to an example of an embodiment, the memory 104 may include a random access memory (RAM), a read-only memory (ROM), a cache, a flash memory, a hard disk, and / or any other suitable storage component. As shown in FIG. 1, in one embodiment, the memory 104 may be a separate component that communicates with the processor 102, but in a further embodiment, the memory 104 may be incorporated into the processor 102.
[0030] The memory 104 may store various software application programs executed by the processor 102 to control the operation of the device 100. Such application programs may include, for example, a wake word spotter 112 for detecting a wake word in a received utterance and a speaker verification engine 114 for verifying a speaker. The speaker verification engine 114 is necessary for verifying the speaker ID when the matched wake word is private. This will be described in more detail below.
[0031] The memory 104 may also store various data records including, for example, one or more wake words 120 and one or more user voiceprints 122. Each of these will be described in more detail.
[0032] The device 100 may further include a communication circuit such as a network interface 124 for connecting to various cloud resources 130 via the Internet. One such resource may be one or more speech recognition and spotter construction servers 150, which are also simply referred to as servers 150 herein. Here, an example of the server 150 will be described with reference to FIG. 2.
[0033] Figure 2 is a schematic block diagram of an embodiment of server 150. As described above, in further embodiments, server 150 may be composed of a plurality of co-located or otherwise configured servers. A more detailed description of sample server 150 is provided below with reference to FIG. 9, but generally, server 150 may include a processor 152 configured to control the operation of server 150 and facilitate communication between various components within server 150. Processor 152 may include a standardized processor, a specialized processor, a microprocessor, etc., capable of executing instructions for controlling server 150.
[0034] Server 150 may further include a memory 154 capable of storing algorithms executable by processor 152. According to an example of an embodiment, memory 154 may include RAM, ROM, cache, flash memory, hard disk, and / or any other suitable storage component. As shown in FIG. 2, in one embodiment, memory 154 may be a separate component communicating with processor 152, but in further embodiments, memory 154 may be incorporated into processor 152.
[0035] Memory 154 may store various software application programs executed by processor 152 to control the operation of server 150. Such application programs may include, for example, a speech recognition engine 162 for converting speech. The application program may further include a general parser 155 and a general query execution engine 157 for processing general (non-WAD) queries. The application program may further include a wake word assignment directive processor (WAD processor 164) for processing dynamic wake word assignment requests. Although speech recognition is complex, many techniques are well established and need not be described in detail here. In most embodiments, it is sufficient to know that the speech recognition engine 162 has a front end that can generate a speech transcription of its input, while the full ASR engine 162 generates a text transcription of the input. When the present disclosure refers to the ASR engine 162, this refers to either the ASR front end or the entire ASR engine, depending on whether a speech output or a text output is required.
[0036] WAD processor 164 may have software components including a WAD parser 166, a spotter builder 168, and a registration engine 170. The components of WAD processor 164 will be described in more detail below. In particular, spotter builder 168 may build a wake word spotter according to one of the methods described below. Registration engine 170 reads and writes wake word data record 120 representing the wake words used by device 100, along with the associated spotters and parameters.
[0037] Server 150 may further include a communication circuit such as network interface 156 for connecting to cloud resources 130, such as client device 100, via the Internet. As described above, server 150 may communicate with a plurality of client devices 100, and each of the plurality of client devices 100 is configured as shown in FIG. 1 and as described below.
[0038] Here, the operations and interactions of client device 100 and server 150 for recognizing wake words and special queries and setting new wake words will be described with reference to the flowchart of FIG. 3. This figure is divided into a left column showing modules operating locally on device 100 and a right column showing modules operating on remote server 150 in the illustrated embodiment. In other embodiments, some or all of the components shown in the right column may actually operate locally on device 100.
[0039] In step 200, device 100 is in an idle state and cannot process a speech request. In this state, it continuously attempts to recognize a wake word using wake word spotter 112. The device may be in a low power mode to conserve energy while in the idle state. Device 100 remains in the idle state until wake word spotter 112 recognizes a wake word in the incoming audio stream. One or more spotters may be used to continuously test one or more wake words against the audio input. The current wake word (and corresponding spotter) is found in wake word data structure 120 that is local to the device.
[0040] It is desirable to recall that when the speaker verification test fails, the private wake word is not considered a match in the voice input. When the wake word matches (Yes in step 202), the device exits the idle state and enters the active state (step 206). The voice input that starts from the end of the matched wake word and ends with the end of utterance (EOU) is an oral query. The EOU may be a pause during speech, or a tap, button press or release. The oral query is sent from the client device to the server (step 212). Thus, the oral query is provided as input to the speech recognition engine 162, and the speech recognition engine 162 creates a transcription of the oral query in step 216. In embodiments, the ASR engine 162 may operate locally on the device 100 or remotely on the server 150. In some examples, the speech recognition engine 162 may generate one or more speech and / or text transcriptions of the wake word or query and a score for each one indicating the confidence level of the respective transcription. The ASR algorithm may utilize any combination of signal processing, hidden Markov models, Viterbi search, speech dictionaries and (possibly recursive) neural networks to generate the transcriptions and their confidence scores.
[0041] Generally, a virtual assistant may process various queries that include information and requests for commands to instruct a client device 100 to perform some action. The way in which a virtual assistant understands queries varies considerably among known embodiments. In the illustrated embodiment, at step 218, non-WAD queries are recognized and parsed and interpreted by a general query parser 155. Specifically, after a speech recognition engine 162 generates a transcription of the query, the general query parser 155 determines the structure (syntax) and meaning (semantics) of the query. In an embodiment, this task may be performed remotely on a server 150 or locally on the device 100. Parsing and processing spoken queries may utilize known algorithms for processing queries. Such systems are disclosed, for example, in U.S. Patent No. 10,217,453 entitled “Virtual Assistant Composed of Selection of Launch Phrases” and U.S. Patent No. 10,347,245 entitled “Natural Language Grammar Validation by Utterance Characterization,” both of which are assigned to SoundHound Inc., headquartered in Santa Clara, California, and the entire contents of which are hereby incorporated by reference herein.
[0042] One particular type of query related to this technology is a query that requests the assignment of a new wake word to device 100. For the purposes of FIG. 3, at step 218, these particular queries, herein referred to as wake word assignment directives or WADs, are processed (parsed) by a special purpose parser, namely the WAD parser 166. As noted above, the speech recognition engine 162 provides both speech transcription and text transcription. The text transcription may be used by the WAD parser 166 according to known NLU algorithms to identify the syntax of the directive. The NLU algorithms may utilize one or more grammar patterns to identify the meaning of the directive portion of the query and, in some cases, the meaning of the wake word portion as well. However, generally, the wake word may be any word or phrase. In special cases, it is also possible for the wake word to be a common word or phrase, or a known name. In some embodiments, the speech recognition engine 162 may use a language model to increase the transcription score of likely wake words. The NLU algorithms may also play a role in parsing the wake word. However, the wake word may be any utterance segment, i.e., a speech wildcard, and is ultimately delimited (segmented) by the directive syntax of the words surrounding the wake word. It is understood that various schemes may be used to determine the presence of the wake word as well as the presence and meaning of directive words (phrases) before or after the wake word. In all such schemes, it is noteworthy that once a wake word segment is determined, its speech transcription becomes available for further processing by the spotter builder 168.
[0043] In FIG. 3, the query recognized by the general query parser 155 is passed for further processing (i.e., execution) by the general query execution engine 157 in step 220, and the general query execution engine 157 operates in a manner expected of a particular virtual assistant. There is no need to change the "host" virtual assistant to implement the dynamic wake word. In the exemplary embodiment shown in FIG. 3, after a normal query is processed by the general query execution engine 157 in step 220, the device returns to the idle state. In a variation of the embodiment not shown, the device remains active for a short period (i.e., several seconds) before returning to the idle state (i.e., does not require a wake word to receive a query). One way to implement this variation is to return to the active state but ensure that a timeout is set when entering the active state, and at the end of the timeout, the device returns to the idle state.
[0044] If the query is a WAD, the general query parser 155 cannot recognize it in step 218. Instead, in step 224, the WAD parser 166 can analyze it to determine any parameters associated with the new wake word and directive. In a variation of the embodiment, the WAD parser 166 may be executed before the general query parser 155, or alternatively, both parsers 155 and 166 may be part of a single integrated parser. These are minor changes to the control flow shown in FIG. 3.
[0045] Insertion phrase in case of failure: If neither parser 155 nor 166 can recognize the query, the device returns to the idle state, most likely after issuing an appropriate error message. Failure can occur, for example, when the speech recognition unit 162 cannot reliably determine the transcription of the query. This is less likely when the system admits multiple transcription hypotheses with different scores. Failure can also occur after an excellent transcription of the query is obtained when the query is not syntactically correct, which can be a syntactic failure when neither parser 155 nor 166 can recognize the query. Other failures can occur when the query is syntactically correct but its interpretation (meaning) cannot be reliably determined, which is a semantic failure. Furthermore, a correctly interpreted query may fail during its execution, which is called query execution. Normal query execution is performed by the normal query execution engine 157. The execution of the WAD consists of the spotter builder and the registration engine 170, and both the spotter builder and the registration engine 170 can indicate their own failures.
[0046] The WAD parser 166 collects information including directives and their parameters from the query. In addition to adding wake words, there are other wake-word related actions such as deleting wake words, listing wake words, and restoring wake words to their previous state stored in the wake-word record 120. This disclosure focuses on adding wake words. This is because it is the most technically difficult part, and the other related actions are easy to explain. Parameters (e.g., public / private and exclusive / inclusive) and wake words are also determined by the WAD parser 166. The spotter builder 168 checks whether the wake word is usable. For example, if the requested wake word is too short, or too ambiguous, or too close to an element of the list of undesirable wake words such as offensive words, or is a wake word that has existed before for the device 100, a new wake-word spotter cannot be constructed. In such cases, the spotter builder 168 exits in a failure state, and a message may be output to the user via the device 100 to convey the fact that the request cannot be satisfied.
[0047] When a new wake word is received, the spotter builder 168 then constructs a spotter for the wake word in step 226. Since the spotter builder 168 constructs a spotter in response to a single reception (receipt) of a new wake word, the spotter can be constructed promptly from the user input of the new wake word. For further details of various embodiments for constructing a new wake word spotter by the spotter builder 168, reference will be made to and described below with reference to FIGS. 5 - 8. When the spotter builder is successful, the wake word, its wake word spotter, and associated parameters are passed to the registration engine 170. When the spotter builder operates on the server 150, this data (including the new spotter) is downloaded to the device 100. The registered spotter, wake word, and parameters are stored in a data structure 120 on the device 100 that includes one or more wake words and associated data.
[0048] As an example of the operation of the flowchart of FIG. 3, the client device 100 may receive the utterance "in response to Okay Agent, Okay Service", where "Okay Agent" is the current wake word. The remainder of the utterance following the wake word "Okay Agent" is the query "in response to Okay Service". This query may be uploaded to the server 150 and converted using the speech recognition engine 162 in step 216. In step 224, the WAD parser 166 recognizes the query as a directive and extracts the wake word. In this example, the query "in response to Okay Service" is a wake word assignment directive that requests that the wake word utterance segment "Okay Service" be assigned as a new wake word. After successfully parsing this wake word assignment directive into its wake word and parameters in step 224, in step 226, the wake word spotter builder 168 must obtain a spotter for the new wake word, either by constructing a brand new wake word spotter or by placing an existing spotter for that wake word. This can be done in a number of ways as will be described later. Next, in step 228, the new wake word spotter, the new wake word and the associated parameters are registered by the registration engine 170, which modifies the data structure of the wake word 120 that holds the current wake word, spotter and associated parameters.
[0049] FIG. 3 shows an embodiment of a process executed on client device 100 or server 150. In the illustrated embodiment, all processes in the left column (labeled "client device") are executed on device 100, and all processes in the right column (labeled "server") are executed on server 150. In an alternative embodiment, some or all of the processes shown in the right column may actually be executed on the client side instead. For example, steps 216 and 224 may be executed locally on device 100 (WAD is recognized locally), while the general query analysis in step 218 and the general query processing in step 220 may be executed on server 150. In such a case, WAD parser 166 operates before general query parser 155. Under these conditions, the spotter builder is provided to operate partially or entirely on device 100. In a first embodiment, the spotter builder has access to the computing power of server 100 and the labeled audio database when constructing the spotter. This is illustrated by the segmentation approach for spotter construction described later. In a second embodiment, the spotter builder is entirely local, which is illustrated by the conversion approach for spotter construction described later.
[0050] As described above, the assignment of wake words requires directives, utterance audio segments of new wake words, and any parameters that control the processing of new wake words. In one example, the parameters may be related to whether the newly assigned wake word is public or private. A public wake word can be used by anyone to start up device 100 and gain access to the resources of device 100. On the other hand, since a private wake word is personal to the user who created it, device 100 will only activate when the wake word is spoken by the wake word creator and will remain idle when spoken by others. In another example, the parameters may be related to whether the new wake word is added to an existing set of wake words or replaces some or all of the existing wake words.
[0051] Referring now to the flowchart of FIG. 4, at step 260, the WAD processor 164 may check whether the wake word assignment directive includes a parameter indicating whether the new wake word was public or private. The following are some examples of parameters that may indicate to the WAD processor 164 that the wake word is public. In the following examples of wake word assignment directives, the wake word itself is omitted and only the query portion of the utterance remains.
[0052] "Respond to Okay Victoria" (assuming the default is "public") "Respond to Okay Madeline, which is a public wake word" "Respond to the public wake word Okay Madeline" "We will call you 'Hey, Jackson'" In the first example, since no parameters are provided, the WAD processor 164 by default assigns the new wake word "Okay Victoria" as the public wake word. In the second example, the parameter phrase "public wake word" is clearly stated before or after the spoken voice segment, and the corresponding characteristics are set. The WAD parser 166 may search for such predefined parameter phrases so as not to consider them as part of the spoken voice segment. In the fourth example, the "public" parameter setting is not explicit and may be inferred from the use of the plural subject "we" (as opposed to the singular "I") in the wake word assignment directive, and the new wake word "Hey, Jackson" is implicitly designated as the public wake word. One can imagine various other examples of wake word assignment syntax where explicit or implicit phrases of parameters in the verbal wake word assignment directive indicate that the new wake word is public.
[0053] Alternatively, the wake word assignment directive may include parameters indicating that the wake word is private. The following are some examples.
[0054] "In response to Okay Victoria" (assuming the default is "private") "In response to Okay Josephine Privately " " Privately "In response to Okay Josephine" " It is a private wake word "In response to Hey, Christopher" " I "will call you Tabatha" " Your nickname "is Penelope" The first example shows that when no parameters are provided, the default may instead be to make the wake word private. In the next four examples, before or after the spoken voice segment, parameter phrases such as "privately" or "private wake word" or "nickname" are used as part of the query to convey private parameter settings. In the last example, the private setting may be inferred from the use of the first-person singular subject pronoun "I" at the beginning of the query. Explicit phrases or contextual parameters in the verbal wake word assignment directive can create various other examples indicating that the new wake word is intended to be private.
[0055] When a public / private parameter is detected at step 260, at step 262, the WAD processor 164 may check whether a private parameter exists. If not, the WAD processor 164 may consider the new wake word to be public when storing it. On the other hand, if the parameter indicates that the wake word is private at step 262, the WAD processor 164 may perform step 264 of creating a voiceprint of the speaker's voice and possibly additional speaker verification data. The user voiceprint information may be stored in the voiceprint data structure 122 on the device 100 for subsequent speaker verification, or the device 100 may be able to access it from the cloud user record.
[0056] At step 268, the WAD processor 164 may associate the speaker verification data with the new wake word and spotter by calculating a voiceprint. The WAD processor 164 may store the speaker verification data considering the new wake word to be private. Such verification data, such as a voiceprint, may be stored in the memory 104 for individual users in the user voiceprint 122.
[0057] During operation, when the processor 102 receives speech and determines the presence of a private wake word, the processor 102 may further confirm whether the speaker is the same as the speaker who created the private wake word using data from the user voiceprint 122 associated with the private wake word. If a match exists, the processor 102 may notify the power circuit 108 to power on the device 100. If no match exists, the processor may ignore the wake word and remain in the idle state.
[0058] Instead of, or in addition to, the public / private parameter, the wake word assignment directive may include parameters regarding whether the new wake word is added to a wake word set of one or more existing wake words or replaces one or more existing wake words in the wake word set. In step 272, the WAD processor 164 checks whether the wake word assignment directive includes parameters regarding whether the new wake word is added to or replaces an existing wake word. The following are some examples of parameters that may indicate to the WAD processor 164 that the wake word is added to an existing wake word.
[0059] "In response to Okay Frederick" "In response to Esteban, the added wake word" "Also in response to Okay Samuel" In the first example, since no parameters are provided, the WAD processor 164, by default, adds the new wake word "Frederick" as an additional wake word to the other existing wake words 120. Thus, the user can wake up the device 100 with the wake word "Okay Frederick" in addition to one or more of the previously existing wake words 120 in the memory 104. Here, the new wake word is said to be additive with respect to the existing wake words. In the second example, the parameter "added wake word" may be a predefined phrase that is explicitly stated before or after the utterance audio segment to make the new wake word additive with respect to the existing wake words. When parsing the received wake word assignment directive, the WAD processor 164 may look for such predefined parameters so as not to consider them as part of the utterance audio segment. In the third example, the parameter "also" (part of "in response to also" or "in response to... also") explicitly requests adding the new wake word to the existing wake words. Various other examples can be envisioned where clear and / or contextual parameters in the verbal wake word assignment directive indicate that the new wake word is additive to the wake word set of one or more existing wake words.
[0060] Alternatively, the wake word assignment directive may include a parameter indicating that the wake word replaces one or more of the current wake words. The following are some examples (underlined to emphasize the parameters).
[0061] "Okay Brittany only in response to" "Marina only in response to" The first example shows that when no parameters are provided, the default can instead exclusively use the wake word and remove previous wake words. In the second example, the parameter "only" may be a predefined phrase that is clearly stated after the uttered voice segment. In the third example, the parameter "only to" may be a predefined phrase that is clearly stated before the uttered voice segment. Instead of the predefined phrases indicating the exclusivity of the new wake word in the second and third examples, the exclusivity of the new wake word may be inferred from the context of the parameters in the wake word assignment directive. Various other examples can be considered where the fact that the new wake word is private is indicated by explicit and / or contextual parameters in the verbal wake word assignment directive.
[0062] If an exclusive or additional parameter is detected in step 272, in step 274, the WAD processor 164 may check whether the parameter indicates that the new wake word is additional. If so, in step 276, a flag may be set to store the new wake word in addition to the existing wake words when the new wake word is stored. On the other hand, if the parameter indicates in step 274 that the wake word replaces one or more existing wake words, in step 280, the WAD processor 164 may check whether multiple wake words are stored. If so, in step 282, the processor 102 may generate a query as to which of the multiple wake words is to be replaced. Alternatively, steps 280 and 282 may be skipped, and by default all existing wake words may be replaced.
[0063] If one or more wake words are replaced, in step 286, when the new wake word is stored, a flag indicating which of the stored wake words is replaced may be set. It may happen that the user does not have the authority to replace one or more of the wake words, which can be determined, for example, from the data stored in the user voiceprint 122 in the memory 104. In this case, the processor may replace only the wake words that the user has the authority to replace.
[0064] It is further understood that parameters related to private / personal or additionally / exclusively may be given in a single wake word assignment directive. The following are some examples.
[0065] "I only call you 'Okay Robert'." "We will also call you 'Okay Natalia'." In the first example, by using the first-person singular subject pronoun "I" at the beginning of the wake word assignment directive, the new wake word "Okay Robert" is specified as a private wake word, and by using the word "only", it is made an exclusive wake word. In the second example, by using the first-person plural subject pronoun "We" at the beginning of the wake word assignment directive, the new wake word "Okay Natalia" is specified as a public wake word, and by using the phrase "also", it is made an additional wake word. The wake words in the above examples are common names for people, but it is understood that the wake word 120 may include any word or phrase that is meaningless or otherwise. In a further embodiment, it is conceivable that the new wake word is formed from a non-audio sound such as, for example, a doorbell, an alarm, or a drumbeat.
[0066] The above description of the parameters relates to two specific aspects (private / personal and additional / exclusive), but it is understood that in addition to, or instead of, the above examples, parameters related to other aspects for generating new wake words and spotters may also be provided.
[0067] Furthermore, in the above example, the parameters were given in a single utterance that included a wake word assignment directive. In further embodiments, the parameters may be established in a modal discourse between the user and the device 100. In particular, the user may first utter a wake word assignment directive without parameters. Thereafter, the WAD processor 164 may prompt the user to provide further parameters and information by the processor 102 that generates text, which text is converted to speech by the TTS algorithm and reproduced by the speaker 136. The following is an example of how parameters may be requested and provided in such a modal discourse.
[0068] (U) In response to Jarvis (D) Understood. Can you say that three times? (U) Jarvis... Jarvis... Jarvis (D) Okay, from now on I will respond to Jarvis (D) Should Jarvis be private or public? (U) Public (D) Should the previous wake word be left? (U) No (D) Do you really want to forget Okay Agent, which is the default public wake word? (U) No, say the wake word (D) Understood. Jarvis is a public wake word and Okay Agent is a public wake word.
[0069] (U) Disable the previous wake word or (U) Disable the okay agent (D) The okay agent has been disabled The above is an example of modal discourse between a user (U) and a device 100 (D). As can be seen above, the device 100 may receive a wake word assignment directive and prompt the user to repeat the wake word several times. Thereafter, the device 100 may prompt the user to specify whether the new wake word is public or private and whether it is additional or exclusive.
[0070] In the case of a device 100 controlled by a single owner, the device may allow the device owner to set its (private or public) wake word so that other users cannot control it. However, if the device 100 is a shared device controlled by multiple users, it may be desirable to save the public wake word after replacement. In particular, when the directive replaces all previous wake words on the device, there is a risk that users (in some cases, even the user who issued the directive) may be locked out and unable to activate or access the resources of the device 100.
[0071] Therefore, a recovery procedure may be provided to avoid the bad consequences of such a lockout. In this case, the device 100 may be restored to its previous state and the previous wake word access may be restored. This enables a minimum level of default functionality for users who did not issue a directive for the new wake word and may not even notice that the wake word they were using has been replaced with the new wake word. In some systems, this is achieved by a factory reset to the factory default settings or a previously saved functional configuration of the wake word set.
[0072] In an embodiment, both the wake word replacement policy, wake word addition policy, and / or a composite mixture of procedures for saving and restoring the wake word set configuration may be supported by using a single directive and / or modal dialog. In a device having a GUI such as a dedicated reset button in addition to voice, these other interfaces may be used instead of or in addition to the voice signal for wake word recovery. In a device without a GUI other than voice, a full reset of the device state is achievable via the voice interface (microphone and speaker) and / or by power cycling.
[0073] The following are further examples of modal dialogs for setting dynamic wake words, including restoring replaced wake words.
[0074] (U) In response to "Butler 3" which is a public name (D) Understood. Is "Butler 3" correct? (U) Yes (D) Shall I leave "Jarvis" which is my private name? (U) No (D) Shall I leave "Okay Agent" which is my public wake word? (U) Disable it As time passes (U) What is the wake word, Butler 3? (D) "Butler 3" is the active public wake word and "Okay Agent" is the disabled public wake word. As time passes (U) Restore the previous wake word, Butler 3 (D) Now "Okay Agent" is the public wake word. Again, other examples of modal dialogs are conceivable.
[0075] The above describes a procedure for setting a new wake word using voice directives via one or more microphones 106. In an alternative embodiment, device 100 may have other interfaces, such as a GUI or a physical (touch) interface, configured to fully or partially implement a wake word assignment directive. For example, the device may have a reserved button that is pressed to enter a new wake word and released at the end of the uttered voice segment. Further, by releasing the button shortly after pressing it, the device may be made to state its current wake word and corresponding status (e.g., public or private).
[0076] As noted above, step 226 (FIG. 3) includes constructing a wake word spotter for the dynamic wake word. As described in the background section, an effective wake word spotter may be designed for the wake words installed at factory shipment by collecting a large dataset related to pre-defined wake words and then using this large dataset to train an acoustic model such as a neural network. Such a dataset generally does not exist for dynamically selected wake words, and thus this technique is not feasible with dynamic wake words.
[0077] This technique overcomes this problem by using one or more methods for rapidly constructing a wake word spotter for any valid dynamic wake word. As noted above, the use of the term "rapidly" here means that a new wake word spotter can be constructed within seconds (e.g., 2 - 5 seconds) after the user has completed issuing a wake word assignment directive. In another example, the term "rapidly" applied to the time taken to construct the spotter may mean the time taken for a confirmation response to the WAD to complete. For example, a spotter is considered to be constructed "rapidly" if it is constructed by the time a response such as "From now on, I will respond to Jarvis" or "For now, Josephine is the public wake word" is provided.
[0078] In an embodiment, a wake word spotter for a new wake word is constructed by a spotter builder 168 of a WAD processor 164 (FIG. 2). As noted above, FIG. 2 shows the spotter builder 168 on the server 150, but the components of the spotter builder 168 may exist in and be implemented in the server 150, the device 100, or a combination of the server 150 and the device 100.
[0079] Wake Word Sampling Approach for Spotter Construction In one embodiment, the spotter builder 168 may construct a wake word spotter using what is referred to herein as a wake word sampling approach. In this approach, multiple sample utterances of a new wake word are collected from the user and then used to construct a new wake word spotter by locally training a classifier. Here, this approach will be described with reference to the flowchart of FIG. 5.
[0080] In steps 290 and 292, the spotter builder 168 uses a modal dialog to ask the user to provide additional audio samples of the new wake word. In some embodiments, one wake word may be requested at a time, and in other embodiments, it may be left open-ended so that the user can provide multiple samples. Receiving a sufficient number (e.g., four or more) of samples of the new wake word may be determined in step 292. In further embodiments, it may be less than four. In step 296, the first audio sample of the wake word is used together with the additional audio samples collected by steps 290 and 292 to build a classifier (such as a neural network (NN) classifier) that serves as a wake word spotter. The collected samples serve as positive examples of the wake word. To avoid false positives, it may be useful to add negative examples during training. Negative examples, if used, can be generated in multiple ways. In an embodiment, the audio samples of the wake word may be converted to phoneme sequences by an ASR front end, and then these sequences may be perturbed slightly to cause near misses. Using speech synthesis, negative audio samples may be obtained from the near-miss phoneme sequences. In a simpler variation of the previous embodiment, a single phoneme sequence from the speech recognition step 162 is used. In this variation, no further conversion step is required. This is particularly convenient when step 162 is executed on the server. Because the device 100 no longer needs to support the ASR function 162.
[0081] In summary, the wake word sampling approach proceeds in the following three steps: 1. Collect positive audio samples of the wake word 2. Optionally, generate some negative audio samples of the wake word 3. Train a classifier to be the desired wake word spotter.
[0082] These steps may be executed locally on device 100 or on a server. It is computationally feasible to locally implement a wake word sampling approach on a device using limited resources because the set of training samples is very small. Since the training dataset is small, its reliability is limited. Since the positive samples are from a single speaker, it is found that the wake word sampling approach is most reliable when combined with a private wake word used only by the specific user who created it. Therefore, device 100 may use the speaker verification engine 114 to verify that the voice of the current speaker matches that of the speaker who created the private wake word. Device 100 stores the user voiceprint 122 together with the speaker verification engine 114 to support this function.
[0083] Continuous transformation approach for spotter construction In another embodiment, the spotter builder 168 may implement a spotter using what is herein referred to as the continuous conversion approach. In this approach, the spotter algorithm relies on the embedded speech recognition engine 162 (or more precisely, the ASR front end) to generate a continuous speech transcription of the input speech. To achieve the low latency required by the spotter, the speech recognition front end may operate locally on the device 100. The flowchart of FIG. 6 illustrates such an embodiment. At step 310, the speech recognition front end maps the incoming speech stream to a phoneme stream, which is the continuous speech transcription of the input. At step 312, the wake word spotter attempts to match consecutive segments of the incoming phoneme stream with the phoneme sequence of the active wake word. This is done continuously and incrementally for each possible alignment of the phoneme sequence of the wake word against the phoneme stream. Once an alignment hypothesis is started, it remains active as long as the acoustic match is maintained. If there are several active wake words in the data structure of the wake words 120 stored in the memory 104, the incoming phoneme stream is thus compared in parallel with each of the stored wake words. The above steps may be applied incrementally each time a new phoneme appears in the incoming phoneme stream. The first match detected at step 316 (the alignment hypothesis is complete with a match) triggers the wake word spotter at step 318. If the acoustic match fails at any time during the alignment hypothesis, the flow returns to step 312 to search for a new phoneme sequence.
[0084] In an embodiment, for example, using a phoneme lattice data structure, multiple speech transcriptions of an incoming speech stream may be considered in parallel. Similarly, in an embodiment, for example, using a phoneme lattice data structure, multiple speech transcriptions of a wake word may be considered in parallel. Each of the multiple hypotheses may have an associated probability or score each time they are considered. In an embodiment, the wake word orthographic sequence is associated with a time component for each phoneme, and the orthographic sequence receives a score associated with the amount of temporal spread between the wake word phonemes and the incoming phonemes. In an embodiment, a low array score, or a low probability or score of an alternative hypothesis, may lead to discarding the array hypothesis.
[0085] This approach of providing a wake word spotter is advantageous in that it does not require training based on such a stored dataset, whether the stored dataset is a large remote dataset on a server or a small set of samples of dynamic wake words collected when needed. This approach is well-suited for the use of public wake words by various people and robustness to noise. This is because the speech recognition engine 162 that generates a phoneme stream from the incoming speech stream, or more precisely its ASR front end, is pre-trained for a wide range of speakers and conditions. Depending on the battery technology, this approach may be optimal for a device 100 plugged into a power outlet or a device 100 that can draw the battery power required to perform continuous speech conversion, although this may not be the case in further embodiments. Matching the continuous speech conversion input to the stored wake word phoneme sequence can be done quite efficiently.
[0086] A main variant of the approach just described is to use continuous text conversion instead of continuous voice conversion. Generally, the full speech recognition module 162 may include a large speech dictionary and a large language model, both of which consume a significant amount of memory. Currently, it is possible to completely ignore the language model (substantially reducing the space) and use a reduced speech dictionary, i.e., a default phoneme / text transducer that does not take exceptions into account.
[0087] The conversion approach does not include the training itself, such as the training of the NN, but "constructs" a spotter composed of a phoneme sequence matching algorithm and one or more target phoneme sequences.
[0088] Wake word segmentation approach for spotter construction In a further embodiment, the spotter builder 168 may construct a wake word spotter using what is referred to herein as a wake word segmentation approach. This approach starts with the speech transcription of the new wake word available from the speech recognition step 216 prior to step 226. Here, with reference to the flowchart of FIG. 7 and the diagram of FIG. 8, such an approach will be described.
[0089] After the analysis of the utterance received by the WAD parser 166, at step 320, the wake word speech transcription of the analyzed utterance may be tested to determine whether there is a wake word spotter that has already been trained and cached for the entire wake word. Such a spotter may be cached on the server 150, but alternatively may be cached on the device 100 or a third-party server.
[0090] For example, FIG. 8 shows the speech transcription of a wake word from a WAD. In this example, the wake word "Hey, Christopher" may be acoustically written as "HH EY1 K R IH1 S T AH0 F ER0" using CMU phonetic characters. CMUP is the standard phonetic character in English. Other phonetic characters (such as the International Phonetic Alphabet (IPA)) may be used to define a sequence of phonemes for a wake word spotter for English or other languages. In step 320, it is tested whether there is a cached spotter that has already been trained for the entire sequence of phonemes. If so, in step 322, the cached spotter is downloaded and used as a new wake word spotter for the new wake word "Hey, Christopher".
[0091] If there is no cached spotter for the entire new wake word, in step 326, the wake word may be split into a plurality of phonetic segments. Splitting the wake word portion into a sequence of phonetic segments may be done in any of a variety of ways, including segmenting the wake word portion into groups of words or syllables, individual syllables, or even finer divisions. As a simple example, FIG. 8 shows a root segmentation where the entire wake word is a single element, i.e., segmentation 1, and a second segmentation, i.e., segmentation 2, where the wake word is divided into phonetic segments from the separate words "Hey" and "Christopher".
[0092] In step 328, the spotter builder 168 checks whether there is a spotter that already exists and is cached for each of the phonetic segments in the current segmentation. If so, in step 348, these spotters for each of the phonetic segments are assembled in the order of the phonetic segments that are serially consecutive. These consecutive spotters are then downloaded and used as a new wake word spotter.
[0093] At any time, if the spotter builder 168 determines that a phonetic segment in a given segmentation does not have a corresponding cached spotter, then, in step 330, the engine 168 checks whether there can be a further split of the wake word into the phonetic segment. For example, FIG. 8 shows a further segmentation in which the wake word is further split into syllables, i.e., segmentation 3. There are syllabification algorithms that automatically segment a valid phonetic sequence into syllables. Further splitting is possible as will be described later. If such another segmentation is possible in step 330, a new split step is taken in step 334, resulting in a new segmentation, which is then retested in step 328 to check whether there is a cached spotter for each segment in the new instance.
[0094] Each time the splitting (segmentation) of the wake word phoneme sequence is completed, the wake word segmentation approach constructs a new spotter for any phonetic segment that does not already have a cached spotter in memory. To construct a new spotter for a phonetic segment, this technique relies on access to a labeled audio database in the memory 154 on the server 150. In particular, a set of audio samples corresponding to a particular phoneme sequence can be retrieved from a database of audio segments labeled by their audio transcriptions. In some embodiments, this search may be optimized by the use of a precomputed index such as a structure (a "trie") where nodes are associated with corresponding locations in the audio segment corpus.
[0095] Next, the spotter may be trained based on the searched matching segments for the positive examples of the wake word. Negative examples useful for training the yes / no classifier can be obtained in a plurality of ways. "Near matches" are useful for avoiding false positives, and false positives can be obtained by confusing the segment phonetic sequences (e.g., by using closely related phonetic sequences). Random audio samples may also be used, which will contribute to improving the output probability of the classifier.
[0096] Since it is most efficient to have as few spotters to build as possible, in step 336, the spotter builder 168 may select an instance that already has the most cached spotter for its phonetic segment. Next, in step 338, the spotter builder 168 searches for a subset of the data from the database in the memory 154 as described above to be used for training a spotter for a phonetic segment that does not have a cached spotter. In step 340, this subset of data is used to train a spotter for this phonetic segment. When the spotter for this phonetic segment is trained, in step 344, the spotter may be added to the cache.
[0097] In step 346, the spotter builder 168 may check whether there are any instances in which a further phonetic segment that does not yet have a cached spotter has been selected. If so, a new phonetic segment is selected and steps 338, 340, and 344 are repeated for this new phonetic segment. This process continues until in step 346 all phonetic segments have a cached spotter. At that point, in step 348, all the cached spotters for each of the phonetic segments are assembled in the order of the sequential phonetic segments in series. Then these sequential spotters are downloaded and used as a new wake word spotter.
[0098] In an embodiment, the training of the spotters for the phonetic segments in this approach (steps 338 and 340) may be performed on server 150 for computational reasons or due to the storage capacity required for the labeled audio database and / or the segment spotter cache. The step 348 of assembling spotters for various phonetic segments into a wake word spotter may be done on server 150 and downloaded to device 100, or may be executed on device 100 itself.
[0099] In an embodiment, the consecutive spotters assembled for each consecutive phonetic segment in step 348 may be regarded as a yes / no classifier that determines the probability of yes (matching that phonetic segment) and no (not matching) when an input stream is provided. The successful path of the spotter is the all-yes path, having the probability that all classifier steps exceed a threshold, and the spotter is successful when the overall probability of the path (the product of the probabilities or the sum of the log probabilities over the entire path) exceeds the threshold. The spotter is applied to the audio in many consecutive sequences, such as for each frame at a given frame rate.
[0100] In the above example, this method tests at various levels whether a spotter for the phonetic segments of a new wake word exists in the cache. The algorithm of FIG. 7 performs a progressive deepening in which the phonetic segments from the wake word are continuously split. Thereby, this method can utilize an existing large-scale wake word segment spotter. However, there are simpler algorithms where a specific level of speech segmentation is assumed. In a variant, it is assumed that in addition to the phonetic sequence, word segmentation of the wake word is available, and the wake word may be split into words. In another variant, the wake word phonetic sequence may be segmented into syllables. Thus, if necessary, a syllable-level spotter may be constructed and cached. However, it can be imagined that spotters are pre-computed for all possible syllables. In an embodiment, the wake word spotter may be pre-defined and stored on server 150 for a limited enumeration of all necessary phonetic segments (such as for all syllables). This is in principle achievable and should work well, making it possible to complete the task without the need to train any new segment spotters in step 340. However, there are a very large number of possible syllables in English or other languages.
[0101] A syllable is formed by three clusters, namely an onset, a nucleus, and a coda. The nucleus cluster is composed of vowels, and the onset and coda clusters are composed of consonants. For example, using the CMU phonetic script, the syllable "S T R EE T S" ("streets") has a three-consonant cluster "S T R" preceding a single vowel cluster "EE" and a two-consonant coda "T S".
[0102] One further variant is to perform sub - syllable segmentation as a further phonetic segment classification on which all wake - word spotters can be trained. For example, each vowel can be split into a word - initial part that includes the word - initial consonant cluster and at least the word - initial vowel, and a word - final part that includes the word - final vowel and the word - final consonant cluster. When the vowel cluster has a length of 1, the vowel is also the word - initial vowel and the word - final vowel. When the vowel cluster has a length of 2, it consists of the word - initial vowel and the word - final vowel. When the vowel cluster has a length of 3, some further rules may be utilized to determine how to split the consonant cluster.
[0103] In any of the above embodiments, once created, the wake - word spotter may be registered by the registration engine 170 even after the wake - word of the spotter has been changed, which includes storing or caching the wake - word and the wake - word spotter in memory. In that way, when the previous wake - word is reused, the cached wake - word spotter can be quickly retrieved.
[0104] In the above embodiments, even when the wake - word is a phrase containing multiple words, a single spotter may be used to spot the wake - word. In a further embodiment, a “multi - spotter” may be used to spot a wake - word containing multiple words. If N words are included in the wake - word, the activation module may depend on running N spotters (each of the N spotters having a binary output of match or no - match) in parallel for the multiple words in the wake - word. In some cases, a joint spotter (such as a classifier having N + 1 results (each possible match result and no - match result)) may be trained, or it may be a combination of these two.
[0105] FIG. 9 is a diagram showing an exemplary computing system 900 that may be a device 100 or a server used to implement an embodiment of the present technology. The computing system 900 of FIG. 9 includes one or more processors 910 and a main memory 920. The main memory 920 stores some instructions and data for execution by the processor unit 910. The main memory 920 can store executable code when the computing system 900 is operating. The computing system 900 of FIG. 9 may further include a mass storage device 930, a portable storage media drive 940, an output device 950, a user input device 960, a display system 970, and other peripheral devices 980.
[0106] The components shown in FIG. 9 are shown as being connected via a single bus 990. These components may be connected via one or more data transmission means. The processor unit 910 and the main memory 920 may be connected via a local microprocessor bus, and the mass storage device 930, the peripheral devices 980, the portable storage media drive 940, and the display system 970 may be connected via one or more input / output (I / O) buses.
[0107] The mass storage device 930, which may be implemented as a magnetic disk drive or an optical disk drive, is a non-volatile storage device for storing data and instructions for use by the processor unit 910. The mass storage device 930 can store system software for implementing embodiments of the present invention for the purpose of loading the system software into the main memory 920.
[0108] The portable memory media drive 940 operates with a portable non-volatile memory media such as a floppy (registered trademark) disk, a compact disk, or a digital video disk, and inputs and outputs data and code to and from the computing system 900 of FIG. 9. System software for implementing embodiments of the present invention may be stored on such a portable media and input to the computing system 900 via the portable memory media drive 940.
[0109] The input device 960 provides a part of the user interface. The input device 960 may include an alphanumeric keypad (such as a keyboard) for inputting alphanumeric and other information, or a pointing device (such as a mouse, a trackball, a stylus, or cursor direction keys). Also, the system 900 shown in FIG. 9 includes an output device 950. Suitable output devices include speakers, printers, network interfaces, and monitors. When the computing system 900 is part of a mechanical client device, the output device 950 may further include a servo controller for a motor within the mechanical device.
[0110] The display system 970 may include a liquid crystal display (LCD) or other suitable display device. The display system 970 receives text and graphics information, processes this information, and outputs it to the display device.
[0111] The peripheral device 980 may include any type of computer support device for adding further functionality to the computing system. The peripheral device 980 may include a modem or a router.
[0112] The components included in the computing system 900 of FIG. 9 are components typically found in a computing system that may be suitable for use in combination with embodiments of the present invention, and are intended to represent a wide variety of such computer components well-known in the art. Thus, the computing system 900 of FIG. 9 may be a personal computer, a handheld computing device, a telephone, a mobile computing device, a workstation, a server, a minicomputer, a mainframe computer, or any other computing device. The computer may also include various bus configurations, networked platforms, multiprocessor platforms, etc. A variety of operating systems can be used, including UNIX®, Linux®, Windows®, Macintosh OS, Palm OS, and other suitable operating systems.
[0113] Some of the above functions may be constituted by instructions stored in a storage medium (e.g., a computer-readable medium). These instructions may be retrieved and executed by a processor. Some examples of storage media are memory devices, tapes, disks, etc. When these instructions are executed by a processor, they operate to instruct the processor to operate in accordance with the present invention. Those skilled in the art are proficient in instructions, processors, and storage media.
[0114] It is worth noting that any hardware platform suitable for executing the processes described herein is also suitable for use in combination with the present invention. The terms "one computer-readable storage medium" and "plural computer-readable storage media" as used herein mean any one or more media involved in providing instructions to a CPU for execution. Such media can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Examples of non-volatile media include optical or magnetic disks (such as fixed disks). Examples of volatile media include dynamic memory (such as system RAM). Examples of transmission media include coaxial cables, copper wire, and fiber optics, including, in particular, wires including one embodiment of a bus. Transmission media can also take the form of acoustic or light waves (such as those generated during radio frequency (RF) and infrared (IR) data communications). Common forms of computer-readable media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROM disks, digital video disks (DVDs), any other optical media, any other physical media having patterns of marks or holes, RAM, PROM, EPROM, EEPROM, FLASHEPROM, any other memory chip or cartridge, carrier waves, or any other media readable by a computer.
[0115] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a CPU for execution. The bus carries data to system RAM, from which the CPU retrieves and executes instructions. Instructions received by system RAM may optionally be stored on a fixed disk before or after execution by the CPU.
[0116] In summary, the present technology relates to a method for changing a set of one or more wake words of a voice-responsive device, the method comprising receiving an oral request from a user and parsing the request into a natural language request and an uttered voice segment, the natural language request instructing the device to accept the uttered voice segment as a new wake word, and the method further comprising constructing a new wake word spotter to recognize the new wake word as an activation trigger.
[0117] In another example, the present technology relates to a method for changing a set of one or more wake words of a voice-responsive device, the method comprising receiving an oral request from a user and parsing the request into a natural language request and an uttered voice segment, the natural language request instructing the device to accept the uttered voice segment as a new wake word, and the method further comprising defining a new wake word spotter to recognize the new wake word as an activation trigger, the step of defining the new wake word spotter being performed by splitting the uttered voice segment into consecutive phonetic segments, comparing the parsed voice segments with a dataset of segments used to train an existing ASR algorithm to find a match between the parsed voice segments and the segments in the dataset of segments, and training the spotter using one or more matching phonetic segments from the existing ASR algorithm.
[0118] In a further example, the technology relates to a method of changing a set of one or more wake words of a voice-responsive device, the method comprising receiving an oral request from a user and parsing the request into a natural language request and an utterance voice segment, the natural language request instructing the device to accept the utterance voice segment as a new wake word, the method further comprising defining a new wake word spotter to recognize the new wake word as an activation trigger, the defining of the new wake word spotter being performed by obtaining further utterance voice samples of the utterance voice segment, converting the utterance voice segment and the further utterance voice samples into a phoneme sequence, and defining the wake word spotter based on one or more of the phoneme sequences of the phoneme sequence.
[0119] The foregoing description is illustrative and not restrictive. Many variations of the invention will be apparent to those skilled in the art upon a review of the disclosure. Accordingly, the scope of the invention should not be determined with reference to the foregoing description, but instead should be determined with reference to the appended claims, along with their full scope of equivalents. Although the invention has been described in connection with a series of embodiments, these descriptions are not intended to limit the scope of the invention to the specific forms described herein. It is further understood that the methods of the invention are not necessarily limited to the individual steps or order of steps described. On the contrary, this description is intended to include such alternatives, modifications and equivalents as may be within the spirit and scope of the invention as defined by the appended claims and otherwise understood by those skilled in the art.
[0120] One skilled in the art will recognize that an Internet service may be configured to provide Internet access to one or more computing devices coupled to the Internet service, and that the computing devices may include one or more processors, buses, memory devices, display devices, input / output devices, etc. Further, one skilled in the art will understand that the Internet service may be coupled to one or more databases, repositories, servers, etc. that may be utilized to implement any of the embodiments of the present invention described herein.
Claims
1. A method performed by a computer for changing a set of one or more wake words of a voice-responsive device, the method comprising: receiving an oral request from a user; and analyzing the oral request into a natural language request and an utterance voice segment, the natural language request instructing the device to accept the utterance voice segment as a new wake word, the method further comprising: constructing a new wake word spotter to recognize the new wake word as an activation trigger; if the oral request specifies that the new wake word is private, the natural language request instructs the device to activate when the new wake word is spoken by the user who created the new wake word and not to activate when spoken by others; if the oral request specifies that the new wake word is public, the natural language request instructs the device to activate when the new wake word is spoken by the user who created the new wake word and by others who did not create the new wake word.
2. The method according to claim 1, wherein the new wake word spotter is constructed in response to one receipt of the oral request.
3. The method according to claim 1 or claim 2, further comprising providing the user with an oral response to the natural language request, the oral response confirming the new wake word, and the new wake word spotter being constructed before the end of the oral response.
4. The method according to any one of claims 1 to 3, wherein the oral request is received at any time during an oral dialogue between the user and the device.
5. The method according to any one of claims 1 to 4, further comprising defining the new wake word spotter, the new wake word spotter being constructed locally on the device.
6. The method according to any one of claims 1 to 4, further comprising training the new wake word spotter, the training of the new wake word spotter being performed remotely on a server connected to the device by a network.
7. The method according to any one of claims 1 to 6, further comprising the step of adding the new wake word to a set of previous wake words that includes at least the previous wake word, wherein the new wake word spotter activates the device when receiving an utterance voice segment including the new wake word or a wake word from the set of previous wake words.
8. The method according to any one of claims 1 to 7, further comprising the step of replacing one or more previous wake words with the new wake word, so that after the replacement, the device activates when receiving an utterance voice segment that matches the new wake word and does not activate when receiving an utterance voice segment that matches the previous wake word.
9. The method according to claim 8, further comprising the step of later resetting the wake word set of the device to include the previous wake word, so that when the new wake word spotter receives an utterance voice segment including the previous wake word, the device is activated.
10. The method according to claim 8, further comprising the step of later resetting the wake word set of the device to the factory shipment setting.
11. The method according to any one of claims 1 to 9, wherein the new wake word spotter is trained based on the utterance voice segment, and the training is speaker-dependent, so that the device activates when the new wake word is spoken by the user who created the new wake word and does not activate when spoken by others.
12. The method according to claim 11, further comprising the step of generating a model of the user's voice for the purpose of speaker verification.
13. The method according to any one of claims 1 to 10, wherein the new wake word spotter is trained based on the utterance voice segment, and the training is speaker-independent, so that the device activates when the new wake word is spoken by the user who created the new wake word and by others who did not create the new wake word.
14. The method according to any one of claims 1 to 13, wherein the device requests further spoken voice samples of the new wake word from the user, and the new wake word spotter is trained using the spoken voice segment and the further spoken voice samples.
15. The method according to claim 14, wherein the training of the new wake word spotter is performed locally on the device.
16. A method executed by a computer for changing a set of one or more wake words of a voice-responsive device, comprising: receiving an oral request from a user; analyzing the oral request into a natural language request and a spoken voice segment, wherein the natural language request instructs the device to accept the spoken voice segment as a new wake word, and the method further comprises: constructing a new wake word spotter to recognize the new wake word as an activation trigger; The step of constructing the new wake word spotter comprises training the new wake word spotter using a cached spotter for at least a part of the phonetic segments of the new wake word.
17. A method executed by a computer for changing a set of one or more wake words of a voice-responsive device, comprising: receiving an oral request from a user; analyzing the oral request into a natural language request and a spoken voice segment, wherein the natural language request instructs the device to accept the spoken voice segment as a new wake word, and the method further comprises: constructing a new wake word spotter to recognize the new wake word as an activation trigger; The step of constructing the new wake word spotter comprises splitting the new wake word into phonetic segments and using a cached spotter for at least a part of the phonetic segments.
18. The step of constructing the novel wake word spotter comprises the step of splitting the novel wake word into phonetic segments and constructing a spotter for the phonetic segments, according to any one of claims 1 to 15.
19. The novel wake word spotter is formed from phonetic segments that are streamed from a remote server and compared with the phonetic segments of the wake words stored in the device, according to any one of claims 1 to 18.
20. A method executed by a computer for changing a set of one or more wake words of a voice-responsive device, comprising: receiving an oral utterance from a user; analyzing the oral utterance into a natural language request and an utterance voice segment, wherein the natural language request instructs the device to accept the utterance voice segment as a novel wake word, and the method further comprises: using automatic speech recognition to map the novel wake word to a novel wake word phonetic sequence; constructing a novel wake word spotter to recognize the novel wake word phonetic sequence as an activation trigger, and the step of constructing the novel wake word spotter is executed by: splitting the novel wake word phonetic sequence into a sequence of two or more consecutive partial phonetic segments; providing a corresponding partial wake word spotter for each partial phonetic segment; sequentially assembling the provided partial wake word spotters into the novel wake word spotter for the entire novel wake word phonetic sequence; when the oral utterance specifies that the novel wake word is private, the natural language request instructs the device to activate when the novel wake word is spoken by the user who created the novel wake word and not to activate when spoken by others. When the oral utterance specifies that the new wake word is public, the natural language request instructs the device to activate when the new wake word is spoken by the user who created the new wake word and by others who did not create the new wake word.
21. The method according to claim 20, wherein the step of splitting the new wake word phonetic sequence into phonetic segments comprises the step of splitting the new wake word phonetic sequence into words.
22. The method according to claim 20, wherein the step of splitting the new wake word phonetic sequence into phonetic segments comprises the step of splitting the new wake word phonetic sequence into separate syllables.
23. The method according to claim 20, wherein the step of splitting the new wake word phonetic sequence into phonetic segments comprises the step of splitting the new wake word phonetic sequence into phonetic segments smaller than syllables.
24. The step of providing a partial wake word spotter for partial phonetic segments comprises identifying a dataset of acoustically labeled audio segments, searching the dataset to collect audio segments whose audio labels match the partial phonetic segments, and training the partial wake word spotter based on the collected audio segments. The method according to any one of claims 20 to 23.
25. A method executed by a computer for changing a set of one or more wake words of an audio-responsive device, comprising receiving an oral utterance from a user, and analyzing the oral utterance into a natural language request and an utterance audio segment. The natural language request instructs the device to accept the utterance audio segment as a new wake word. The method further comprises using automatic speech recognition to map the new wake word to a new wake word phonetic sequence, and constructing a new wake word spotter to recognize the new wake word phonetic sequence as an activation trigger. The step of constructing the new wake word spotter comprises The step of splitting the novel wake word phonetic sequence into a sequence of two or more consecutive partial phonetic segments; For each partial phonetic segment, the step of providing a corresponding partial wake word spotter; It is executed by the steps of sequentially assembling the provided partial wake word spotters into the novel wake word spotter for the entire novel wake word phonetic sequence; The step of providing a partial wake word spotter for a partial phonetic segment; The step of identifying a set of cached wake word spotters indexed by wake words; A method comprising the step of searching for a cached wake word spotter for the partial phonetic segment.
26. A method executed by a computer for changing one or more sets of wake words of a voice-responsive device, comprising: The step of receiving an oral utterance from a user; The step of parsing the oral utterance into a natural language request and an uttered voice segment, wherein the natural language request instructs the device to accept the uttered voice segment as a novel wake word, and the method further comprises: The step of mapping the novel wake word to a novel wake word phonetic sequence using automatic speech recognition; The step of constructing a novel wake word spotter to recognize the novel wake word phonetic sequence as an activation trigger, and the step of constructing the novel wake word spotter: The step of splitting the novel wake word phonetic sequence into a sequence of two or more consecutive partial phonetic segments; For each partial phonetic segment, the step of providing a corresponding partial wake word spotter; It is executed by the steps of sequentially assembling the provided partial wake word spotters into the novel wake word spotter for the entire novel wake word phonetic sequence; The step of providing a corresponding partial wake word spotter for each partial phonetic segment; The step of searching a cache of wake word spotters for a partial wake word spotter for a consecutive phonetic segment of the novel wake word phonetic sequence; A method comprising the step of assembling the partial wake word spotter into the new wake word spotter.
27. The step of providing a corresponding partial wake word spotter for each of the partial phonetic segments comprises a step of checking memory for a cached wake word spotter for the phonetic segment, and if the phonetic segment does not have a cached wake word spotter in the memory, a step of constructing a wake word spotter for the phonetic segment, the method according to any one of claims 20 to 26.
28. The method further comprises the step of adding the new wake word to a set of previous wake words comprising at least the previous wake word, and the new wake word spotter activates the device when receiving an utterance voice segment comprising the new wake word or a wake word from the set of previous wake words, the method according to any one of claims 20 to 27.
29. A method executed by a computer for changing a set of one or more wake words of a voice-responsive device, comprising a step of receiving an oral request, and a step of parsing the oral request into a natural language request and an utterance voice segment, the natural language request instructs the device to accept the utterance voice segment as a new wake word, and the method further comprises a step of defining a new wake word spotter to recognize the new wake word as an activation trigger, the step of defining the new wake word spotter is executed by a step of requesting further utterance voice samples of the utterance voice segment, a step of converting the utterance voice segment and the further utterance voice samples into a phoneme sequence, and a step of defining the new wake word spotter based on one or more phoneme sequences of the phoneme sequence, and if the oral request specifies that the new wake word is private, the natural language request instructs the device to activate when the new wake word is spoken by the user who created the new wake word and not to activate when spoken by others. When the oral request specifies that the new wake word is public, the natural language request instructs the device to activate when the new wake word is spoken by the user who created the new wake word and by others who did not create the new wake word.
30. The method according to claim 29, wherein the new wake word is a private wake word for the person providing the oral request, whereby the device activates when receiving the new wake word from the person rather than from others.
31. The method according to claim 29 or claim 30, further comprising the step of receiving feedback from the user to verify the accuracy of the phoneme sequence when two or more of the phoneme sequences are different from each other.
32. The method according to any one of claims 29 to 31, further comprising the step of adding the new wake word to a set of previous wake words comprising at least the previous wake word, and the new wake word spotter activates the device when receiving an utterance voice segment comprising the new wake word or a wake word from the set of previous wake words.
33. The method according to any one of claims 29 to 32, further comprising the step of replacing the previous wake word with the new wake word so that the new wake word spotter activates the device when receiving an utterance voice segment comprising the new wake word and does not activate the device when receiving an utterance voice segment comprising the previous wake word.
34. The method according to any one of claims 29 to 33, further comprising the step of resetting the wake word set of the device to include the previous wake word so that the new wake word spotter activates the device when receiving an utterance voice segment comprising the previous wake word.
35. A program that, when executed by at least one processor of a computer, causes the computer to execute the method according to any one of claims 1 to 34.
Citation Information
Patent Citations
Method and device for waking up equipment
CN109887505A
Computer system for voice recognition of utterance input, and its method and computer program
JP2010072098A
Dictionary creation device, dictionary creation method, and dictionary creation program
JP2010097239A
A unified framework for device configuration, interaction and control, and related methods, devices and systems
JP2016502137A
Keyword model generation for detecting user-defined keywords
JP2017515147A