Attenuation of automatic speech recognition processing results
The method and system dynamically adjust ASR processing levels based on microphone duration and user interaction probability to enhance user experience and conserve resources on voice-enabled devices.
Patent Information
- Application Number
- JP2024015483
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-03
- Filing Date
- 2024-02-05
- Publication Date
- 2025-08-14
- Estimated Expiration
- 2041-11-16
AI Technical Summary
It is cost-prohibitive to continuously perform speech recognition on resource-constrained voice-enabled devices like smartphones or smartwatches, making it difficult for them to consistently respond to verbal requests, and requiring repeated invocation phrases or gestures becomes tedious for users.
A method and system that attenuate the level of automatic speech recognition (ASR) processing based on the duration of an open microphone window and user interaction probability, allowing the microphone to remain open for additional queries without requiring repeated invocation phrases, while conserving power and resources.
Improves user experience by allowing natural conversational interaction without repeated invocation, while optimizing power consumption and resource usage by adjusting ASR processing levels dynamically.
Smart Images

Figure 0007723771000001 
Figure 0007723771000002 
Figure 0007723771000003
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates to attenuating automatic speech recognition results. [Background technology]
[0002] Users frequently interact with voice-enabled devices such as smartphones, smartwatches, and smart speakers through digital assistant interfaces that enable users to complete tasks and get answers to questions they have, all through natural conversational interaction.
[0003] Ideally, when conversing with a digital assistant interface, a user should be able to communicate as if the user were speaking with another person by verbal requests directed to the user's voice-enabled device running the digital assistant interface. The digital assistant interface provides these verbal requests to an automatic speech recognizer to process and recognize the verbal requests so that an action can be taken. However, in practice, it is cost-prohibitive to continuously perform speech recognition on a resource-constrained voice-enabled device such as a smartphone or smartwatch, making it difficult for the device to always respond to these verbal requests. Summary of the Invention [Means for solving the problem]
[0004] One aspect of the present disclosure provides a method for damping automatic speech recognition processing. The method includes receiving, at data processing hardware of a voice-enabled device, an indication of a microphone trigger event indicating a possible user interaction with the voice-enabled device by speech, the voice-enabled device having a microphone, the microphone configured to capture speech for recognition by an automatic speech recognition (ASR) system when open. In response to receiving the indication of the microphone trigger event, the method further includes instructing, by the data processing hardware, the microphone to open or remain open for an open microphone duration window to capture an audio stream within an environment of the voice-enabled device, and providing, by the data processing hardware, the audio stream captured by the open microphone to the ASR system for performing ASR processing on the audio stream. While the ASR system is performing ASR processing on the audio stream captured by the open microphone, the method further includes attenuating, by the data processing hardware, a level of ASR processing that the ASR system performs on the audio stream based on a function of the open microphone duration window, and instructing, by the data processing hardware, the ASR system to use the attenuated level of ASR processing on the audio stream captured by the open microphone.
[0005] In some examples, while the ASR system is performing ASR processing on the audio stream captured by the open microphone, the method also includes determining, by the data processing hardware, whether voice activity is detected in the audio stream captured by the open microphone. In these examples, attenuating a level of ASR processing performed by the ASR system on the audio stream is further based on determining whether any voice activity is detected in the audio stream. In some implementations, the method further includes obtaining, by the data processing hardware, a current context when an indication of a microphone trigger event is received. In these implementations, instructing the ASR system to use the attenuated level of ASR processing includes instructing the ASR system to bias the speech recognition results based on the current context. In some configurations, after instructing the ASR system to use an attenuated level of ASR processing on the audio stream, the method further includes receiving, by the data processing hardware, an indication that the confidence of the speech recognition result for the speech query output by the ASR system fails to meet a confidence threshold, and instructing, by the data processing hardware, the ASR system to increase the level of ASR processing from the attenuated level and reprocess the speech query using the increased level of ASR processing. In some examples, while the ASR system is performing ASR processing on the audio stream captured by the open microphone, the method further includes determining, by the data processing hardware, when the attenuated level of ASR processing that the ASR performs on the audio stream is equal to zero based on a function of the duration of the open microphone, and instructing, by the data processing hardware, the microphone to close when the attenuated level of ASR processing is equal to zero.Optionally, the method may also include displaying, by the data processing hardware, a graphical indicator in a graphical user interface of the voice-enabled device that indicates the attenuated level of ASR processing performed by the ASR system on the audio stream.
[0006] Another aspect of the present disclosure provides a system for attenuating automatic speech recognition processing. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving, at a voice-enabled device, an indication of a microphone trigger event indicating a possible user interaction with the voice-enabled device through speech, the voice-enabled device having a microphone, the microphone configured, when open, to capture speech for recognition by an automatic speech recognition (ASR) system. In response to receiving the indication of the microphone trigger event, the operations further include instructing the microphone to open or remain open for an open microphone duration window to capture an audio stream within an environment of the voice-enabled device, and providing the audio stream captured by the open microphone to the ASR system for performing ASR processing on the audio stream. While the ASR system is performing ASR processing on the audio stream captured by the open microphone, the operations further include attenuating a level of ASR processing that the ASR system performs on the audio stream based on a function of the open microphone duration window, and instructing the ASR system to use the attenuated level of ASR processing on the audio stream captured by the open microphone.
[0007] This aspect may include one or more of any of the following features. In some examples, while the ASR system is performing ASR processing on the audio stream captured by the open microphone, the operations also include determining whether voice activity is detected in the audio stream captured by the open microphone. In these examples, the operations of attenuating a level of ASR processing that the ASR system performs on the audio stream are further based on determining whether any voice activity level is detected in the audio stream. In some implementations, the operations further include obtaining a current context when an indication of a microphone trigger event is received. In these implementations, the operations of instructing the ASR system to use the attenuated level of ASR processing include instructing the ASR system to bias the speech recognition results based on the current context. In some configurations, after instructing the ASR system to use the attenuated level of ASR processing on the audio stream, the operations further include receiving an indication that the confidence of the speech recognition result for the speech query output by the ASR system fails to meet a confidence threshold, and instructing the ASR system to increase the level of ASR processing from the attenuated level and reprocess the speech query using the increased level of ASR processing. In some implementations, while the ASR system is performing ASR processing on the audio stream captured by the open microphone, the operations further include determining when the attenuated level of ASR processing that the ASR performs on the audio stream equals zero based on a function of the duration of the open microphone, and instructing the microphone to close when the attenuated level of ASR processing equals zero. Optionally, the operations may also include displaying a graphical indicator in a graphical user interface of the voice-enabled device that indicates the attenuated level of ASR processing performed by the ASR system on the audio stream.
[0008] Implementations of the system or method may include one or more of any of the following features. In some implementations, the ASR system initially uses a first processing level to perform ASR processing on the audio stream at the start of an open microphone duration window, the first processing level being associated with a maximum processing capacity of the ASR system. In these implementations, attenuating the level of ASR processing that the ASR system performs on the audio stream based on a function of the open microphone duration window includes determining whether a first time interval has elapsed since the start of the open microphone duration window, and attenuating the level of ASR processing that the ASR system performs on the audio stream when the first time interval has elapsed by reducing the level of ASR processing from the first processing level to a second processing level, the second processing level being lower than the first processing level. In some examples, instructing the ASR system to use the attenuated level of ASR processing includes instructing the ASR system to switch from performing ASR processing on a remote server in communication with the voice-enabled device to performing ASR processing on data processing hardware of the voice-enabled device. In some configurations, instructing the ASR system to use an attenuated level of ASR processing includes instructing the ASR system to switch from using a first ASR model to using a second ASR model to perform ASR processing on the audio stream, where the second ASR model includes fewer parameters than the first ASR model. Instructing the ASR system to use an attenuated level of ASR processing may include instructing the ASR system to reduce the number of ASR processing steps performed on the audio stream. Instructing the ASR system to use an attenuated level of ASR processing may include instructing the ASR system to adjust beam search parameters to reduce a decoding search space of the ASR system.Additionally or alternatively, instructing the ASR system to use an attenuated level of ASR processing may include instructing the ASR system to perform quantization and / or sparsification on one or more parameters of the ASR system. In some configurations, instructing the ASR system to use an attenuated level of ASR processing may include instructing the ASR system to switch from system-on-chip (SOC-based) processing for performing ASR processing on the audio stream to digital signal processor-based (DSP-based) processing for performing ASR processing on the audio stream. While the ASR system is using the attenuated level of ASR processing on the audio stream captured by the open microphone, the ASR system is configured to generate speech recognition results for the audio data corresponding to a query spoken by a user and provide the speech recognition results to an application to perform an action specified by the query.
[0009] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0010] [Figure 1A] 1 is a schematic diagram of an exemplary speech environment for attenuating audio processing; [Figure 1B] FIG. 1B is a schematic diagram of an exemplary timeline of the speech environment of FIG. 1A. [Figure 2A] FIG. 1 is a schematic diagram of an exemplary interaction analyzer for attenuating voice processing. [Figure 2B] FIG. 1 is a schematic diagram of an exemplary interaction analyzer for attenuating voice processing. [Figure 3] 1 is a flow diagram of an exemplary sequence of operations for a method of attenuating audio processing. [Figure 4]FIG. 1 is a schematic diagram of an example computing device that may be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0011] Like reference symbols in the various drawings indicate like elements.
[0012] Ideally, when conversing with a digital assistant interface, a user should be able to communicate as if the user were speaking with another person by verbal requests directed to the user's voice-enabled device running the digital assistant interface. The digital assistant interface provides these verbal requests to an automatic speech recognizer to process and recognize the verbal requests so that an action can be taken. However, in practice, it is cost-prohibitive to continuously perform speech recognition on a resource-constrained voice-enabled device such as a smartphone or smartwatch, making it difficult for the device to always respond to these verbal requests.
[0013] A conversation with a digital assistant interface is typically initiated by a convenient user interaction, such as the user saying a set phrase (e.g., a hot word / keyword / wake word) or using some predefined gesture (e.g., lifting or squeezing the voice-enabled device). However, once a conversation begins, requiring the user to say the same set phrase or use the same predefined gesture for each successive verbal request / query can become tedious and inconvenient for the user. To alleviate this requirement, the microphone of the voice-enabled device may be left open for some predefined amount of time immediately after the interaction to allow the microphone to capture additional queries spoken by the user in a much more natural way. However, there are some trade-offs regarding how long the microphone should be left open immediately after the interaction. For example, leaving the microphone open for too long may unnecessarily consume power to perform voice recognition and increase the likelihood of capturing unintentional speech in the voice-enabled device's environment. On the other hand, closing the microphone too quickly creates a poor user experience, as the user is required to restart the conversation with a set phrase, gesture, or other means, which may be inconvenient.
[0014] Implementations herein are directed to initiating a speech recognizer to perform speech recognition in response to an event by opening a microphone of a speech-enabled device and gradually attenuating both the responsiveness and processing capabilities of the speech recognizer. More particularly, implementations herein include attenuating the level of speech recognition processing based on the probability of a user interaction with the speech-enabled device, such as an additional query after an initial query. As opposed to making a binary decision of either closing the microphone or preventing further processing by the speech recognizer at any given time, attenuating speech recognizer processing over time improves the user experience by keeping the microphone open longer to capture additional speech directed to the speech-enabled device, and improves power conservation by allowing the speech recognizer to run in different power modes depending on the confidence level of the upcoming user interaction.
[0015] 1A and 1B, in some implementations, a system 100 includes a user 10 providing a user interaction 12 to interact with a voice-enabled device 110 (also referred to as a device 110 or a user device 110), where the user interaction 12 is a verbal utterance 12, 12U, corresponding to a query or command to solicit a response from the device 110 or to cause the device 110 to perform a task specified by the query. In this sense, the user 10 may have a conversational interaction with the voice-enabled device 110 to perform a computing activity or find an answer to a question.
[0016] The device 110 is configured to capture user interactions 12, such as speech, from one or more users 10 in a speech environment. Utterances 12U spoken by the users 10 may be captured by the device 110 and may correspond to queries or commands for a digital assistant interface 120 executing on the device 110 to perform an action / task. The device 110 may correspond to any computing device associated with the user 10 and capable of receiving audio signals. Some examples of the user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, e-book readers, etc.), computers, wearable devices (e.g., smart watches), music players, casting devices, smart appliances (e.g., smart TVs) and Internet of Things (IoT) devices, remote controls, smart speakers, etc. The device 110 includes data processing hardware 112 and memory hardware 114 that communicates with the data processing hardware 112 and stores instructions that, when executed by the data processing hardware 112, cause the data processing hardware 112 to perform one or more operations related to voice processing.
[0017] Device 110 further includes an audio subsystem having an audio capture device (e.g., an array of one or more microphones) 116 for capturing audio data in the speech environment and converting it into an electrical signal. While device 110 implements audio capture device 116 (also generally referred to as microphone 116) in the illustrated example, audio capture device 116 may not be physically present on device 110 and may communicate with an audio subsystem (e.g., peripherals of device 110). For example, device 110 may correspond to a vehicle infotainment system that utilizes an array of microphones located throughout the vehicle.
[0018] A voice-enabled interface (e.g., a digital assistant interface) 120 may process queries or commands conveyed in verbal utterances 12U captured by device 110. The voice-enabled interface 120 (also referred to as interface 120 or assistant interface 120) generally receives audio data 124 corresponding to the utterances 12U and facilitates coordinating voice processing on the audio data 124 or other activity derived from the utterances 12U to generate responses 122. The interface 120 may execute on the data processing hardware 112 of the device 110. The interface 120 may stream the audio data 124 including the utterances 12U to various systems related to voice processing. For example, FIG. 1 shows that the interface 120 communicates with a voice recognition system 150. Here, the interface 120 receives the audio data 124 corresponding to the utterances 12U and provides the audio data 124 to the voice recognition system 150. In some configurations, the interface 120 serves as an open communication channel between the microphone 116 of the device 110 and the speech recognition system 150. In other words, the microphone 116 captures the speech 12U in the audio stream 16, and the interface 120 conveys audio data 124 corresponding to the speech 12U converted from the audio stream 16 to the speech recognition system 150 for processing. More specifically, the speech recognition system 150 may process the audio data 124 to generate a transcription 152 of the speech 12U and perform semantic interpretation on the transcription 152 to identify appropriate actions to perform. The interface 120 used to interact with the user 10 at the device 110 may be any type of program or application configured to perform the functions of the interface 120. For example, the interface 120 may be an application programming interface (API) that interfaces with other programs hosted on or communicating with the device 110.
[0019] 1A , a first utterance 12U, 12Ua by a user 10 states, “Hey computer, who is the president of France?” Here, the first utterance 12U includes a hot word 14 that, when detected in audio data 124, triggers the interface 120 to open the microphone 116 and relay subsequently captured audio data corresponding to the query “who is the president of France?” to the speech recognition system 150 for processing. That is, the device 110 may be in a sleep or hibernation state and execute a hot word detector to detect the presence of the hot word 14 in the audio stream 16. Acting as a wake-up phrase, the hot word 14, when detected by the hot word detector, triggers the device 110 to wake up and begin speech recognition for the hot word 14 and / or one or more words following the hot word 14. The hot word detector may be a neural network-based model configured to detect acoustic features indicative of hot words without performing speech recognition or semantic analysis.
[0020] In response to the hot word detector detecting the hot word 14 in the audio stream 16, the interface 120 relays audio data 124 corresponding to the utterance 12Ua to a speech recognition system 150, which performs speech recognition on the audio data 124 to generate a speech recognition result (e.g., a transcription) 152 of the utterance 12Ua. The speech recognition system 150 and / or the interface 120 perform semantic interpretation on the speech recognition result 152 to determine that the utterance 12Ua corresponds to a search query regarding the identity of the president of France. Here, the interface 120 may send the transcription 152 to a search engine 160, which searches for and returns search results 162 of "Emmanuel Macron" for the query "Who is the president of France?" The interface 120 receives this search result 162 of "Emmanuel Macron" from the search engine 160 and further communicates "Emmanuel Macron" to the user 10 as a response 122 to the query of the first utterance 12Ua. In some examples, the response 122 includes synthesized speech that is audibly output from the device 110.
[0021] To perform the functions of assistant interface 120, interface 120 may be configured to control one or more peripherals (e.g., one or more components of an audio subsystem) of device 110. In some examples, interface 120 controls microphone 116 to direct when microphone 116 is open, i.e., actively receiving audio data 124 for purposes of some audio processing, or closed, i.e., not receiving audio data 124 or receiving a limited amount of audio data 124 for purposes of audio processing. Here, whether the microphone 116 is “open” or “closed” may refer to whether the interface 120 communicates audio data 124 received at the microphone 116 to the speech recognition system 150 such that the interface 120 has an open channel of communication with the speech recognition system 150 to enable speech recognition of any utterances 12U included in the audio stream 16 received at the microphone 116, or a closed channel of communication with the speech recognition system 150 to prevent the speech recognition system 150 from performing speech recognition on the audio stream 16. In some implementations, the interface 120 specifies or indicates whether such a channel is open or closed based on whether the interface 120 receives or has received an interaction 12 with an interaction 12 and / or trigger 14. For example, when interface 120 receives an interaction 12 with a trigger 14 (e.g., an utterance 12U with a hotword), interface 120 instructs microphone 116 to open and relays audio data 124 converted from the audio captured by microphone 116 to speech recognition system 150. After the utterance is completed, interface 120 may instruct microphone 116 to close to prevent speech recognition system 150 from processing additional audio data 124.
[0022] In some implementations, device 110 communicates with remote system 140 over network 130. Remote system 140 may include remote resources 142, such as remote data processing hardware 144 (e.g., a remote server or CPU) and / or remote memory hardware 146 (e.g., a remote database or other storage hardware). Device 110 may utilize remote resources 142 to perform various functions related to speech processing. For example, search engine 160 may reside on remote system 140, and / or some of the functionality of speech recognition system 150 may reside on remote system 140. In one example, speech recognition system 150 may reside on device 110 to perform on-device automatic speech recognition (ASR). In another example, speech recognition system 150 resides on a remote system to provide server-side ASR. In yet another example, functionality of speech recognition system 150 is split between device 110 and server 140. For example, FIG. 1A shows speech recognition system 150 and search engine 160 in dotted boxes to indicate that these components may reside on device 110 or on the server side (i.e., on remote system 140).
[0023] In some configurations, different types of speech recognition models reside in different locations (e.g., on-device or remote) depending on the model. Similarly, end-to-end or streaming-based speech recognition models may reside on device 110 due to their space-efficient size, while larger, more traditional speech recognition models built from multiple models (e.g., acoustic models (AM), pronunciation models (PM), and language models (LM)) are server-based models that reside on remote system 140 rather than on-device. In other words, depending on the desired level of speech recognition and / or the desired speed at which to perform the speech recognition, speech recognition may reside on-device (i.e., user-side) or remotely (i.e., server-side).
[0024] When a user 10 conversationally engages with an interface 120, it can become quite inconvenient for the user 10 to repeatedly say the same invocation phrase (e.g., hot word) 14 for each interaction 12 in which the user 10 desires some feedback (e.g., a response 122) from the interface 120. In other words, requiring the user 10 to say the same hot word for each successive query by the user 10 can become very tedious and inconvenient. Further unfortunately, having the device 110 continuously perform speech recognition on all audio data 124 received at the microphone 116 can also be a waste of computing resources and can be computationally expensive.
[0025] To address the inconvenience of requiring user 10 to repeatedly say hotword 14 each time user 10 desires to communicate a new query 14 while user 10 is actively engaged in a conversation (or interaction 12 session) with interface 120, interface 120 may allow user 10 to provide an additional query after interface 120 has responded to a previous query without requiring user 10 to say hotword 14. That is, a response 122 to a query may serve as an indication of a microphone trigger event 202 ( FIG. 1B ) indicating a possible user interaction 12 with voice-enabled device 110 through speech, thereby causing device 110 to instruct microphone 116 to open or remain open to capture audio stream 16 and provide the captured audio stream 16 to speech recognition system 150 for performing speech recognition processing on audio stream 16. Here, the microphone 116 may be open to accept queries from the user 10 without requiring the user 10 to first say the hotword 14 again to trigger the microphone 116 to open to accept queries.
[0026] A microphone trigger event 202 (also referred to as a trigger event 202) generally refers to the occurrence of an event that indicates the possibility that user 10 may interact with device 110 through speech and therefore requires activation of microphone 116 to capture all speech for voice processing. Here, because trigger event 202 indicates a possible user interaction 12, trigger event 202 may range from a gesture to a recognized user characteristic (e.g., a pattern of behavior) to any action by user 10 that device 110 may discern as a potential interaction 12. For example, user 10 may have a weekday routine of querying device 110 about the weather and whether there are any events on the user's calendar when entering the kitchen. Due to this pattern of behavior, device 110 may recognize that user 10 enters the kitchen around a certain time (e.g., by listening for movement in the direction of the kitchen entrance) and treat the action of user 10 entering the kitchen in the morning as a trigger event 202 for a possible user interaction 12. In the case of gestures, the trigger event 202 may be an interaction 12 in which the user 10 lifts the device 110, squeezes the device 110, presses a button on the device 110, taps the screen of the device 110, moves their hand in a predefined manner, or any other type of pre-programmed gesture to indicate that the user 10 may intend to engage in a conversation with the assistant interface 120. Another example of a trigger event 202 is when the device 110 (e.g., the interface 120) communicates a response 122 to a query from the user 10. In other words, when the interface 120 relays the response 122 to the user 10, the response 122 acts as a communication interaction for the user 10 on behalf of the device 110; that is, additional queries spoken by the user 10 are likely to follow after receiving the response 122, or may occur simply because the user 10 is currently conversing with the device 110.Based on this possibility, the response 122 may be considered a trigger event 202 such that after the device 110 outputs the response 122, the microphone 116 either opens or remains open to capture and enable audio processing of any additional queries spoken by the user 10 without requiring the user 10 to prefix the additional queries with the hotword 14.
[0027] To implement the microphone trigger event 202 and provide a conversation between the device 110 and the user 10 without sacrificing processing resources, the device 110 deploys an interaction analyzer 200 (also referred to as analyzer 200) that recognizes when the user 10 is speaking to the assistant interface 120 while also maintaining voice recognition throughout the conversation with the user 10. In other words, the analyzer 200 may provide endpointing functionality by identifying when the conversation between the user 10 and the interface 120 begins and when it is best to consider the conversation ended in order to determine the endpoints of the interaction 12 and deactivate voice recognition. Furthermore, in addition to determining the endpoints, the analyzer 200 may modify the voice recognition process by lowering the voice recognition processing level 222 (FIGS. 2A and 2B) (i.e., using the voice recognition system 150) depending on the nature of one or more interactions 12 by the user 10. In particular, the analyzer 200 may be able to instruct the speech recognition system 150 to use an attenuated level 222 of speech recognition for the audio stream 16 captured by the microphone 116. For example, when the device 110 instructs the microphone 116 to open or remain open for an open microphone duration window to capture and provide the audio stream 16 to the speech recognition system 150 in response to receiving an indication of the microphone trigger event 202, the analyzer 200 may attenuate the level 222 of speech processing that the speech recognition system 150 performs on the audio stream 16 as a function of the open microphone duration window 212. Here, the analyzer 200 operates in conjunction with the interface 120 to control speech recognition and / or speech-related processing.
[0028] 1B , when device 110 (e.g., interface 120) receives an indication of a trigger event 202, interface 120 instructs microphone 116 to open or remain open and transmits audio data 124 associated with audio stream 16 captured by microphone 116 to speech recognition system 150 for processing. Here, when interface 120 instructs microphone 116 to open, analyzer 200 may be configured to specify that microphone 116 remain open for some open microphone duration window 212 after initiation by trigger event 202. For example, open microphone duration window 212 specifies a period of time during which interface 120 transmits audio stream 16 of audio data 124 captured by microphone 116 to speech recognition system 150. Here, microphone duration window 212 refers to a set duration that begins upon receipt of trigger event 202 and ends after the set duration (e.g., following a microphone close event). In some examples, the analyzer 200 may instruct the interface 120 to extend the microphone duration window 212 or refresh (i.e., refresh) the microphone duration window 212 when the interface 120 receives another trigger event 202 (e.g., a subsequent trigger event 202) during the microphone duration window 212. By way of example, FIGS. 1A and 1B show a user 10 generating a first utterance 12Ua stating, "Hey computer, who is the president of France?" and a second utterance 12U, 12Ub, "how old is he?" as a follow-up question to a response 122 that the president of France is Emmanuel Macron.1B , the hotword 14 “hey computer” corresponds to a phrase that initiates speech recognition of the subsequent portion of the utterance, “who is the president of France?” When the interface 120 responds “Emmanuel Macron,” the analyzer 200 establishes the response 122 as a microphone trigger event 202 that initiates a first microphone duration window 212, 212a. Here, the window 212 may have a duration defined by a start point 214 and an end point 216, at which point the microphone duration window 212 expires and the interface 120 and / or the analyzer 200 execute a microphone close event. Furthermore, in this example, the user 10 asks a follow-up question, “how old is he?” prior to the originally specified end point 216a, 216b of the first microphone duration window 212a. The follow-up question (i.e., another user interaction 12) of the second utterance 12Ub causes the interface 120 to generate a second response 122, 122b stating that Emmanuel Macron's age is 47. Like the first response 122, 122a, "Emmanuel Macron," the second response 122, 122b is a second trigger event 202, 202b that starts a new microphone duration window 212, 212b as the second microphone duration window 212b, even though the first microphone duration window 212a may or may not have expired (e.g., FIG. 1B shows that the first window 212a expired before the second response 122b). When the interface 120 starts a new microphone duration window 212 while the microphone duration window 212 has not yet ended, this may be considered as extending the unended window 212 or keeping the unended window 212 open.In some implementations, when a duration window 212 has not yet ended, a trigger event 202 during that window 212 may extend the unended duration window 212 by some specified amount of time, rather than starting an entirely new duration window 212 (i.e., starting the duration window anew). For example, if the duration window 212 is 10 seconds and the duration window 212 is not currently ending, a trigger event 202 during this unended duration window 212 will extend the unended window 212 by an additional 5 seconds, rather than a new additional full 10 seconds. In FIG. 1B , the first window 212a ends before the second response 122b, such that the second trigger event 202b starts an entirely new duration window 212, 212b at a second starting point 214, 214b, and this duration window 212, 212b ends at a second ending point 216, 216b (unless, for example, one or more additional trigger events 202 occur).
[0029] 1B , analyzer 200 is further configured to generate attenuated state 204 during a portion of window 212. For example, by transitioning to attenuated state 204, speech recognition processing level 222 may be altered (e.g., lowered by a certain amount) during window 212. In other words, the level of speech processing 222 performed by speech recognition system 150 on audio data 124 may be based on a function of window 212. By basing the level of speech processing (e.g., for speech recognition) 222 on window 212, the amount of speech processing performed on audio stream 16 of audio data 124 is a function of time. That is, as time passes and it becomes less likely that a conversation between interface 120 and user 10 is still ongoing (e.g., audio data 124 does not include user speech or speech is not directed at device 110), analyzer 200 recognizes this potential decrease and lowers the speech recognition accordingly. 1B shows that during the last half of the first microphone duration window 212a, the audio processing transitions to an attenuated state 204, where the level of audio processing 222 is reduced. This approach may allow the analyzer 200 and / or interface 120 to reduce the amount of computing resources being used to perform audio processing at the device 110. Similarly, in this example of FIG. 1B, when the analyzer 200 transitions the speech recognition to the attenuation state 204, even if the user 10 generates an additional question, "how old is he?" during the last third of the first window 212a, the analyzer 200 determines that the attenuation state 204 in the second window 212b should occur more quickly because the speech recognition was already in the attenuation state 204 in the first window 212 and the likelihood that the user 10 would perform the second interaction 12Ub was already somewhat low; therefore, the analyzer 200 determines that the attenuation state 204 should occur for a longer period in the second duration window 212b, for example, for approximately 90% of the second duration window 212b.By utilizing analyzer 200, device 110 and / or interface 120 strives to strike a balance between maintaining the ability to perform speech recognition without hotwords 14 and trying to avoid actively performing speech recognition when user 10 is unlikely to intend to further interact with assistant interface 120.
[0030] 2A and 2B , the analyzer 200 generally includes a window generator 210 and a reducer 220. The window generator 210 (also referred to as the generator 210) is configured to generate open microphone duration windows 212 that determine the duration for which the microphone 116 is open to capture and provide the audio stream 16 to the speech recognition system 150 for processing. Each microphone duration window 212 includes a start point 214 that specifies the time at which the microphone duration window 212 begins when the audio stream 16 received at the microphone 116 begins to be transmitted to the speech recognition system 150 for processing, and an end point 216 that specifies the time at which the open microphone duration window 212 ends, after which the audio stream 16 is no longer transmitted to the speech recognition system 150 for processing. The window generator 210 initially generates the open microphone duration window 212 when it is determined that the interface 120 has received a trigger event 202 from the user 10.
[0031] In some implementations, the window generator 210 is configured to intelligently generate windows 212 of different sizes (i.e., windows 212 having different lengths of time) based on the configuration of the analyzer 200 or based on conditions occurring at the microphone 116. For example, an administrator of the analyzer 200 sets a default duration for the windows 212 generated by the window generator 210 (e.g., each window 212 is 10 seconds long). In contrast, the generator 210 may recognize patterns of behavior by the user 10 or aspects / characteristics of the utterances 12U by the user 10 and intelligently generate windows 212 having sizes that match or correspond to these recognized characteristics. For example, it may be common for a particular user 10 to engage in multiple questions each time they generate a trigger 14 to begin a conversation with the interface 120. Here, a voice processing system associated with device 110 may determine the identity of user 10, and based on this identity, generator 210 may generate window 212 having a size that corresponds to or is responsive to the frequency of interactions 12 of that user 10 during an interaction session (i.e., a conversation with interface 120). For example, instead of generating a default window 212 having a duration of 5 seconds, generator 210 may generate a window 212 having a duration of 10 seconds because user 10 who sent hotword 14 tends to generate a higher frequency of interactions 12 with interface 120. On the other hand, if user 10 who sent hotword 14 tends to ask only a single query whenever interacting with interface 120, generator 210 may receive this behavioral information about user 10 and shorten the default window 212 from 5 seconds to a custom window 212 of 3 seconds.
[0032] Generator 210 may generate custom open microphone duration window 212 based on aspects or characteristics of utterance 12U or response 122 to utterance 12U. As a basic example, user 10 utters utterance 12U: "hey computer, play the new Smashing Pumpkins album, Cyr." However, when interface 120 attempts to respond to this command, interface 120 determines that the new Smashing Pumpkins album, Cyr, is scheduled to be released later that month and provides response 122 indicating that the album is not currently available. Response 122 may further ask whether user 10 would like to listen to something else. Here, following response 122 by interface 120, generator 210 may generate a larger window 212 (i.e., a window 212 that lasts for a longer duration) or may automatically extend the initially generated window 212 based on the fact that interface 120 and / or device 110 determine that there is a high probability of additional interaction 12 by user 10 when interface 120 generates response 122 that the requested album is not available. In other words, interface 120 and / or device 110 intelligently recognize that due to the results of the command by user 10, user 10 is likely to send another music request following response 122.
[0033] 1B shows, generator 210 may not only generate window 212 when trigger event 202 is first received (e.g., when interface 120 communicates response 122), but may also generate a new window 212 or extend a currently open window 212 when a trigger event 202 is received by interface 120 during an open window 212. By generating a new window 212 when a trigger event 202 is received during an open window 212 or extending an open window 212, generator 210 functions to keep window 212 open as the conversation between user 10 and interface 120 continues.
[0034] 2B , when the generator 210 extends a window 212 or generates a new window 212 in an ongoing conversation, the generator 210 is configured to discount the duration of the window 212. For example, the generator 210 is configured to generate subsequent windows 212 of shorter duration after each subsequent interaction 12 by the user 10. This approach may take into account the fact that after a first interaction 12, there is a first probability that a second or additional interaction 12 will occur, but thereafter the probability of further interactions 12 occurring decreases. For example, after an initial inquiry in a first utterance 12Ua, an additional inquiry in a second utterance 12Ub is likely to occur with approximately a 20% probability during an interaction session between the user 10 and the interface 120, but a third utterance 12U after the additional inquiry in the second utterance 12Ub will occur with only approximately a 5% probability. Based on this pattern, the generator 210 may reduce the size of the window 212 or the length by which the open window 212 is extended as a function of this probability that another trigger event 202 (i.e., a possible interaction 12) will occur. For example, FIG. 2B shows three trigger events 202, 202a-c, each corresponding to subsequent responses 122 (e.g., three responses 122a-c) from the interface 120, such that the generator 210 generates a first window 212a for the first trigger event 202a, followed by a second window 212b for the second trigger event 202b that occurs after the first trigger event 202a, and then a third window 212c for the third trigger event 202c that occurs after the second trigger event 202b. In this example, the generator 210 shortens the generated windows 212 so that the third window 212c has a shorter duration than the second window 212b, which in turn has a shorter duration than the first window 212a.Additionally or alternatively, the generator 210 may analyze the audio data 124 of the audio stream 16 to determine whether the level of voice activity in the audio data 124 provides any indication that a trigger event 202 is likely to occur. With this information, the generator 210 may modify the size of the currently open window 212 or modify the size of any subsequently created / extended windows 212.
[0035] The reducer 220 is configured to specify a processing level 222 of speech recognition to be performed on the audio stream 16 of the audio data 124. The processing level 222 may be based on several variables, including, but not limited to, the type of speech recognition model performing the speech recognition for the speech recognition system 150, the location where the speech recognition is performed, the speech recognition parameters used to perform the speech recognition, whether the speech model is specified to operate at full capacity or at some lower level of capacity, etc. The processing level 222 may generally refer to the amount of computing resources (e.g., local resources such as data processing hardware and memory hardware, or remote resources) dedicated to or consumed by speech processing such as speech recognition at any given time. This means, for example, that when the amount of computing resources or computing power dedicated to speech processing in a first state is less than in a second state, a first processing level 222 in a first state is lower than a second processing level 222 in a second state.
[0036] In some examples, when speech recognition is performed at a reduced processing level 222 (e.g., when compared to maximum processing capacity), the reduced processing level 222 may function as a first pass to recognize speech, but may then result in a second pass having a higher processing level 222 than the first pass. For example, when speech recognition system 150 is operating at a reduced processing level 222, speech recognition system 150 identifies a low-confidence speech recognition result (e.g., a low-confidence hypothesis) that audio data 124 includes utterance 12U to command or query interface 120. This low-confidence speech recognition result may cause reducer 220 to increase processing level 222 so that it may be determined at the higher processing level 222 whether this low-confidence speech recognition result actually corresponds to a higher-confidence speech recognition result (e.g., a high-confidence hypothesis) that audio data 124 includes utterance 12U to command or query interface 120. In some examples, a low-confidence speech recognition result is a speech recognition result that fails to meet a confidence threshold during speech recognition. In other words, the reducer 220 may change the processing level 222 based not only on the variables described above, but also on the results obtained during speech recognition at a particular processing level 222.
[0037] In some implementations, the reducer 220 specifies the processing level 222 as a function of the window 212. That is, the reducer 220 can lower the processing level 222 for speech recognition when an open microphone duration window 212 exists. For example, when the window 212 corresponds to a duration of 10 seconds (e.g., as shown in FIG. 2A ), for the first 5 seconds of the window 212, the reducer 220 instructs the speech recognition system 150 to operate at maximum power / processing (e.g., a first processing level 222, 222a) for speech recognition. After these 5 seconds at maximum processing occur, the speech processing transitions to the attenuation state 204, and the processing level 222 is lowered to a level below maximum processing. As an example, from 5 seconds to 7 seconds into the duration of the window 212, the reducer 220 instructs the speech recognition system 150 to operate at a second processing level 222, 222b, which corresponds to 50% of the maximum processing of the speech recognition system 150. Then, after the seventh second of the duration, until the end of the duration of the window 212, the reducer 220 instructs the speech recognition system 150 to operate at a third processing level 222, 222c, which corresponds to 25% of the maximum processing of the speech recognition system 150. Thus, the reducer 220 controls the processing level 222 of the speech recognition system 150 during the open microphone duration window 212.
[0038] In some implementations, the generator 210 and the reducer 220 work together such that the duration of the window 212 depends on the attenuation state 204. The generator 210 may not generate the end point 216 at a specified time, but instead allow the reducer 220 to reduce the processing level 222 to a level that has the same effect as closing the microphone 116. For example, after a third processing level 222c, which corresponds to 25% of the maximum processing of the speech recognition system 150, a further 25% reduction in processing would reduce the processing level 222 to 0% of the maximum processing of the speech recognition system 150 (i.e., no processing or “closed”), causing the reducer 220 to close the microphone 116. While this particular example is a stepwise approach that discretely and gradually reduces the processing level 222, other types of attenuation to result in a closed microphone 116 are also possible. For example, the processing level 222 may decay linearly at any point during the open microphone duration window 212. By allowing the reducer 220 to lower the processing level 222 until the microphone 116 is closed, this technique may promote a continuous attenuation of audio processing (e.g., when no interaction 12 is occurring and the microphone 116 is open).
[0039] 2B , the processing level 222 within the window 212 is based on the time when the interface 120 received the last trigger event 202 by the user 10. In these configurations, the reducer 220 may determine whether a time period 224 from when the interface 120 received the last trigger event 202 to the current time satisfies a time threshold 226. When the time period 224 meets the time threshold 226 (i.e., no interaction 12 has occurred for the threshold time), the reducer 220 may generate a processing level 222 at the current time. For example, if the interface 120 sends a response 122 to an utterance 12U by the user 10, the window 212 begins when the interface 120 communicates the response 122. When the time threshold 226 is set to 5 seconds, the reducer 220 determines whether 5 seconds have passed since the interface 120 communicated the response 122. When the reducer 220 determines that five seconds have elapsed, it may, for example, specify some processing level 222 of speech recognition to begin at the current time that the reducer 220 determines that five seconds have elapsed or at some particular time thereafter.
[0040] In some configurations, the processing level 222 is changed by adjusting one or more parameters of the speech recognition system 150. In one approach to changing the processing level 222 of speech recognition in the speech recognition system 150, the speech recognition is changed depending on where it occurs. The location of the speech recognition may change from occurring server-side (i.e., remote) to occurring on-device (i.e., local). In other words, the first processing level 222a corresponds to remote speech recognition using a server-based speech recognition model, while the second processing level 222b corresponds to local speech recognition occurring on-device. When the speech recognition system 150 is hosted “on-device,” the device 110 receives the audio data 124 and uses its processor (e.g., data processing hardware 112 and memory hardware 114) to perform the functions of the speech recognition system 150. Because a server-based model may utilize a greater number of remote processing resources (e.g., along with other costs such as bandwidth and transmission overhead), speech recognition using a server-based model may be considered to have a higher processing level 222 than an on-device speech recognition model. Due to the greater amount of processing, server-based models may potentially be larger in size than on-device models and / or may perform decoding using a larger search graph than on-device models. For example, a server-based speech recognition model may utilize multiple larger models (e.g., an acoustic model (AM), a pronunciation model (PM), and a language model (LM)) trained specifically for dedicated speech recognition purposes, while an on-device model must often integrate these different models into a smaller package to operate effectively and space-efficiently with the finite processing resources of device 110. Thus, when reducer 220 lowers processing level 222 to attenuation state 204, reducer 220 may modify speech recognition that occurs remotely using a server-based model to occur locally using an on-device model.
[0041] In some situations, there may be two or more devices 110 near a user 110 generating an interaction 12, such as a verbal utterance 12U. When multiple devices 110 are in the user 10's vicinity, each device 110 may be capable of performing some aspect of speech recognition. Because multiple devices 110 performing speech recognition on the same verbal utterance 12U may overlap, the reducer 220 may know that another device 110 is processing or configured to process speech recognition and may lower the processing level 222 in that device 110 by closing the microphone 116 of that device. As an example, when a user 10 has a mobile device and a smartwatch, both devices 110 may be capable of performing speech recognition. Here, the reducer 220 may close the microphone 116 for speech recognition on the mobile device to conserve the mobile device's processing resources for a wide range of other computing tasks that the mobile device may need to perform. For example, the user 10 may want to conserve the battery of their mobile device over the battery of their smartwatch. In some examples, when multiple devices are present, the reducer 220 may attempt to determine the characteristics of each device 110 (e.g., current processor drain, current battery life, etc.) to identify which devices 110 are best suited to have their microphones 116 remain open and which devices are best suited to have their microphones 116 closed.
[0042] In addition to differences in processing levels between the remote speech recognition system 150 and the on-device speech recognition system 150, there may be different versions of the on-device or server-side speech recognition model. With different versions, the reducer 220 may change the processing level 222 by changing the model or model version being used for speech recognition. Broadly speaking, a model may have a large version with a high processing level 222, a medium version with a medium processing level 222, and a small version with a low processing level 222. In this sense, if the reducer 220 wants to lower the processing level 222, the reducer 220 may transition the speech recognition being performed on a first version of the model to a second version of the model that has a lower processing level 222 than the first version of the model. In addition to changing between versions of either the on-device or server-side model, the reducer 220 may also change from one version of the server-side model to a particular version of the on-device model. By having models and versions of these models, the reducer 220 has a greater number of processing levels 222 at its disposal to attenuate the speech recognition processing levels 222 during an open microphone window 212 .
[0043] In some implementations, versions of on-device speech recognition models have different processing demands, such that the reducer 220 may designate such versions for different processing levels 222. Some examples of on-device speech recognition models include sequence-to-sequence models, such as recurrent neural network transducer (RNN-T) models, listen-attend-spell (LAS) models, neural transducer models, monotonic alignment models, and recurrent neural alignment (RNA) models. On-device models may also exist that are hybrids of these models, such as a two-pass model that combines an RNN-T model and an LAS model. Given these different versions of the on-device model, the reducer 220 may rank or identify the processing requirements of each of these versions to generate different processing levels 222. For example, a two-pass model includes a first pass of an RNN-T network followed by a second pass of an LAS network. Because the two-pass model includes multiple networks, the reducer 220 may designate the two-pass model as a large-scale, on-device model with a relatively high processing level 222 for speech recognition. To lower the processing level 222 for speech recognition from that of the two-pass model, the reducer 220 may change from the two-pass model to an LAS model, where the LAS model is an attention-based model that performs attention during its decoding process to generate the strings that form the transcription 152. In general, attention-based approaches tend to be more computationally intensive because they focus on specific features for a given speech input.For comparison, the RNN-T model does not employ an attention mechanism and performs its beam search using a single neural network instead of a large decoder graph. Therefore, the RNN-T model may be more compact and computationally intensive than the LAS model. For these reasons, the reducer 220 may switch between a two-pass model for speech recognition, an LAS model, and an RNN-T model to lower the processing level 222. That is, by employing an RNN-T network as the first pass and rescoring the first pass with an LAS network as the second pass, the two-pass model has a higher processing level 222 than either the LAS model or the RNN-T model alone, while the LAS model, as an attention-based model, has a higher processing level 222 than the RNN-T model. By identifying the processing requirements for speech recognition for different versions of the on-device speech recognition model, the reducer 220 can attenuate (or increase) speech processing by switching between different versions of the on-device model. Furthermore, when reducer 220 combines the processing level options of the on-device model with the processing level options of the server-side model, the attenuation of the audio processing by reducer 220 has many potential processing level steps.
[0044] To further expand the potential number of processing level stages, reducer 220 may be configured to modify the speech processing steps or parameters of a given model to change the processing level of that particular model. For example, a particular model may include one or more layers of a neural network (e.g., a recurrent neural network with a long-short-term memory (LSTM) layer). In some examples, an output layer may receive information from past states (backward) and future states (forward) to generate its output. A layer is considered bidirectional when it receives backward and forward states. In some configurations, reducer 220 is configured to modify the processing steps of the model to change the speech recognition model from bidirectional (i.e., forward and backward) operation to solely unidirectional (e.g., forward) operation. Additionally or alternatively, reducer 220 may reduce the number of layers of the neural network that the model uses to perform speech recognition to change the processing level 222 of a particular model.
[0045] The reducer 220 may lower the processing level 222 of the speech recognition model by changing the model's beam search parameters (or other pruning / search mode parameters). Generally, a beam search includes a beam size or beam width parameter that specifies how many of the best potential solutions (e.g., hypotheses or candidates) to evaluate to generate a speech recognition result. Thus, the beam search process performs a kind of pruning of potential solutions to reduce the number of solutions evaluated to form a speech recognition result. That is, the beam search process can limit the computation involved by using a limited number of active beams to search for the most likely sequences of words spoken in utterance 12U to generate a speech recognition result (e.g., transcription 152 of utterance 12U). Here, the reducer 220 may adjust the beam size to reduce the number of best potential solutions to evaluate, which in turn reduces the amount of computation involved in the beam search process. For example, the reducer 220 changes the beam size from 5 to beam size 2 to have the speech recognition model evaluate two best candidates instead of five best candidates.
[0046] In some examples, the reducer 220 performs quantization or sparsification on one or more parameters of the speech recognition model to generate lower processing levels 222 of the model. When generating a speech recognition result, the speech recognition model typically generates a large number of weights. For example, the speech recognition model weights different speech parameters and / or speech-related features to output a speech recognition result (e.g., transcription 152). Due to this large number of weights, the reducer 220 may discretize the values of these weights by performing quantization. For example, the quantization process converts floating-point weights into weights represented as fixed-point integers. This quantization process loses some information or quality, but may allow resources to process these quantized parameters to use less memory and allow more efficient operations (e.g., multiplication) to be performed on certain hardware.
[0047] In a similar vein, sparsification also aims to reduce the amount of processing required to run a model. Here, sparsification refers to the process of removing redundant parameters or features in a speech recognition model to focus on more relevant features. For example, while determining a speech recognition result, a speech model may determine probabilities for all speech-related features (e.g., characters, symbols, or words) even though not all speech-related features are associated with a particular speech input. By using sparsification, the model may expend fewer computational resources by determining probabilities for speech-related features associated with the input instead of all speech-related features, allowing the sparsification process to ignore speech-related features not associated with the input.
[0048] Optionally, reducer 220 may generate a lower processing level 222 of the speech recognition by determining the context of interaction 12 (e.g., spoken utterance 12U) that originally generated or resulted in open microphone duration window 212. Once reducer 220 identifies the context, reducer 220 may use the context to lower processing level 222 of speech recognition system 150 by biasing the speech recognition results based on the context. In some implementations, reducer 220 biases the speech recognition results based on the context by restricting speech recognition system 150 to vocabulary relevant to the context. As an example, user 10 may ask interface 120, "how do you tie a prusik hitch?" From this question, reducer 220 determines that the prusik knot is primarily used in mountaineering or rock climbing. In other words, reducer 220 identifies the context of interaction 12 as mountaineering. In this example, when reducer 220 begins to lower speech recognition processing level 222, reducer 220 limits the speech recognition output to mountain climbing vocabulary. Thus, if user 10 subsequently asks a follow-up question about the Appalachian Mountains, speech recognition system 150 may generate possible speech recognition results that include the terms “Application” and “Appalachian,” but reducer 220 ensures that speech recognition system 150 is biased toward the term “Appalachian” because “Appalachian” is associated with the “Mountaineering” context. For example, reducer 220 instructs speech recognition system 150 to increase the probability score of potential results related to the identified context (e.g., mountain climbing). In other words, speech recognition system 150 increases the probability score of potential results that include mountain climbing-related vocabulary.
[0049] When the speech recognition system 150 is running on the device 110, the speech recognition system 150 may generally use a system-on-chip (SOC-based) processor to perform speech recognition. A system-on-chip (SOC) processor refers to a general-purpose processor, a signal processor, and additional peripherals. The reducer 220 may generate a reduced processing level 222 when the speech recognition uses SOC-based processing by instructing the speech recognition system 150 to change from SOC-based processing to a digital signal processor (DSP). Here, this change results in a lower processing level 222 because DSPs tend to consume less power and memory than SOC-based processing.
[0050] When the reducer 220 attenuates the speech recognition processing level 222, it may be advantageous to provide the user 10 with some indication that the attenuation is occurring. To provide this indication, a graphical user interface (GUI) associated with the device 110 may include a graphical indicator to show the current processing level 222 of the speech recognition system 150. In some examples, the graphical indicator has a brightness level configured to decrease proportionally to the attenuation of the processing level 222. For example, the screen of the device 110 may include a GUI showing a red microphone dot indicating that the microphone 116 is open (i.e., listening to the interaction 12), which gradually fades out as the reducer 220 attenuates the speech recognition processing level 222. Here, when the microphone 116 closes, the red microphone dot disappears. Additionally or alternatively, the indicator that indicates that the microphone 116 is open and / or the degree of attenuation of the speech recognition processing level 222 may be a hardware indicator such as a light (e.g., a light emitting diode (LED)) on the device 110. For example, the LED may fade toward off as the processing level 222 decreases or flash at a decreasing rate (e.g., slower and slower) until the microphone 116 is closed.
[0051] 3 is a flow diagram of an example sequence of operations for a method 300 of attenuating audio processing. At operation 302, method 300 receives, at a voice-enabled device 110, an indication of a microphone trigger event 202 indicating a possible user interaction with the voice-enabled device 110 by speech, the voice-enabled device 110 having a microphone 116 that, when open, is configured to capture speech for recognition by an automatic speech recognition (ASR) system 150. Operation 304 includes two sub-operations 304, 304a-b that occur in response to receiving the indication of the microphone trigger event 202. At operation 304a, method 300 instructs microphone 116 to open or remain open for an open microphone duration window 212 to capture an audio stream 16 within the environment of voice-enabled device 110. At operation 304b, method 300 provides audio stream 16 captured by open microphone 116 to ASR system 150 for performing ASR processing on audio stream 16. Operation 306 includes two sub-operations 306, 306a-b, that occur while ASR system 150 is performing ASR processing on audio stream 16 captured by open microphone 116. At operation 306, method 300 attenuates a level 222 of ASR processing that ASR system 150 performs on audio stream 16 based on a function of open microphone duration window 212. At operation 306b, method 300 instructs ASR system 150 to use the attenuated level 204, 222 of ASR processing on audio stream 16 captured by open microphone 116.
[0052] 4 is a schematic diagram of an exemplary computing device 400 that may be used to implement the systems (e.g., device 110, interface 120, remote system 140, speech recognition system 150, search engine 160, and / or analyzer 200) and methods (e.g., method 300) described herein. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functionality are intended to be merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0053] Computing device 400 includes a processor 410, memory 420, a storage device 430, a high-speed interface / controller 440 that connects to memory 420 and a high-speed expansion port 450, and a low-speed interface / controller 460 that connects to a low-speed bus 470 and storage device 430. Each of components 410, 420, 430, 440, 450, and 460 are interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as appropriate. Processor 410 can process instructions for execution within computing device 400, including instructions stored in memory 420 or on storage device 430, to display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 480 coupled to high-speed interface 440. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as appropriate. Additionally, multiple computing devices 400 may be connected together (eg, as a server bank, a group of blade servers, or a multi-processor system) with each device providing a portion of the required operations.
[0054] The memory 420 stores information non-transiently within the computing device 400. The memory 420 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 420 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 400. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0055] The storage device 430 can provide mass storage for the computing device 400. In some implementations, the storage device 430 is a computer-readable medium. In various different implementations, the storage device 430 can be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 420, the storage device 430, or memory on the processor 410.
[0056] The high-speed controller 440 manages bandwidth-intensive operations for the computing device 400, while the low-speed controller 460 manages less bandwidth-intensive operations. Such assignment of roles is merely exemplary. In some implementations, the high-speed controller 440 is coupled to the memory 420, to the display 480 (e.g., through a graphics processor or accelerator), and to a high-speed expansion port 450, which may accept various expansion cards (not shown). In some implementations, the low-speed controller 460 is coupled to the storage device 430 and to a low-speed expansion port 490. The low-speed expansion port 490, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or may be coupled to a network device, such as a switch or router, for example, via a network adapter.
[0057] Computing device 400 may be implemented in several different forms, as shown in the figure. For example, computing device 400 may be implemented as a standard server 400a, or multiple times within a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.
[0058] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0059] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0060] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by dedicated logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. A processor typically receives instructions and data from a read-only memory or a random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. A computer typically also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operatively coupled to receive data from or transfer data to such mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0061] To provide for interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0062] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0063] 10 users 12 User Interaction 12U Oral Speech 12Ua First utterance 12Ub Second utterance 14 Hotwords, Triggers 16 audio streams 100 systems 110 Voice-enabled devices, devices, user devices 112 Data Processing Hardware 114 Memory Hardware 116 Audio capture device, microphone 120 Digital Assistant Interface, Voice-Enabled Interface, Interface, Assistant Interface 122, 122a~c Responses 122a First Response 122b Second Response 124 audio data 130 Network 140 Remote Systems 142 Remote Resources 144 Remote Data Processing Hardware 146 Remote Memory Hardware 150 Voice Recognition System 152 transcriptions and speech recognition results 160 search engines 162 results 200 Interaction Analyzer, Analyzer 202 Microphone trigger event, trigger event 202a First trigger event 202b Second trigger event 202c Third trigger event 204 Decaying State 210 Window Generator, Generator 212 Open Microphone Duration Window 212a First microphone duration window, first window 212b Second microphone duration window, new microphone duration window, second window 212c Third Window 214 Starting point 214b Second Starting Point 216 End Point 216b Second End Point 220 Reducer 222 Processing Level 222a First Processing Level 222b Second Processing Level 222c Third Processing Level 224 hour period 226 hour threshold 300 ways 400 computing devices 400a Standard Server 400b laptop computer 400c Rack Server System 410 processor 420 memory 430 Storage Devices 440 High-Speed Interface / Controller 450 High-Speed Expansion Port 460 Low-Speed Interface / Controller 470 Slow Bus 480 display 490 Low-Speed Expansion Port
Claims
1. When executed by data processing hardware, the data processing hardware receiving audio data corresponding to the speech and captured by the open microphone; generating first-pass speech recognition results for the audio data by processing the audio data using a first level of automatic speech recognition (ASR) processing; determining a confidence level of the first-pass speech recognition result; generating second-pass speech recognition results for the audio data by processing the audio data using a second level of ASR processing based on the confidence level of the first-pass speech recognition results, the second level of ASR processing being greater than the first level of ASR processing, and the second level of ASR processing decaying over time when processing the audio data using the second level of ASR processing; instructing the microphone to close when the second level of ASR processing is equal to zero; A computer-implemented method for performing operations including:
2. The operation is determining that the confidence of the first-pass speech recognition result does not satisfy a confidence threshold; increasing the first level of ASR processing to the second level of ASR processing based on determining that the confidence of the first-pass speech recognition result does not satisfy the confidence threshold; The computer-implemented method of claim 1 further comprising:
3. The computer-implemented method of claim 1 , wherein the first level of ASR processing comprises a partial processing capacity of an ASR system.
4. 2. The computer-implemented method of claim 1, wherein generating the first-pass speech recognition results for the audio data using the first level of ASR processing comprises performing speech recognition on a voice-enabled device.
5. The computer-implemented method of claim 1 , wherein the second level of ASR processing comprises all of the processing power of an ASR system.
6. 2. The computer-implemented method of claim 1, wherein generating the second-pass speech recognition results for the audio data using the second level of ASR processing includes performing speech recognition at a remote server in communication with a voice-enabled device.
7. 2. The computer-implemented method of claim 1, wherein the operations further include receiving an indication of a microphone trigger event indicating a possible user interaction with a voice-enabled device through speech, the voice-enabled device having a microphone, the microphone configured in an open state to capture speech for recognition by an ASR system.
8. 8. The computer-implemented method of claim 7, wherein the actions further include instructing the microphone to open or remain open for an open microphone duration window to capture the speech.
9. 9. The computer-implemented method of claim 8, wherein the operations further comprise attenuating the second level of ASR processing to a third level of ASR processing lower than the second level of ASR processing after generating the second-pass speech recognition result.
10. 2. The computer-implemented method of claim 1, wherein the first and second levels of ASR processing respectively correspond to respective amounts of computing resources used to process the audio data.
11. data processing hardware; memory hardware in communication with the data processing hardware, the memory hardware, when executed on the data processing hardware, receiving audio data corresponding to the speech and captured by the open microphone; generating first-pass speech recognition results for the audio data by processing the audio data using a first level of automatic speech recognition (ASR) processing; determining a confidence level of the first-pass speech recognition result; generating second-pass speech recognition results for the audio data by processing the audio data using a second level of ASR processing based on the confidence level of the first-pass speech recognition results, the second level of ASR processing being greater than the first level of ASR processing, and the second level of ASR processing decaying over time when processing the audio data using the second level of ASR processing; instructing the microphone to close when the second level of ASR processing is equal to zero; and memory hardware that stores instructions for performing operations including: A system including:
12. The operation is determining that the confidence of the first-pass speech recognition result does not satisfy a confidence threshold; increasing the first level of ASR processing to the second level of ASR processing based on determining that the confidence of the first-pass speech recognition result does not satisfy the confidence threshold; The system of claim 11 further comprising:
13. The system of claim 11 , wherein the first level of ASR processing comprises a partial processing capacity of an ASR system.
14. 12. The system of claim 11, wherein generating the first-pass speech recognition results for the audio data using the first level of ASR processing comprises performing speech recognition on a voice-enabled device.
15. The system of claim 11 , wherein the second level of ASR processing comprises all of the processing capabilities of an ASR system.
16. 12. The system of claim 11, wherein generating the second-pass speech recognition results for the audio data using the second level of ASR processing includes performing speech recognition at a remote server in communication with a voice-enabled device.
17. 12. The system of claim 11, wherein the operations further include receiving an indication of a microphone trigger event indicating a possible user interaction with a voice-enabled device by speech, the voice-enabled device having a microphone, the microphone configured in an open state to capture speech for recognition by an ASR system.
18. 20. The system of claim 17, wherein the action further comprises instructing the microphone to open or remain open for an open microphone duration window to capture the speech.
19. 20. The system of claim 18, wherein the operations further include attenuating the second level of ASR processing to a third level of ASR processing lower than the second level of ASR processing after generating the second-pass speech recognition result.
20. 12. The system of claim 11, wherein the first and second levels of ASR processing each correspond to a respective amount of computing resources used to process the audio data.
Citation Information
Patent Citations
Multi-pass vehicle voice recognition systems and methods
US20150379987A1
Trigger Word Detection With Multiple Digital Assistants
US20190251960A1
Noise cancellation for open microphone mode
US9646628B1
Information processing device, information processing method, and program
WO2020004213A1