Speech recognition using word or phoneme time markers based on user input - Patent Application 20070122997

User-input time markers improve ASR performance by correlating with audio signals to separate target speech from noise and competing voices, enhancing transcription accuracy in noisy conditions.

JP2026041784APending Publication Date: 2026-03-10GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Automatic speech recognition (ASR) systems face significant degradation in performance due to background noise and competing voices, which existing modeling strategies struggle to address effectively, particularly in low signal-to-noise ratio conditions.

Method used

The use of user-input time markers to correlate with audio signals, allowing for the separation of target speech from background noise and competing voices by calculating word timestamps and generating enhanced audio features for improved transcription accuracy.

Benefits of technology

Enhances speech recognition accuracy by effectively separating target speech from background interference, enabling accurate transcription even in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041784000001_ABST
    Figure 2026041784000001_ABST
Patent Text Reader

Abstract

A speech recognition system and method are provided. The method includes receiving an input audio signal captured by a user device, the input audio signal corresponding to a target speech of a plurality of words spoken by a target user and including background noise where the user device is located while the target user is speaking the plurality of words in the target speech. The method also includes receiving a sequence of time markers input by the target user as the target user speaks the plurality of words in the target speech, and correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal. The method also includes processing the enhanced audio features using a speech recognition model to generate a transcription of the target speech.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE This disclosure relates to speech recognition using word or phoneme time markers based on user input. [Background technology]

[0002] An automatic speech recognition (ASR) system may operate on a computing device to recognize and transcribe speech spoken by a user who queries a digital assistant to perform an action. The robustness of automatic speech recognition (ASR) systems has improved significantly over the years with the advent of neural network-based end-to-end models, large-scale training data, and improved strategies for expanding the training data. However, various conditions, such as stronger background noise and competing voices, significantly degrade the performance of ASR systems. Summary of the Invention

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations, the operations including receiving an input audio signal captured by a user device, the input audio signal corresponding to a target speech of a plurality of words spoken by a target user and including background noise at which the user device is located while the target user is speaking the plurality of words in the target speech, receiving a sequence of time markers input by the target user as the target user speaks the plurality of words in the target speech, correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal, and processing the enhanced audio features using a speech recognition model to generate a transcription of the target speech.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, correlating the sequence of time markers with the input audio signal includes: using the sequence of time markers to calculate a sequence of word timestamps, each word timestamp designating a respective time corresponding to one of a plurality of words of the target speech spoken by the target user; and using the calculated sequence of word timestamps to separate the target speech from background noise of the input audio signal and generate an extended audio feature. In these implementations, separating the target speech from background noise of the input audio signal may include removing the background noise from inclusions in the extended audio feature. Additionally or alternatively, separating the target speech from background noise of the input audio signal includes assigning the sequence of word timestamps to corresponding audio segments of the extended audio feature to distinguish the target speech from the background noise.

[0005] In some examples, receiving the sequence of time markers input by the target user includes receiving each time marker of the sequence of time markers in response to the target user touching or pressing a predefined area of ​​the user device or other device in communication with the data processing hardware. Here, the predefined area of ​​the user device or other device may include a physical button located on the user device or other device. Additionally or alternatively, the predefined area of ​​the user device or other device includes a graphical button displayed on a graphical user interface of the user device. In some implementations, receiving the sequence of time markers input by the target user includes receiving each time marker of the sequence of time markers in response to a sensor in communication with the data processing hardware detecting that the target user is performing a predefined gesture.

[0006] In some examples, the number of time markers in the sequence of time markers entered by the target user is equal to the number of words spoken in the target voice by the target user. In some embodiments, the data processing hardware resides in a user device associated with the target user. Additionally or alternatively, the data processing hardware resides in a remote server that communicates with a user device associated with the target user. In some examples, the background noise included in the input audio signal includes competing voices spoken by one or more other users. In some embodiments, the target voice spoken by the target user includes a query directed to a digital assistant running on the data processing hardware. Here, the query specifies an action for the digital assistant to perform.

[0007] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations including receiving an input audio signal captured by a user device. The input audio signal corresponds to a target speech of a plurality of words spoken by a target user and includes background noise where the user device is located while the target user is speaking the plurality of words in the target speech. The operations further include receiving a sequence of time markers input by the target user as the target user speaks the plurality of words in the target speech, correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal, and processing the enhanced audio features using a speech recognition model to generate a transcription of the target speech.

[0008] This aspect may include one or more of the following optional features. In some implementations, correlating the sequence of time markers with the input audio signal includes: using the sequence of time markers to calculate a sequence of word timestamps, each word timestamp designating a respective time corresponding to one of a plurality of words of the target speech spoken by the target user; and using the calculated sequence of word timestamps to separate the target speech from background noise of the input audio signal and generate an enhanced audio feature. In these implementations, separating the target speech from background noise of the input audio signal may include removing the background noise from inclusions of the enhanced audio feature. Additionally or alternatively, separating the target speech from background noise of the input audio signal includes assigning the sequence of word timestamps to corresponding audio segments of the enhanced audio feature to distinguish the target speech from the background noise.

[0009] In some examples, receiving the sequence of time markers input by the target user includes receiving each time marker of the sequence of time markers in response to the target user touching or pressing a predefined area of ​​the user device or other device in communication with the data processing hardware. Here, the predefined area of ​​the user device or other device may include a physical button located on the user device or other device. Additionally or alternatively, the predefined area of ​​the user device or other device includes a graphical button displayed on a graphical user interface of the user device. In some implementations, receiving the sequence of time markers input by the target user includes receiving each time marker of the sequence of time markers in response to a sensor in communication with the data processing hardware detecting that the target user is performing a predefined gesture.

[0010] In some examples, the number of time markers in the sequence of time markers entered by the target user is equal to the number of words spoken in the target voice by the target user. In some embodiments, the data processing hardware resides in a user device associated with the target user. Additionally or alternatively, the data processing hardware resides in a remote server that communicates with a user device associated with the target user. In some examples, the background noise included in the input audio signal includes competing voices spoken by one or more other users. In some embodiments, the target voice spoken by the target user includes a query directed to a digital assistant running on the data processing hardware. Here, the query specifies an action for the digital assistant to perform.

[0011] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram of an exemplary system for improving automatic speech recognition accuracy for transcribing target speech using a sequence of time markers entered by a user. [Figure 2] 1 is an exemplary plot showing correlations between sequences of input time markers and word time markers spoken in a target and competing speech spoken by different speakers; [Figure 3] 1 is a flowchart of an exemplary sequence of operations of a method for using input time markers to enhance / improve speech recognition for a target speech contained in a noisy audio signal. [Figure 4] FIG. 1 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0013] Like reference symbols in the various drawings indicate like elements.

[0014] The robustness of automatic speech recognition (ASR) systems has improved significantly over the years with the advent of neural network-based end-to-end models, large-scale training data, and improved strategies for expanding training data. Nevertheless, background interference can significantly degrade an ASR system's ability to accurately recognize speech directed at it. Background interference can be broadly categorized as background noise and competing voices. While separate ASR models can be trained to treat each of these background interference groups in isolation, maintaining multiple task / condition-specific ASR models and switching between them on the fly during use is difficult and impractical.

[0015] Background noise with non-speech characteristics is typically adequately handled using data augmentation strategies such as multi-style training (MTR) of ASR models. Here, noise is added to training data using a room simulator, which is then carefully weighted with clean data during training to balance performance between clean and noisy conditions. As a result, large-scale ASR models are robust to moderate levels of non-speech noise. However, in the presence of low signal-to-noise ratio (SNR) conditions, background noise can still affect the performance of back-end speech systems.

[0016] Unlike non-speech background noise, competing speech presents significant challenges for ASR models trained to recognize a single speaker. Training an ASR model with multiple-speaker speech can be problematic in itself, as it can be difficult to disambiguate which speaker to focus on during inference. Using a model that recognizes multiple speakers is also suboptimal, since it is difficult to know in advance how many users it will support. Furthermore, such multi-speaker models typically perform poorly in single-speaker settings, making them undesirable.

[0017] These aforementioned classes of background interference are typically addressed separately, each using a separate modeling strategy. In recent literature, speech separation using speaker embeddings, along with techniques such as deep clustering and permutation-invariant learning, has received much attention. When using speaker embeddings, the target speaker of interest is assumed to be known a priori. When the target speaker embedding is unknown, blind speaker separation involves using the input audio waveforms of the speech mixture itself and performing unsupervised learning techniques by clustering similar-looking audio features into the same bucket. However, due to the severe ill-posedness of this inverse problem, audio-only compensated speaker sets are often noisy and still contain inter-speech artifacts. For example, the audio-only speech of a speaker set associated with a first speaker may contain speech spoken by another second speaker. Techniques developed for speaker separation are also applied to remove non-speech noise by modifying the training data.

[0018] 1 , in some embodiments, system 100 includes a user 10 in an audio environment communicating a spoken target speech 12 to a speech-enabled user device 110 (also referred to as device 110 or user device 110). The target user 10 (i.e., the speaker of the speech 12) may speak the target speech 12 as a query or command that solicits a response from a digital assistant 105 running on user device 110. User device 110 is configured to capture sounds from one or more users 10, 11 within the audio environment. Here, audio sounds may refer to the speech 12 spoken by user 10 that functions as an audible query, a command / action to digital assistant 105, or an audible communication captured by device 110. Digital assistant 105 may process the query for the command by answering the query and / or causing the command to be executed.

[0019] The device 110 may correspond to any computing device associated with the user 10 and capable of receiving the noisy audio signal 202. Some examples of the user device 110 include, but are not limited to, a mobile device (e.g., a mobile phone, a tablet, a laptop, etc.), a computer, a wearable device (e.g., a smart watch), a smart appliance, an Internet of Things (IoT) device, a smart speaker, and the like. The device 110 includes data processing hardware 112 and memory hardware 114 in communication with the data processing hardware 112, which stores instructions that, when executed by the data processing hardware 112, cause the data processing hardware 112 to perform one or more operations. In some examples, an automatic speech recognition (ASR) system 120 executes on the data processing hardware 112. Additionally or alternatively, the digital assistant 105 may execute on the data processing hardware 112.

[0020] Device 110 further includes an audio subsystem with an audio capture device (e.g., a microphone) 116 for capturing and converting spoken speech 12 within the audio environment into electrical signals, and an audio output device (e.g., an audio speaker) 117 for communicating audible audio signals. Although device 110 implements a single audio capture device 116 in the illustrated example, device 110 may implement an array of audio capture devices 116 without departing from the scope of this disclosure, whereby one or more audio capture devices 116 in the array may communicate with the audio subsystem (e.g., peripherals of device 110) rather than being physically present on device 110. For example, device 110 may correspond to a vehicle infotainment system that utilizes an array of microphones located throughout the vehicle.

[0021] In the illustrated example, the ASR system 120 uses an ASR model 160 that processes the extended speech features 145 to generate a speech recognition result (e.g., a transcription) 160 for the target speech 12. As described in more detail below, the ASR system 120 may derive the extended speech features 145 from a noisy input audio signal 202 corresponding to the target speech 12 based on a sequence of input time markers 204 (interchangeably referred to as an “input time marker sequence 204” or a “time marker sequence 204”). The ASR system 120 may further include a natural language understanding (NLU) module 170 that performs semantic interpretation on the transcription 165 of the target speech 12 to identify queries / commands directed to the digital assistant 105. Accordingly, the output 180 from the ASR system 120 may include instructions for executing the queries / commands identified by the NLU module 170.

[0022] In some examples, device 110 is configured to communicate with remote system 130 (also referred to as remote server 130) over a network (not shown). Remote system 130 may include remote resources 132, such as remote data processing hardware 134 (e.g., a remote server or CPU) and / or remote memory hardware 136 (e.g., a remote database or other storage hardware). User device 110 may utilize remote resources 132 to perform various functions related to voice processing and / or query fulfillment. ASR system 120 may reside within device 110 (referred to as an on-device system) or may reside remotely while in communication with device 110 (e.g., may reside in remote system 130). In some examples, one or more components of ASR system 120 reside locally or within the device, while one or more other components of ASR system 120 reside remotely. For example, if a model or other component of ASR system 120 is significantly larger in size or involves increased processing requirements, the model or component may reside in remote system 130. Furthermore, if device 110 can support the size or processing requirements of a given model and / or component of ASR system 120, those models and / or components may reside within device 110 using data processing hardware 112 and / or memory hardware 114. Optionally, components of ASR system 120 may reside both locally / on-device and remotely. For example, a first-pass ASR model 160 capable of performing streaming speech recognition may run on user equipment 110, while a second-pass ASR model 160, which is more computationally intensive than the first-pass ASR model 160, may run on remote system 130 to rescore the speech recognition results generated by the first-pass ASR model 160.

[0023] Various types of background interference can interfere with the ability of the ASR system 120 to process the target voice 12 specifying a query or command to the device 110. As previously mentioned, background interference can include competing voices 13, such as speech 13 other than the target voice 12 spoken by one or more other users 11 not directed at the user device 110, and background noise having non-voice characteristics. In some cases, background inference can also be caused by device echo corresponding to played audio output from the user device (e.g., smart speaker) 110, such as media content or synthesized speech from a digital assistant 105 conveying a response to a query spoken by the target user 10.

[0024] In the illustrated example, user device 110 captures a noisy audio signal 202 (also referred to as audio data) of target speech 12 spoken by user 10 in the presence of background interference emanating from one or more sources other than user 10. Thus, audio signal 202 also includes background noise / interference that was present at user device 110 while target user 10 was speaking target speech 110. Target speech 12 may correspond to a query directed to digital assistant 105 that specifies an action for digital assistant 105 to perform. For example, target speech 12 spoken by target user 10 may include the query, "What is the weather today?", requesting digital assistant 105 to obtain today's weather forecast. The presence of background interference due to at least one of competing voices 13, device echo, and / or non-voice background noise interfering with the target voice 12 may make it difficult for the ASR model 140 to recognize an accurate transcription of the target voice 12 corresponding to the query "What is the weather today?" in the noisy audio signal 202. As a result of the inaccurate transcription 165 output by the ASR model 160, the NLU module 170 may not be able to grasp the actual intent / content of the query spoken by the target user 10, thereby resulting in the digital assistant 105 being unable to obtain an appropriate response (e.g., today's weather forecast) or being unable to obtain a response at all.

[0025] To combat background noise / interference in a noisy audio signal 202 that may degrade the performance of the ASR model 160 to accurately transcribe the target speech 12 of the words spoken by the target user 10, embodiments herein are directed to using a sequence of time markers 204 input by the target user 10 as the target user 10 speaks the words of the target speech 12 as additional context to enhance / improve the accuracy of the transcription 165 output by the ASR model 160. That is, the target user 10 may input a time marker 204 via the user input source 115 each time the user 10 speaks one of the words, such that each time marker 204 specifies the time at which the user 10 spoke a corresponding one of the words of the target speech 12. Thus, the number of time markers 204 input by the user may be equal to the number of words spoken by the target user 10 in the target speech 12. In the illustrated example, the ASR system 120 may receive a sequence of five time markers 204 associated with five words "what," "is," "the," "weather," and "today" spoken by the target user 10 in the target speech 12. In some examples, the target user 10 inputs a time marker 204 via the user input source 115 each time the user 10 speaks a syllable that occurs in a corresponding one of the words, such that each time marker 204 designates the time at which the user 10 spoke a corresponding syllable of the words in the target speech 12. Here, the number of time markers 204 input by the user may be equal to the number of syllables occurring in the words spoken by the target user in the target speech 12. The number of syllables may be equal to or greater than the number of words spoken by the target user 10.

[0026] An embodiment includes the ASR system 120 employing an input-to-speech (ITS) correlator 140 configured to receive as input an input time marker sequence 204 input by the target user 10 and initial features 154 extracted from the input audio signal 202 by the feature extraction module 150. The initial features 154 may correspond to parameterized acoustic frames, such as Mel frames. For example, the parameterized acoustic frames may correspond to log Mel filter bank energy (LFBE) extracted from the input audio signal 202 and may be represented by a series of time windows / frames. In some examples, the time windows representing the initial features 154 may include a fixed size and a fixed overlap. An embodiment is directed to the ITS correlator 140 correlating the received sequence of time markers 204 with the initial features 154 extracted from the input audio signal 202 to generate enhanced audio features 145 that separate the target speech 12 from background interference (i.e., competing speech 13) in the input audio signal 202. The ASR model 160 may then process the enhanced audio features 145 generated by the ITS correlator 140 to generate a transcription 165 of the target speech 12. In some configurations, the user device 110 displays the transcription 165 on the graphical user interface 118 of the user device 110.

[0027] In some implementations, the ITS correlator 140 correlates the sequence of time markers 204 with the initial features 154 extracted from the input audio signal 202 by using the sequence of time markers 204 to calculate a sequence of word timestamps 144, each specifying a respective time at which a corresponding one of a plurality of words of the target speech 12 was spoken by the target user 10. In these implementations, the ITS correlator 140 may use the calculated sequence of word timestamps 144 to separate the target speech 12 from background interference (i.e., competing speech 13) in the input audio signal 202.

[0028] In some examples, the ITS correlator 140 separates the target speech 12 from the background interference (i.e., competing speech 13) by removing the background interference (i.e., competing speech 13) from the inclusions of the extended audio features 145. For example, the background interference may be removed by filtering a time window from the initial features 154 to which no word timestamps 144 are assigned / assigned from the sequence of word timestamps 144. In other examples, the ITS correlator 140 separates the target speech 12 from the background interference (i.e., competing speech 13) by assigning a sequence of word timestamps 144 to corresponding audio segments of the extended audio features 145, thereby distinguishing the target speech 12 from the background interference. For example, the extended audio features 145 may correspond to initial features 154 extracted from the input audio signal 202, represented by time windows / frames, whereby the word timestamps 144 are assigned to corresponding time windows from the initial features 154 to specify the timing of each word.

[0029] In some examples, the ITS correlator 140 may perform blind speaker diagnosis on the raw initial features 154 by extracting speaker embeddings and grouping similar speaker embeddings extracted from the initial features 154 into clusters. Each cluster of speaker embeddings may be associated with a corresponding speaker label associated with a different respective speaker. The speaker embedding may include a d-vector or an i-vector. Thus, the ITS correlator 140 can assign each extended audio feature 145 to a corresponding speaker label 146 to convey speech spoken by different speakers in the input audio signal 202.

[0030] The ITS correlator 140 may correlate each input time marker 204 with each word spoken by the target user 10 in the target voice using the following formula: speaker_1_word_data = SPEECH_TO_WORD(speaker_1_waveform)(1) Here, SPEECH_TO_WORD describes an ordered set of time markers 204 (eg, in milliseconds) that correlate with words spoken by the target user 10 in the target voice 12 .

[0031] In a multi-speaker scenario, the ITS correlator 140 may receive respective input time marker sequences 204 from at least two different users 10, 11 and initial features 154 extracted from an input audio signal 202 containing speech spoken by each of the at least two different users 10, 11. Here, the ITS correlator 140 may apply Equation 1 to correlate each input time marker 204 from each input time marker sequence with words spoken by the different users 10, 11, thereby generating enhanced audio features 145 that effectively separate the speech spoken by the at least two different users 10, 11. That is, the ITS correlator 140 identifies an optimal correlation between the sequence of input time markers 204 and the initial features 154 extracted from the noisy input audio signal 202 by performing point-by-point time difference matching. Of course, only one of the different users 10, 11 may input their respective sequences of time markers 204, so that the ITS correlator 140 may use the sequence of time markers 204 input by the target user 10 to generate enhanced audio features 145 that separate the target speech 12 from the other speech 13, and then interpolate which words were spoken in the speech 13 from the user 11 who did not input any time markers.

[0032] 2 is a plot 200 illustrating how the ITS correlator 140 performs point-by-point time difference matching between a sequence of time markers 204 input by a first speaker (Speaker #1) and a noisy audio signal 202 containing a target speech 12 of multiple words spoken by the first speaker in the presence of multiple other competing speeches 13 spoken by a different second speaker (Speaker #2). The x-axis of the plot 200 indicates increasing time from left to right. The ITS correlator 140 may calculate the average time difference between each time marker 204 in the sequence input by Speaker #1 and the nearest word timestamp 144 associated with the position of the spoken word in the mixed audio signal 202 containing both the target speech 12 and the competing speeches 13. Plot 200 shows that word timestamps 144 indicating words spoken by speaker #1 in target speech 12 are associated with a smaller average time difference between the input sequences of time markers 204 than word timestamps indicating other words spoken by speaker #2 in competing speech 13. Thus, ITS correlator 140 can use any of the above techniques to generate extended speech features 154 that effectively separate target speech 12 from competing speech 13, allowing downstream ASR model 160 to generate a transcription 165 of target speech 12 in noisy input audio signal 202 while ignoring competing speech 13.

[0033] The target user 10 may provide the sequence of time markers 204 using various techniques. The target user 10 may opportunistically provide the time markers 204 when in a noisy environment (e.g., a subway, an office, etc.) where competing voices 13 are susceptible to being captured by the user device 110 while the target user 10 is calling the digital assistant 105 via the target voice 12. The ASR system 120, and more specifically the ITS correlator 140, may receive the sequence of time markers 204 entered by the user 10 via a user input source 115. The user input source 115 may allow the user 10 to enter the sequence of time markers 204 along with multiple words and / or syllables spoken by the user 10 in the target voice 12 that the ASR model 140 is to recognize.

[0034] In some examples, the ASR system 120 receives each time marker 204 in response to the target user 10 touching or pressing a predefined region 115 of the user device 110 or another device communicating with the user device 110. Here, the predefined region touched / pressed by the user 10 serves as a user input source 115 and may include a physical button 115a (e.g., a power button) on the user device 110 or a physical button 115a on a peripheral device (e.g., a steering wheel button if the user device 110 is a vehicle infotainment system) (e.g., a keyboard button / key if the user device 110 is a desktop computer). The predefined region 115 may also include a capacitive touch region (e.g., a capacitive touch region on headphones) located on the user device 110. Without departing from the scope of the present disclosure, the predefined area 115 that is touched / pressed by the user 10 and serves as a user input source 115 may include a graphical button 115b displayed on the graphical user interface 118 of the user device 110. For example, in the illustrated example, a graphical button 115b labeled "Touch While Speaking" is displayed on the graphical user interface 118 of the user device 110 to allow the user 10 to tap the graphical button 115b for each word that the user 10 speaks in the target voice 12.

[0035] The ASR system 120 may also receive each time marker 204 in response to the sensor 113 detecting that the target user 10 is performing a predefined gesture. The sensor 113 may include an array of one or more sensors of or in communication with the user device 110. For example, the sensor 113 may include an image sensor (e.g., a camera), a radar / lidar sensor, a motion sensor, a capacitance sensor, a pressure sensor, an accelerometer, and / or a gyroscope. For example, the image sensor 113 may capture streaming video of the user 10 performing a predefined gesture, such as clenching a fist, each time the user 10 utters a word / syllable in the target voice 12. Similarly, the user may perform other predefined gestures, such as squeezing the user device, which can be detected by a capacitance or pressure sensor, or shaking the user device, which can be detected by accelerometers and / or a gyroscope.

[0036] In some additional examples, the sensor includes a microphone 116b for detecting audible sounds corresponding to input time markers 204 input by the user 10 while speaking each word in the target voice 12. For example, the user 10 may clap their hands, snap their fingers, knock on the surface supporting the user device 110, or make some other audible sound that can be captured by the microphone 116b in conjunction with speaking a word in the target voice.

[0037] 3 illustrates an exemplary sequence of operations for a method 300 for enhancing / improving speech recognition accuracy by using a sequence of time markers 204 entered by the target user 10 as the target user 10 speaks each of a plurality of words in the target voice 12. The target voice 12 may include an utterance of a query directed to the digital assistant application 105, requesting the digital assistant application 105 to perform an action specified in the query in the target voice 12. The method 300 includes a computer-implemented method executed on the data processing hardware 410 (FIG. 4) by communicating with the data processing hardware 410 and executing an exemplary sequence of actions stored in the memory hardware 420 (FIG. 4). The data processing hardware 410 may include the data processing hardware 112 of the user device 110 or the data processing hardware 134 of the remote system 130. The memory hardware 420 may include the memory hardware 114 of the user device 110 or the memory hardware 136 of the remote system 130.

[0038] At operation 302, the method 300 includes receiving an input audio signal 202 captured by a user device 110, where the input audio signal 202 corresponds to a target speech 12 of multiple words spoken by a target user 10 and includes background noise in the presence of the user device 110 while the target user 10 is uttering the multiple words of the target speech 12. The background noise may include competing speeches 13 spoken by one or more other users 11 in the presence of the user device 110 when the target speech 12 is spoken by the target user 10.

[0039] At operation 304, the method 300 includes receiving a sequence of time markers 204 entered by the target user 10 as the target user 10 speaks a plurality of words in the target voice 12, where each time marker 204 in the sequence of time markers 204 may be received in response to the target user 10 touching or pressing a predefined region 115 of the user device 110 or other device. In some examples, the number of time markers 204 in the sequence of time markers 204 entered by the target user 10 is equal to the number of words spoken by the target user 10 in the target voice 12.

[0040] At operation 306, the method 300 also includes correlating the sequence of time markers 204 with the input audio signal 202 to generate enhanced audio features 145 that separate the target speech 12 from background noise 13 in the input audio signal 202. At operation 308, the method 300 includes processing the enhanced audio features 145 using the speech recognition model 160 to generate a transcription 165 of the target speech 12.

[0041] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0042] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0043] 4 is a schematic diagram of an exemplary computing device 400 that can be used to implement the systems and methods described herein. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functionality are for illustrative purposes only and are not intended to limit the embodiments of the present disclosure described and / or claimed herein.

[0044] Computing device 400 includes a processor 410, memory 420, a storage device 430, a high-speed interface / controller 440 connecting to memory 420 and a high-speed expansion port 450, and a low-speed interface / controller 460 connecting to a low-speed bus 470 and storage device 430. Each of the components 410, 420, 430, 440, 450, and 460 is interconnected using various buses and may reside on a common motherboard or exist in other ways as needed. Processor 410 processes instructions for execution within computing device 400, including instructions stored in memory 420 or storage device 430, and can display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 480 connected to high-speed interface 440. In other implementations, multiple processors and / or multiple buses may be used as needed, along with multiple memories and memory types. Multiple computing devices 400 may also be connected, each providing a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0045] The memory 420 stores information non-transiently within the computing device 400. The memory 420 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 420 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 400. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0046] Storage device 430 can provide mass storage for computing device 400. In some embodiments, storage device 430 is a computer-readable medium. In various different implementations, storage device 430 can be a device array, including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In additional embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 420, storage device 430, or memory on processor 410.

[0047] High-speed controller 440 manages bandwidth-intensive operations for computing device 400, while low-speed controller 460 manages lower-bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, high-speed controller 440 is coupled to memory 420, display 480 (e.g., via a graphics processor or accelerator), and high-speed expansion port 450, which can accept various expansion cards (not shown). In some implementations, low-speed controller 460 is coupled to storage device 430 and low-speed expansion port 490. Low-speed expansion port 490 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet, etc.) and can connect to one or more input / output devices, such as a keyboard, pointing device, scanner, or network devices, such as a switch or router, via a network adapter or the like.

[0048] The computing device 400, as shown, can be implemented in many different forms. For example, it may be implemented as a standard server 400a, or multiple times within a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.

[0049] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, which may be specialized or general-purpose, at least one input device, and at least one output device, coupled to receive data and instructions from and transmit data and instructions to the storage system.

[0050] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic circuit (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0051] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating output. The processors and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operably coupled to receive data from or transfer data to them, or both. However, a computer need not include such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0052] To provide for interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor) or touch screen for displaying information to the user, and optionally a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide for interaction with a user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, spoken, or tactile input. Additionally, a computer can interact with a user by sending and receiving documents to a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0053] While several implementations have been described, it will, of course, be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method (300) that, when executed on data processing hardware (410), causes the data processing hardware (410) to perform operations, the operations comprising: receiving an input audio signal (202) captured by a user device (110), the input audio signal (202) corresponding to a target speech (12) of a plurality of words spoken by a target user (10) and including background noise in which the user device (110) is located while the target user (10) is speaking the plurality of words of the target speech (12); receiving a sequence of time markers (204) entered by the target user (10) in time with the target user (10) speaking the plurality of words in the target voice (12); correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal; processing the extended audio features (145) using a speech recognition model (160) to generate a transcription (165) of the target speech (12); A computer-implemented method (300) comprising:

2. Correlating the sequence of time markers (204) with the input audio signal (202) comprises: using the sequence of time markers (204) to calculate a sequence of word timestamps (144), each word timestamp designating a respective time corresponding to one of the plurality of words of the target speech (12) spoken by the target user (10); using the sequence of calculated word timestamps (144) to separate the target speech (12) from the background noise of the input audio signal (202) and generate the enhanced audio features (145); The computer-implemented method (300) of claim 1, comprising:

3. 3. The computer-implemented method of claim 2, wherein separating the target voice from the background noise of the input audio signal comprises removing the background noise from inclusions of the extended audio features.

4. 4. The computer-implemented method of claim 2, wherein separating the target speech from the background noise of the input audio signal comprises assigning a sequence of the word timestamps to a corresponding audio segment of the extended audio features to distinguish the target speech from the background noise.

5. 5. The computer-implemented method of claim 1, wherein receiving the sequence of time markers input by the target user comprises receiving each time marker in the sequence of time markers in response to the target user touching or pressing down a predefined area of ​​the user device or other device in communication with the data processing hardware.

6. 6. The computer-implemented method of claim 5, wherein the predefined area of ​​the user device includes a physical button located on the user device.

7. 7. The computer-implemented method of claim 5, wherein the predefined area of ​​the user device includes a graphical button displayed on a graphical user interface of the user device.

8. 8. The computer-implemented method (300) of claim 1, wherein receiving the sequence of time markers (204) input by the target user (10) comprises receiving each time marker (204) of the sequence of time markers (204) in response to a sensor (113) in communication with the data processing hardware (410) detecting that the target user (10) is performing a predefined gesture.

9. 9. The computer-implemented method of claim 1, wherein the number of time markers in the sequence of time markers input by the user is equal to the number of words spoken by the target user in the target voice.

10. The computer-implemented method (300) of any of claims 1 to 9, wherein the data processing hardware (410) resides in the user device (110) associated with the target user (10).

11. The computer-implemented method (300) of any of claims 1 to 10, wherein the data processing hardware (410) resides in a remote server (130) that communicates with the user device (110) associated with the target user (10).

12. 12. The computer-implemented method (300) of any of claims 1 to 11, wherein the background noise included in the input audio signal (202) includes competing voices (13) spoken by one or more other users (11).

13. The computer-implemented method (300) of any one of claims 1 to 12, wherein the target speech (12) spoken by the target user (10) includes a query directed to a digital assistant (105) running on the data processing hardware (410), the query specifying an action for the digital assistant (105) to perform.

14. data processing hardware (410); memory hardware (420) in communication with the data processing hardware (410), the memory hardware (420) storing instructions that, when executed by the data processing hardware (410), cause the data processing hardware (410) to: receiving an input audio signal (202) captured by a user device (110), the input audio signal (202) corresponding to a target speech (12) of a plurality of words spoken by a target user (10) and including background noise in which the user device (110) is located while the target user (10) is speaking the plurality of words of the target speech (12); receiving a sequence of time markers (204) entered by the target user (10) in time with the target user (10) speaking the plurality of words in the target voice (12); correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal; processing the extended audio features (145) using a speech recognition model (160) to generate a transcription (165) of the target speech (12); the memory hardware (420) causing A system (100) comprising:

15. Correlating the sequence of time markers (204) with the input audio signal (202) comprises: using the sequence of time markers (204) to calculate a sequence of word timestamps (144), each word timestamp designating a respective time corresponding to one of the plurality of words of the target speech (12) spoken by the target user (10); using the sequence of calculated word timestamps (144) to separate the target speech (12) from the background noise of the input audio signal (202) and generate the enhanced audio features (145); The system (100) of claim 14, comprising:

16. 16. The system (100) of claim 15, wherein separating the target voice (12) from the background noise of the input audio signal (202) comprises removing the background noise from inclusions of the extended audio features (145).

17. 17. The system (100) of claim 15 or claim 16, wherein separating the target speech (12) from the background noise of the input audio signal (202) comprises assigning a sequence of the word timestamps (144) to corresponding audio segments of the extended audio features (145) to distinguish the target speech (12) from the background noise.

18. The system (100) of any one of claims 14 to 17, wherein receiving the sequence of time markers (204) input by the target user (10) comprises receiving each time marker (204) of the sequence of time markers (204) in response to the target user (10) touching or pressing down a predefined area (115) of the user device (110) or another device in communication with the data processing hardware (410).

19. 20. The system (100) of claim 18, wherein the predefined area (115) of the user device (110) or the other device includes a physical button (115a) located on the user device (110) or the other device.

20. 20. The system (100) of claim 18 or claim 19, wherein the predefined area (115) of the user device (110) or the other device includes a graphical button (115b) displayed on a graphical user interface (118) of the user device (110).

21. The system (100) of any one of claims 14 to 20, wherein receiving the sequence of time markers (204) input by the target user (10) comprises receiving each time marker (204) of the sequence of time markers (204) in response to a sensor (113) in communication with the data processing hardware (410) detecting that the target user (10) is performing a predefined gesture.

22. 22. The system (100) of claim 14, wherein the number of time markers (204) in the sequence of time markers (204) input by the user is equal to the number of the plurality of words spoken by the target user (10) in the target voice (12).

23. The system (100) of any of claims 14 to 22, wherein the data processing hardware (410) resides in the user device (110) associated with the target user (10).

24. The system (100) of any of claims 14 to 23, wherein the data processing hardware (410) resides in a remote server that communicates with the user device (110) associated with the target user (10).

25. The system (100) of any one of claims 14 to 24, wherein the background noise included in the input audio signal (202) includes competing voices (13) spoken by one or more other users (11).

26. The system (100) of any one of claims 14 to 25, wherein the target speech (12) spoken by the target user (10) includes a query directed to a digital assistant (105) running on the data processing hardware (410), the query specifying an action for the digital assistant (105) to perform.