Adaptive Batch Processing for Reducing Recognition Latency

By dynamically controlling batch sizes in ASR systems, the method addresses the challenge of reducing latency without parallel models, enhancing both user experience and processing efficiency.

JP7693988B2Active Publication Date: 2025-06-18MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022542461
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2020-12-15
Publication Date
2025-06-18
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

Conventional automatic speech recognition (ASR) systems face challenges in reducing latency without relying on parallel models, which increases training and deployment costs and introduces undesirable user-perceived latency due to batching requirements.

Method used

The implementation of initial latency-sensitive adaptive batching, where the batch size is dynamically controlled during ASR processing, allowing for a small batch size initially to minimize latency and then increasing the batch size after word hypotheses are generated to improve processing efficiency.

Benefits of technology

This approach enables quick presentation of word hypotheses with reduced user-perceived latency while maintaining the advantages of larger batch sizes, such as better CPU and cache efficiency, thereby improving overall ASR performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693988000001
    Figure 0007693988000001
  • Figure 0007693988000002
    Figure 0007693988000002
  • Figure 0007693988000003
    Figure 0007693988000003
Patent Text Reader

Abstract

An embodiment may include collecting a first batch of acoustic feature frames of the audio signal, where the number of acoustic feature frames in the first batch is equal to a first batch size; inputting the first batch into a speech recognition network; collecting a second batch of acoustic feature frames of the audio signal in response to detection of word hypotheses output by the speech recognition network, where the number of acoustic feature frames in the second batch is equal to a second batch size that is greater than the first batch size; and inputting the second batch into the speech recognition network.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Neural-network-based models are commonly used to perform automatic speech recognition (ASR). In some examples, a neural-network-based acoustic model extracts senone-discriminative features from an input audio frame and is trained to classify senones based on the extracted features. The decoder generates word hypotheses based on the classification and outputs the corresponding text.

[0002] Input audio frames can be batched into batches of two or more frames in order to enable joint processing and recognition for the purpose of improving accuracy and performance. However, batching requires the system to wait until all frames of the batch have been received before submitting the batch to the ASR. This waiting can introduce an undesirable latency that is perceived by the user and is independent of the processing speed of the ASR.

[0003] Conventional ASR systems can use several physically different models to satisfy different latency requirements within a given deployment. This approach increases the training and deployment costs associated with the processor. The system is desired to improve latency without relying on parallel models.

Brief Description of the Drawings

[0004]

Figure 1A

[0005]

Figure 1B

[0006]

Figure 1C

[0007]

Figure 1D

[0008]

Figure 2

[0009]

Figure 3A

[0010]

Figure 3B

[0011]

Figure 4A

Figure 4B

Figure 4C

[0012]

Figure 5A

[0013]

Figure 5B

[0014]

Figure 5C

[0015]

Figure 5D

[0016]

Figure 6A

[0017]

Figure 6B

[0018]

Figure 6C

[0019]

Figure 6D

[0020]

Figure 7A

Figure 7B

[0021]

Figure 8

[0022]

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0023] The following description is provided to enable one of ordinary skill in the art to make and use the described embodiments. However, various modifications will be readily apparent to one of ordinary skill in the art.

[0024] According to some embodiments, the number of frames simultaneously input to the acoustic model (i.e., the batch size) is dynamically controlled during ASR processing. For example, a small batch size is used until word hypotheses are generated. Using a small batch size may result in less latency than in the case of a larger batch size because collecting a small batch requires less time than collecting a larger batch of frames of the same size. Then, the batch size may be increased after word hypotheses are generated, which may improve processing efficiency because simultaneous processing of a large number of frames has better CPU efficiency and cache efficiency than processing a smaller number of frames. Thus, the user can receive the presentation of word hypotheses quickly while maintaining most of the advantages of processing with a larger batch size.

[0025] Some embodiments are operable with a model that utilizes look-ahead frames within each batch of input frames, where the number of look-ahead frames is related to the degree of overlap between consecutive batches of input frames. These embodiments can also dynamically control the batch size as described above to provide initially low user-perceived latency and, after the first hypothesis, further provide improved processing efficiency.

[0026] Some models are operable to receive input batches of various sizes containing various numbers of look-ahead frames. By dynamically increasing the batch size in response to the first hypothesis, embodiments using such models provide initially low user-perceived latency, and then, as described above, the processing efficiency may be improved. Further, an improvement in recognition accuracy may be achieved by increasing the number of look-ahead frames in response to the first hypothesis.

[0027] In this regard, such an acoustic model can be trained considering a specific number of look-ahead frames that is greater than the number of look-ahead frames initially used. Therefore, although the initial operation is sub-optimal, it is possible to improve the recognition accuracy by changing the number of look-ahead frames to the number of look-ahead frames on which the acoustic model was trained. Also, embodiments may allow changing the number of look-ahead frames while maintaining a constant batch size.

[0028] Additionally or alternatively, some embodiments monitor the input frames independently of hypothesis generation to detect an end-of-speech state. When this state is detected, the batch size can be reduced to avoid waiting for non-speech frames that would otherwise be added to the current batch of input frames. The batch size may also be reduced in response to detection of the end of a stream, a target maximum processing time, or other conditions. Thus, such embodiments can support early exit from the batch process and reduce latency in the generation of the final hypothesis.

[0029] FIG. 1A shows a speech recognition system 100 that utilizes initial latency-sensitive adaptive batching during operation according to some embodiments. System 100 can be implemented using any suitable combination of hardware and software components. Each of the illustrated functions of system 100, or the functions otherwise described herein, can be realized by one or more computing devices (e.g., computer servers), storage devices (e.g., hard or solid state disk drives), and other hardware known in the art. The components can be located remotely from each other and can be elements of one or more cloud computing platforms including, but not limited to, software-as-a-service, platform-as-a-service, and infrastructure-as-a-service platforms. According to some embodiments, one or more components are implemented by one or more dedicated virtual machines that execute one or more trained neural network models.

[0030] The frame generation unit 110 receives an audio signal and generates a frame-level amplitude spectrum corresponding to each frame of the audio signal, or a raw acoustic feature frame, as is known in the art. For example, assuming an audio frame size of 25 ms, the frame generation unit 110 generates the first raw acoustic feature frame s0 based on the first 25 ms of the incoming audio signal. In this regard, FIG. 1A shows the start of ASR based on the incoming audio signal.

[0031] The next raw acoustic feature frame can be generated based on the next 25 ms of the incoming audio signal. In some embodiments, temporally adjacent raw acoustic feature frames can represent overlapping frames of the audio signal. For example, the first raw acoustic feature frame s0 can be generated based on the first 25 ms of the incoming audio signal, the second raw acoustic feature frame can be generated based on the 15 ms to 40 ms frame of the incoming audio signal, and the third raw acoustic feature frame can be generated based on the 30 ms to 55 ms frame of the incoming audio signal.

[0032] According to some embodiments, the generated raw acoustic feature frames are Mel-Frequency Cepstral Coefficients or log-filterbank features, as known in the art. The frame generation unit 110 can perform any suitable signal processing to determine the raw acoustic feature frames corresponding to each frame of the incoming audio signal. For example, the frame generation unit 110 may implement a custom-designed fast Fourier transform and custom filterbank parameters.

[0033] The batching unit 120 operates to batch the raw acoustic feature frames and provide the batch to the trained acoustic model 130. The trained acoustic model 130 can include any acoustic model capable of processing batches of input frames having various batch sizes. For example, the trained acoustic model 130 may include a one-directional long short-term memory model (LSTM), a bi-directional LSTM (BLSTM) model, or a layer trajectory (ltLSTM) model, but the embodiments are not limited thereto.

[0034] According to some embodiments described below, the trained acoustic model 130 can also process a batch of input frames including look-ahead frames. The trained acoustic model 130 may support only a fixed number of look-ahead frames (e.g., a contextual layer-trajectory LSTM (cltLSTM) model), or may support a batch including a varying number of look-ahead frames (e.g., a latency-controlled BLSTM (LC-BLSTM) model).

[0035] The decoder 140 forms word hypotheses based on the output of the trained acoustic model 130. According to some non-exhaustive examples, the model 130 outputs senone posteriors in response to the received batch of input frames. Based on the posterior hidden Markov model, the decoder 140, under a given input frame S = s0…s n outputs the most likely word sequence W ^ = w0...w m (W ^ represents Ŵ).

[0036] The batching unit 120 can determine the number of look-ahead frames and / or the batch size that serve as the criteria for collecting batches. As described herein, the determination may be based on the word hypotheses generated by the decoder 140 and / or the speech end state detected by other methods. Accordingly, the batching unit 120 may implement a function of determining the batch size (and, in some embodiments, the number of look-ahead frames) based on the received word hypotheses and / or speech end signals.

[0037] Figure 1B shows the ongoing operation of the system 200 according to this embodiment. As shown, the frame generation unit 110 has generated frame s1 based on a second (possibly overlapping) frame of the input audio signal. Assume that the initial batch size is set to 2.

[0038] As shown in FIG. 1A, the batching unit 120 has already received frame s0. Due to the initial batch size of 2, the batching unit 120 combines frames s0 and s1 into batch 150 and provides batch 150 to the trained acoustic model 130. The model 130 outputs the corresponding posterior probability based on the model architecture and the trained parameter values, which are known in the art. The posterior probability is received by the decoder 140, and the decoder 140 attempts to form a hypothesis.

[0039] According to some embodiments, the above process continues as described for subsequent frames of the audio signal. For example, FIG. 1C shows the generation of frame s 10 based on each frame of the audio signal. At this point, the batching unit 120 has already collected frames s8 and s9 into batch 152 and passes these frames to the model 130.

[0040] It is assumed that the decoder 140 generates a non - silent and non - noise hypothesis based on the batches received up to the current point. As shown, the hypothesis is returned to the batching unit 120. Based on the generation of the hypothesis, the batching unit 120 decides to increase the batch size.

[0041] Changing the batch size and / or determining a new batch size based on a hypothesis may be performed by a component different from the batching unit 120. Thus, such a component can send the new batch size to the batching unit 120, so that the batching unit 120 can adapt subsequent batch processing to the new batch size.

[0042] Determining a new batch size can be triggered by the initially generated word hypothesis, but embodiments are not limited thereto. The determination may be triggered by the nth generated word hypothesis, by the generation of one (or more) of a particular subset of word hypotheses, or by any other function of the word hypothesis output by the decoder 140.

[0043] FIG. 1D shows the implementation of a new batch size according to some embodiments. Although the new batch size is assumed to be 4, embodiments are not limited thereto. Thus, as shown in FIG. 1C, after the output of batch 152, the batching unit 120 collects frames s 10 through s 14 and puts them into batch 154 and provides batch 154 to model 130. In contrast, if the batch size had remained 2, the batching unit 120 would have collected only frames s10 and s11, put them into a batch, and provided that batch to model 130.

[0044] The illustrated process can continue using batches consisting of four frames until ASR ends (e.g., until the end of the audio signal is reached). According to some embodiments, the batch size is reduced to an initial batch size (e.g., 2) when silence is detected based on the output of the decoder 140, and is increased again as described above in response to word hypotheses.

[0045] In some embodiments, the batch size may gradually increase up to a target batch size. Regarding the above example, the batching unit 120 may increase the batch size from 2 to 3 over a specific period and then increase the batch size from 3 to 4. This increase may be achieved based on the creation of further hypotheses, instead of or in addition to the elapsed time.

[0046] FIG. 2 is a flowchart of a process 200 for providing dynamic batch processing according to some embodiments. Process 200 and other processes described herein can be executed using any suitable combination of hardware and software. The software program code embodying these processes is stored on any non-transitory tangible medium including a fixed disk, volatile or non-volatile random access memory, DVD, flash drive, or magnetic tape, and can be executed by any number of processing units including, but not limited to, processors, processor cores, and processor threads. Embodiments are not limited to the examples described herein.

[0047] Process 200 may be initiated in response to a request to perform ASR on an audio signal. Process 200 can be executed, for example, by a cloud-based ASR service that receives requests and audio signals from a cloud-based meeting provider. The audio signal may be received as a near real-time stream according to some embodiments. In some embodiments, process 200 operates continuously to process any signals present in the incoming audio channel.

[0048] The initial batch size for processing the audio signal is determined at S210. The initial batch size may be a predetermined experimentally determined value. The initial batch size may be selected from among several values based on a desired initial latency, where a lower latency corresponds to a smaller batch size. Referring to the examples of FIGS. 1A through 1D, an initial batch size of two frames is determined at S210.

[0049] The raw acoustic feature frames of the audio signal are collected at S220. The number of collected raw acoustic feature frames is equal to the determined initial batch size. Thus, in the above example, while the flow stops at S220, the batching unit 120 waits until it receives frames s0 and s1 from the frame generation unit 110.

[0050] The collected frames are input as a batch to the speech recognition network at S230. The speech recognition network can include any system capable of generating word hypotheses based on a batch of acoustic frames. The trained acoustic model 130 and decoder 140 of FIGS. 1A through 1D comprise a speech recognition network according to some embodiments.

[0051] Next, at S240, it is determined whether a word hypothesis has been generated by the speech recognition network. More specifically, S240 may include a determination of whether the speech recognition network has generated a non - silent, non - noise hypothesis based on the batch received up to the current point in time. If not, the flow returns to S220 to collect another batch of frames that matches the initial batch size. Thus, the flow circulates among S220, S230, and S240 until a word hypothesis is generated by the speech recognition network. As described above, the determination at S240 is not limited to the identification of the first generated word hypothesis.

[0052] Although shown as a linear flow, the determination at S240 may be made independently and in parallel to the flow circulation between S220 and S230. For example, the collection and input of a batch of frames may proceed until interrupted by a separate determination of the generated word hypotheses, after which the flow proceeds to S250.

[0053] At S250, the next batch size for processing the audio signal is determined. The next batch size may be predetermined based on experiments and / or may be selected from among several possible batch sizes. According to some embodiments, the next batch size is 80 frames.

[0054] As described with respect to FIG. 1D, the raw acoustic feature frames of the next batch size are collected at S260 and input to the speech recognition network at S270. According to some embodiments, the flow circulates between S260 and S270, collecting and inputting batches of frames, until silence is detected at S280. The silence detection at S280 may include detecting the absence of word hypotheses generated by the speech recognition network over a certain period of time and / or may be based on other heuristic methods.

[0055] According to the example shown, the detection of silence at S280 causes the process 200 to return to S210 and reset the batch size to the initial batch size. Thus, if the subsequent frames of the audio signal contain speech, the perceived latency for the user to process that speech can be made smaller than the latency resulting from a larger batch size.

[0056] Switching from an initial batch size to the next batch size can introduce significant latency into the process. In some embodiments, in response to the determination in S240, the batch size may be gradually increased to the target batch size. For example, the batch size may be increased by a predetermined amount for each of the n iterations of S260 until the target batch size is reached.

[0057] Process 200 may end when it reaches the end of the audio signal. In some embodiments, process 200 ends in response to detection of an end-of-speech state. FIGS. 3A and 3B illustrate such processing according to some embodiments.

[0058] System 300 adds a speech end detector 360 to the components of system 100 described above. Speech end detector 360 directly receives raw acoustic frames from frame generation unit 310 and determines based thereon whether speech has ended. Speech end detector 360 can include any system for detecting an end-of-speech state based on the input frames and can include one or more trained neural networks. In one example, detector 360 is an LSTM classifier formulated and trained independently of the decoder hypothesis and predicts the end of a query based only on past frames. Detector 360 may operate in relation to operational heuristics.

[0059] In an example of operation, components 310, 320, 330, and 340 operate as described above to modify the initial batch size to the next batch size of 4 frames. Thus, batching unit 320 collects batch 350 of frames s 10 to s 13 for input to model 330. FIG. 3A also shows frames s from frame generation unit 310 to batching unit 320 and detector 360 16shows the transmission. Therefore, the frame generation unit 310 has already generated frames s 14 and s 15 and these have been collected by the batching unit 320 in anticipation of generating the next batch of 4 frames.

[0060] The speech end detector 360 is assumed to detect the speech end state in response to the reception of frame s 16 . Accordingly, the speech end detector 360 notifies the detected state to the batching unit 320. Then, the batching unit 320 collects frame s 16 as the last frame of the current batch. FIG. 3B shows the above operation, where the batching unit 320 generates a batch 352 including frames s 14 , s 15 , s 16 and the batch 352 is input to the model 330. Therefore, the batch size of the last batch is reduced to avoid waiting for non-speech frames, and non-speech frames would have been added to the current batch of input frames if not as in this example. Thus, such an embodiment can support early withdrawal from batch processing and reduce latency in the generation of the final hypothesis. As described above, the final batch size may be reduced according to other conditions in some embodiments, and other conditions include, but are not limited to, detection of the end of the stream state and detection of the elapse of the maximum recognition time.

[0061] FIGS. 4A to 4C include a flowchart of a process 400 for providing dynamic batch processing according to some embodiments. The process 400 may be implemented by the system 300 of FIG. 3, but the embodiments are not limited thereto.

[0062] Steps S405, S410, S415, S420, and S425 can proceed in the same manner as steps S210, S220, S230, S240, and S250 of process 200, respectively. As described, an initial batch size for processing the incoming audio signal is determined at S405, and a corresponding batch of raw acoustic feature frames of the audio signal is collected at S410.

[0063] The batch of collected frames is input to the speech recognition network at S415. Then, at S420, it is determined whether an appropriate word hypothesis has been generated by the speech recognition network. If not, the flow returns to S410 to collect another batch of frames that matches the initial batch size. If so, the next batch size for processing the audio signal is determined at S425 as described above.

[0064] Next, at S430, it is determined whether an end-of-speech state has been detected based on the next raw acoustic feature frame. S430 may be performed by an end-of-speech detector that operates independently of the speech recognition network as described with respect to FIGS. 3A and 3B. Referring to FIG. 3A, S430 may include receiving a frame from frame generation unit 310 and evaluating an end-of-speech state based on the received frame (and, for example, based on previously received frames and / or other heuristic methods). If the end-of-speech state is not detected, the flow continues to S435, where the next raw acoustic frame is collected. The frame can be collected at S435 for inclusion within the next batch of frames. In this regard, at S440, it is determined whether a batch that conforms to the next batch size has been collected. If not, the flow returns to S430.

[0065] The determination in S430 is assumed to continue to be negative such that the flow between S430, S435, and S440 circulates until a desired number of frames are collected. Then, the flow proceeds to S445, where the collected frames of the current batch are input into the speech recognition network. If no silence is detected in S450, the flow returns to S430. Thus, the flow circulates between S430 and S450, collecting and inputting frame batches of the current batch size, until silence is detected at S450 or an end-of-speech state is detected at S430. Detection of silence at S430 causes the flow to return to S405 and reset the batch size to the initial batch size. As described above, when the next received frame of the audio signal contains speech, the latency perceived by the user when processing that speech can be made lower than the latency caused by a larger batch size.

[0066] During collection of the next batch size batch, if an end-of-speech state is detected at S430, the flow proceeds from S430 to S455. At S455, the next raw acoustic frame based on the detected end-of-speech state is collected. Referring to the examples of FIGS. 3A and 3B, an end-of-speech state is detected at S430 based on frame s 16 , and then frame s 16 is collected by the batching unit 320 as the last frame of the current batch at S455. Then, even if the size of the current batch is still smaller than the next batch size determined at S425, at S460, the collected frames (s 14 , s 15 , s 16 ) of the current batch are input into the model 330.

[0067] The determination in S420 may occur independently of and in parallel with the loop of the flow between S410 and S415. Similarly, the speech end determination in S430 may be performed independently of and in parallel with the cycle of the flow from S435 to S450, or alternatively.

[0068] Figures 5A to 5D show a look-ahead voice recognition system 500 using initial latency-sensitive adaptive batch processing during operation according to some embodiments. The system 500 includes a trained acoustic model 530 that supports a "look-ahead (or look-ahead)" input frame known in the art. The trained acoustic model 530 may include any model architecture that uses look-ahead frames, and the model architecture is known or becoming known, including but not limited to a latency-controlled bidirectional long short-term memory (LC-BLSTM) acoustic model and a context layer trajectory type long short-term memory (cltLSTM) acoustic model. As is known in the art, the cltLSTM model can support only a fixed number of look-ahead frames, while the LC-BLSTM model can support a batch including a variable number of look-ahead frames.

[0069] Figure 5A shows an example where the batching unit 520 collects a batch consisting of six raw acoustic feature frames, four of which (i.e., frames s2, s3, s4, s5) are look-ahead frames. In some embodiments, the batch size is defined as a window including a start frame and an end frame, and the look-ahead frames are determined by an overlap value. For example, the batch 550 may be defined by the window {s0, s5} and an overlap value of 4 (frames).

[0070] FIG. 5B shows collecting and inputting the next batch 552 according to this example. Batch 552 overlaps with the last 4 frames (s2, s3, s4, s5) of batch 550 and includes look-ahead frames s4, s5, s6, s7. Thus, the batching unit 520 must wait for frames s6 and s7 before generating batch 552.

[0071] As described above, the model 130 outputs a posterior probability corresponding to the received batch based on its architecture and trained parameter values, as is known in the art. The posterior probability is received by a decoder (not shown), and the decoder attempts to form a word hypothesis based on the received posterior probability.

[0072] Here, it is assumed that the decoder generates a non-silent, non-noise hypothesis based on the batch received up to the point shown in FIG. 5B. Based on the generation of the hypothesis, the batching unit 520 (or another component) determines to increase the batch size.

[0073] FIG. 5C shows the implementation of a new batch size according to some embodiments. According to the example of FIG. 5C, the number of look-ahead frames is fixed, but the embodiments are not limited thereto. Thus, this example can show the operation of using a cltLSTM model that uses a fixed number of look-ahead frames.

[0074] As shown in FIG. 5C, the new batch size is 8 frames. Thus, the next batch 554 includes 8 frames and overlaps with the last 4 frames of the preceding batch 552. However, the embodiments are not limited thereto. Batch 554 includes look-ahead frames s8, s9, s 10 , s 11 . Thus, as shown in FIG. 1B, after the output of batch 552, the batching unit 520 collects frames s8 to s 11 and adds them to the already received frames s4 to s7 to generate batch 554.

[0075] FIG. 5D shows operations based on a modified batch size. The batching unit 520 collects frames s 12 through s 15 and adds these frames to the already received frames s8 through s 11 to generate a batch 556. The batch 556 has a batch size of 8 frames and includes four look-ahead frames.

[0076] As described above, the batch size may be gradually increased to a target batch size in some embodiments. Such an increase may be, for example, in response to the creation of further hypotheses and / or the elapsed time since the first hypothesis.

[0077] FIGS. 5A through 5D show operations of changing the batch size when the number of look-ahead frames remains fixed. In contrast, FIGS. 6A through 6D show a look-ahead responsive speech recognition system 600 that uses an initial latency-sensitive adaptive look-ahead.

[0078] The frame generation unit 610 receives an audio signal and generates a frame-level amplitude spectrum or raw acoustic feature frame corresponding to each frame of the audio signal. The batching unit 620 collects the generated raw acoustic feature frames into a batch of an initial specified batch size (or window) having a specified number of look-ahead frames (or specified overlap number). The trained acoustic model 630 receives the batch and outputs, for example, a senone posterior probability, which is used by the decoder 640 to generate word hypotheses. The trained acoustic model 630 supports batches containing different numbers of look-ahead frames and thus may include an LC-BLSTM model in some embodiments.

[0079] In the example of FIG. 6A, the initial batch size is 3, and each batch contains two look-ahead frames (i.e., frames s1 and s2 in batch 650). FIG. 6B shows the output of the next batch 652, which contains three frames and overlaps with two look-ahead frames in batch 650. Here, it is assumed that the decoder 640 generates non-silent, non-noise hypotheses based on the batches received up to the current point in time.

[0080] In response, as shown in FIGS. 6C and 6D, the batch size is changed to 4, and the number of look-ahead frames per batch is changed to 3. Thus, batch 654 overlaps with look-ahead frames s2 and s3 of batch 652 and contains four frames. The four frames of batch 654 include three look-ahead frames s3, s4, s5. Currently, considering the changed batch size and the number of look-ahead frames, the batching unit 620 collects the next batch 656 that contains four frames. The four frames of batch 656 include frames that overlap with look-ahead frames s3, s4, s5 of batch 654, and also include look-ahead frames s4, s5, s6.

[0081] Accordingly, the embodiment can increase or decrease the batch size and / or the number of look-ahead frames per batch. The increase or decrease for any parameter may occur once or gradually according to the detected hypothesis.

[0082] According to some embodiments, the acoustic model 630 and the decoder 640 are trained using the batch sizes and the number of look-ahead frames of FIGS. 6C and 6D, and thus operate optimally after detection of the first word hypothesis. Thus, as shown in FIGS. 6A and 6B, when operating in relation to the initial batch size and the initial number of look-ahead frames, the accuracy of the acoustic model 630 and the decoder 640 may not be optimal as a trade-off for the user to perceive only a shorter latency.

[0083] FIGS. 7A and 7B include a flowchart of a process 700 for providing dynamic batching according to some embodiments. The process 700 may be implemented by the system 600, but the embodiments are not limited thereto.

[0084] An initial batch size for processing the input audio signal is determined at S710. Also at S710, the initial number of look-ahead frames included in each batch is determined. The initial batch size may be specified as a window size, and the number of look-ahead frames may be defined as a specific fraction (e.g., 1 / 2) of the window size.

[0085] Raw acoustic feature frames of the audio signal are collected at S720 based on the initial batch size and the initial number of look-ahead frames. As shown in the example of FIG. 6A, assuming an initial batch size of 3 frames, while the flow remains at S720, the batching unit 620 waits until it receives frames s0, s1, s2 from the frame generation unit 610.

[0086] The collected frames are input as a batch to the speech recognition network at S730. The speech recognition network may include any system capable of generating word hypotheses based on a batch of acoustic frames including look-ahead frames.

[0087] Next, in S740, raw acoustic feature frames of the audio signal are collected based on the initial batch size and the initial number of look-ahead frames. The frames collected in S740 overlap with the look-ahead frames of the previously collected frames. Referring to FIG. 6B, batch 652 can be collected in S740, which consists of three frames and overlaps with the look-ahead frames s1 and s2 of batch 650. The frames collected in S740 are input as a batch to the speech recognition network in S750.

[0088] Collecting frames that overlap with the look-ahead frames of the previously collected frames can include setting the start of the window of the current batch of frames to the start of the overlapping frames of the previous batch. For example, if the overlap is half of the window size, the window of the current batch of frames is set to start from the center of the window of the previous batch.

[0089] S760 includes determining whether a word hypothesis has been generated by the speech recognition network and may be implemented as described above. If the word hypothesis has not been generated yet, the flow returns to S740 to collect another batch of frames that matches the initial batch size, includes look-ahead frames, and overlaps with the look-ahead frames of the previous batch. Thus, the flow circulates among S740, S750, and S760 until a word hypothesis is generated by the speech recognition network.

[0090] If the determination at S760 is affirmative, the next batch size for processing the audio signal is determined at S770. The determination of the next batch size may include determining the next number of look-ahead frames. Accordingly, S770 may include determining a new batch size per batch and / or a new number of look-ahead frames. In some embodiments, S770 includes determining the next window size, and the next look-ahead size is determined as a function of the next window size. As described above, the new number of look-ahead frames can be determined only if the process 700 is executed in connection with an acoustic recognition network capable of processing batches containing any number of look-ahead frames.

[0091] As described in connection with FIGS. 6C and 6D, the raw acoustic feature frames of the next batch size that overlap with the previous look-ahead frame are collected at S775 and input to the acoustic recognition network at S780. The flow circulates between S775 and S780 until silence is detected at S785, collecting and inputting batches of frames. The flow then returns to S710, resets the batch size and the number of look-ahead frames to their initial values, and then reduces the latency perceived by the user when processing the received speech.

[0092] As described in connection with process 400, embodiments of process 700 may further incorporate a speech end detector for independently detecting a speech end state. Based on the detection, the size of the final batch may be reduced (e.g., the window is shortened) and may include only those up to the speech end frame.

[0093] Neural networks (e.g., deep learning, deep convolutional, or recurrent types) according to some embodiments include a series of "neurons" configured in the network, such as LSTM nodes. Neurons are architectures used in data processing and artificial intelligence, particularly machine learning, and include memory capable of determining when to "recall" and when to "forget" based on the weights of the inputs given to a given neuron. Each of the neurons used herein is configured to receive a predetermined number of inputs from other neurons within the network and provide relational and sub-relational outputs regarding the content of the analyzed frame. Individual neurons can be connected together and / or organized in a tree structure in various configurations of the neural network to provide a learning model of how each frame during speech is related to each other in terms of interaction and involvement.

[0094] For example, an LSTM functioning as a neuron includes several gates for processing an input vector, a memory cell, and an output vector. The input gate and the output gate respectively control the information flowing into and out of the memory cell, while the forget gate selectively removes information from the memory cell based on the input from cells previously linked in the neural network. The weights and bias vectors of the various gates are adjusted during the training phase, and once the training phase is completed, those weights and biases are finalized for normal operation. Neurons and neural networks may be constructed programmatically (e.g., via software instructions) or by special hardware that links each neuron to form a neural network.

[0095] FIG. 8 shows a distributed transcription system 800 according to some embodiments. System 800 may be cloud-based, and its components may be implemented using on-demand virtual machines, virtual servers, and cloud storage instances.

[0096] The speech-to-text service 810 may be implemented as a cloud service that provides a transcription of a speech audio signal received via the cloud 820. The speech-to-text service 810 includes components of a batching unit, an ASR model, and a decoder that operate as described above, and provides ASR using latency-sensitive adaptive batching as described in this case.

[0097] Each of the client devices 830, 832 can be operated to request services such as a search service 840 and a voice assistant service 850. Then, the services 840, 850 can request a speech-to-text function from the speech-to-text service 810. According to some embodiments, a client device (e.g., a smart speaker, a smartphone) includes all components of a system such as the system 100 or the system 300, and the operation does not require communication with a cloud service. In some embodiments, the client device includes at least one of a component for frame generation, a component of a batching unit, an ASR model, a decoder, and optionally a speech end detector, and the client device calls a remote system (e.g., a cloud service) to access the functions of any of these components not implemented in the client device.

[0098] FIG. 9 is a block diagram of a system 900 according to some embodiments. The system 900 may include a general-purpose computing device and may execute program code to provide ASR using latency-sensitive adaptive batching as described in this case. According to some embodiments, the system 900 may be implemented by a stand-alone device such as a personal computer, a smart speaker, and a smartphone, or a cloud-based virtual server.

[0099] System 900 includes a processing unit 910 operably coupled to a communication device 920, a persistent data storage system 930, one or more input devices 940, one or more output devices 950, and a memory 960. The processing unit 910 may include one or more processors, processing cores, etc. for executing program code. The communication interface 920 is capable of facilitating communication with an external network such as the Internet. The input device 940 may include, for example, a keyboard, keypad, mouse or other pointing device, microphone, touch screen, and / or eye-tracking device. The output device 950 may include, for example, a display (e.g., a display screen), speaker, and / or printer.

[0100] The data storage system 930 may include any number of suitable persistent storage devices, including combinations of magnetic storage devices (e.g., magnetic tape, hard disk drive, and flash memory), optical storage devices, read-only memory (ROM) devices, etc. The memory 960 may include random access memory (RAM), storage class memory (SCM), or any other high-speed access memory.

[0101] The batching unit 932 and the decoder 936 may include program code executable to provide the functions as described in this case. The batching unit 932 and / or the decoder 936 may include one or more trained neural networks according to some embodiments. The node operator library 934 may include program code executable to provide an acoustic model defined by hyperparameters 938 and trained parameter values 939, as known in the art. Also, the data storage device 930 may store data and other program code to provide additional functions such as device drivers, operating system files, etc., and / or necessary for the operation of the system 900.

[0102] The individual elements and processes of each function described in this case can be implemented, at least in part, in computer hardware, in program code, and / or in one or more computing systems that execute the program code, as known in the art. Such a computing system can include one or more processing units that execute processor-executable program code stored in a memory system.

[0103] The processor-executable program code embodying the processes described above is stored by any non-transitory tangible medium including a fixed disk, volatile or non-volatile random access memory, DVD, flash drive, or magnetic tape, and can be executed by any number of processing units including, but not limited to, processors, processor cores, and processor threads. Embodiments are not limited to the examples described below.

[0104] The foregoing figures represent a logical architecture for explaining a system according to some embodiments, and actual embodiments may include more or different components arranged in other ways. Other topologies may be used with other embodiments. Further, each component or device described herein may be implemented by any number of devices communicating via any number of other public and / or private networks. Two or more of such computing devices may be located remotely from each other and be capable of communicating with each other via any known method of network and / or dedicated connection. Each component or device may include any number of hardware and / or software elements suitable for providing the functions described herein and any other functions. For example, any computing device used in the implementation of a system according to some embodiments may include a processor for executing program code, such that the computing device operates as described herein.

[0105] The figures described herein do not imply a fixed order to the illustrated methods, and embodiments may be implemented in any order that is executable. Further, any of the methods described herein may be implemented by hardware, software, or any combination of these approaches. For example, a computer-readable storage medium can store instructions that, when executed by a machine, result in performance of a method according to any of the embodiments described herein.

[0106] Those skilled in the art will understand that various adaptations and modifications of the above-described embodiments can be constructed without departing from the claims. Accordingly, it should be understood that the claims may be practiced otherwise than as specifically set forth herein.

Claims

1. A system comprising a processing unit and a storage device containing program code, wherein when the program code is executed by the processing unit, the system is caused to: Determine a first batch size for processing an audio signal; Collect a first batch containing a first number of raw acoustic feature frames of the audio signal, wherein the first number is equal to the first batch size; Input the first batch into an automatic speech recognition network; Determine a second batch size based on a word hypothesis output by the automatic speech recognition network, wherein the second batch size is larger than the first batch size; Collect a second batch containing a second number of acoustic feature frames of the audio signal, wherein the second number is equal to the second batch size; and Input the second batch into the automatic speech recognition network; A system that causes the above to be executed.

2. The system according to claim 1, wherein determining the second batch size based on the word hypothesis includes detecting a first non - silent word hypothesis output by the automatic speech recognition network and determining the second batch size in response to the detection.

3. The system according to claim 1, wherein when the program code is executed by the processing unit, the system is caused to: Collect a plurality of batches, each of the plurality of batches containing the first number of acoustic feature frames of the audio signal; Input the plurality of batches into the automatic speech recognition network until the word hypothesis is output by the automatic speech recognition network; A system that causes the above to be executed.

4. In the system according to claim 1, when the program code is executed by the processing unit, the system is caused to: Detect a speech end frame; and Reduce the size of the next batch of acoustic feature frames so as to include the speech end frame as the last frame of the next batch to be collected; A system that causes the above to be executed.

5. In the system according to claim 1, the first batch size includes a first number of look-ahead frames, the first batch includes a first look-ahead frame, the number of first look-ahead frames in the first batch is equal to the first number of look-ahead frames, and when the program code is executed by the processing unit, the system is caused to: Execute a step of collecting a third batch including a first number of acoustic feature frames of the audio signal, and a period associated with a plurality of frames in the third batch overlaps with a period associated with the first look-ahead frame in the first batch; Determining the second batch size includes determining a second number of look-ahead frames; and The second batch includes a second look-ahead frame, and the number of second look-ahead frames in the second batch is equal to the second number of look-ahead frames. A system.

6. The system according to claim 5, wherein the speech recognition network includes a latency-controlled bidirectional long-short term memory module.

7. In the system according to claim 1, the first batch size includes a first number of look-ahead frames, the first batch includes a first look-ahead frame, the number of first look-ahead frames in the first batch is equal to the first number of look-ahead frames, and when the program code is executed by the processing unit, the system is caused to: execute a step of collecting a third batch including a first number of acoustic feature frames of the audio signal, and a period associated with a plurality of frames in the third batch overlaps with a period associated with the first look-ahead frame of the first batch; The second batch includes a second look-ahead frame, and the number of second look-ahead frames in the second batch is equal to the first number of look-ahead frames.

8. In the system according to claim 7, the speech recognition network includes a context layer trajectory type long short-term memory module.

9. A method executed by a computer, comprising: a step of collecting a first batch of acoustic feature frames of an audio signal, wherein the number of acoustic feature frames in the first batch is equal to a first batch size; a step of inputting the first batch into a speech recognition network; a step of collecting a second batch of acoustic feature frames of the audio signal in response to detection of a word hypothesis output by the speech recognition network, wherein the number of acoustic feature frames in the second batch is equal to a second batch size, and the second batch size is larger than the first batch size; and a step of inputting the second batch into the speech recognition network. A method including the above steps.

10. The method according to claim 9, wherein the detection of the word hypothesis includes the detection of a first non-silent word hypothesis output by the speech recognition network.

11. In the method according to claim 9, the step of detecting a speech end frame; a step of collecting a third batch of acoustic feature frames of the audio signal in response to detecting the speech end frame, the third batch being collected such that the speech end frame is included as the last frame of the third batch; and the step of inputting the third batch into the speech recognition network; A method further comprising.

12. In the method according to claim 9, the first batch size includes a first number of look-ahead frames, the first batch includes a first look-ahead frame, the number of first look-ahead frames in the first batch is equal to the first number of look-ahead frames, and the method comprises: executing a step of collecting a third batch including a first number of acoustic feature frames of the audio signal, a period associated with a plurality of frames of the third batch overlapping a period associated with the first look-ahead frame of the first batch; The second batch includes a second look-ahead frame, and the number of second look-ahead frames in the second batch is equal to the second number of look-ahead frames.

13. The method according to claim 12, wherein the speech recognition network includes a latency-controlled bidirectional long-short term memory module.

14. In the method according to claim 9, the first batch size includes a first number of look-ahead frames, the first batch includes a first look-ahead frame, the number of the first look-ahead frames in the first batch is equal to the first number of the look-ahead frames, and the method comprises: further comprising collecting a third batch including a first number of acoustic feature frames of the audio signal, wherein a period associated with a plurality of frames of the third batch overlaps with a period associated with the first look-ahead frame of the first batch; wherein the second batch includes a second look-ahead frame, and the number of the second look-ahead frames in the second batch is equal to the first number of the look-ahead frames.

15. In the method according to claim 14, the speech recognition network includes a context layer trajectory type long-short term memory module.

Citation Information

Patent Citations

  • Interactive device

    JP2012226068A

  • Speech recognition system, speech recognition method, and speech recognition program

    WO2006075648A1