Speech recognition two-stage decoding acceleration method based on acoustic clustering

By adopting a two-stage decoding acceleration method based on acoustic clustering in the automatic speech recognition system, the size of the decoding vocabulary is reduced, and the problem that the system's decoding speed is difficult to meet real-time response is solved, and a significant acceleration effect is achieved.

CN120220679APending Publication Date: 2025-06-27XIAONIU FANYI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375802.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing automatic speech recognition system based on deep learning is difficult to meet the demand for real-time response in terms of decoding speed, mainly due to the huge decoding vocabulary, the inference speed is limited.

Method used

The two-stage decoding acceleration method of speech recognition based on acoustic clustering is adopted. By extracting the acoustic information sequence of audio, binary data of is constructed, text-to-sound unit mapping model is trained, and clustered using the KMeans method to obtain a sub-vocabulary set, thereby reducing the size of the target decoded vocabulary list.

Benefits of technology

While ensuring model performance, the decoding speed is significantly improved, an acceleration ratio of nearly 10% on average is obtained, and orthogonal to other acceleration methods, which can be used simultaneously to achieve a acceleration ratio of 90%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220679A_ABST
    Figure CN120220679A_ABST
Patent Text Reader

Abstract

The invention discloses a speech recognition two-stage decoding acceleration method based on acoustic clustering, and the method comprises the steps: obtaining an acoustic information sequence corresponding to an audio according to a pre-trained sound unit extraction model, and constructing lt; a text and acoustic information sequence gt; training a text-to-sound unit mapping model according to the binary data; converting the text into a corresponding acoustic information sequence, and clustering by using a KMeans method to obtain a sub-word list set; constructing an automatic speech recognition model, screening speech recognition training data from audio to text, and extracting an audio file into an fbank feature sequence; performing first-stage decoding to obtain a corresponding target sub-word list; and according to the target sub-word list, calculating probability distribution under the sub-word list in second-stage decoding, and selecting a word with the highest probability as an identification result. According to the method, on the basis of the latest implementation of rapid reasoning, 1.08 times of speed-up ratio can be continuously obtained, and meanwhile, the model performance is almost not reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for accelerating speech recognition decoding, specifically a two-stage decoding acceleration method for speech recognition based on acoustic clustering. Background Art

[0002] Generally speaking, automatic speech recognition (ASR) is a process of converting the speech content of a language into the corresponding text output by a computer. As an important branch of natural language processing (NLP) and artificial intelligence (AI), ASR has a wide range of requirements in application scenarios such as voice assistants, real-time captions, meeting records, and voice input. For example, major technology companies have launched ASR-related products, such as Apple's Siri, Google's Google Assistant, and Amazon's Alexa.

[0003] Looking at the development history of ASR, its methods can be roughly divided into two categories: rule-based speech recognition and data-driven speech recognition. Specifically, data-driven ASR methods can be further divided into statistical-based methods and deep learning-based methods. Early ASR mainly relied on manually defined speech rules for recognition. From the 1980s to the 1990s, the hidden Markov model (HMM) became the mainstream technology in the ASR field. The combination of HMM and Gaussian mixture model (HMM-GMM) greatly improved the accuracy of speech recognition, enabling ASR to enter the commercial application stage. However, statistical methods still rely on a large amount of feature engineering and assume that the speech signal has a certain implicit structure, resulting in limited performance in complex speech environments. In addition, the HMM-GMM method has insufficient modeling ability for the mutual dependence of long speech sequences and is difficult to effectively process the context information in speech. Furthermore, researchers have proposed automatic speech recognition methods based on deep learning, which directly model speech recognition with neural networks, and the model learning is completed in an end-to-end manner without the need for the design of artificial features.

[0004] Compared with traditional statistical-based speech recognition methods, automatic speech recognition systems based on deep learning have attracted the attention of many researchers due to their high recognition quality. However, due to the characteristics of neural networks themselves, they are more time-consuming in the process of use. This problem is particularly prominent in practical speech recognition systems because they generally have more stringent requirements for response time. Therefore, the decoding speed of speech recognition systems has also become the key to their practicality, and whether it is possible to optimize the speed on the basis of existing deep learning-based automatic speech recognition systems has become an extremely important topic.

[0005] Since the neural network based on deep learning involves a large number of matrix operations and takes up a lot of decoding time, people have begun to try methods such as knowledge distillation and attention acceleration to optimize the efficiency. Existing automatic speech recognition systems based on self-attention mechanisms abandon the use of traditional neural network structures (such as recurrent neural networks, convolutional neural networks, etc.). In their structures, except for simple feed-forward networks, almost all model the transformation of sequences through attention mechanisms. Inside the encoder and decoder, the audio and text information are respectively modeled through the self-attention mechanism, and this part of the overhead becomes the primary difficulty in solving model acceleration. However, after using methods such as attention acceleration and knowledge distillation, the overhead of attention calculation is significantly reduced. Instead, the huge decoding vocabulary severely limits the overall inference speed of the model, accounting for more than 30% of the overall inference time.

[0006] It can be seen that the size of the decoding vocabulary is still the key factor restricting the further improvement of the decoding speed. Since the automatic speech recognition model based on deep learning needs to perform decoding prediction in an overall overly large vocabulary space during the calculation process, the overly large vocabulary search space greatly reduces the decoding efficiency, resulting in the decoding speed of this automatic speech recognition method being difficult to meet the requirements of real-time response in actual use. How to compress the decoding vocabulary and improve the decoding efficiency has become the key issue for the low-latency implementation of automatic speech recognition. Summary of the Invention

[0007] Aiming at the deficiencies that the decoding speed of the existing automatic speech recognition method is difficult to meet the requirements of real-time response in actual use, etc., the technical problem to be solved by the present invention is to provide a two-stage decoding acceleration method for speech recognition based on acoustic clustering, which can improve the real-time response speed on the basis of the latest implementation of fast inference and with almost no decrease in model performance.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is:

[0009] The present invention provides a two-stage decoding acceleration method for speech recognition based on acoustic clustering, including the following steps:

[0010] 1) Obtain the acoustic information sequence corresponding to the audio according to the pre-trained sound unit extraction model, construct the binary data of <text, acoustic information sequence>, and use the binary data to train the text-to-sound unit mapping model;

[0011] 2) Based on the sound unit mapping model, transform the text into the corresponding acoustic information sequence and use the KMeans method for clustering to obtain the sub-vocabulary set;

[0012] 3) Build an automatic speech recognition model, screen the speech recognition training data from audio to text, and extract the audio files into fbank feature sequences for training the automatic speech recognition model;

[0013] 4) According to the output of the decoding layer of the automatic speech recognition model in step 3), perform the first-stage decoding to obtain the corresponding target sub-word table;

[0014] 5) According to the target sub-word table predicted by the first-stage decoding in step 4), calculate the probability distribution under this sub-word table in the second-stage decoding, and select the word with the highest probability as the recognition result.

[0015] In step 1), according to the pre-trained voice unit extraction model, obtain the acoustic information sequence corresponding to the audio, and construct the binary data of <text, acoustic information sequence>, and use the binary data to train the text-to-voice unit mapping model. Specifically:

[0016] 101) For each audio data, use the pre-trained voice unit extraction model to obtain its corresponding acoustic information sequence;

[0017] 102) Combine the text with the extracted acoustic information sequence to obtain the binary training data of <text, acoustic information sequence>;

[0018] 103) Use the standard Transformer, adopt the Encoder-Decoder architecture setting, and train according to the text-to-acoustic unit mapping training data to obtain the text-to-voice unit mapping model of text-to-acoustic information.

[0019] In step 3), build an automatic speech recognition model, screen the speech recognition training data from audio to text, and extract the audio files into fbank feature sequences for training the automatic speech recognition model. The specific steps are as follows:

[0020] 301) Adopt the Conformer-Encoder architecture and the Transformer-Decoder architecture to build the framework of the speech recognition model;

[0021] 302) Filter and clean the speech recognition training data from audio to text, screen out high-quality training data, and extract the audio files as fbank feature representations;

[0022] 303) Use the SpecAug method to perform data augmentation on the audio data, and use the convolutional layer for downsampling to shorten the length of the audio data to ensure the robustness and stability of training;

[0023] 304) Input the extracted audio fbank features into the speech recognition model for training to obtain the output result of the decoding layer of the automatic speech recognition model.

[0024] In step 4), according to the output of the decoding layer of the automatic speech recognition model in step 3), the first-stage decoding is performed to obtain the corresponding target sub-word list. First, the output result of the decoding layer of the automatic speech recognition model is obtained, and then, using the clustered sub-word list set in step 2) as the classification target, the Softmax method is used to predict the most likely target sub-word list. Specifically:

[0025] 401) According to the clustered sub-word list set, set the number of all target sub-word lists for the first-stage decoding classification prediction, and construct the corresponding Softmax layer;

[0026] 402) According to the output result of the decoding layer of the automatic speech recognition model, use the Softmax method for classification prediction to obtain the most likely target sub-word list. The specific formula is:

[0027]

[0028] where p n is the normalized prediction probability value of the nth target sub-word list; n is the current target sub-word list serial number; N is the number of all target sub-word lists; i = [1, N] represents any sub-word list serial number; c i is the output probability of any target sub-word list in the automatic speech recognition decoding layer; c n is the output probability of the nth target sub-word list in the automatic speech recognition decoding layer; τ is the temperature parameter used to control the transition of the distribution from smooth to sharp; g i is the additional noise corresponding to any sub-word list; g n is the additional noise corresponding to the target sub-word list.

[0029] In step 5), the second-stage decoding process decodes according to the obtained target sub-word list, calculates the probability distribution of each word under this sub-word list, and selects the word with the highest probability as the recognition result. The calculation formula is:

[0030]

[0031] where p n,k is the normalized prediction probability of the target word; V n is the size of the number of words in the nth sub-word list; k is the position of the target word in this sub-word list; l k is the probability corresponding to the target word; i is any word in the target sub-word list; l i is the probability of any word in the target sub-word list; T is the temperature parameter used to control the probability smoothness.

[0032] The present invention has the following beneficial effects and advantages:

[0033] 1. The present invention proposes a two-stage decoding acceleration method for speech recognition based on acoustic clustering, which can greatly improve the efficiency of the system during the inference process by reducing the size of the target decoding vocabulary. This method achieves an average acceleration ratio of nearly 10% in inference speed, while the model performance does not decline. It is orthogonal to other acceleration methods and can be used simultaneously to achieve an acceleration ratio of 90%. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the mapping from text to acoustic information sequence in the present invention;

[0035] Figure 2 It is a schematic diagram of clustering the vocabulary text according to acoustic information in the present invention;

[0036] Figure 3 It is a schematic diagram of the automatic speech recognition model architecture in the present invention;

[0037] Figure 4 It is a schematic diagram of the two-stage decoding acceleration method in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The present invention will be further described below with reference to the accompanying drawings of the specification.

[0039] The present invention provides a two-stage decoding acceleration method for speech recognition based on acoustic clustering, including the following steps:

[0040] 1) Obtain the acoustic information sequence corresponding to the audio according to the pre-trained sound unit extraction model, construct the binary data of <text, acoustic information sequence>, and use the binary data to train the text-to-sound unit mapping model;

[0041] 2) Based on the sound unit mapping model, transform the text into the corresponding acoustic information sequence and perform clustering using the KMeans method to obtain the sub-vocabulary set;

[0042] 3) Construct an automatic speech recognition model, screen the speech recognition training data from audio to text, and extract the audio file into an fbank feature sequence for training the automatic speech recognition model;

[0043] 4) According to the output of the decoding layer of the automatic speech recognition model in step 3), perform the first-stage decoding to obtain the corresponding target sub-vocabulary;

[0044] 5) According to the target sub-vocabulary predicted by the first-stage decoding in step 4), calculate the probability distribution under this sub-vocabulary in the second-stage decoding, and select the word with the highest probability as the recognition result.

[0045] The method of the present invention optimizes the decoding speed of an automatic speech recognition system based on a deep neural network from the perspective of decoding vocabulary compression, aiming to significantly improve the decoding speed of the automatic speech recognition system at the cost of a small performance loss, achieving a balance between performance and speed.

[0046] In step 1), an acoustic information sequence corresponding to the audio is obtained according to a pre-trained sound unit extraction model, and a binary data of <text, acoustic information sequence> is constructed, and the text-to-sound unit mapping model is trained using the binary data. Specifically:

[0047] 101) For each audio data, use the pre-trained sound unit extraction model to obtain its corresponding acoustic information sequence; such as the Hubert model, which is trained in an unsupervised manner and can map an audio data to a corresponding Unit sequence as its acoustic information unit. For example, “It’s a good day.” can be mapped to “001_123_342_012_132_067”;

[0048] 102) Combine the text with the extracted acoustic information sequence to obtain the binary training data of <text, acoustic information sequence>; such as <“It’s a good day.”, “001_123_342_012_132_067”>;

[0049] 103) Use a standard Transformer, adopt the Encoder-Decoder architecture setting, and train according to the text-to-acoustic unit mapping training data to obtain the text-to-sound unit mapping model from text to acoustic information; in this embodiment, 6 layers of Encoder and 6 layers of Decoder are set, and the constructed mapping training data from text to acoustic information is used for training to obtain the mapping model from text to acoustic information, so as to realize converting each text token in the vocabulary into its corresponding acoustic unit sequence, as Figure 1 shown.

[0050] In step 2), use the mapping model to map the text in the vocabulary to the corresponding acoustic information sequence, embed the acoustic information sequences of all vocabulary texts into Embedding representations, and then use the Kmeans method for clustering to obtain the set of all target sub-vocabularies, as Figure 2 shown.

[0051] 3) Construct an automatic speech recognition model, screen the speech recognition training data from audio to text, and extract the audio file into an fbank feature sequence for training the automatic speech recognition model;

[0052] In this step, a standard speech recognition model is constructed based on the Conformer-Encoder architecture and the Transformer-Decoder. The speech recognition training data from audio to text is filtered, and the audio files are extracted into 80-dimensional fbank feature sequences for training. Specifically:

[0053] 301) Use a 12-layer Conformer-Encoder architecture and a 6-layer Transformer-Decoder architecture to build the framework of the speech recognition model, as Figure 3 shown;

[0054] 302) Filter and clean the speech recognition training data from audio to text, select high-quality training data, and extract the audio files into 80-dimensional fbank feature representations;

[0055] 303) Use the SpecAug method to perform data augmentation on the audio data, and use convolutional layers for downsampling to shorten the length of the audio data to ensure the robustness and stability of training;

[0056] In this step, the speech recognition training data from audio to text is filtered and cleaned to remove data with abnormal ratios of audio to text lengths, such as data with a ratio greater than 5, and audio data with too long or too short durations, such as data less than 1 second or exceeding 30 seconds, so as to select high-quality training data. Then, the audio files are extracted into 80-dimensional fbank feature representations for training;

[0057] 304) Input the extracted audio fbank features into the speech recognition model for training, and obtain the output result of the last layer Decoder of the automatic speech recognition model for decoding and prediction.

[0058] In step 4), according to the output of the decoding layer of the automatic speech recognition model in step 3), the first-stage decoding is performed to obtain the corresponding target sub-word list. First, obtain the output result of the decoding layer of the automatic speech recognition model, and then use the clustered sub-word list set in step 2) as the classification target, and use the Softmax method to predict the most likely target sub-word list, as Figure 4 shown.

[0059] In this step, according to the output result of the last layer Decoder of the automatic speech recognition model, the first-stage decoding is performed. The clustered sub-word list set is used as the classification target, and Gumbel Softmax is used for classification prediction to obtain the most likely target sub-word list. Specifically:

[0060] 401) Set the number of all target sub - vocabulary lists for the first - stage decoding classification prediction according to the clustered sub - vocabulary list set; construct the corresponding Softmax layer. For example, if there are 10 candidate target sub - vocabulary lists in the first stage, set the output dimension size of the classification linear layer of Softmax to 10;

[0061] 402) According to the output result of the decoding layer of the automatic speech recognition model, use the Softmax method for classification prediction to obtain the most likely target sub - vocabulary list. The specific formula is:

[0062]

[0063] where, p n is the normalized prediction probability value of the nth target sub - vocabulary list; n is the current target sub - vocabulary list serial number; N is the number of all target sub - vocabulary lists; i = [1, N] represents any sub - vocabulary list serial number; c i is the output probability of any target sub - vocabulary list in the automatic speech recognition decoding layer; c n is the output probability of the nth target sub - vocabulary list in the automatic speech recognition decoding layer; τ is the temperature parameter used to control the transition of the distribution from smooth to sharp; g i is the additional noise corresponding to any sub - vocabulary list; g n is the additional noise corresponding to the target sub - vocabulary list.

[0064] The second - stage decoding process in step 5) decodes according to the obtained target sub - vocabulary list, calculates the probability distribution of each word under this sub - vocabulary list, and selects the word with the highest probability as the recognition result, as Figure 4 shown; if the third sub - vocabulary list is predicted as the target sub - vocabulary list in the first stage, then select this sub - vocabulary list as the vocabulary list for decoding in the second stage, calculate the probability distribution of each candidate word in this sub - vocabulary list through the Softmax function, and take the word with the highest probability as the final result of this decoding process; its calculation formula is:

[0065]

[0066] where, p n,k is the normalized prediction probability of the target word; V n is the size of the number of words in the nth sub - vocabulary list; k is the position of the target word in this sub - vocabulary list; l k is the probability corresponding to the target word; i is any word in the target sub - vocabulary list; l i is the probability of any word in the target sub - vocabulary list; T is the temperature parameter used to control the probability smoothness.

[0067] In this embodiment, the two-stage decoding acceleration method proposed by the present invention is verified through an English automatic speech recognition task. The open English automatic speech recognition dataset LibriSpeech is used as training data, with a total of 960 hours of audio data and corresponding texts. The audio data is extracted to obtain 80-dimensional fbank feature data. At the same time, the text data is segmented into sub-words to construct an overall vocabulary with a size of 10,000. Then, the size of the sub-vocabulary set is set to 10, and the Kmeans clustering method is used to cluster the original sub-table to obtain 10 different target sub-vocabularies.

[0068] By reducing the size of the target decoding vocabulary, the efficiency of the system during the inference process can be greatly improved. As shown in Table 1, in this embodiment, the clean and other datasets in LibriSpeech are used for performance testing, comparing the performance and efficiency of the two-stage decoding method in the present invention with the traditional direct decoding method. While ensuring the performance of the automatic speech recognition model on datasets of different sizes, significant efficiency improvements can be achieved. The base model is a single-stage decoding model, and base+2stage is the model using the two-stage decoding method of the present invention. It is found that the overall FLOPs overhead is reduced by 57M, the decoding speed is increased by 5 tokens / s, and an average acceleration ratio of nearly 8% is obtained in terms of the acceleration rate. At the same time, the model performance does not decline. Moreover, the method of the present invention is orthogonal to other acceleration methods. Especially when combined with other model acceleration methods, such as Flash-attention and knowledge distillation, the model acceleration efficiency can be further improved. small * is a single-stage decoding model using Flash-attention and knowledge distillation based on base, small * +2stage is the two-stage decoding model using the present invention. Compared with base, the overall FLOPs overhead is reduced by 367M, the decoding speed is increased by 58 tokens / s, and an average acceleration ratio of nearly 97% is obtained in terms of the acceleration rate. This can bring greater benefits in automatic speech recognition scenarios with high real-time requirements.

[0069] Table 1: Comparison results of acceleration rates between the two-stage decoding method and the single-stage decoding

[0070]

Claims

1. A two-stage decoding acceleration method for speech recognition based on acoustic clustering, characterized in that The following steps are involved: 1) Obtain the acoustic information sequence corresponding to the audio according to the pre-trained sound unit extraction model, and construct binary data of <text, acoustic information sequence>, and use the binary data to train the text to sound unit mapping model; 2) Based on the sound unit mapping model, the text is converted into a corresponding acoustic information sequence and clustered using the KMeans method to obtain a subword list set; 3) Build an automatic speech recognition model, filter audio-to-text speech recognition training data, and extract audio files into fbank feature sequences for training the automatic speech recognition model; 4) performing first-stage decoding according to the output of the automatic speech recognition model decoding layer in step 3) to obtain a corresponding target subword list; 5) According to the target sub-word table predicted by the first stage decoding in step 4), the probability distribution under the sub-word table is calculated in the second stage decoding, and the word with the highest probability is selected as the recognition result.

2. The method for accelerating speech recognition decoding in two stages based on acoustic clustering according to claim 1, characterized in that: In step 1), the acoustic information sequence corresponding to the audio is obtained according to the pre-trained sound unit extraction model, and binary data of <text, acoustic information sequence> is constructed, and the text to sound unit mapping model is trained using the binary data, specifically: 101) For each audio data, use the pre-trained sound unit extraction model to obtain its corresponding acoustic information sequence; 102) combining the text with the extracted acoustic information sequence to obtain binary training data of <text, acoustic information sequence>; 103) Using a standard Transformer and an Encoder-Decoder architecture, training is performed based on text-to-acoustic unit mapping training data to obtain a text-to-sound unit mapping model of text-to-acoustic information.

3. The method for accelerating speech recognition decoding in two stages based on acoustic clustering according to claim 1, characterized in that: In step 3), an automatic speech recognition model is constructed, audio-to-text speech recognition training data is screened, and the audio file is extracted into an fbank feature sequence for training the automatic speech recognition model. The specific steps are as follows: 301) Use the Conformer-Encoder architecture and Transformer-Decoder architecture to build a speech recognition model framework; 302) filtering and cleaning the audio-to-text speech recognition training data, selecting high-quality training data, and extracting the audio file as fbank feature representation; 303) Use the SpecAug method to enhance the audio data, and use the convolutional layer to downsample and shorten the length of the audio data to ensure the robustness and stability of the training; 304) The extracted audio fbank features are input into the speech recognition model for training to obtain the decoding layer output result of the automatic speech recognition model.

4. The method for accelerating speech recognition decoding in two stages based on acoustic clustering according to claim 1, characterized in that: In step 4), according to the output of the automatic speech recognition model decoding layer in step 3), the first stage decoding is performed to obtain the corresponding target sub-word table. First, the output result of the automatic speech recognition model decoding layer is obtained, and then the clustered sub-word table set in step 2) is used as the classification target, and the Softmax method is used to predict the most likely target sub-word table, specifically: 401) According to the clustered sub-word table set, set the number of all target sub-word tables for the first stage decoding classification prediction, and construct the corresponding Softmax layer; 402) According to the output result of the automatic speech recognition model decoding layer, the Softmax method is used to perform classification prediction to obtain the most likely target subword list. The specific formula is: Among them, p n is the normalized predicted probability value of the nth target sub-word list; n is the number of the current target sub-word list; N is the number of all target sub-word lists; i = [1, N] represents the number of any sub-word list; c i is the output probability of any target subword list in the automatic speech recognition decoding layer; c n is the output probability of the nth target subword list in the automatic speech recognition decoding layer; is the temperature parameter, which is used to control the transition of the distribution from smooth to sharp; g i is the additional noise corresponding to any sub-vocabulary; g n is the additional noise corresponding to the target sub-word list.

5. The method for accelerating speech recognition decoding in two stages based on acoustic clustering according to claim 1, characterized in that: The second stage decoding process in step 5) is to decode according to the obtained target sub-word table, calculate the probability distribution of each word under the sub-word table, and select the word with the highest probability as the recognition result. The calculation formula is: Among them, p n,k is the normalized predicted probability of the target word; V n is the number of words in the nth sub-word list; k is the position of the target word in the sub-word list; l k is the probability of the corresponding target word; i is any word in the target subword list; l i is the probability of any word in the target subword list; T is the temperature parameter used to control the probability smoothness.