Hot word sorting and selecting method for context speech recognition
By designing a scoring network, using the TTS model and cross attention mechanism to process the audio characteristics of hot words, the performance bottleneck of the context ASR model when processing a large number of hot words is solved, and the accuracy and efficiency of hot words recognition are improved.
Patent Information
- Application Number
- CN202510444856.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-27
AI Technical Summary
The contextual automatic speech recognition (ASR) model performs performance bottlenecks when processing a large number of hot words, resulting in a decrease in recognition accuracy and reduced computational efficiency.
A scoring network is designed to convert hot words into audio through the TTS model, combine pre-trained audio encoder and cross-attention mechanism to extract the characteristics of voice and hot words audio, and use CNN and softmax layers to score hot words, filter and sort hot words.
It significantly reduces the hot word error rate (B-WER), improves the accuracy and computing efficiency of the model to recognize hot words, and enhances the scalability and performance of the context ASR model.
Smart Images

Figure CN120220688A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a method for hot word sorting and selection for context speech recognition. Background Art
[0002] With the continuous development of speech recognition technology, end-to-end automatic speech recognition (ASR) systems have achieved remarkable results, mainly including three categories: connectionist temporal classification (CTC) models, attention-based encoder-decoder models, and transformer-based models, which are widely used in various ASR tasks. However, standard ASR systems have great difficulties in recognizing low-frequency words such as rare words and proper nouns. The main reason is that low-frequency words in the training data show a long-tail distribution, resulting in inaccurate transcription results.
[0003] To solve these problems, context hot word technologies have been applied, such as shallow fusion and deep fusion technologies. By integrating context information into the ASR process, the performance of ASR has been effectively improved. Shallow fusion combines a pre-trained language model (LM) and an acoustic model during decoding. First, the acoustic model generates candidate transcripts, and then the LM re-scores them according to language likelihood. Deep fusion is to jointly train the acoustic model and the LM, making them interact more deeply during the inference stage. By merging the intermediate representations before the final prediction layer, the fusion of acoustic and language information is strengthened, thereby improving the ASR accuracy. In recent years, many studies have focused on integrating large-scale base models with context ASR technologies, hoping to better recognize rare words and domain-specific terms in context scenarios and more accurately handle various language details with the capabilities of these advanced models.
[0004] Although context automatic speech recognition (ASR) systems have made great progress, they still face challenges when dealing with a large number of hot words. When the number of hot words is large (such as more than 1000), context ASR models often have difficulty coping and cannot process efficiently. Especially models built based on large-scale base models are very sensitive to the number of hot words. This is because the context length is limited, which restricts the model's ability to process and integrate a large number of hot words. At the same time, the limitation in computational efficiency also makes it difficult for the model to handle the exponentially growing complexity caused by a large number of hot words, ultimately affecting the overall performance of the context ASR system.
[0005] Therefore, those skilled in the art are committed to developing a method for hot word sorting and selection for context speech recognition. A scorer network is proposed, which comprehensively utilizes technologies such as TTS models, audio encoders, cross-attention mechanisms, and CNN (convolutional neural network) to accurately screen and sort hot words, improving the model's ability to recognize hot words. Summary of the Invention
[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is the performance bottleneck problem of the context ASR model when processing a large number of hot words.
[0007] To achieve the above object, the present invention provides a hot word sorting and selection method for context speech recognition, including a scorer network, which screens hot words to reduce the total number of hot words before integrating the hot words into the ASR system.
[0008] Further, the hot words are converted into hot word audio by means of a TTS model and fused with the speech audio; the pre-trained audio encoder extracts features from the speech and hot word audio respectively, and captures cross-modal relationships through a cross-attention mechanism; then the local features are extracted by a CNN, and the global features are obtained through a global pooling layer; finally, the softmax layer scores the hot words, and the hot words are screened according to the scores.
[0009] Further, different hot word permutation methods are set, the scorer network is used to generate hot word scores, the hot words are input into the model in different orders, the performance changes of the model are observed, and the best hot word sorting method is selected.
[0010] Further, the hot word permutation methods include a random order permutation method, an ascending order permutation method, and a descending order permutation method.
[0011] Further, in the ascending order permutation method, the high-probability hot words are placed at the end.
[0012] Further, in the descending order permutation method, the high-probability hot words are placed at the beginning.
[0013] Further, a named entity recognition model is used to generate a list of proper noun hot words close to the real-world scenario.
[0014] Further, the named entity recognition model screens each word in the text, identifies the proper nouns therein, and constructs a comprehensive list of hot words.
[0015] Further, the proper nouns include contact names, phone numbers, personal names, and location names.
[0016] Further, it includes the following steps:
[0017] Step 1, data preparation;
[0018] Step 2, model construction and training;
[0019] Step 3, hot word sorting and selection;
[0020] Step 4, comparison and analysis.
[0021] When the existing context ASR model faces a large number of hot words, it is difficult to effectively process them due to the limitations of the context length and computational efficiency, resulting in a decline in overall performance. The present invention designs a new scoring network to screen hot words and reduce the total number of hot words before integrating them into the ASR system. The present invention uses a TTS model to convert hot words into hot word audio and fuse it with the speech audio. The pre-trained audio encoder is used to extract features from the speech and hot word audio respectively, and the cross-attention mechanism is used to capture cross-modal relationships, enabling the model to better associate hot words with speech content. Then, a CNN is used to extract local features, and global features are obtained through a global pooling layer. Finally, a softmax layer scores the hot words, and the hot words are screened according to the scores. The present invention is tested on the LibriSpeech dataset combined with the IS21 hot word list, and the hot word error rate (B-WER) is relatively reduced by more than 40%, improving the model's performance in recognizing hot words, enhancing the scalability and efficiency of the context ASR model in processing a large number of hot words, and having good generalization in different models and hot word lists.
[0022] The existing technology has not explored the impact of the order of hot words when inputting into the model on the performance of context ASR, and lacks a method to optimize the input order of hot words. The present invention studies the impact of hot word sorting on the model performance and compares the performance of the model under different sorting methods. For the IS21 hot word list, the present invention sets three permutation methods: random order, ascending order (high-probability hot words are placed at the end), and descending order (high-probability hot words are placed at the beginning). The proposed scoring network is used to generate hot word scores, and the hot words are input into the Whisper model in different orders to observe the changes in model performance. The present invention finds that when the real hot words are input into the Whisper model in ascending order, the model performance is the best. It provides a reference for optimizing the order of hot words input into the model and helps to improve the performance of the context ASR system.
[0023] The existing hot word list construction method is not close enough to the actual application scenario, resulting in poor performance of the ASR system in processing hot words in real scenarios. The present invention uses a named entity recognition (NER) model to generate a more realistic hot word list of proper nouns. The present invention uses the NER model to screen each word in the LibriSpeech text one by one, identify the proper nouns among them, such as contact names, phone numbers, personal names, location names, etc., and construct a comprehensive hot word list. The experimental results of the present invention show that by using this hot word list combined with the proposed method and selecting the top 50 hot words with the highest scores in the Whisper-turbo model, the B-WER can be significantly reduced by 30%, more effectively improving the model's ability to process hot words in real scenarios.
[0024] Compared with the prior art, the present invention has the following obvious substantial features and remarkable advantages:
[0025] 1. Technical advantages: Through innovative hot word sorting and selection technology, the present invention effectively solves the performance bottleneck problem of the context ASR model when processing a large number of hot words. The proposed scorer network comprehensively utilizes technologies such as the TTS model, audio encoder, cross-attention mechanism, and CNN, and can accurately screen and sort hot words, significantly improving the model's ability to recognize hot words. Compared with traditional methods, when dealing with the same hot word task, the B-WER is significantly reduced, which means that in practical applications, the accuracy of speech recognition is greatly improved, effectively reducing information errors caused by hot word recognition errors, and providing more reliable technical support for the speech interaction-related industries.
[0026] 2. In terms of metrics: The experimental results strongly prove the superiority of the technical solution of the present invention. On the LibriSpeech dataset, whether using the IS21 hot word list or the hot word list generated by NER, a significant reduction in B-WER can be achieved, with the highest relative reduction exceeding 40%. At the same time, in different context ASR models, such as Whisper and TCPGen-based BiasingWhisper, the present invention can achieve good results and improve the model performance. This shows that the technical solution of the present invention has wide applicability and stability on different datasets and models, providing a solid performance guarantee for its industrial application.
[0027] 3. From the implementation perspective: The technical components adopted by the present invention, such as the TTS model (edge-tts), ASR model (Whisper-turbo), etc., all have mature open-source implementations, reducing the threshold and cost of technology implementation. The model parameter settings detailed in the experiment, such as the projection dimension of the linear layer, the number of heads and dropout rate of the cross-attention mechanism, the output channels and kernel size of each layer of the CNN, etc., provide clear guidance for model construction and optimization in practical applications, facilitating enterprises and developers to quickly integrate this technology into existing speech recognition systems, accelerating product iteration and upgrading, and having high implementability.
[0028] 4. The technical solution of the present invention has significant technical advantages, excellent metric performance, and good implementation feasibility, and has broad industrial application prospects in speech recognition-related industries, such as intelligent voice assistants, speech transcription, intelligent customer service, etc., and has extremely high conversion value.
[0029] The following will further illustrate the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings to fully understand the purpose, features, and effects of the present invention. Description of the Drawings
[0030] Figure 1 It is a flowchart of hot word processing in a preferred embodiment of the present invention. Detailed Embodiments
[0031] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0032] In the drawings, components with the same structure are denoted by the same numerical labels, and components with similar structures or functions everywhere are denoted by similar numerical labels. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of some parts in the drawings is appropriately exaggerated.
[0033] As Figure 1 shown, the present invention is a method for ranking and selecting bias words (hot words) for context-biased speech recognition:
[0034] 1. Data preparation stage: The LibriSpeech dataset is selected to verify the effectiveness of the method. This dataset contains approximately 1000 hours of English audiobook speech, divided into training, validation, and test subsets. In this experiment, train-clean-100 is used as the training set, dev-clean as the validation set, and test-clean and test-other as the test sets. Two types of hot word lists are determined. One is the IS21 deep bias words list. By removing the 5000 most common words in the vocabulary, the remaining rare words are used as hot words, and different numbers of interference words are introduced (N = {100, 500, 1000, 2000}); the other is the hot word list generated from the LibriSpeech text using the NER model, which is closer to the real scenario. After screening, a hot word list containing 4365 words is obtained.
[0035] 2. Model construction and training stage: Use edge-tts as the TTS model to convert hot words into corresponding hot word audio, realizing the unified representation of hot words and speech audio in the same audio domain for facilitating subsequent similarity calculation. Use Whisper-turbo as the ASR model. Project the 1280-dimensional features extracted from the Whisper-large-v3 model to 368 dimensions using a linear layer to reduce the dimension and improve the calculation efficiency. Set up a cross-attention mechanism with 8 heads and use a dropout rate of 0.1 in the training stage to enhance the generalization ability of the model. Construct a CNN consisting of four layers, and the number of output channels of each layer is set to 32, 64, 128, and 256 respectively, with a unified kernel size of 3, for extracting local patterns in the speech and hot word audio features. Aggregate the features into a single 256-dimensional vector through an adaptive average pooling layer for convenient subsequent processing.
[0036] 3. Hot Word Ranking and Selection Phase: The pre-trained audio encoder is used to extract features from the speech audio and hot word audio respectively, and then a linear projection layer is used for dimensionality reduction to improve computational efficiency. The bidirectional cross-attention mechanism is adopted. First, the speech audio features are used as queries, and the hot word audio features are used as keys and values. Then, vice versa, the hot word audio features are used as queries, and the speech audio features are used as keys and values to learn the bidirectional relationship between speech and hot words, highlighting the hot words most relevant to the speech input. Multiple cosine similarity matrices are calculated, including the similarity between speech features and hot word features, the similarity between speech features and the cross-attention features from hot words to speech, etc., to measure the relationship between speech and hot word audio from multiple perspectives. These similarity matrices are stacked and input into a series of CNN layers to extract powerful feature representations, capturing the local and global dependencies between speech and hot words. All the features output by the CNN layers are merged into a set of global features through a global pooling layer, and then through a linear layer and a softmax layer, the final score of each hot word relative to the entire speech input is calculated. The higher the score, the more likely the hot word is the true hot word. The hot words are ranked according to the scores, and the top k hot words with the highest scores are selected.
[0037] 4. Experimental Comparison and Analysis Phase: Compare the performance of different methods when using the IS21 deep bias words list, including DB-RNNT, DB-RNNT+NNLM, BPB, Whisper-turbo, etc., as well as the effects when using different prompt types (naive method, colloquial method) and hot word orders (random, ascending, descending). For the NER hot word list, test the performance changes of the model when selecting different numbers (k = {10, 50, 100, 200}) of hot words. By comparing the overall word error rate (WER) and the hot word error rate (B-WER), evaluate the performance of the model under different settings and analyze the influence of each factor on the model performance.
[0038] In the entire hot word processing flow of hot word ranking and selection, through a series of carefully designed steps, such as hot word audio conversion, feature extraction and dimensionality reduction, cross-attention mechanism to calculate similarity, multi-matrix fusion and feature extraction, and finally score calculation and ranking selection, the efficient screening and ranking of a large number of hot words are achieved. This not only improves the accuracy of the context ASR model in recognizing hot words, reduces the B-WER, but also shows good generalization and adaptability on different models and hot word lists, providing an effective means to improve speech recognition performance in practical applications.
[0039] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. A method for hot word sorting and selection for contextual speech recognition, characterized in that: It includes a scorer network to filter hot words and reduce the total number of hot words before integrating them into the ASR system.
2. The hot word sorting and selection method for contextual speech recognition according to claim 1, characterized in that: With the help of the TTS model, hot words are converted into hot word audio and fused with the speech audio; the pre-trained audio encoder is used to extract features from the speech and hot word audio respectively, and the cross-modal relationship is captured through the cross-attention mechanism; CNN is then used to extract local features, and the global features are obtained through the global pooling layer; finally, the softmax layer scores the hot words and selects the hot words based on the scores.
3. The hot word sorting and selection method for contextual speech recognition according to claim 1, characterized in that: Set different hot word arrangements, use the scorer network to generate hot word scores, input hot words into the model in different orders, observe the changes in model performance, and select the best hot word sorting method.
4. The hot word sorting and selection method for contextual speech recognition according to claim 3, characterized in that: The hot word arrangement method includes a random order arrangement method, an ascending order arrangement method, and a descending order arrangement method.
5. The hot word sorting and selection method for contextual speech recognition according to claim 4, characterized in that: In the ascending order, high-probability hot words are placed at the end.
6. The method for hot word sorting and selection for contextual speech recognition according to claim 4, characterized in that: In the descending order, high-probability hot words are placed at the beginning.
7. The method for hot word sorting and selection for contextual speech recognition according to claim 1, characterized in that: Use the named entity recognition model to generate a hot word list of proper nouns that is close to real-world scenarios.
8. The hot word sorting and selection method for contextual speech recognition according to claim 7, characterized in that: The named entity recognition model screens the words in the text one by one, identifies the proper nouns therein, and constructs a comprehensive hot word list.
9. The method for hot word sorting and selection for contextual speech recognition according to claim 8, characterized in that: The proper nouns include contact names, phone numbers, personal names, and location names.
10. The method for hot word sorting and selection for contextual speech recognition according to claim 1, characterized in that: The following steps are involved: Step 1: Data preparation; Step 2: Model construction and training; Step 3: Hot word sorting and selection; Step 4: Compare and analyze.