A speech recognition method supporting a dual-line scenario
By preprocessing and feature extraction of voice signals, combined with dynamic encoder and full-sequence encoder for online and offline processing, the problem of difficulty in achieving high accuracy in both online and offline scenarios in the prior art is solved, and the high accuracy of two-line scene speech recognition is achieved.
Patent Information
- Application Number
- CN202310041299.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-01-12
AI Technical Summary
Existing voice recognition technologies are difficult to achieve high-accuracy recognition in both online real-time and offline scenarios.
A speech recognition method that supports two-line scenes is adopted to pre-process the speech signal, extract acoustic features, and capture local features using 18-layer ResNet. Then, dynamic encoder and decoder are used for online processing, combined with full-sequence encoder, text encoder and decoder for offline processing, and finally identification is achieved through model training and deployment.
It realizes high accuracy speech recognition in both online real-time scenarios and offline scenarios, reducing the usage threshold and maintenance costs of the model.
Smart Images

Figure CN116312479B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech recognition, and specifically provides a speech recognition method supporting a dual-line scenario. Background Art
[0002] Speech recognition is a technology for extracting text from audio. Online speech recognition technology processes audio files sent in a streaming manner in real time and returns recognition results. Since the integrity of sentences is not considered during recognition, problems such as recognition errors and reduced accuracy may occur. Common scenarios include speech input methods, real-time reminders for conversation intelligence, etc.; compared with real-time processing, offline speech recognition technology can better consider context information and correct recognition results in terms of sentence integrity and semantics, which can reduce recognition errors and improve recognition accuracy, but it cannot return recognition results in real time. Common scenarios include conference audio transcription, automatic subtitle generation, etc.
[0003] In the technical field of speech recognition, the early popular stage was the generative model stage, such as methods based on hidden Markov models. In the early days of the explosion of deep learning, the CTC method based on deep neural networks greatly improved the corresponding recognition effect on the original technology. However, the outputs of such methods are independent of each other, which makes the ability to utilize context information for each frame slightly insufficient. Subsequently, many methods based on RNNs emerged, such as the DeepSpeech series of methods. After the Transformer became extremely popular, many methods based on the Attention mechanism emerged. Its encoder can more effectively model the context of speech features, and at the same time, the autoregressive decoder method can make the modeling ability stronger. However, the drawback of such methods is that the processing effect for streaming, that is, online real-time scenarios, is very poor.
[0004] To better handle the speech recognition problem in online real-time scenarios, Transformer-transducer changes the global attention mechanism to a local attention mechanism and uses a predictor to enable the decoder to generate a streaming output, which can slightly alleviate the problem of poor recognition of online speech data.
[0005] In summary, the vast majority of existing technologies cannot handle the problem of online speech recognition well; for methods that emphasize improving the online recognition problem, they usually adopt the method based on the local attention mechanism. Although this can slightly alleviate the problem of poor online speech recognition scenarios, it cannot handle the problem of offline scenario recognition. For many application scenarios, high accuracy requirements are often imposed on both online and offline recognition. Summary of the Invention
[0006] The purpose of this application is to provide a speech recognition method supporting dual-line scenarios to solve the problem that although the problem of the poor online speech recognition scenario is alleviated, the problem of being unable to handle the offline scenario recognition in the above-mentioned background technology.
[0007] To achieve the above purpose, this application provides the following technical solutions: A speech recognition method supporting dual-line scenarios, including the following steps:
[0008] Step 1: Speech processing: First, preprocess the speech signal and output acoustic features;
[0009] Step 2: Local feature extraction: Based on the neural network structure of vision that is good at processing local features, input the acoustic features in Step 1 into an 18-layer ResNet to reduce the length of the sequence and capture local features;
[0010] Step 3: Online processing: Includes a dynamic encoder and a decoder. The features output in Step 2 are used as the input of this module. Each segment frame is first input into the dynamic encoder, and then the corresponding output is input into the decoder to generate an intermediate value. When all acoustic segments are processed by the online module, all intermediate values are connected as the input of the offline processing module;
[0011] Step 4: Offline processing: Adopt a fully sequential encoder, a text encoder, and a fully sequential decoder; The result output in Step 3 is used as the input of offline processing. First, input it into the fully sequential encoder in the offline module to generate a new intermediate vector, and at the same time input it into the text encoder to generate the corresponding semantic intermediate vector. Then, splice these two intermediate vectors and input them into the fully sequential decoder to output the final result;
[0012] Step 5: Model training: Based on the AISHELL dataset, the ratio of training, validation, and evaluation is 8:1:1. The optimizer uses AdamOptimizer, where β1 is 0.8, β2 is 0.98, the warm-up times is 8000, the learning rate is 0.00003, and the number of iterations is 6000 times;
[0013] Step 6: Model deployment and service: Deploy the model trained in Step 5 to the server, and recognize the incoming speech in real-time or batch mode and convert it into the corresponding text for return.
[0014] Preferably, the preprocessing of the speech signal in Step 1 includes pre-emphasis, framing, and windowing.
[0015] Preferably, in step 3, the encoder used is a 20-layer Transformer encoder, with residual connections based on the multi-head attention sub-layer and layer normalization operations in the feed-forward network sub-layer, and the decoder used is a 16-layer Transformer decoder.
[0016] Preferably, in step 4, the full-order encoder used is a 20-layer Transformer encoder, the text encoder used is a 10-layer Transformer encoder, and the full-order decoder used is a 10-layer Transformer decoder.
[0017] Preferably, in step 4, the attention mechanism in the full-order decoder is used to calculate the context vector and semantic context vector in the sound, and these two vectors will be linked together as the attention vector.
[0018] Compared with the prior art, the beneficial effects of this application are:
[0019] Based on a novel speech recognition system supporting dual-line scenarios, the present invention can simultaneously achieve the response speed of online real-time and the recognition accuracy of offline. One system meets all speech recognition scenarios, reducing the usage threshold and maintenance cost of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of the process of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0022] Embodiment:
[0023] Please refer to Figure 1 , this application provides a technical solution: a speech recognition method supporting dual-line scenarios, including the following steps:
[0024] Step 1: Speech processing: First, preprocess the speech signal. The preprocessing of the speech signal includes performing conventional operations such as pre-emphasis, framing, and windowing, which belong to the prior art and will not be elaborated here, and output the acoustic features;
[0025] Step 2: Local feature extraction: Since the features output in Step 1 are generally quite long, and although the Transformer based on the attention mechanism captures and processes global feature patterns, it does not handle local features well. Therefore, in this application, based on a neural network structure for vision that is good at processing local features, the acoustic features in Step 1 are input into an 18-layer ResNet to reduce the length of the sequence and capture local features;
[0026] Step 3: Online processing: It includes a dynamic encoder and a decoder. The encoder uses a 20-layer Transformer encoder, with residual connections based on the multi-head attention sub-layer and layer normalization operations based on the feed-forward network sub-layer. The decoder uses a 16-layer Transformer decoder. The features output in Step 2 are used as the input to this module. Each segment frame is first input into the dynamic encoder, and then the corresponding output is input into the decoder to generate an intermediate value. After all the acoustic segments are processed by the online module, all the intermediate values are concatenated as the input to the offline processing module;
[0027] Step 4: Offline processing: It uses a fully sequential encoder, a text encoder, and a fully sequential decoder. The fully sequential encoder uses a 20-layer Transformer encoder, the text encoder uses a 10-layer Transformer encoder, and the fully sequential decoder uses a 10-layer Transformer decoder. The attention mechanism in the fully sequential decoder is used to calculate the context vector and semantic context vector in the sound, and these two vectors are concatenated as the attention vector. The result output in Step 3 is used as the input for offline processing. It is first input into the fully sequential encoder in the offline module to generate a new intermediate vector, and at the same time, it is input into the text encoder to generate the corresponding semantic intermediate vector. Then these two intermediate vectors are concatenated and input into the fully sequential decoder to output the final result;
[0028] Step 5: Model training: Based on the AISHELL dataset, the ratio of training, validation, and evaluation is 8:1:1. The optimizer uses AdamOptimizer, where β1 is 0.8, β2 is 0.98, the warm-up times is 8000, the learning rate is 0.00003, and the number of iterations is 6000;
[0029] Step 6: Model deployment and service: Deploy the model trained in Step 5 to the server, and identify the incoming speech in real-time or batch mode and convert it into the corresponding text for return.
[0030] The following is the prior art mentioned in the background art of this application. By comparing the methods in the references of the prior art in this field with this application, the advancement of the technology of this application can be better reflected.
[0031] [1] Hannun A, Case C, Casper J, et al. Deep speech: Scaling up end-to-end speech recognition[J]. arXiv preprint arXiv:1412.5567, 2014.
[0032] [2] Amodei D, Ananthanarayanan S, Anubhai R, et al. Deep speech2: End-to-end speech recognition in english and mandarin[C] / / International conference on machine learning. PMLR, 2016:173-182.
[0033] [3] Zhang Q, Lu H, Sak H, et al. Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss[C] / / ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2020:7829-7833.
[0034] The above has shown and described the basic principles, main features and advantages of this application. For those skilled in the art, it is obvious that this application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of this application; therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of this application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in this application, and any reference signs in the claims should not be regarded as limiting the claims involved.
[0035] Although embodiments of the present application have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A speech recognition method for supporting a dual-line scenario, characterized in that: It includes the following steps: Step 1: Speech processing: First, preprocess the speech signal and output acoustic features; Step 2: Local feature extraction: Based on the neural network structure of vision that is good at processing local features, input the acoustic features in Step 1 into an 18-layer ResNet to reduce the length of the sequence and capture local features; Step 3: Online processing: It includes a dynamic encoder and a decoder. The features output in Step 2 are used as the input of this module. Each segment frame is first input into the dynamic encoder, and then the corresponding output is input into the decoder to generate an intermediate value. After all the acoustic segments are processed by the online module, all the intermediate values are concatenated as the input of the offline processing module; Step 4: Offline processing: Adopt a fully sequential encoder, a text encoder, and a fully sequential decoder; the result output in Step 3 is used as the input of the offline processing. First, input it into the fully sequential encoder in the offline module to generate a new intermediate vector, and at the same time input it into the text encoder to generate the corresponding semantic intermediate vector. Then concatenate these two intermediate vectors and input them into the fully sequential decoder to output the final result; Step 5: Model training: Based on the AISHELL dataset, the ratio of training, validation, and evaluation is 8:1:
1. The optimizer uses AdamOptimizer, where β1 is 0.8, β2 is 0.98, the warm-up times is 8000, the learning rate is 0.00003, and the number of iterations is 6000 times; Step 6: Model deployment and service: Deploy the model trained in Step 5 to the server, and identify the incoming speech in real-time or batch mode and convert it into the corresponding text for return.
2. The voice recognition method supporting a dual-line scenario according to claim 1, wherein: The preprocessing of the speech signal in Step 1 includes pre-emphasis, framing, and windowing.
3. The voice recognition method supporting a dual-line scenario according to claim 1, characterized in that: In Step 3, the encoder uses a 20-layer Transformer encoder, performs residual connection based on the multi-head attention sublayer and layer normalization operation based on the feed-forward network sublayer, and the decoder uses a 16-layer Transformer decoder.
4. A speech recognition method for supporting a dual-line scenario according to claim 1, characterized in that: In Step 4, the fully sequential encoder uses a 20-layer Transformer encoder, the text encoder uses a 10-layer Transformer encoder, and the fully sequential decoder uses a 10-layer Transformer decoder.
5. A speech recognition method for supporting a dual-line scenario according to claim 1, characterized in that: The attention mechanism in the fully sequential decoder in Step 4 is used to calculate the context vector and semantic context vector in the sound, and these two vectors will be concatenated as the attention vector.
Citation Information
Patent Citations
Intelligent robot semantic interactive system and method based on white light communication and brain-similar cognition
CN108717852A
Voice recognition method, equipment thereof and storage medium
CN109036379A