Speech Recognition Acoustic Model Using Dynamic Blank Units
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies based on state modeling, such as Hidden Markov Models, face challenges with confusion between pronunciation units, leading to poor recognition performance and slow decoding speeds.
Innovation Solution
A speech recognition method and device that uses connectionist temporal classification to establish an acoustic model and decoding network, dynamically adding a blank unit during the decoding process to improve accuracy and speed by selecting the optimum decoding path.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If Hidden Markov Model based state modeling is used for speech recognition, then the recognition system can process speech signals, but confusion occurs between pronunciation units leading to poor recognition performance
Solution Approach 1:
The patent segments the speech recognition process by introducing blank units that separate and delimit pronunciation units in the decoding network. This segmentation allows the system to clearly distinguish between different pronunciation units by inserting blank symbols at their boundaries, thereby preventing confusion and improving recognition accuracy.
Solution Approach 2:
The blank unit acts as an intermediary element in the decoding network. It mediates between different pronunciation units by serving as a delimiter that prevents direct confusion between adjacent pronunciation units. The blank unit is inserted dynamically during decoding to mark boundaries without interfering with the actual pronunciation recognition.
2Productivity
If traditional decoding networks are used without dynamic blank unit addition, then the decoding process is simpler, but the number of possible decoding paths increases leading to slower decoding speed
Solution Approach 1:
The patent applies dynamics by making the decoding network structure adaptable during the decoding process. Blank units are dynamically added or removed based on the specific speech signal being processed and the current decoding state. This dynamic adjustment optimizes the number of decoding paths for each specific case, improving overall decoding speed while maintaining necessary complexity only where needed.
Solution Approach 2:
The system changes the structural parameters of the decoding network dynamically by adding or removing blank units based on the input speech signal characteristics. This parameter change allows the system to optimize the decoding path count adaptively, reducing unnecessary computational paths and improving decoding speed without permanently increasing system complexity.
3Reliability
If more decoding paths are considered to improve recognition accuracy, then better recognition performance is achieved, but the decoding time increases
Solution Approach 1:
The patent applies partial action by selectively adding blank units only where necessary to resolve ambiguity between pronunciation units, rather than adding them uniformly throughout all decoding paths. This selective approach maintains recognition accuracy where needed while avoiding unnecessary expansion of the decoding search space, thereby reducing overall decoding time.
Data Source
AI summary
The present disclosure provides a speech recognition method and device. The method includes: receiving a speech signal; decoding the speech signal according to an acoustic model, a language model and a decoding network established in advance, and dynamically adding a blank unit in a decoding process to obtain an optimum decoding path with the added blank unit, in which the acoustic model is obtained based on connectionist temporal classification training, the acoustic model includes basic pronunciation units and the blank unit, and the decoding network includes a plurality of decoding paths consisting of the basic pronunciation units; and outputting the optimum decoding path as a recognition result of the speech signal.


