Speech Recognition Acoustic Model Using Dynamic Blank Units

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies based on state modeling, such as Hidden Markov Models, face challenges with confusion between pronunciation units, leading to poor recognition performance and slow decoding speeds.

Innovation Solution

A speech recognition method and device that uses connectionist temporal classification to establish an acoustic model and decoding network, dynamically adding a blank unit during the decoding process to improve accuracy and speed by selecting the optimum decoding path.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If Hidden Markov Model based state modeling is used for speech recognition, then the recognition system can process speech signals, but confusion occurs between pronunciation units leading to poor recognition performance

Engineering Contradiction:
Improverecognition performanceVSAvoidpronunciation unit distinction
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the speech recognition process by introducing blank units that separate and delimit pronunciation units in the decoding network. This segmentation allows the system to clearly distinguish between different pronunciation units by inserting blank symbols at their boundaries, thereby preventing confusion and improving recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The blank unit acts as an intermediary element in the decoding network. It mediates between different pronunciation units by serving as a delimiter that prevents direct confusion between adjacent pronunciation units. The blank unit is inserted dynamically during decoding to mark boundaries without interfering with the actual pronunciation recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional decoding networks are used without dynamic blank unit addition, then the decoding process is simpler, but the number of possible decoding paths increases leading to slower decoding speed

Engineering Contradiction:
Improvedecoding speedVSAvoiddecoding process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies dynamics by making the decoding network structure adaptable during the decoding process. Blank units are dynamically added or removed based on the specific speech signal being processed and the current decoding state. This dynamic adjustment optimizes the number of decoding paths for each specific case, improving overall decoding speed while maintaining necessary complexity only where needed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the structural parameters of the decoding network dynamically by adding or removing blank units based on the input speech signal characteristics. This parameter change allows the system to optimize the decoding path count adaptively, reducing unnecessary computational paths and improving decoding speed without permanently increasing system complexity.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If more decoding paths are considered to improve recognition accuracy, then better recognition performance is achieved, but the decoding time increases

Engineering Contradiction:
Improverecognition accuracyVSAvoiddecoding time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively adding blank units only where necessary to resolve ambiguity between pronunciation units, rather than adding them uniformly throughout all decoding paths. This selective approach maintains recognition accuracy where needed while avoiding unnecessary expansion of the decoding search space, thereby reducing overall decoding time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10650809B2Speech recognition method and device
Publication Date: 2020.05.12 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10650809B2 patent drawing
  • US10650809B2 patent drawing
  • US10650809B2 patent drawing

AI summary

The present disclosure provides a speech recognition method and device. The method includes: receiving a speech signal; decoding the speech signal according to an acoustic model, a language model and a decoding network established in advance, and dynamically adding a blank unit in a decoding process to obtain an optimum decoding path with the added blank unit, in which the acoustic model is obtained based on connectionist temporal classification training, the acoustic model includes basic pronunciation units and the blank unit, and the decoding network includes a plurality of decoding paths consisting of the basic pronunciation units; and outputting the optimum decoding path as a recognition result of the speech signal.