Speech Recognition System Using Pronunciation Difference Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Chinese speech recognition systems face challenges in efficiently recognizing Cantonese language due to limited training resources, leading to long project cycles, high resource consumption, and complex system integration, especially when translating results from Cantonese to simplified Chinese.

Innovation Solution

A method that utilizes existing system resources by generating differential pronunciation pairs and using these pairs, along with acoustic and language features, to directly recognize Cantonese speech as simplified Chinese text, eliminating the need for additional translator training and reducing resource labeling costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a translator is trained to translate Cantonese recognition results into simplified Chinese, then recognition accuracy is improved, but project cycle and resource consumption increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidproject cycle
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges the translator training process into the existing speech recognition system by integrating pronunciation difference statistics. Instead of training a separate translator model, the system incorporates pronunciation variations directly into the recognition pipeline, combining multiple functions into a unified model that reduces project cycles while maintaining accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The speech recognition system is enhanced to perform multiple functions: it not only recognizes Cantonese speech but also handles translation to simplified Chinese through the integrated pronunciation difference statistics. This multi-functional approach eliminates the need for separate translator training, reducing resource consumption and time while improving overall system efficiency

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If a translator is trained to translate Cantonese recognition results into simplified Chinese, then recognition accuracy is improved, but resource consumption and development costs increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent merges the translator training process into the existing speech recognition system by integrating pronunciation difference statistics. Instead of training a separate translator model, the system incorporates pronunciation variations directly into the recognition pipeline, combining multiple functions into a unified model that reduces project cycles while maintaining accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The speech recognition system is enhanced to perform multiple functions: it not only recognizes Cantonese speech but also handles translation to simplified Chinese through the integrated pronunciation difference statistics. This multi-functional approach eliminates the need for separate translator training, reducing resource consumption and time while improving overall system efficiency

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If existing recognition systems are used without pronunciation difference statistics, then system complexity is reduced, but recognition results require additional translation processing

Engineering Contradiction:
Improvesystem complexityVSAvoidrecognition efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-calculating and storing pronunciation difference statistics before the recognition process. These statistics are prepared in advance and integrated into the recognition system, allowing the model to directly output simplified Chinese results without requiring subsequent translation processing, thus improving efficiency while maintaining manageable complexity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12033615B2Method and apparatus for recognizing speech, electronic device and storage medium
Publication Date: 2024.07.09 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12033615B2 patent drawing
  • US12033615B2 patent drawing
  • US12033615B2 patent drawing

AI summary

The disclosure provides a method and an apparatus for recognizing a speech, an electronic device and a storage medium. A speech to be recognized is obtained. An acoustic feature of the speech to be recognized and a language feature of the speech to be recognized are obtained. The speech to be recognized is input to a pronunciation difference statistics to generate a differential pronunciation pair corresponding to the speech to be recognized. The text information of the speech to be recognized is generated based on the differential pronunciation pair, the acoustic feature and the language feature.