Speech Recognition System Using Pronunciation Difference Statistics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Chinese speech recognition systems face challenges in efficiently recognizing Cantonese language due to limited training resources, leading to long project cycles, high resource consumption, and complex system integration, especially when translating results from Cantonese to simplified Chinese.
Innovation Solution
A method that utilizes existing system resources by generating differential pronunciation pairs and using these pairs, along with acoustic and language features, to directly recognize Cantonese speech as simplified Chinese text, eliminating the need for additional translator training and reducing resource labeling costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a translator is trained to translate Cantonese recognition results into simplified Chinese, then recognition accuracy is improved, but project cycle and resource consumption increase
Solution Approach 1:
The patent merges the translator training process into the existing speech recognition system by integrating pronunciation difference statistics. Instead of training a separate translator model, the system incorporates pronunciation variations directly into the recognition pipeline, combining multiple functions into a unified model that reduces project cycles while maintaining accuracy
Solution Approach 2:
The speech recognition system is enhanced to perform multiple functions: it not only recognizes Cantonese speech but also handles translation to simplified Chinese through the integrated pronunciation difference statistics. This multi-functional approach eliminates the need for separate translator training, reducing resource consumption and time while improving overall system efficiency
2Measurement precision
If a translator is trained to translate Cantonese recognition results into simplified Chinese, then recognition accuracy is improved, but resource consumption and development costs increase
Solution Approach 1:
The patent merges the translator training process into the existing speech recognition system by integrating pronunciation difference statistics. Instead of training a separate translator model, the system incorporates pronunciation variations directly into the recognition pipeline, combining multiple functions into a unified model that reduces project cycles while maintaining accuracy
Solution Approach 2:
The speech recognition system is enhanced to perform multiple functions: it not only recognizes Cantonese speech but also handles translation to simplified Chinese through the integrated pronunciation difference statistics. This multi-functional approach eliminates the need for separate translator training, reducing resource consumption and time while improving overall system efficiency
3Device complexity
If existing recognition systems are used without pronunciation difference statistics, then system complexity is reduced, but recognition results require additional translation processing
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing pronunciation difference statistics before the recognition process. These statistics are prepared in advance and integrated into the recognition system, allowing the model to directly output simplified Chinese results without requiring subsequent translation processing, thus improving efficiency while maintaining manageable complexity
Data Source
AI summary
The disclosure provides a method and an apparatus for recognizing a speech, an electronic device and a storage medium. A speech to be recognized is obtained. An acoustic feature of the speech to be recognized and a language feature of the speech to be recognized are obtained. The speech to be recognized is input to a pronunciation difference statistics to generate a differential pronunciation pair corresponding to the speech to be recognized. The text information of the speech to be recognized is generated based on the differential pronunciation pair, the acoustic feature and the language feature.


