English Speech Recognition Using Language-Specific Phonetic Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies are inaccurate when applied to English speech due to differences in pronunciation and phonetics from Chinese speech, leading to erroneous recognition results.

Innovation Solution

A method and apparatus that utilize a deep learning algorithm to determine a target speech recognition model specific to English speech, identifying original phonemes and applying a phonetic model generated by pre-training English texts to match and convert English speech into text, distinguishing between standard English and Chinese accents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing speech recognition technologies are applied to English speech, then the system can process speech input, but the recognition accuracy deteriorates due to pronunciation and phonetic differences from Chinese speech

Engineering Contradiction:
Improverecognition accuracyVSAvoidadaptability to different languages
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by creating language-specific phonetic models and speech recognition models tailored to English pronunciation characteristics. The system uses distinct phoneme sets (IPA for English,汉语拼音 for Chinese) and trains separate acoustic models for different languages, allowing each language to be processed with its own optimized parameters and features rather than a universal model

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the speech recognition system into distinct language-specific components: English speech is processed through English phoneme identification and English language models, while Chinese speech follows a separate Chinese phoneme identification path. This segmentation allows the system to maintain high accuracy for each language by treating them as independent processing streams with dedicated resources

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If a universal speech recognition model is used for both Chinese and English, then the device complexity is reduced, but the recognition precision deteriorates due to phonetic differences

Engineering Contradiction:
Improverecognition precisionVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements universality by designing a multi-functional speech recognition system that can handle multiple languages through a unified architecture. The system uses a common front-end for speech signal processing and a common decision-making framework, while incorporating language-specific phoneme sets and acoustic models. This allows the same device to accurately process both Chinese and English speech without requiring separate hardware systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10755701B2Method and apparatus for converting English speech information into text
Publication Date: 2020.08.25 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10755701B2 patent drawing
  • US10755701B2 patent drawing
  • US10755701B2 patent drawing

AI summary

The present disclosure proposes a method and an apparatus for converting English speech information into a text. The method may include: receiving the English speech information inputted by a user, determining a target speech recognition model according to a preset algorithm, and identifying original phonemes of the English speech information by applying the target speech recognition model; performing a matching on the original phonemes by applying a phonetic model generated by pre-training English texts and a preset probability model, and determining a target phoneme matched successfully; and acquiring a target English text corresponding to the target phoneme, and displaying the target English text on a speech conversion textbox.