Language Adaptivity in Speech Recognition Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multilingual speech recognition systems face low accuracy due to confusion between pronunciations in different languages, leading to a worse user experience and reduced processing efficiency.

Innovation Solution

A method and apparatus for speech recognition based on language adaptivity, which involves extracting phoneme features from voice data, using a pre-trained language discrimination model to determine the language, and switching to the corresponding language acoustic model for accurate recognition, thereby avoiding confusion between languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a hybrid acoustic pronunciation unit set including multiple languages is used to train an acoustic model, then the speech recognition system can support multiple languages, but the recognition accuracy in different languages is greatly affected due to pronunciation confusion

Engineering Contradiction:
Improvemultilingual supportVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the multilingual speech recognition system into separate language-specific acoustic models and a language discrimination model. Instead of using a single hybrid acoustic model that mixes pronunciations from multiple languages, the system divides the recognition process into two stages: first identifying the language through the discrimination model, then routing to the appropriate language-specific acoustic model. This segmentation eliminates pronunciation confusion while maintaining multilingual support.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a language discrimination model as an intermediary component between the voice input and the acoustic models. This intermediary first analyzes the phoneme features of the input voice data to determine the language type, then directs the recognition process to the corresponding language-specific acoustic model. This intermediary structure prevents direct mixing of different language pronunciations while enabling seamless multilingual recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual language selection is required for speech recognition, then accurate language-specific recognition can be achieved, but unnecessary user operations reduce processing efficiency

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent enables the speech recognition system to automatically identify and select the appropriate language without user intervention. The language discrimination model autonomously analyzes the input voice data's phoneme features to determine the language type, then automatically routes to the corresponding acoustic model. This self-service mechanism eliminates the need for manual language selection while maintaining accurate language-specific recognition, thereby improving processing efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs language identification as a preliminary action before executing the main speech recognition task. By first analyzing the phoneme features and determining the language type through the discrimination model, the system prepares the appropriate language-specific acoustic model in advance. This preliminary language identification action enables subsequent accurate recognition without requiring user input, streamlining the overall processing workflow.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12033621B2Method for speech recognition based on language adaptivity and related apparatus
Publication Date: 2024.07.09 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12033621B2 patent drawing
  • US12033621B2 patent drawing
  • US12033621B2 patent drawing

AI summary

A method for speech recognition based on language adaptivity comprises obtaining voice data of a user. The method also comprises extracting, based on the obtained voice data, a phoneme feature representing pronunciation phoneme information. The phoneme feature is input to a pre-trained language discrimination model that is pre-trained based on a multilingual corpus. A language discrimination result corresponding to the phoneme feature and in accordance with the language discrimination model is obtained. The method also comprises obtaining a speech recognition result of the voice data based on a language acoustic model of a language corresponding to the language discrimination result. The method further comprises determining a speech recognition result of the voice data based on a language acoustic model of a language corresponding to the language discrimination result.