Multilingual Speech Recognition via Automatic Language Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems face challenges in handling multilingual users, requiring users to select a single language and involving inconvenient setup processes, as more than 50% of people use multiple languages in their daily lives.

Innovation Solution

The development of speech recognition systems that can recognize multiple languages simultaneously by using language recognition modules for different languages, identifying the candidate language through phonological analysis, and selecting recognition candidates based on recognition scores and user input, allowing for seamless switching between languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single language model is used for speech recognition, then the system is simple and easy to operate, but it cannot recognize multiple languages

Engineering Contradiction:
Improvemultilingual recognition capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by creating a single speech recognition system that can handle multiple languages through a unified architecture. The system uses a common acoustic model that works across languages, combined with language-specific language models and dictionaries, allowing one system to serve multiple linguistic functions without requiring separate recognition systems for each language.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the speech recognition system into distinct modular components: a universal acoustic model, separate language models for each language, and language-specific dictionaries. This segmentation allows the system to maintain simplicity in the acoustic processing while managing linguistic complexity through separate, manageable modules that can be independently selected based on the input language.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If multiple language models are used simultaneously, then multilingual recognition is enabled, but the setup process becomes complex and inconvenient

Engineering Contradiction:
Improvelanguage recognition capabilityVSAvoiduser configuration convenience
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements self-service by enabling the system to automatically detect the language of the input audio and select the appropriate language model without requiring user intervention. The system analyzes the acoustic characteristics of the speech signal to identify the language and autonomously configures which language model to use, eliminating the need for users to manually select or configure language settings.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by pre-configuring the system with multiple language models and dictionaries ready for immediate use. The system prepares language detection capabilities in advance, allowing it to quickly identify and switch between languages based on the input audio without requiring users to perform setup or configuration actions during actual use.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If language identification is performed separately from recognition, then the recognition process is simple, but the overall process time increases

Engineering Contradiction:
Improvelanguage identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges language identification and speech recognition into a single integrated processing step. The system simultaneously performs acoustic analysis for both language detection and speech recognition, using the same acoustic model and processing pipeline to accomplish both tasks concurrently, thereby eliminating the time penalty associated with separate sequential processing while maintaining high accuracy in language identification.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9129591B2Recognizing speech in multiple languages
Publication Date: 2015.09.08 GOOGLE LLC
  • US9129591B2 patent drawing
  • US9129591B2 patent drawing
  • US9129591B2 patent drawing

AI summary

Speech recognition systems may perform the following operations: receiving audio; recognizing the audio using language models for different languages to produce recognition candidates for the audio, where the recognition candidates are associated with corresponding recognition scores; identifying a candidate language for the audio; selecting a recognition candidate based on the recognition scores and the candidate language; and outputting data corresponding to the selected recognition candidate as a recognized version of the audio.