Speech Recognition Adaptation via Language Skill Profiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in achieving accurate speaker adaptation due to limited availability of speaker-specific speech data, which is costly in terms of computational complexity and memory consumption, and often does not accurately represent voice characteristics of each user.

Innovation Solution

The system employs information indicative of language skills of users to adapt and improve speech recognition performance by building user-specific acoustic models, utilizing a client-server architecture that collects and utilizes user data to adjust acoustic models and pronunciation rules based on native and non-native language proficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker adaptation is performed using speaker-specific speech data, then speech recognition accuracy is improved, but computational complexity and memory consumption increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates simplified copies of speaker-specific acoustic models by using language skill information as a proxy. Instead of building full speaker adaptation models requiring extensive speaker-specific data, the system generates simplified acoustic models based on language skill profiles, which capture essential speaker characteristics with reduced computational requirements.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameters used for speaker adaptation from traditional speaker-specific acoustic features to language skill information. By representing speakers through language skill parameters (native language, proficiency levels) rather than extensive acoustic measurements, the system reduces the dimensionality and complexity of the adaptation process while maintaining recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If speaker adaptation is performed using speaker-specific speech data, then speech recognition accuracy is improved, but memory consumption increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidmemory consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system creates compact representations of speaker characteristics by copying only the essential language skill information rather than storing extensive speaker-specific speech data. This allows the acoustic models to be adapted to individual speakers using minimal memory resources while preserving recognition accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts only the critical language skill information from speaker profiles, separating essential adaptation parameters (language skills) from non-essential data. This extraction process reduces memory consumption by retaining only the most relevant features for speaker adaptation.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If speaker adaptation is performed for each speaker, then speech recognition accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary adaptation by pre-processing language skill information and creating baseline acoustic models based on language proficiency before actual speech recognition occurs. This preliminary preparation reduces the processing time required during real-time recognition while maintaining speaker-specific accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By changing the adaptation parameters to language skill information, which can be processed more quickly than extensive acoustic data, the system reduces processing time. Language skill parameters require less computational processing while still enabling effective speaker adaptation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3097553B1Method and apparatus for exploiting language skill information in automatic speech recognition
Publication Date: 2022.06.01 NUANCE COMMUNICATIONS INC
  • EP3097553B1 patent drawingFigure 1
  • EP3097553B1 patent drawingFigure 2
  • EP3097553B1 patent drawingFigure 3

AI summary

Typical speech recognition systems usually use speaker-specific speech data to apply speaker adaptation to models and parameters associated with the speech recognition system. Given that speaker-specific speech data may not be available to the speech recognition system, information indicative of language skills is employed in adapting configurations of a speech recognition system. According to at least one example embodiment, a method and corresponding apparatus, for speech recognition comprise maintaining information indicative of language skills of users of the speech recognition system. A configuration of the speech recognition system for a user is determined based at least in part on corresponding information indicative of language skills of the user. Upon receiving speech data from the user, the configuration of the speech recognition system determined is employed in performing speech recognition.