Speech Recognition Using Visual Biometric Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems face challenges in adapting to speakers with different vocal characteristics and environments with varying noise and reverberation conditions, requiring extensive voice samples for accurate adaptation, which can lead to frustration and inefficiency.

Innovation Solution

The system uses visual information to extract biometric and environmental features, estimating adaptation parameters before speech recognition, allowing for prompt dereverberation and noise cancellation, and refining these parameters with voice samples for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition systems use extensive voice samples for adaptation, then recognition accuracy improves, but system complexity and user time investment increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidtime investment for adaptation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary adaptation by extracting biometric features (gender, age, ethnicity) from visual images of the speaker before any voice samples are collected. This allows the acoustic model to be pre-adapted to the speaker's characteristics, eliminating the need for extensive voice sample collection and achieving accurate recognition with minimal or no voice samples.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces visual information (images) as an intermediary to bridge the gap between speaker characteristics and acoustic model adaptation. By extracting biometric features from images, the system creates a visual-to-acoustic mapping that enables adaptation without requiring extensive direct acoustic data (voice samples).

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the system collects extensive voice samples for adaptation, then adaptation accuracy improves, but ease of operation deteriorates

Engineering Contradiction:
Improveadaptation accuracyVSAvoiduser convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs adaptation准备工作 in advance by processing visual images to extract biometric features and generating adapted acoustic models before the user needs to speak. This preliminary adaptation eliminates the burden of requiring users to record extensive voice samples, making the system easy to operate while maintaining high adaptation accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces the traditional acoustic-only adaptation mechanism with a visual-acoustic hybrid approach. By substituting some of the acoustic data collection requirement with visual information processing, the system maintains high adaptation accuracy while significantly improving ease of operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If acoustic models are adapted using traditional methods, then recognition accuracy improves, but processing time increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidadaptation processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs acoustic model adaptation in advance by processing visual images to extract biometric features and generating speaker-specific acoustic models before speech recognition is needed. This preliminary adaptation eliminates the need for time-consuming real-time adaptation processing, enabling rapid recognition while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system segments the adaptation process into distinct stages: visual feature extraction from images, biometric parameter estimation, and acoustic model adaptation. By segmenting and pre-processing these steps, the system reduces the computational burden during actual speech recognition, achieving both high accuracy and fast processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8660842B2Enhancing speech recognition using visual information
Publication Date: 2014.02.25 HONDA MOTOR CO LTD
  • US8660842B2 patent drawing
  • US8660842B2 patent drawing
  • US8660842B2 patent drawing

AI summary

Speech recognition device uses visual information to narrow down the range of likely adaptation parameters even before a speaker makes an utterance. Images of the speaker and/or the environment are collected using an image capturing device, and then processed to extract biometric features and environmental features. The extracted features and environmental features are then used to estimate adaptation parameters. A voice sample may also be collected to refine the adaptation parameters for more accurate speech recognition.