Multi-Modal Speech Input via Audio and Image Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for inputting commands on portable devices face challenges due to their small form factor, with voice recognition being inaccurate due to background noise and requiring lengthy training processes, and lacking efficient multi-modal input solutions.

Innovation Solution

A computing device that combines audio and image capture to identify user inputs through speech recognition and image analysis, using state-specific dictionaries and gesture recognition to enhance accuracy and reduce resource intensity, allowing for immediate use without extensive training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice recognition is used for input on portable devices, then hands-free operation is enabled, but accuracy decreases due to background noise and device proximity

Engineering Contradiction:
Improvehands-free operationVSAvoidvoice input accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent combines audio data from microphones with image data from cameras to create a multi-modal input system. The speech recognition module processes audio while the facial expression recognition module processes visual data, and their results are integrated to improve overall input accuracy, especially in noisy environments where pure voice recognition fails.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary multi-modal processing layer that mediates between the raw audio/image inputs and the final command interpretation. This layer fuses speech recognition results with facial expression analysis results, allowing the system to disambiguate commands and improve accuracy when one modality is degraded by background noise.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conventional speech recognition is implemented, then voice input capability is provided, but lengthy training processes are required

Engineering Contradiction:
Improvevoice input capabilityVSAvoidtraining process duration
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs self-training by automatically collecting and analyzing user speech and facial expression data during normal device operation. The machine learning models are continuously refined through self-service learning from real-world usage patterns, eliminating the need for separate lengthy training sessions while maintaining personalized recognition accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates preliminary action by pre-training the speech and facial expression recognition models with large datasets before deployment. This pre-training enables the system to achieve functional capability immediately upon first use, with only minimal adaptation needed for individual user patterns, thereby reducing the perceived training time for end users.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multiple input modalities are combined, then input accuracy is improved, but device complexity increases

Engineering Contradiction:
Improveinput accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the multi-modal processing system into distinct modular components: audio capture module, image capture module, speech recognition module, facial expression recognition module, and fusion module. Each module operates independently with well-defined interfaces, allowing the complex system to be developed, tested, and maintained as separate manageable units while achieving high input accuracy through their coordinated operation.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If audio and image processing are performed simultaneously, then multi-modal recognition accuracy is enhanced, but computational resources are consumed

Engineering Contradiction:
Improverecognition accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system employs periodic action by processing audio and image data at different frequencies and only performing full multi-modal fusion when necessary. For example, audio may be processed continuously at a lower computational level, while full facial expression analysis is performed periodically or triggered by specific events, reducing overall energy consumption while maintaining recognition accuracy when needed.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS8700392B1Speech-inclusive device interfaces
Publication Date: 2014.04.15 AMAZON TECH INC
  • US8700392B1 patent drawing
  • US8700392B1 patent drawing
  • US8700392B1 patent drawing

AI summary

A user can provide input to a computing device through various combinations of speech, movement, and/or gestures. A computing device can analyze captured audio data and analyze that data to determine any speech information in the audio data. The computing device can simultaneously capture image or video information which can be used to assist in analyzing the audio information. For example, image information is utilized by the device to determine when someone is speaking, and the movement of the person's lips can be analyzed to assist in determining the words that were spoken. Any gestures or other motions can assist in the determination as well. By combining various types of data to determine user input, the accuracy of a process such as speech recognition can be improved, and the need for lengthy application training processes can be avoided.