Image-Aided Voice Recognition for Multi-User Command Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice recognition technologies face challenges in adapting to multiple-user environments, leading to increased instances of false positives and reduced efficiency in voice command recognition.

Innovation Solution

The use of image data to capture and analyze the vicinity of a device, adjusting voice recognition parameters such as trigger thresholds and algorithms based on the presence of individuals, and employing gaze detection to authenticate and prioritize voice inputs from authorized users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice recognition is performed without image data, then the system is simpler and faster, but accuracy in multiple-user environments deteriorates

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines voice recognition with image recognition by merging the voice processing module and image processing module into an integrated system. The voice recognition module processes audio input while the image recognition module processes camera input, and both modules work together to identify the active user through multiple modalities, thereby improving accuracy without significantly increasing perceived system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an active user identification module that acts as an intermediary between the voice recognition module and the device control system. This module uses both voice and image data to determine which user is actively speaking and intends to interact with the device, then provides this information to adjust voice recognition parameters accordingly.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If voice recognition parameters are fixed, then the system is more stable, but adaptability to multiple-user environments deteriorates

Engineering Contradiction:
Improveadaptation to multiple usersVSAvoidparameter stability
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent implements dynamic adjustment of voice recognition parameters based on real-time identification of the active user. When a different user is detected through image or voice analysis, the system automatically adjusts parameters such as sensitivity thresholds, accepted voice commands, and language models to match the newly identified user's profile, enabling seamless adaptation while maintaining operational stability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes multiple parameters including voice recognition sensitivity, acceptable dialect ranges, and command interpretation thresholds based on the identified user's characteristics. These parameter changes allow the same device to serve multiple users with different voice patterns and preferences while maintaining stable performance for each individual user.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If image processing is added to aid voice recognition, then user differentiation improves, but processing time increases

Engineering Contradiction:
Improveuser differentiation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously capturing and pre-processing image data in the background even before voice input is received. Facial recognition features are extracted and stored in advance, so when voice recognition is triggered, the system can quickly match the current speaker against pre-processed user profiles without performing full image analysis from scratch, significantly reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a selective processing approach where the system skips full image analysis in favor of quicker verification methods when possible. For example, if basic facial recognition quickly identifies the active user, the system skips more time-consuming detailed analysis and proceeds directly to voice recognition parameter adjustment, thereby maintaining accuracy while minimizing time loss.

Inventive Principle:
Principle #21Skipping (Rushing through)

Data Source

PatentUS20240221745A1Method and Apparatus for Using Image Data to Aid Voice Recognition
Publication Date: 2024.07.04 GOOGLE TECHNOLOGY HOLDINGS LLC
  • US20240221745A1 patent drawing
  • US20240221745A1 patent drawing
  • US20240221745A1 patent drawing

AI summary

A device performs a method for using image data to aid voice recognition. The method includes the device capturing image data of a vicinity of the device and adjusting, based on the image data, a set of parameters for voice recognition performed by the device. The set of parameters for the device performing voice recognition include, but are not limited to: a trigger threshold of a trigger for voice recognition; a set of beamforming parameters; a database for voice recognition; and/or an algorithm for voice recognition. The algorithm may include using noise suppression or using acoustic beamforming.