Voice Recognition Device Using Visual Trigger Events

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition devices face challenges in accurately determining the desired utterance section in noisy environments, as they struggle to differentiate between the user's voice and ambient noise, and rely on unreliable methods such as lip motion analysis or user operation, which fail when the user is not directly interacting with the device.

Innovation Solution

A voice recognition system that uses a combination of microphone arrays and camera inputs to determine the voice source direction and section by analyzing phase differences in sound signals and visual triggers, such as face direction and posture, to enhance the target sound and reduce ambient noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If noise reduction techniques such as beam forming or echo cancellation are used, then voice recognition accuracy is improved, but it is still difficult to achieve sufficient voice recognition accuracy in noisy environments

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidambient noise impact
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces image information from a camera as an intermediary to assist in determining the voice section. The image processing unit detects face direction and posture, which serve as visual cues to identify when the user is actively speaking. This visual intermediary complements the audio-based beam forming technique, allowing the system to more accurately distinguish the user's voice from ambient noise by cross-referencing both audio and visual data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If lip motion detection is used to determine utterance section, then voice section identification is improved, but inaccurate detection occurs when unrelated motions such as gum chewing are made

Engineering Contradiction:
Improveutterance section determination accuracyVSAvoiddetection reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system employs feedback by continuously monitoring both audio and visual data streams and adjusting the voice section determination based on multiple indicators. The voice section determination unit receives input from both the audio processing unit (which analyzes sound characteristics) and the image processing unit (which analyzes face direction and posture). This feedback mechanism allows the system to distinguish between genuine speaking motions and unrelated motions like gum chewing by looking for consistent patterns across both modalities.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If user operation is required to determine voice section, then accurate voice section identification is achieved, but the method becomes unusable when the user is apart from the device

Engineering Contradiction:
Improvevoice section determination accuracyVSAvoiduser interaction requirement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system implements self-service by automatically detecting the voice section through analysis of audio and visual data without requiring any user operation. The voice section determination unit autonomously identifies when the user is speaking by analyzing sound characteristics from the microphone array and visual cues from the camera, such as face direction and posture changes. This eliminates the need for users to manually press buttons or interact with the device, allowing accurate voice section identification even when the user is at a distance.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach significantly improves voice recognition accuracy by accurately identifying the voice source direction and section, even in noisy conditions, reducing the impact of ambient noise and eliminating the need for direct user interaction.

Implementation Method 1

analyzing phase differences in sound signals

Methodology Applied
Scientific EffectPhase difference analysis:

Implementation Method 2

acquisition sound acquired by a microphone includes various kinds of noises

Methodology Applied
Scientific EffectAcoustic wave propagation: Sound

Data Source

PatentEP2956940B1Voice recognition device, voice recognition method, and program
Publication Date: 2019.04.03 SONY GROUP CORP
  • EP2956940B1 patent drawingFigure 1
  • EP2956940B1 patent drawingFigure 2
  • EP2956940B1 patent drawingFigure 3

AI summary

By recognizing visual trigger events to determine start points and/or end points of voice data signals, the negative effects of noise on voice recognition may be significantly minimized. The visual trigger events may be predetermined gestures and/or predetermined postures of a user captured by a camera, which allow a system to appropriately focus attention on a user to optimize the receipt of a voice command in a noisy environment. This may be accomplished through the assistance of visual feedback complementing the voice feedback provided to the system by the user. Since the visual trigger events are predetermined gestures and/or postures, the system may be able to distinguish which sounds produced by a user are voice commands and which sounds produced by the user is noise that in unrelated to the operation of the system.