Video Display Voice Control via Spatial Microphone Echo Cancellation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video display apparatuses face difficulties in discriminating sound signals intended for operation from the sounds emitted by the apparatus itself and background noise, making voice control ineffective due to echo and noise interference.

Innovation Solution

A video display apparatus equipped with at least two spatially separated microphones, an audio signal processing unit, a sound source locating unit, a voice recognition unit, and a voice command execution unit, allowing for the separation of sound emitted by the loudspeaker from ambient sounds and precise localization of voice commands for executing control functions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice control is implemented using a single microphone, then the apparatus can receive voice commands, but the sound signals intended for operation cannot be discriminated from sounds emitted by the loudspeaker and background noise

Engineering Contradiction:
Improvevoice control operationVSAvoidsound source discrimination accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent divides the sound reception function into multiple spatially separated microphones (at least two microphones positioned at different locations). This segmentation allows the system to capture sound signals from different spatial perspectives, enabling discrimination between voice commands and other sounds (loudspeaker output, background noise) through spatial analysis and signal processing.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If multiple spatially separated microphones are used, then sound source discrimination becomes possible, but the device complexity increases

Engineering Contradiction:
Improvesound source discrimination accuracyVSAvoidmicrophone array and signal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent integrates multiple functions into a single apparatus: the video display apparatus simultaneously performs display, sound output, and voice control reception. The microphones serve dual purposes - capturing both ambient sound for voice commands and monitoring loudspeaker output for echo cancellation. This multi-functionality reduces the need for separate dedicated control devices, offsetting the complexity increase through functional consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements feedback mechanisms where the microphones capture both voice commands and loudspeaker output, and the processor analyzes these signals to distinguish between them. The echo cancellation feature uses feedback from the loudspeaker signal to subtract echoed sound from the microphone input, improving voice command recognition accuracy while managing the complexity through systematic signal processing.

Inventive Principle:
Principle #23Feedback

3Use of energy by moving object

If the apparatus emits sound through loudspeaker, then audio output is provided, but echoes and background noise interfere with voice command recognition

Engineering Contradiction:
Improveaudio output functionVSAvoidecho and noise interference
Core Design Contradiction:
Use of energy by moving objectVSObject-affected harmful factors

Solution Approach 1:

The patent converts the harmful effect of loudspeaker sound (which creates echoes that interfere with voice recognition) into a beneficial element. The system uses the same loudspeaker output signal as a reference to actively cancel echoes from the microphone input through adaptive filtering. By utilizing the loudspeaker signal itself for cancellation purposes, the system transforms the source of interference into a tool for eliminating that interference, enabling reliable voice control despite the presence of audio output.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables reliable voice-controlled operation of video display apparatuses without the need for separate control devices, enhancing user interaction through hands-free multimedia and gaming applications by accurately distinguishing and executing voice commands amidst background noise and echoes.

Implementation Method 1

at least two spatially separated microphones... receiving a sound

Methodology Applied
Scientific EffectSound propagation: Sound

Implementation Method 2

an audio signal processing unit configured to separate the sound emitted by the loudspeaker and received by the microphones from sound received by the microphones

Methodology Applied
Scientific EffectEcho cancellation: Echo

Data Source

PatentEP3349480B1Video display apparatus and method of operating the same
Publication Date: 2020.09.02 VESTEL ELEKTRONIK SANAYI & TICARET ANONIM SIRKETI
  • EP3349480B1 patent drawingFigure 1
  • EP3349480B1 patent drawingFigure 2
  • EP3349480B1 patent drawingFigure 3~4

AI summary

The present invention provides a video display apparatus (100) at least comprising a display screen (10), at least one loudspeaker for emitting a sound in association with at least one still or moving image displayed on the display screen (10), at least two spatially separated microphones (1, 2, 3), and an audio signal processing unit (20) configured to separate the sound emitted by the loudspeaker from a sound received by the microphones. The present invention also provides a method of operating a video display apparatus, wherein the method at least comprises displaying on a display screen of the apparatus at least one still or moving image, emitting a sound from a loudspeaker of the apparatus in association with displaying the at least one still or moving image, receiving a sound by at least two spatially separated microphones of the apparatus, and separating the sound emitted from the loudspeaker from the sound received by the microphones. Such a method allows for three-dimensional localization and separation of sound sources to receive and execute voice commands for control of a video display apparatus, such as a television, without the need for a remote control.