TV Remote Wake Word Detection for AI Speaker Interference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The speech recognition performance of AI speakers is degraded in large living spaces due to audio outputs from nearby devices, such as TVs, making it difficult to control connected home appliances effectively.
Innovation Solution
A method that uses a TV and a remote control to automatically recognize a specific AI speaker and its wake-up word, by outputting a second wake-up word through the TV speaker, and adjusting the audio signal output based on the TV's volume level or ambient noise to improve speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the TV outputs audio signals during content playback, then the TV provides entertainment functionality, but the speech recognition performance of the AI speaker is degraded due to audio interference
Solution Approach 1:
The TV proactively detects the wake-up word and takes preliminary action to reduce or mute audio output before the AI speaker's speech recognition is affected by the audio interference. This preventive measure eliminates the harmful effect before it degrades recognition performance.
Solution Approach 2:
The system uses the TV's microphone to detect wake-up words and provides feedback by automatically adjusting the audio output level. This closed-loop feedback mechanism ensures that audio interference is reduced precisely when speech recognition is needed, improving reliability without permanently sacrificing entertainment functionality.
2Reliability
If the TV manually reduces audio output for AI speaker control, then speech recognition improves, but user convenience deteriorates due to manual intervention required
Solution Approach 1:
The TV system automatically detects wake-up words through its microphone and self-adjusts the audio output level without requiring user intervention. The system serves itself by autonomously managing the conflict between entertainment audio and speech recognition needs, thereby maintaining both reliability and ease of operation.
Solution Approach 2:
The TV dynamically changes the audio output parameter (volume level) based on detected wake-up words. This automatic parameter adjustment eliminates the need for manual user intervention while ensuring optimal speech recognition conditions, resolving the contradiction between reliability and ease of operation.
3Reliability
If the TV automatically detects and responds to wake-up words, then speech recognition performance improves, but device complexity increases due to additional processing requirements
Solution Approach 1:
The TV's existing microphone, which is already part of the system for other functions, is utilized to detect wake-up words. This multi-functional use of existing components improves speech recognition performance without significantly increasing device complexity, as the same hardware serves multiple purposes.
Solution Approach 2:
The TV acts as an intermediary between the user and the AI speaker. Instead of requiring direct complex interaction between the user and AI speaker, the TV mediates by detecting wake-up words and adjusting audio levels, thereby simplifying the overall system architecture while improving speech recognition reliability.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution enhances the speech recognition performance of AI speakers by reducing interference from TV audio outputs, allowing for quicker and more reliable control of connected devices using a TV remote control.
Implementation Method 1
receiving a first wake-up word corresponding to an AI speaker of the STB through a microphone of the remote control
Implementation Method 2
transmitting the first wake-up word to a transceiver of the TV through a wireless communication transceiver of the remote control
Implementation Method 3
outputting a second wake-up word through the speaker of the TV
Data Source
AI summary
A control method for a system comprising a TV and a remote control, according to an embodiment of the present invention, comprises the steps of: outputting, through a screen of the TV, a video signal of content received from an STB; outputting, through a speaker of the TV, an audio signal of content received from the STB; receiving, through a microphone of the remote control, a first wake word corresponding to an AI speaker of the STB; transmitting, through a wireless communication transceiver of the remote control, the first wake word to a transceiver of the TV; and outputting, through the speaker of the TV, a second wake word.


