Voice Detection Using Time-Gated Audio Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice detection systems in electronic devices suffer from deteriorated keyword detection performance due to ambient noise and echo issues, particularly when specific sound sources like music or TTS are reproduced, causing keyword detection malfunctions.
Innovation Solution
An electronic device with a processor and memory configured to temporarily store messages, output them through a speaker, and process input sounds using time information to detect keywords while ignoring parts of the input sound, especially during sound source reproduction, thereby preventing keyword distortion and improving detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If keyword detection is performed using pattern matching technology, then voice detection function is enabled, but keyword detection performance deteriorates when ambient noise or echo is present
Solution Approach 1:
The system performs preliminary analysis to determine whether a sound source is being reproduced before executing keyword detection. This preliminary action allows the system to prepare appropriate processing strategies in advance, preventing keyword detection malfunctions by anticipating potential echo interference scenarios.
Solution Approach 2:
The keyword detection process is made dynamic by adjusting its execution based on real-time conditions. When sound source reproduction is detected, the system dynamically modifies the keyword detection behavior, such as temporarily suspending detection or adjusting sensitivity thresholds, to adapt to changing acoustic environments.
2Ease of operation
If sound source reproduction is performed, then audio output function is enabled, but echo interferes with microphone input causing keyword detection malfunction
Solution Approach 1:
The system extracts and identifies the specific time periods when sound sources are reproduced by analyzing audio output patterns. By separating the sound source reproduction periods from other time periods, the system can apply different processing strategies during these extracted intervals, such as ignoring microphone input or adjusting detection thresholds.
Solution Approach 2:
The system introduces an intermediary analysis layer that monitors both speaker output and microphone input simultaneously. This intermediary mechanism compares the timing and characteristics of sound source reproduction with captured audio, enabling the system to identify and filter out echo interference before it affects keyword detection.
3Adaptability or versatility
If keyword detection is performed during sound source reproduction, then voice activation is maintained, but keyword distortion occurs reducing detection accuracy
Solution Approach 1:
The system dynamically adjusts keyword detection parameters based on whether sound source reproduction is occurring. During reproduction periods, detection sensitivity and threshold levels are dynamically modified to account for the changed acoustic conditions, maintaining adaptability while preserving detection accuracy.
Solution Approach 2:
The system changes detection parameters such as sensitivity thresholds, matching criteria, and processing gain when sound source reproduction is detected. These parameter changes allow the keyword detection algorithm to maintain accuracy despite the presence of echo and altered acoustic conditions during audio output.
Data Source
AI summary
An electronic device is provided, which includes a housing; a microphone located on or within a predetermined distance of a first portion of the housing; a speaker located on or within a predetermined distance of a second portion of the housing; a communication circuit; a processor electrically connected to the microphone, the speaker, and the communication circuit; and a memory electrically connected to the processor configured to store a message to be provided as a voice through the speaker, wherein the memory stores instructions, wherein the processor is configured to execute the instructions to perform operations comprising: determining time information corresponding to a first part of the message if providing of the message is necessary, outputting the message through the speaker, receiving an input sound through the microphone while at least a part of the message is output, and processing the input sound using the time information to detect at least one word or sentence from the input sound, and the processing the input sound includes processing the input sound by ignoring at least a part of the input sound using the time information.


