Electronic Device Wake-Up Detection Without Acoustic Echo Cancellation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition devices face challenges in distinguishing user speech from device output due to the limitations of acoustic echo cancellation (AEC) technology, which requires synchronized voice input and output and increases manufacturing costs.
Innovation Solution
An electronic device with a speaker, microphone, memory, and processor that uses a wake-up word detection model to identify and activate voice recognition models based on the presence of predefined wake-up words in input and output sound signals, allowing for efficient voice recognition without relying on AEC technology.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If AEC technology is used to distinguish user speech from device output, then voice recognition accuracy is improved, but computational complexity and manufacturing cost increase
Solution Approach 1:
The patent segments the voice recognition process into two distinct stages: wake-up word detection and voice recognition. The wake-up word detection model processes only the speaker output signal to determine if the device should be activated, while the full voice recognition model processes the microphone input. This segmentation reduces the computational burden on the wake-up word detection stage and simplifies the overall system architecture compared to applying AEC throughout the entire voice recognition process.
Solution Approach 2:
The patent applies preliminary action by using a dedicated wake-up word detection model to first determine whether the device should be activated before engaging the full voice recognition model. This preliminary detection stage processes only the speaker output signal to identify wake-up words, avoiding the need for complex AEC calculations during the activation decision-making process. Only after this preliminary action does the system proceed to full voice recognition if activation is confirmed.
2Measurement precision
If AEC technology is used to synchronize voice signals, then voice recognition accuracy is improved, but manufacturing cost increases
Solution Approach 1:
The patent extracts the wake-up word detection function from the main voice recognition process and implements it as a separate, simpler model that processes only the speaker output signal. This extraction eliminates the need for complex AEC synchronization during the wake-up word detection stage, reducing computational requirements and manufacturing costs while maintaining voice recognition accuracy through the dedicated wake-up word detection model.
Solution Approach 2:
The patent employs a simpler, less expensive wake-up word detection model for the initial activation decision, rather than using a complex AEC-based voice recognition model throughout. This wake-up word detection model is designed to be computationally lightweight and cost-effective, processing only the speaker output signal to determine device activation. Only when activation is confirmed does the system engage the more expensive full voice recognition model, thereby reducing overall manufacturing costs.
3Loss of time
If wake-up word detection model is used to identify wake-up words in speaker output, then response time is improved, but computational resources are consumed
Solution Approach 1:
The patent applies local quality by using different processing approaches for different signals: the speaker output signal is processed by the wake-up word detection model to identify wake-up words and determine device activation, while the microphone input signal is processed by the full voice recognition model only after activation is confirmed. This localized application of processing models ensures that computational resources are consumed only when necessary, improving response time for wake-up word detection while managing computational resource usage efficiently.
Data Source
AI summary
An electronic device including a speaker; a microphone; a memory configured to store a voice recognition model; and a processor configured to: identify, while a first sound signal is input through the microphone, whether a wake-up word is included in a second sound signal that is output through the speaker by inputting the second sound signal into a wake-up word detection model, and identify, based on identifying that the wake-up word is not included in the second sound signal, whether the wake-up word is included in the first sound signal by inputting the first sound signal into the wake-up word detection model.


