Wakeup Model Generation Using Microphone Transfer Functions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face challenges in efficiently recognizing a wakeup word across different devices due to variations in microphone characteristics, leading to inconsistent performance and user reluctance to use voice agents due to repeated learning requirements.
Innovation Solution
An electronic device method that involves obtaining audio data from an external device, converting it using a transfer function specific to the device's microphone, and generating a wakeup model to verify the wakeup word, allowing for seamless recognition across various devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio data is directly used from external devices without conversion, then the process is simple and fast, but the wakeup word recognition accuracy deteriorates due to microphone characteristic variations
Solution Approach 1:
The patent introduces a transfer function as an intermediary mathematical model that mediates between the external device's audio data and the target device's microphone characteristics. This transfer function acts as a bridge, transforming audio data from one microphone's characteristic space to another, thereby resolving the mismatch without requiring physical re-recording or complex adaptive learning processes.
Solution Approach 2:
The patent changes the parameters of the audio data by applying a transfer function that adjusts frequency response, gain, and other acoustic parameters. This parameter transformation aligns the external device's audio characteristics with the target device's microphone profile, enabling accurate wakeup word recognition across different devices without retraining the recognition model.
2Reliability
If a wakeup model is trained for each specific device, then recognition accuracy is high, but the user experience deteriorates due to repeated learning requirements when switching devices
Solution Approach 1:
The patent creates a universal transfer function that can be applied across multiple devices with different microphone characteristics. Instead of training separate wakeup models for each device, the system uses a single wakeup model combined with device-specific transfer functions, making the system universally applicable across different electronic devices while maintaining high recognition accuracy.
Solution Approach 2:
The patent performs preliminary action by pre-calculating and storing transfer functions for different device types before actual wakeup word recognition is needed. These transfer functions are prepared in advance based on microphone characteristics, so when a user switches devices, the appropriate transfer function is already available for immediate application, eliminating the need for on-the-spot relearning or retraining.
3Reliability
If audio data is converted using transfer function, then microphone characteristic variations are compensated, but processing time increases
Solution Approach 1:
The transfer functions are pre-calculated and stored based on microphone characteristics of different device types. This preliminary preparation means that during actual wakeup word recognition, the system only needs to apply the pre-existing transfer function rather than performing complex real-time calculations, significantly reducing processing time while maintaining recognition consistency across devices.
Data Source
AI summary
In accordance with an aspect of the disclosure, an electronic device comprises a first audio receiving circuit; a communication circuit; at least one processor operatively connected to the first audio receiving circuit and the communication circuit; and a memory operatively connected to the at least one processor, wherein the memory stores one or more instructions that, when executed, cause the at least one processor to: obtain first audio data, wherein the first audio data is based on a user utterance recorded by an external electronic device, through the communication circuit; convert the first audio data into second audio data, using a first transfer function of the first audio receiving circuit; and generate a wakeup model using the second audio data, the wakeup model configured to verify a wakeup word associated with the first audio data.


