Voice Extraction Filter for AI Voice Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition technologies face challenges in accurately distinguishing between users in noisy environments and ensuring secure voice control, particularly in systems that require precise user identification and differentiated authority settings.
Innovation Solution
An AI apparatus that registers users with their voice information, uses a voice extraction filter to isolate the registered user's voice, and performs control operations based on activated user registration information and authority levels, enabling precise voice recognition and secure voice control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice recognition technology is used in noisy environments, then voice control function is provided, but voice recognition accuracy deteriorates due to surrounding noise and multiple users
Solution Approach 1:
The patent segments the voice recognition process by introducing user-specific voice extraction filters. Each registered user has a dedicated filter that segments and extracts only their voice from the mixed audio signal, separating the target voice from background noise and other users' voices, thereby maintaining high recognition accuracy in noisy environments
Solution Approach 2:
The patent introduces voice extraction filters as intermediary components between the microphone input and the voice recognition system. These filters act as mediators that pre-process the audio signal by extracting the target user's voice before it reaches the recognition engine, improving accuracy without compromising voice control availability
2Measurement precision
If user-specific voice extraction filters are introduced, then voice recognition accuracy improves, but device complexity increases
Solution Approach 1:
The patent implements a universal voice extraction filter framework that serves multiple functions: noise filtering, user identification, and voice separation. This multi-functional approach achieves high voice recognition accuracy without requiring separate complex systems for each function, thereby limiting the increase in device complexity
Solution Approach 2:
The patent performs preliminary voice extraction and user identification before the main voice recognition process. By pre-processing the audio signal with user-specific filters and identifying the active user in advance, the system simplifies the subsequent recognition task, achieving high accuracy without proportionally increasing overall system complexity
3Reliability
If multiple users are registered with different authority levels, then security is enhanced, but ease of operation decreases due to activation state management
Solution Approach 1:
The patent implements automatic user identification and authority verification through the voice extraction filter system. The system automatically determines which user is speaking and applies their corresponding authority level without requiring manual intervention, maintaining security while simplifying operation by eliminating the need for users to manually manage activation states
Solution Approach 2:
The patent introduces dynamic activation states for user registration information that can be automatically adjusted based on context. The system can dynamically enable or disable specific user profiles based on time, location, or other factors, maintaining security requirements while reducing the operational burden of manual user management
Data Source
AI summary
According to an embodiment of the present invention, an artificial intelligence (AI) apparatus for performing voice control, includes a memory configured to store a voice extraction filter for extracting a voice of a registered user, and a processor to receive identification information of a user and a first voice signal of the user, to register the user using the received identification information, to extract a voice of the registered user from the received second voice signal by using the voice extraction filter corresponding to the registered user, when a second voice signal is received, and to proceed a control operation corresponding to intention information of the extracted voice of the registered user. The voice extraction filter is generated by using the received first voice signal of the registered user.


