Voice Registration Neural Network Feature Vector Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic apparatuses face challenges in registering high-quality wake-up voice inputs efficiently, often requiring multiple utterances and risking activation by similar texts.
Innovation Solution
An electronic apparatus equipped with a microphone, communication interface, memory, and processor that uses a trained neural network model to process user voice inputs, compares feature vectors with verification data sets from an external server, and registers the wake-up voice input based on similarity thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the electronic apparatus requires multiple utterances (e.g., 5 or more times) to register a wake-up voice input, then the recognition accuracy is improved, but the usability and user convenience deteriorate
Solution Approach 1:
The patent introduces a server as an intermediary to perform wake-up voice verification. The server receives the user's voice input, compares it with pre-stored verification voices, and determines whether to register the wake-up voice. This external intermediary handles the complex verification process, allowing the electronic apparatus to register wake-up voices with only one utterance while maintaining high accuracy through the server's sophisticated comparison algorithms.
2Ease of operation
If the electronic apparatus registers a wake-up voice input after a single utterance, then the usability is improved, but the reliability and accuracy of wake-up recognition deteriorate
Solution Approach 1:
The patent implements preliminary action by pre-storing multiple verification voices corresponding to potential wake-up words on the server before actual wake-up recognition occurs. When a user attempts to register or use a wake-up voice, the system compares it against these pre-prepared verification voices. This advance preparation enables accurate single-utterance recognition without requiring multiple trial utterances during the actual wake-up process.
Solution Approach 2:
The server acts as an intermediary that performs sophisticated voice comparison algorithms, receiving the user's single utterance and comparing it with pre-stored verification voices. This external verification mechanism ensures high reliability of wake-up recognition while allowing the process to complete in a single utterance, thus maintaining both usability and accuracy.
3Ease of operation
If the electronic apparatus uses a simple wake-up voice registration process, then the ease of operation is improved, but the quality and distinctiveness of registered wake-up voices deteriorate
Solution Approach 1:
The patent employs a server as an intermediary to perform comprehensive voice quality verification. The server compares the registered wake-up voice against multiple pre-stored verification voices using sophisticated algorithms, ensuring high voice quality and distinctiveness. This external verification process maintains strict quality standards while keeping the user-facing registration process simple and requiring only a single utterance.
Solution Approach 2:
The patent replaces traditional mechanical repetition-based verification (requiring multiple utterances) with an information-processing approach using neural networks and voice recognition algorithms on the server. This substitution allows the system to achieve high voice quality verification through intelligent processing rather than requiring multiple physical utterances from the user.
Data Source
Figure 1A
Figure 1B
Figure 2
AI summary
An electronic apparatus is disclosed. The electronic apparatus may include a microphone; a communication interface; a memory configured to store at least one instruction; and a processor configured to execute the at least one instruction to: obtain a user voice input for registering a wake-up voice input via the microphone; input the user voice input into a trained neural network model to obtain a first feature vector corresponding to text included in the user voice input; receive a verification data set determined based on information related to the text included in the user voice input from an external server via the communication interface; input a verification voice input included in the verification data set into the trained neural network model to obtain a second feature vector corresponding to the verification voice input; and identify whether to register the user voice input as the wake-up voice input based on a similarity between the first feature vector and the second feature vector.