Voice Color Conversion for Additional Wake-up Word Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition devices can only recognize a preset wake-up word and do not allow users to set additional wake-up words of their preference, limiting their functionality and convenience, especially for those unfamiliar with the preset word.
Innovation Solution
The implementation of a speech recognition method that enables users to set an additional wake-up word, using voice color conversion technology to generate and learn various spoken utterances for the additional wake-up word, allowing for accurate recognition and activation of the speech recognition function.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If only preset wake-up words are recognized, then device complexity is reduced and processing is simpler, but adaptability and user convenience deteriorate
Solution Approach 1:
The system performs preliminary voice color conversion to generate multiple utterance variations of the additional wake-up word before recognition. This pre-processing step creates a comprehensive set of training data that enables the recognition algorithm to accurately identify the wake-up word across different speaking conditions, thereby improving adaptability without requiring complex real-time processing
Solution Approach 2:
The system creates synthetic copies of the additional wake-up word through voice color conversion, generating multiple utterance variations that simulate different speaking styles and conditions. These copied utterances are then used to train the recognition algorithm, enabling flexible wake-up word recognition without directly processing diverse real-world variations during operation
2Measurement precision
If voice color conversion is used to generate various utterances, then recognition accuracy improves, but processing time and computational load increase
Solution Approach 1:
Voice color conversion and utterance generation are performed in advance during a setup phase, creating a comprehensive training dataset before actual wake-up word recognition is needed. This preliminary processing eliminates the need for real-time voice conversion during operation, maintaining high recognition accuracy while avoiding time loss during critical recognition moments
Solution Approach 2:
The system dynamically adjusts the number and type of utterance variations generated based on the specific characteristics of the additional wake-up word and the recognized speaking environment. This dynamic approach ensures sufficient recognition accuracy while optimizing processing time by avoiding unnecessary generation of excessive utterance variations
3Ease of operation
If multiple additional wake-up words are allowed, then user convenience and satisfaction improve, but system complexity and configuration difficulty increase
Solution Approach 1:
The system automatically performs voice color conversion and generates training data for additional wake-up words without requiring manual configuration of complex parameters. Users simply need to provide their preferred wake-up words, and the system handles the sophisticated processing automatically, thereby improving ease of operation while managing system complexity through automation
Data Source
AI summary
Disclosed are a speech recognition device and a speech recognition method, which perform speech recognition by executing an artificial intelligence (AI) algorithm and/or a machine learning algorithm provided therein, and which can communicate with other electronic devices and an external server in a 5G communication environment. According to an embodiment, the speech recognition method includes setting an additional wake-up word target capable of activating a speech recognition function in addition to a preset wake-up word, generating a plurality of additional wake-up word utterances formed on the basis of the additional wake-up word target being uttered under various conditions, learning a wake-up word recognition algorithm by using each of the spoken utterances of the additional wake-up word to generate an additional wake-up word recognition algorithm, and executing the additional wake-up word recognition algorithm upon receiving a select word uttered by a user to determine whether to activate the speech recognition function.


