Robot Voice Trigger Model Generation via Environmental Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems for robots struggle to detect various utterances of a wake word, especially in different tones and environments, leading to failed activation of voice recognition services and lack of consideration for environmental factors and mechanism characteristics.
Innovation Solution
A trigger recognition model generating method that phonetically synthesizes input text to create learning data, applying filters based on environmental factors and mechanism characteristics to enhance recognition precision, allowing for easy change of voice triggers and improved productivity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a wake word is extracted from input voice information and registered as a user wake word, then the voice recognition service can be activated in response to the wake word, but the system fails to detect various utterances for the wake word in different tones and environments
Solution Approach 1:
The patent applies parameter changes by generating multiple voice triggers from a single input text through phonetic synthesis with varying parameters (tones, speeds, pitches). This creates diverse training data that enables the recognition model to handle different utterance variations while maintaining reliable activation across various environments.
Solution Approach 2:
The patent implements preliminary action by pre-generating multiple voice trigger variations through phonetic synthesis before actual voice recognition occurs. The system creates a comprehensive set of training data in advance, including environmental noise and mechanism characteristics, so that the recognition model is already prepared to handle diverse utterances when activation is needed.
2Ease of manufacture
If voice data is generated by combining synthesized voice and recorded voice, then voice data can be created, but the voice data does not reflect actual environmental factors or mechanism characteristics
Solution Approach 1:
The patent applies feedback by incorporating environmental noise and mechanism characteristics into the voice trigger generation process. The system uses feedback from environmental sensors and robot mechanism data to adjust and refine the synthesized voice triggers, ensuring that the training data accurately reflects real-world conditions while maintaining ease of data generation.
Solution Approach 2:
The patent uses composite materials by combining multiple data sources (phonetic synthesis, environmental noise, mechanism characteristics) into a unified training dataset. This composite approach integrates diverse information types to create comprehensive voice trigger data that simultaneously reflects environmental factors and mechanism characteristics while remaining manufacturable.
3Productivity
If a trigger recognition model is trained with limited voice trigger data, then the model can be generated, but it cannot recognize various utterances for the same trigger in different environments
Solution Approach 1:
The patent implements preliminary action by pre-generating a comprehensive set of voice trigger variations through phonetic synthesis before model training. This preliminary data preparation expands the training dataset to include multiple utterance variations, enabling the model to achieve high adaptability while maintaining efficient training processes and productivity.
Solution Approach 2:
The patent applies parameter changes by systematically varying phonetic synthesis parameters (tone, speed, pitch) and environmental conditions to generate diverse training data. This parameter expansion creates a rich dataset that enables the recognition model to learn and recognize various utterances for the same trigger across different environments while keeping the training process manageable and productive.
Data Source
AI summary
Provided are a trigger recognition model generating method for a robot and a robot to which the method is applied. A trigger recognition model generating method comprises obtaining an input text which expresses a voice trigger, obtaining a first set of voice triggers by voice synthesis from the input text, obtaining a second set of voice triggers by applying a first filter in accordance with an environmental factor to the first set of voice triggers, obtaining a third set of voice triggers by applying a second filter in accordance with a mechanism characteristic of the robot to the second set of voice triggers, and applying the first, second, and third sets of voice triggers to the trigger recognition model as learning data for the voice trigger. By doing this, a trigger recognition model which is capable of recognizing a new trigger is generated.


