Robot Voice Data Collection for Owner Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing robot systems struggle to efficiently collect training data for identifying specific users, particularly in situations where sound data is insufficient, limiting their ability to mimic the behavior of a pet-like companion.
Innovation Solution
A robot system that includes sensors to detect external stimuli, processes these stimuli to determine specific conditions for collecting sound data, and adjusts data collection criteria based on detected events to enhance the likelihood of capturing relevant training data, particularly when interacting with its owner.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the robot collects sound data continuously to improve speaker identification accuracy, then the quantity of training data increases, but the energy consumption and operational time increase
Solution Approach 1:
The robot dynamically adjusts the sound data collection condition based on its current state and interaction context. When the robot is actively interacting with a user or when speaker identification confidence is low, the collection condition is relaxed to capture more sound data. When the robot is idle or battery level is low, the collection condition becomes stricter to reduce energy consumption. This dynamic adjustment resolves the contradiction between data quantity and energy usage.
Solution Approach 2:
The system changes the parameters of data collection by adjusting the sensitivity and triggers for recording sound data. Instead of continuous recording, the system modifies collection parameters to record only when specific conditions are met (e.g., detected speech patterns, interaction events), thereby reducing overall energy consumption while still accumulating sufficient training data over time.
2Measurement precision
If the robot uses strict collection conditions to ensure high-quality training data, then the precision of training data improves, but the quantity of collected data decreases
Solution Approach 1:
The robot performs preliminary processing of sound data by detecting speech patterns, voice characteristics, and interaction context before deciding whether to collect the data. This preliminary filtering ensures that only potentially valuable sound data (those with clear speech patterns or relevant to speaker identification) are collected, maintaining high quality while increasing the effective quantity of useful training data.
Solution Approach 2:
The system uses feedback from its speaker identification performance to adjust data collection strategies. When the identification accuracy is low or when certain speakers are poorly represented in the training set, the system relaxes collection conditions to gather more data from those specific sources, thereby improving both quantity and targeted quality of the training data.
3Use of energy by moving object
If the robot collects sound data only during specific interactions to save energy, then the energy consumption decreases, but the diversity of training data is limited
Solution Approach 1:
The robot's sound data collection system serves multiple functions: it collects data for speaker identification, adapts to different user voices and speaking styles, and learns from various interaction contexts. By designing the collection condition to trigger on diverse interaction types (not just speech but also environmental sounds during interactions), the system achieves energy efficiency while maintaining data diversity that supports multiple AI functions.
Data Source
AI summary
A robot includes a sensor that detects an external stimulus, and at least one processor. In response to detection, by the sensor and as the external stimulus, of a sound that satisfies a predetermined collection condition, the at least one processor stores sound data expressing the detected sound in a storage device as training data for identifying a speaker from the sound, and in response to detection by the sensor of an external stimulus that satisfies a specific condition, the at least one processor changes the collection condition such that the sound detected by the sensor more easily satisfies the collection condition for a period from the detection of the external stimulus satisfying the specific condition until a predetermined amount of time elapses.


