Filled Pause Detection Training Data Generation via Domain Randomization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating a filled pause detector with excellent performance is challenging due to the difficulty in acquiring large amounts of speech data in similar noise environments and defining noise environments effectively, which affects the training of filled pause detecting models.
Innovation Solution
A method and device using a domain randomization algorithm to generate training data in a simulation noise environment by acquiring acoustic data, generating noise data, and synthesizing it with speech data to create labeled training data, enhancing filled pause detection performance in actual noise environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If training data is generated without considering noise environment, then data acquisition is simple, but filled pause detection performance deteriorates
Solution Approach 1:
The patent creates synthetic training data by copying and combining clean speech data with simulated noise data. The noise generator creates artificial noise patterns that mimic real noise environments, allowing the model to learn filled pause detection in controlled conditions without requiring actual recorded noise-filled speech data.
Solution Approach 2:
The patent systematically varies noise parameters such as signal-to-noise ratio, noise type, and temporal characteristics when generating training data. By changing these parameters across different training samples, the model learns to detect filled pauses under diverse noise conditions, improving generalization performance.
2Reliability
If training data is generated in simulation noise environment, then detection performance in actual noise environment improves, but data generation complexity increases
Solution Approach 1:
The patent divides the data generation process into separate modules: a clean speech data acquisition module, a noise generation module, and a data synthesis module. This segmentation allows each component to be independently designed and optimized, reducing overall system complexity while maintaining effectiveness.
Solution Approach 2:
The patent introduces a noise generator as an intermediary component that creates artificial noise patterns from simple random processes. This intermediary transforms complex real-world noise into manageable synthetic noise data, simplifying the overall data generation pipeline while still capturing essential noise characteristics.
3Measurement precision
If large amount of speech data in similar noise environment is acquired, then training accuracy improves, but data collection difficulty increases
Solution Approach 1:
The system generates its own training data autonomously by processing clean speech data through the noise generator and synthesis components. This self-service approach eliminates the need for manual collection of noise-filled speech data, allowing unlimited data generation without additional field work or human effort.
Solution Approach 2:
The patent performs preliminary data generation in a controlled environment before actual deployment. By pre-generating diverse training data samples with various noise conditions and filled pause patterns, the system prepares extensive training material in advance, avoiding the need for difficult real-world data collection during deployment.
Data Source
AI summary
Disclosed is a method for generating training data for training a filled pause detecting model and a device therefor, which execute mounted artificial intelligence (AI) and/or machine learning algorithms in a 5G communication environment. The method includes acquiring acoustic data including first speech data including a filled pause, second speech data not including a filled pause, and noise, generating a plurality of noise data based on the acoustic data, and generating first training data including a plurality of filled pauses and second training data not including a plurality of filled pauses by synthesizing the plurality of noise data with the first speech data and the second speech data. According to the present disclosure, training data for training a filled pause detecting model in a simulation noise environment can be generated, and filled pause detection performance for speech data generated in an actual noise environment can be enhanced.


