Acoustic Model Retraining via Heterogeneity-Based Data Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The acoustic model training support device faces issues where unnecessary retraining occurs within already trained ranges and fails to perform retraining in desirable ranges due to the inclusion of all trained data in retraining processes.
Innovation Solution
A data processing device that generates candidate input and intermediate data by combining trained and untrained data, prioritizing heterogeneity for selection in subsequent training, to optimize retraining ranges.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all trained data is used for retraining, then the retraining process is comprehensive, but unnecessary retraining occurs within already trained ranges
Solution Approach 1:
The patent segments the retraining data by generating candidate input data that combines trained and untrained input data, and candidate intermediate data that combines trained and untrained intermediate data. This segmentation allows selective use of only necessary data for retraining, avoiding redundant processing of already-trained data ranges while maintaining comprehensiveness where needed.
2Productivity
If only relationship-based candidate data selection is used, then retraining is promoted in untrained ranges, but retraining may occur unnecessarily within already trained ranges
Solution Approach 1:
The patent employs feedback mechanisms by calculating heterogeneity degrees of candidate intermediate data relative to selected intermediate data, and using this feedback to preferentially select data with greater heterogeneity. This feedback loop ensures that retraining is directed toward untrained ranges while automatically avoiding already-trained ranges, resolving the contradiction between productivity and time loss.
3Productivity
If heterogeneity-based selection is applied, then retraining is optimized for untrained ranges, but the selection process becomes more complex
Solution Approach 1:
The patent changes the selection parameter from simple relationship-based criteria to heterogeneity degree-based criteria. By defining and calculating heterogeneity degrees of candidate intermediate data, the system achieves more precise selection for untrained ranges. The complexity is managed by systematically defining heterogeneity calculation methods and using automated processing.
Data Source
AI summary
A data processing device includes: a first generating unit that generates a plurality of pieces of candidate input data including a plurality of pieces of trained input data and a plurality of pieces of untrained input data; a second generating unit that generates a plurality of pieces of candidate intermediate data including trained intermediate data and untrained intermediate data; a first selection unit that selects one piece of candidate intermediate data from the plurality of pieces of candidate intermediate data, and preferentially selects one piece of candidate intermediate data having a greater degree of heterogeneity when used for second training subsequent to first training as compared with the selected intermediate data; and a second selection unit that selects one piece of candidate input data corresponding to the one piece of candidate intermediate data selected from the plurality of pieces of candidate input data to be used in the second training.


