Learning Data Generation for End-of-Talk Prediction Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The generation of learning data for end-of-talk prediction models in dialog systems is costly due to the manual appending of training data, which is time-consuming and inefficient.
Innovation Solution
A learning data generation device and method that predicts end-of-talk utterances using a combination of an end-of-talk prediction model and predefined rules, and automatically generates learning data by detecting interruption utterances within a prescribed time frame to determine when an utterance is not an end-of-talk, thereby reducing the need for manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual appending of training data is performed to create learning data for end-of-talk prediction models, then the accuracy and quality of learning data is improved, but the cost and time consumption increase significantly
Solution Approach 1:
The system automatically generates learning data by having the end-of-talk prediction model predict its own training data. The model predicts whether utterances are end-of-talk utterances, and these predictions are automatically used as training data without requiring manual annotation, thus the system serves itself
Solution Approach 2:
The system performs preliminary predictions using the end-of-talk prediction model before actual training is needed. By pre-generating predicted results as training data, the system prepares learning data in advance without waiting for manual processing
2Manufacturing precision
If manual appending of training data is performed to create learning data for end-of-talk prediction models, then the accuracy and quality of learning data is improved, but the cost increases
Solution Approach 1:
The system automatically generates learning data by having the end-of-talk prediction model predict its own training data. The model predictions are automatically used as training data without requiring manual annotation, eliminating labor costs
Solution Approach 2:
The manual mechanical process of annotating training data is replaced by an automated computational system. The end-of-talk prediction model automatically generates training data through computational prediction, substituting human manual work
3Loss of time
If the end-of-talk prediction model is used to generate learning data automatically, then the cost and time consumption are reduced, but the accuracy of learning data may be compromised
Solution Approach 1:
The system uses the end-of-talk prediction model to generate predictions, which are then fed back as training data to improve the model. This feedback loop allows the model to learn from its own predictions and improve accuracy over time
Solution Approach 2:
The system performs preliminary predictions using the end-of-talk prediction model before actual training is needed. By pre-generating predicted results as training data, the system prepares learning data in advance without waiting for manual processing
Data Source
AI summary
The learning data generation device (10) of the present invention comprises: an end-of-talk predict unit (11) for performing: a first prediction in which it is predicted, based on utterance information on an utterance in the dialog, using the end-of-talk prediction model (16), whether the utterance is an end-of-talk utterance of the speaker; and a second prediction in which it is predicted, based on one or more prescribed rules, whether the utterance is an end-of-talk utterance; and a training data generate unit (13) for generating, when, in the first prediction it is predicted that the utterance is not an end-of-talk utterance and in the second prediction it is predicted that the utterance is an end-of-talk utterance, for the utterance information on the utterance, learning data to which training data indicating that the utterance is an end-of-talk utterance is appended.


