Utterance Pair Acquisition via Keyword Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The limited amount of manually created training data restricts the ability to handle a wide variety of input utterances, making it challenging to train utterance generation models to produce suitable output utterances.
Innovation Solution
An utterance pair acquisition device and method that extracts keywords from expansion source utterance pairs and comparison data, using statistical analysis to identify characteristic keywords, and then selects utterance pairs that meet predetermined conditions for expansion, thereby increasing the quantity and quality of training data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If training data is collected manually, then the quality of training data is high, but the quantity of training data is limited
Solution Approach 1:
The patent uses keyword extraction from high-quality manually created utterance pairs to identify characteristic keywords, then automatically searches for and extracts utterance pairs containing these keywords from large-scale corpora. This copying approach replicates the patterns and characteristics of manually created data at scale, resolving the contradiction between data quality and quantity.
Solution Approach 2:
The patent introduces keywords as an intermediary between manual utterance pairs and automatic data extraction. By extracting characteristic keywords from manual data and using them as search queries in corpora, the system bridges the gap between small-scale high-quality data and large-scale automated data collection, maintaining quality while increasing quantity.
2Measurement precision
If manual data collection is used, then data quality is maintained, but the variety of input utterances is limited
Solution Approach 1:
The system copies the linguistic patterns and keyword structures from manually created high-quality utterance pairs, then applies these patterns to search diverse corpora. This allows the model to learn from the quality of manual data while encountering a much wider variety of input utterances from the expanded training set.
Solution Approach 2:
The extracted keywords serve multiple functions: they characterize the manual utterance pairs, guide the search in large corpora, and filter relevant utterance pairs. This multi-functional use of keywords enables the system to maintain quality standards while adapting to diverse input types across different corpora.
3Adaptability or versatility
If more training data is obtained manually, then the model can handle more variety, but the time and resources required increase significantly
Solution Approach 1:
The system performs preliminary keyword extraction from a small set of manual utterance pairs before the main data collection process. This preliminary action creates a reusable keyword set that can be applied to multiple corpora without requiring additional manual annotation, significantly reducing the time and resources needed for expanding training data variety.
Solution Approach 2:
Instead of manually creating diverse utterance pairs from scratch, the system copies keywords from manual data and uses them to automatically retrieve diverse utterance pairs from existing corpora. This copying strategy achieves variety expansion without the proportional time and resource investment that manual creation would require.
Data Source
AI summary
Acquisition of an utterance pair for expanding a set of utterance pairs for outputting an output utterance in response to receiving a given utterance is described. A keyword extraction unit is configured to compare a degree of characteristic of a word in expansion source utterance pair data and a degree of characteristics of a word in the given utterance data. The expansion source utterance pair data represents a set of expansion source utterance pairs including an input utterance and an output utterance for the input utterance. The present technology includes extracting, based on a comparison result, a keyword list including a keyword that is characteristic of the expansion source utterance pair data. An utterance pair extraction unit is configured to extract, based on the keyword list, an utterance pair from a set of given utterance pairs as an addition for expanding the set of utterance pairs.


