Utterance Pair Acquisition via Keyword Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The limited amount of manually created training data restricts the ability to handle a wide variety of input utterances, making it challenging to train utterance generation models to produce suitable output utterances.

Innovation Solution

An utterance pair acquisition device and method that extracts keywords from expansion source utterance pairs and comparison data, using statistical analysis to identify characteristic keywords, and then selects utterance pairs that meet predetermined conditions for expansion, thereby increasing the quantity and quality of training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If training data is collected manually, then the quality of training data is high, but the quantity of training data is limited

Engineering Contradiction:
Improvequality of training dataVSAvoidquantity of training data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses keyword extraction from high-quality manually created utterance pairs to identify characteristic keywords, then automatically searches for and extracts utterance pairs containing these keywords from large-scale corpora. This copying approach replicates the patterns and characteristics of manually created data at scale, resolving the contradiction between data quality and quantity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces keywords as an intermediary between manual utterance pairs and automatic data extraction. By extracting characteristic keywords from manual data and using them as search queries in corpora, the system bridges the gap between small-scale high-quality data and large-scale automated data collection, maintaining quality while increasing quantity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual data collection is used, then data quality is maintained, but the variety of input utterances is limited

Engineering Contradiction:
Improvedata qualityVSAvoidvariety of input utterances
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system copies the linguistic patterns and keyword structures from manually created high-quality utterance pairs, then applies these patterns to search diverse corpora. This allows the model to learn from the quality of manual data while encountering a much wider variety of input utterances from the expanded training set.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The extracted keywords serve multiple functions: they characterize the manual utterance pairs, guide the search in large corpora, and filter relevant utterance pairs. This multi-functional use of keywords enables the system to maintain quality standards while adapting to diverse input types across different corpora.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If more training data is obtained manually, then the model can handle more variety, but the time and resources required increase significantly

Engineering Contradiction:
Improvehandling capability of input utterancesVSAvoidtime for data collection
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary keyword extraction from a small set of manual utterance pairs before the main data collection process. This preliminary action creates a reusable keyword set that can be applied to multiple corpora without requiring additional manual annotation, significantly reducing the time and resources needed for expanding training data variety.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of manually creating diverse utterance pairs from scratch, the system copies keywords from manual data and uses them to automatically retrieve diverse utterance pairs from existing corpora. This copying strategy achieves variety expansion without the proportional time and resource investment that manual creation would require.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12019986B2Utterance pair acquisition apparatus, utterance pair acquisition method, and program
Publication Date: 2024.06.25 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12019986B2 patent drawing
  • US12019986B2 patent drawing
  • US12019986B2 patent drawing

AI summary

Acquisition of an utterance pair for expanding a set of utterance pairs for outputting an output utterance in response to receiving a given utterance is described. A keyword extraction unit is configured to compare a degree of characteristic of a word in expansion source utterance pair data and a degree of characteristics of a word in the given utterance data. The expansion source utterance pair data represents a set of expansion source utterance pairs including an input utterance and an output utterance for the input utterance. The present technology includes extracting, based on a comparison result, a keyword list including a keyword that is characteristic of the expansion source utterance pair data. An utterance pair extraction unit is configured to extract, based on the keyword list, an utterance pair from a set of given utterance pairs as an addition for expanding the set of utterance pairs.