Natural Language Training Data Generation via Semantic Re-parametrization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for training natural language processing systems require large numbers of annotated utterances, which is a time-consuming and expertise-dependent process, limiting the pool of annotators and the efficiency of data accumulation.

Innovation Solution

A method is introduced that automatically generates training data by loading transcripts of dialogue events, acquiring and re-parametrizing commands with alternative semantic parameters, and generating alternative ordered subsequences of dialogue events, allowing for the expansion of training data coverage and annotation without the need for extensive annotator expertise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation of training data is used, then training data can be created with expert quality, but the process is time-consuming and limits the pool of annotators

Engineering Contradiction:
Improveannotation qualityVSAvoiddata accumulation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-annotation by automatically generating training data with semantic parameters and annotations without requiring human annotators. The computerized assistant autonomously creates dialogue events, assigns semantic parameters, and generates annotations, eliminating the need for manual expert annotation while maintaining data quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system generates synthetic training data by creating copies of real dialogue patterns with varied semantic parameters. It uses template-based generation to produce multiple variations of dialogue events, effectively copying and adapting existing dialogue structures to create diverse training examples without manual annotation.

Inventive Principle:
Principle #26Copying

2Reliability

If manual annotation by experts is required, then high-quality training data can be produced, but the complexity and expertise requirement limit annotator availability

Engineering Contradiction:
Improvetraining data qualityVSAvoidannotation process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical process of manual expert annotation with an automated computational system. The computerized assistant uses algorithms to generate dialogue events, assign semantic parameters, and create annotations, substituting human cognitive work with automated processing that maintains consistency and quality without requiring expert annotators.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameters of training data generation by using semantic parameters that can be systematically varied. Instead of relying on human experts to manually create diverse examples, the system programmatically generates variations by modifying semantic parameters such as entities, attributes, and relationships, thereby achieving diversity and quality through parameter manipulation rather than human expertise.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If large numbers of annotated utterances are collected manually, then comprehensive training data can be accumulated, but the process is slow and resource-intensive

Engineering Contradiction:
Improvetraining data volumeVSAvoiddata accumulation time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-generating large volumes of training data with complete annotations before they are needed for model training. The computerized assistant creates extensive datasets in advance by systematically generating dialogue events with varied semantic parameters, eliminating the need for time-consuming manual annotation processes later.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables continuous generation of training data through automated processes that can operate without interruption. The computerized assistant continuously generates dialogue events, assigns semantic parameters, and creates annotations in an uninterrupted workflow, whereas manual annotation processes must stop and restart based on annotator availability, creating gaps and delays.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS11145291B2Training natural language system with generated dialogues
Publication Date: 2021.10.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11145291B2 patent drawing
  • US11145291B2 patent drawing
  • US11145291B2 patent drawing

AI summary

A method for generating training data for training a natural language processing system comprises loading, into a computer memory, a computer-readable transcript representing an ordered sequence of one or more dialogue events. The method further comprises acquiring a computer-readable command describing an exemplary ordered subsequence of one or more dialogue events from the computer-readable transcript. The method further comprises re-parametrizing the computer-readable command with an alternative semantic parameter. The method further comprises generating an alternative ordered subsequence of one or more dialogue events based on the re-parametrized computer-readable command. The method further comprises outputting, to a data store, an alternative computer-readable transcript including the alternative ordered subsequence of one or more dialogue events, the alternative computer-readable transcript having a predetermined format usable to train the computerized assistant.