Pre-trained Language Model for Automated Labeled Training Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The generation of labeled training data for text classification tasks is often limited due to the need for manual annotation by expert users, making it time-consuming and costly.

Innovation Solution

A system utilizing a pre-trained language model neural network to automatically generate high-quality labeled training data through few-shot prompts, reducing the reliance on human annotations and enhancing data augmentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation by expert users is used to generate labeled training data, then data quality and accuracy are improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training a language model on large unlabeled corpora before using it to generate labeled training data. The pre-trained model's acquired knowledge is then leveraged to automatically annotate data, eliminating the need for manual annotation while maintaining high quality. This resolves the contradiction by performing the useful work (knowledge acquisition) in advance, allowing rapid high-quality data generation without expert intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service by enabling the language model to automatically generate its own training data without external human annotation. The model uses its pre-acquired linguistic knowledge to autonomously classify and label text data, replacing the manual expert annotation process. This self-service capability maintains data quality while dramatically reducing time consumption and costs.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual annotation by expert users is used to generate labeled training data, then data quality and accuracy are improved, but computational cost and resource requirements increase

Engineering Contradiction:
Improvedata qualityVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs the computationally intensive work of language model training in advance on unlabeled data. Once pre-trained, the model can rapidly generate labeled training data with minimal additional computational cost, compared to the expensive and time-consuming manual annotation process. This preliminary investment resolves the contradiction by shifting computational burden to a one-time pre-training phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses the pre-trained language model as a template or copy of human linguistic knowledge, which can then be applied repeatedly to generate labeled data without requiring repeated human expert involvement. This copying of knowledge reduces both time and computational costs while maintaining data quality, as the model can generate numerous labeled examples from a single pre-training investment.

Inventive Principle:
Principle #26Copying

3Reliability

If large amounts of labeled training data are collected through manual annotation, then model performance is improved, but the complexity and difficulty of data collection increase

Engineering Contradiction:
Improvemodel performanceVSAvoiddata collection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent enables the system to self-generate labeled training data using the pre-trained language model, eliminating the complex manual data collection process. The model autonomously creates high-quality labeled datasets by applying its pre-acquired knowledge to unlabeled text, thereby improving model performance while dramatically simplifying data collection complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The pre-trained language model serves multiple functions: it acts as both the data annotator and the foundation for the target task model. This multi-functionality allows a single pre-trained model to generate diverse labeled training data across different tasks, reducing overall system complexity while maintaining high model performance through quality labeled data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230196105A1Generating labeled training data using a pre-trained language model neural network
Publication Date: 2023.06.22 GOOGLE LLC
  • US20230196105A1 patent drawing
  • US20230196105A1 patent drawing
  • US20230196105A1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating labeled training data using a pre-trained language model neural network. In particular, the language model neural network can generate the text input in a new labeled training example from an input sequence that includes (i) one or more context inputs and (ii) a text label that identifies the ground truth category for the new labeled training example.