Pre-trained Language Model for Automated Labeled Training Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The generation of labeled training data for text classification tasks is often limited due to the need for manual annotation by expert users, making it time-consuming and costly.
Innovation Solution
A system utilizing a pre-trained language model neural network to automatically generate high-quality labeled training data through few-shot prompts, reducing the reliance on human annotations and enhancing data augmentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation by expert users is used to generate labeled training data, then data quality and accuracy are improved, but time consumption and cost increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-training a language model on large unlabeled corpora before using it to generate labeled training data. The pre-trained model's acquired knowledge is then leveraged to automatically annotate data, eliminating the need for manual annotation while maintaining high quality. This resolves the contradiction by performing the useful work (knowledge acquisition) in advance, allowing rapid high-quality data generation without expert intervention.
Solution Approach 2:
The patent implements self-service by enabling the language model to automatically generate its own training data without external human annotation. The model uses its pre-acquired linguistic knowledge to autonomously classify and label text data, replacing the manual expert annotation process. This self-service capability maintains data quality while dramatically reducing time consumption and costs.
2Measurement precision
If manual annotation by expert users is used to generate labeled training data, then data quality and accuracy are improved, but computational cost and resource requirements increase
Solution Approach 1:
The patent performs the computationally intensive work of language model training in advance on unlabeled data. Once pre-trained, the model can rapidly generate labeled training data with minimal additional computational cost, compared to the expensive and time-consuming manual annotation process. This preliminary investment resolves the contradiction by shifting computational burden to a one-time pre-training phase.
Solution Approach 2:
The patent uses the pre-trained language model as a template or copy of human linguistic knowledge, which can then be applied repeatedly to generate labeled data without requiring repeated human expert involvement. This copying of knowledge reduces both time and computational costs while maintaining data quality, as the model can generate numerous labeled examples from a single pre-training investment.
3Reliability
If large amounts of labeled training data are collected through manual annotation, then model performance is improved, but the complexity and difficulty of data collection increase
Solution Approach 1:
The patent enables the system to self-generate labeled training data using the pre-trained language model, eliminating the complex manual data collection process. The model autonomously creates high-quality labeled datasets by applying its pre-acquired knowledge to unlabeled text, thereby improving model performance while dramatically simplifying data collection complexity.
Solution Approach 2:
The pre-trained language model serves multiple functions: it acts as both the data annotator and the foundation for the target task model. This multi-functionality allows a single pre-trained model to generate diverse labeled training data across different tasks, reducing overall system complexity while maintaining high model performance through quality labeled data.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating labeled training data using a pre-trained language model neural network. In particular, the language model neural network can generate the text input in a new labeled training example from an input sequence that includes (i) one or more context inputs and (ii) a text label that identifies the ground truth category for the new labeled training example.


