Language-Model Data Tagging With Sample-Based Error Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual and semi-supervised learning methods for data tagging are inefficient, labor-intensive, and inaccurate, while existing supervised learning approaches based on traditional neural networks require pre-training and consume significant computing resources, limiting tag quality.
Innovation Solution
A data tagging and prompt generation system (TE) that automates the tagging process, using a sample of input data to generate high-quality tags with reduced computing resources and manual labor, incorporating a language model (LM) to infer tags based on metadata, statistics, and user-provided context.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If manual tagging or semi-supervised learning methods are used, then labor and computing resources are reduced, but tagging accuracy and efficiency deteriorate significantly
Solution Approach 1:
The patent segments the tagging process into multiple stages: initial tagging using a first language model, error identification through data sampling and verification, and iterative refinement using a second language model. This segmentation allows the system to achieve high accuracy without requiring full manual verification of all data, thus reducing computing resources while maintaining tagging precision.
Solution Approach 2:
The patent introduces an intermediary verification mechanism that samples tagged data to identify errors without requiring complete manual review. This intermediary step acts as a mediator between automated tagging and full manual verification, achieving high accuracy with reduced resource consumption by focusing verification efforts only on potentially erroneous cases.
2Measurement precision
If traditional neural network supervised learning approaches are used, then tag quality can be improved, but pre-training requirements and computing resource consumption increase
Solution Approach 1:
The patent performs preliminary tagging actions using a first language model before more intensive processing. By pre-tagging data with an initial model and then selectively refining only erroneous cases through sampling and verification, the system achieves high tag quality without requiring all data to undergo complex pre-training and full supervised learning processes, thus reducing overall system complexity.
Solution Approach 2:
Instead of applying intensive supervised learning to the entire dataset, the patent applies partial action by selectively verifying and refining only a sampled portion of tagged data. This partial verification approach achieves high tag quality without the excessive computing resources and system complexity required for complete retraining of neural networks on all data.
3Measurement precision
If complete manual verification of all tagged data is performed, then tagging accuracy is maximized, but processing time and resource consumption increase significantly
Solution Approach 1:
The patent applies partial verification by sampling a subset of tagged data for error identification rather than verifying all tagged data manually. This partial action approach maintains high tagging accuracy by focusing verification resources on potentially erroneous cases while significantly reducing processing time compared to complete manual verification of the entire dataset.
Solution Approach 2:
The system implements self-service through automated error identification mechanisms that use language models to detect and flag potentially erroneous tags without requiring continuous manual intervention. This self-verification capability reduces processing time by automating the accuracy check process while maintaining high tagging standards.
Data Source
AI summary
System, method, and various embodiments for data tagging and prompt generation are described herein. An embodiment operates by receiving input data, identifying metadata, generating one or more statistics based on the input data, calculating a sample size for the input data based on the one or more statistics and extracting a sample of the input data of the sample size. A prompt is generated based on a prompt template, and the prompt is provided to a language model configured to tag the input in accordance with the prompt. The output including tagged input data is received, and a query is executed against the tagged input data.


