Synthetic Text Generation for Numerical Entity Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current entity recognition systems for numerical quantities require significant manual effort and time for data annotation and training, as well as large datasets, making them inefficient and costly.
Innovation Solution
A computerized method and system that automatically generates synthetic text documents by identifying numeric values and unit expressions in a text corpus, auto-labels these documents, and trains a model using the generated dataset, reducing the need for manual annotation and large datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data annotation is used to train entity recognition models, then the model can be trained with real-world data, but the process becomes time-consuming and expensive
Solution Approach 1:
The patent creates synthetic training data by copying and transforming existing labeled data through template-based generation. Synthetic documents are generated by filling templates with randomized values while preserving the structural patterns and linguistic characteristics of real documents, enabling model training without manual annotation of new data
Solution Approach 2:
The patent performs preliminary data preparation by pre-defining templates and patterns from existing labeled data. The synthesis process pre-generates large volumes of training data in advance, eliminating the need for time-consuming manual annotation during the actual model training phase
2Adaptability or versatility
If large datasets are collected and manually annotated for training, then the model can learn from diverse examples, but the cost and complexity increase significantly
Solution Approach 1:
The system copies the essential structural and linguistic patterns from a small set of labeled examples and replicates them extensively through template instantiation. This allows the model to learn from diverse synthetic examples generated from limited source data, maintaining adaptability without requiring large manual datasets
Solution Approach 2:
The patent creates universal templates that can generate multiple variants of training data by substituting different values while maintaining the same structural patterns. A single template can serve multiple training purposes by generating diverse examples through parameter variation, reducing the need for separate data collection for different scenarios
3Measurement precision
If rule-based approaches with handcrafted rules are used, then the system can identify named entities using domain-specific knowledge, but significant manual work is required to generalize the approach
Solution Approach 1:
The patent copies successful pattern-matching rules from domain-specific applications and transforms them into template-based synthesis rules. These synthesized templates automatically generate training data that embodies domain-specific patterns without requiring manual rule crafting for each new application domain
Solution Approach 2:
The patent develops universal template structures that can represent various entity types and patterns across different domains. These multi-functional templates can generate training data for multiple entity recognition tasks by substituting different parameters, eliminating the need to create separate handcrafted rules for each domain
4Reliability
If feature-based supervised learning is used with manually generated training data, then the model can be trained with labeled features, but the process is time-consuming and expensive
Solution Approach 1:
The patent copies labeled feature patterns from existing training data and replicates them through template instantiation. The synthesis process automatically generates corresponding feature labels for synthetic data by applying the same labeling logic used for real data, enabling efficient feature-based supervised learning without manual labeling
Solution Approach 2:
The patent performs preliminary feature extraction and labeling rule definition from existing labeled data. The synthesis process then automatically applies these pre-defined feature labeling rules to generate labeled synthetic data, eliminating the need for time-consuming manual feature annotation during training data preparation
Data Source
AI summary
A computerized method for training a computer executed model for recognizing numerical quantities is provided. An input, at least one unit expression, is received by an input module. The input module may then search for numeric values and the unit expression in a text corpus, wherein, the text corpus comprises sets of words and frequency of occurrence of each of the sets. The input module may identify identified sets, wherein the identified sets may comprise a combination of a numeric value and the unit expression. A synthetic text generation module may then generate sentences from the text corpus by applying the identified sets as input. A training dataset may be generated by a labeling module by auto labelling features in the generated sentences based on the numeric value and the unit expression and further a training module may train the training model by providing input based on the training dataset.


