Synthetic Text Generation for Numerical Entity Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current entity recognition systems for numerical quantities require significant manual effort and time for data annotation and training, as well as large datasets, making them inefficient and costly.

Innovation Solution

A computerized method and system that automatically generates synthetic text documents by identifying numeric values and unit expressions in a text corpus, auto-labels these documents, and trains a model using the generated dataset, reducing the need for manual annotation and large datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data annotation is used to train entity recognition models, then the model can be trained with real-world data, but the process becomes time-consuming and expensive

Engineering Contradiction:
Improvemodel training qualityVSAvoidannotation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates synthetic training data by copying and transforming existing labeled data through template-based generation. Synthetic documents are generated by filling templates with randomized values while preserving the structural patterns and linguistic characteristics of real documents, enabling model training without manual annotation of new data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary data preparation by pre-defining templates and patterns from existing labeled data. The synthesis process pre-generates large volumes of training data in advance, eliminating the need for time-consuming manual annotation during the actual model training phase

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If large datasets are collected and manually annotated for training, then the model can learn from diverse examples, but the cost and complexity increase significantly

Engineering Contradiction:
Improvemodel learning capabilityVSAvoiddata preparation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system copies the essential structural and linguistic patterns from a small set of labeled examples and replicates them extensively through template instantiation. This allows the model to learn from diverse synthetic examples generated from limited source data, maintaining adaptability without requiring large manual datasets

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent creates universal templates that can generate multiple variants of training data by substituting different values while maintaining the same structural patterns. A single template can serve multiple training purposes by generating diverse examples through parameter variation, reducing the need for separate data collection for different scenarios

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If rule-based approaches with handcrafted rules are used, then the system can identify named entities using domain-specific knowledge, but significant manual work is required to generalize the approach

Engineering Contradiction:
Improveentity recognition accuracyVSAvoidsystem generalization effort
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent copies successful pattern-matching rules from domain-specific applications and transforms them into template-based synthesis rules. These synthesized templates automatically generate training data that embodies domain-specific patterns without requiring manual rule crafting for each new application domain

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent develops universal template structures that can represent various entity types and patterns across different domains. These multi-functional templates can generate training data for multiple entity recognition tasks by substituting different parameters, eliminating the need to create separate handcrafted rules for each domain

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If feature-based supervised learning is used with manually generated training data, then the model can be trained with labeled features, but the process is time-consuming and expensive

Engineering Contradiction:
Improvemodel training effectivenessVSAvoiddata generation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent copies labeled feature patterns from existing training data and replicates them through template instantiation. The synthesis process automatically generates corresponding feature labels for synthetic data by applying the same labeling logic used for real data, enabling efficient feature-based supervised learning without manual labeling

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary feature extraction and labeling rule definition from existing labeled data. The synthesis process then automatically applies these pre-defined feature labeling rules to generate labeled synthetic data, eliminating the need for time-consuming manual feature annotation during training data preparation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11915157B2Computerized method of training a computer executed model for recognizing numerical quantities
Publication Date: 2024.02.27 KIRA INC
  • US11915157B2 patent drawing
  • US11915157B2 patent drawing
  • US11915157B2 patent drawing

AI summary

A computerized method for training a computer executed model for recognizing numerical quantities is provided. An input, at least one unit expression, is received by an input module. The input module may then search for numeric values and the unit expression in a text corpus, wherein, the text corpus comprises sets of words and frequency of occurrence of each of the sets. The input module may identify identified sets, wherein the identified sets may comprise a combination of a numeric value and the unit expression. A synthetic text generation module may then generate sentences from the text corpus by applying the identified sets as input. A training dataset may be generated by a labeling module by auto labelling features in the generated sentences based on the numeric value and the unit expression and further a training module may train the training model by providing input based on the training dataset.