Entity Learning Recognition via Geometric Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional entity recognition algorithms require a large number of samples for effective training, especially in applications like surveillance, where labeled data is limited, leading to models with limited expressive power, increased vulnerability to attacks, and impracticality for real-world deployment.

Innovation Solution

A method for data augmentation that generates numerous noisy entities by applying geometric transformations and superimposing structural elements onto original entities, creating an expanded dataset for training models that are robust to noise and attacks, while preserving label information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional entity recognition algorithms are trained with limited labeled data, then training time and resource requirements are reduced, but the model's expressive power and robustness deteriorate

Engineering Contradiction:
Improvevolume of training dataVSAvoidrobustness to attacks
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-processing original entities through geometric transformations (rotation, scaling, translation) and noise addition to create augmented training samples before model training. This preliminary transformation of data ensures the model encounters varied and robust training examples, improving its resistance to attacks without requiring additional collected data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs copying by generating multiple augmented copies of each original entity through structural element transformations. These copied and transformed versions serve as additional training samples, effectively increasing the training data volume while maintaining the original data distribution characteristics, thereby improving model robustness.

Inventive Principle:
Principle #26Copying

2Reliability

If more training samples are used to improve model robustness, then resistance to noise and attacks increases, but data collection and labeling costs increase

Engineering Contradiction:
Improverobustness to noiseVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses copying to generate multiple augmented versions of each original entity through systematic geometric transformations and noise additions. This approach creates diverse training samples without requiring additional data collection or manual labeling, thereby improving robustness while avoiding the complexity and costs associated with expanding the original dataset through traditional means.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies parameter changes by systematically varying geometric parameters (rotation angles, scale factors, translation vectors) and noise parameters during data augmentation. These controlled parameter transformations generate diverse training samples that improve model robustness to noise and attacks without requiring complex data processing pipelines or additional data collection infrastructure.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If data augmentation with geometric transformations is applied, then model robustness to attacks improves, but computational resources required for training increase

Engineering Contradiction:
Improveresistance to attacksVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-computing augmented training samples through geometric transformations and noise additions before model training begins. This preliminary data preparation step ensures the model is trained on robust, attack-resistant samples without requiring additional computational resources during the actual training process, as the augmentation is performed once during data preparation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying to generate augmented training samples that are then used for model training. By creating these transformed copies in advance and using them as the training dataset, the computational cost of transformation is amortized over the training process, improving model robustness while managing computational resource requirements efficiently.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12093839B2Entity learning recognition
Publication Date: 2024.09.17 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12093839B2 patent drawing
  • US12093839B2 patent drawing
  • US12093839B2 patent drawing

AI summary

An entity learning recognition method, system, and computer program product include learning (i.e., in a training phase) from at least one entity to produce augments entities such that an augmented entity is still recognizable as the original entity but differs sufficiently to produce a different feature representation of the entity to create a database for use (i.e., in an implementation phase).