Entity Learning Recognition via Geometric Data Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional entity recognition algorithms require a large number of samples for effective training, especially in applications like surveillance, where labeled data is limited, leading to models with limited expressive power, increased vulnerability to attacks, and impracticality for real-world deployment.
Innovation Solution
A method for data augmentation that generates numerous noisy entities by applying geometric transformations and superimposing structural elements onto original entities, creating an expanded dataset for training models that are robust to noise and attacks, while preserving label information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional entity recognition algorithms are trained with limited labeled data, then training time and resource requirements are reduced, but the model's expressive power and robustness deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-processing original entities through geometric transformations (rotation, scaling, translation) and noise addition to create augmented training samples before model training. This preliminary transformation of data ensures the model encounters varied and robust training examples, improving its resistance to attacks without requiring additional collected data.
Solution Approach 2:
The patent employs copying by generating multiple augmented copies of each original entity through structural element transformations. These copied and transformed versions serve as additional training samples, effectively increasing the training data volume while maintaining the original data distribution characteristics, thereby improving model robustness.
2Reliability
If more training samples are used to improve model robustness, then resistance to noise and attacks increases, but data collection and labeling costs increase
Solution Approach 1:
The patent uses copying to generate multiple augmented versions of each original entity through systematic geometric transformations and noise additions. This approach creates diverse training samples without requiring additional data collection or manual labeling, thereby improving robustness while avoiding the complexity and costs associated with expanding the original dataset through traditional means.
Solution Approach 2:
The patent applies parameter changes by systematically varying geometric parameters (rotation angles, scale factors, translation vectors) and noise parameters during data augmentation. These controlled parameter transformations generate diverse training samples that improve model robustness to noise and attacks without requiring complex data processing pipelines or additional data collection infrastructure.
3Reliability
If data augmentation with geometric transformations is applied, then model robustness to attacks improves, but computational resources required for training increase
Solution Approach 1:
The patent applies preliminary action by pre-computing augmented training samples through geometric transformations and noise additions before model training begins. This preliminary data preparation step ensures the model is trained on robust, attack-resistant samples without requiring additional computational resources during the actual training process, as the augmentation is performed once during data preparation.
Solution Approach 2:
The patent uses copying to generate augmented training samples that are then used for model training. By creating these transformed copies in advance and using them as the training dataset, the computational cost of transformation is amortized over the training process, improving model robustness while managing computational resource requirements efficiently.
Data Source
AI summary
An entity learning recognition method, system, and computer program product include learning (i.e., in a training phase) from at least one entity to produce augments entities such that an augmented entity is still recognizable as the original entity but differs sufficiently to produce a different feature representation of the entity to create a database for use (i.e., in an implementation phase).


