Structured Object Representation for Pixel Labeling Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional pixel labeling techniques rely on manual user input for training data, leading to inaccuracies and inefficiencies due to the reliance on manually drawn borders, resulting in limited and expensive training data sets.
Innovation Solution
The use of structured object representations, such as SVG, PDF, or HTML images, to accurately and efficiently label objects and their corresponding pixels, which are then used to train a machine learning model for improved pixel labeling accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual user input is used to label objects in images, then training data can be obtained, but labeling accuracy deteriorates due to manual dexterity limitations
Solution Approach 1:
The patent uses structured object representations (SVG, PDF, HTML) as templates to generate multiple training examples through transformations such as scaling, rotating, and color changes. This copying approach allows one accurately labeled object to produce many training variants without requiring manual relabeling, thereby increasing the quantity of training data while maintaining labeling accuracy through programmatic precision
Solution Approach 2:
The patent performs preliminary automated labeling by extracting object boundaries and attributes from structured representations before manual review. This preliminary action pre-processes the labeling task, reducing the manual effort required and minimizing errors by establishing accurate ground truth labels before any user interaction occurs
2Quantity of substance
If manual user input is used to label objects in images, then training data can be obtained, but time consumption increases significantly
Solution Approach 1:
The patent generates multiple training examples by programmatically copying and transforming structured object representations through scaling, rotating, and color changes. This automated copying process produces large numbers of training examples instantaneously without requiring proportional manual time investment, thereby increasing training data quantity while minimizing time loss
Solution Approach 2:
The system performs self-service by automatically extracting object boundaries, attributes, and labels from structured representations without requiring manual user input for each training example. The automated pipeline generates and prepares training data independently, eliminating the time-consuming manual border definition process while maintaining high labeling accuracy
3Measurement precision
If manual user input is used to label objects in images, then training data can be obtained, but costs increase due to expense of obtaining sufficient examples
Solution Approach 1:
The patent uses automated copying and transformation of structured object representations to generate large volumes of training data at minimal cost. By programmatically creating variants through scaling, rotating, and color changes, the system eliminates the need to manually commission and pay for numerous individual labeling tasks, thereby reducing costs while maintaining model accuracy through sufficient training examples
Solution Approach 2:
The system performs self-service by automatically generating and preparing training data from structured representations without requiring external human labor. This automated approach eliminates costs associated with manual labeling services while maintaining high data quality, making it economically efficient to obtain sufficient training examples for accurate model training
Data Source
AI summary
Techniques are described to generate improved training data for pixel labeling. To generate training data, objects are displayed in a user interface by a computing device, e.g., iteratively. The objects are taken from a structured object representation associated with a respective one of a plurality of images. The structured object representation defines a hierarchical relationship of the objects within the respective image. Inputs are then received that are originated through user interaction with the user interface. The inputs label respective ones of the iteratively displayed objects, e.g., as text, a graphical element, background, foreground, and so forth. A model is trained by the computing device using machine learning.


