Visual Semi-Automated Data Labeling Platform

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Supervised learning for machine learning models requires accurate and time-consuming labeling of large training data sets, which is resource-intensive and demands significant domain knowledge, limiting the efficiency of the data labeling process.

Innovation Solution

A visual platform for rapid assignment of labels to training data based on ontological classes, utilizing visualizations and user input to iteratively transform and label data points, leveraging weak supervision strategies and visual tools to facilitate semi-automated data labeling, enabling scalable and accurate labeling of data across varying sizes and types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If supervised learning is implemented with accurate manual labeling of training data, then the reliability of the ML model is improved, but the loss of time and resource requirements increase significantly

Engineering Contradiction:
Improveaccuracy of labelsVSAvoidtime for data labeling
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables semi-automated labeling where the ML model performs self-labeling on unlabeled data based on patterns learned from initially labeled data. This self-service mechanism reduces manual intervention while maintaining label quality through iterative refinement and user validation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback loops where user corrections to automatically generated labels are fed back into the system to refine the labeling process. This continuous feedback improves labeling accuracy over time while reducing the time required for manual labeling of new data.

Inventive Principle:
Principle #23Feedback

2Reliability

If supervised learning is implemented with accurate manual labeling of training data, then the reliability of the ML model is improved, but the resource intensity increases significantly

Engineering Contradiction:
Improveaccuracy of labelsVSAvoidresource requirements for labeling
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system enables semi-automated labeling where the ML model performs self-labeling on unlabeled data based on patterns learned from initially labeled data. This self-service mechanism reduces manual intervention while maintaining label quality through iterative refinement and user validation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system introduces an automated labeling intermediary that acts as a bridge between manual labeling and final labeled data output. This intermediary processes data through multiple labeling functions and ensembles, reducing the need for direct human involvement in labeling every data point.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If traditional data labeling processes are used, then domain knowledge requirements are reduced, but the productivity and speed of labeling decrease

Engineering Contradiction:
Improvespeed of labelingVSAvoiddomain knowledge required
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system introduces an automated labeling intermediary that acts as a bridge between manual labeling and final labeled data output. This intermediary processes data through multiple labeling functions and ensembles, reducing the need for direct human involvement in labeling every data point.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces manual mechanical labeling processes with automated computational labeling functions. These functions use algorithms and heuristics to automatically assign labels, substituting human cognitive effort with machine-based automation while maintaining accuracy through user validation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Quantity of substance

If manual labeling of large data sets is performed, then the quantity of labeled training data is sufficient, but the loss of time and resources become prohibitive

Engineering Contradiction:
Improveamount of labeled training dataVSAvoidtime for labeling large data sets
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system implements continuous labeling where the ML model continuously processes and labels data points as they become available. This continuous action allows large quantities of data to be labeled efficiently over time rather than requiring intensive batch processing, maintaining productivity while reducing overall time investment.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system enables semi-automated labeling where the ML model performs self-labeling on unlabeled data based on patterns learned from initially labeled data. This self-service mechanism reduces manual intervention while maintaining label quality through iterative refinement and user validation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11593458B2System for time-efficient assignment of data to ontological classes
Publication Date: 2023.02.28 ACCENTURE GLOBAL SOLUTIONS LTD
  • US11593458B2 patent drawing
  • US11593458B2 patent drawing
  • US11593458B2 patent drawing

AI summary

Implementations are directed to receiving a set of training data including a plurality of data points, at least a portion of which are to be labeled for subsequent supervised training of a computer-executable machine learning (ML) model, providing at least one visualization based on the set of training data, the at least one visualization including a graphical representation of at least a portion of the set of training data, receiving user input associated with the at least one visualization, the user input indicating an action associated with a label assigned to a respective data point in the set of training data, executing a transformation on data points of the set of training data based on one or more heuristics representing the user input to provide labeled training data in a set of labeled training data, and transmitting the set of labeled training data for training the ML model.