RPA Data Labelling from Production Sensor Workflows

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high cost and labor intensity of generating labelled datasets for machine learning models, particularly in industrial production settings, where traditional data labelling methods are time-consuming and prone to human errors, and existing approaches do not efficiently address the challenges of data collection and labelling quality.

Innovation Solution

Implementing a system that uses sensors to capture data during normal production processes, leveraging expert workers' classifications to create labelled datasets, and deploying machine learning models to automate classification tasks, while continuously improving model accuracy through additional training and data collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional manual data labelling methods are used, then data quality can be maintained through human expertise, but the process becomes expensive and labor intensive

Engineering Contradiction:
Improvedata labelling qualityVSAvoidlabelling efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system allows workers to perform their normal duties while sensors automatically capture and label their actions, eliminating the need for dedicated data collection tasks. The production process itself generates the labelled data as a byproduct, making the system self-servicing for data collection purposes

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The sensor system serves multiple functions: it monitors production quality, tracks worker performance, and simultaneously generates labelled training data for machine learning models. This multi-functionality resolves the contradiction by making the data collection process efficient without compromising quality

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If more labelled data is collected to improve model accuracy, then model performance improves, but the cost and time for data collection increases

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Sensors continuously capture data during normal production operations without interruption. The system operates continuously as workers perform their duties, accumulating labelled data over time without dedicated data collection periods, thus improving model accuracy without proportional time loss

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system pre-captures labelled data during production processes before it is needed for model training. By having data already labelled and stored in advance, the system eliminates the time delay that would occur if data were collected after model development begins

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual data labelling is performed to ensure data quality, then labelling accuracy improves, but the process becomes prone to human errors and is time-consuming

Engineering Contradiction:
Improvelabelling accuracyVSAvoidlabelling process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system replaces manual human labelling with automated sensor-based detection and classification. Sensors objectively capture and label worker actions without human intervention in the labelling process itself, eliminating human errors while maintaining accuracy, and reducing process complexity by automating what was previously a manual task

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Quantity of substance

If traditional data collection methods are used, then data can be obtained for training, but the process is expensive and labor intensive

Engineering Contradiction:
Improvetraining data quantityVSAvoiddata generation efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The production system generates labelled training data as a self-service byproduct of normal operations. Workers do not need to be trained as data annotators, and no separate data collection team is required. The system automatically produces the training data needed, increasing efficiency while maintaining data quantity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system merges the data collection function with the existing production process. Instead of separating data collection from production, the system combines them so that production activities simultaneously generate the labelled data needed for training, thereby increasing productivity without sacrificing data quantity

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20230229119A1Robotic process automation (RPA)-based data labelling
Publication Date: 2023.07.20 BAIDU USA LLC
  • US20230229119A1 patent drawing
  • US20230229119A1 patent drawing
  • US20230229119A1 patent drawing

AI summary

One application of deep learning methods and labelled data is for industrial production or work applications. For such applications implemented with machine learning applications, massive amounts of data are required to train, validate, and/or tune models for better fitting the requirements. However, obtaining such data has typically be costly and difficult. Embodiments provide adaptable processes that provide data labelling methods for work settings. Embodiments take advantage of the work or production processes to label and collect data, which save time and money and improves accuracy. Embodiments prevent or reduce the need for worker training costs and human mistake-triggered data labelling problems. Embodiments also improve data labelling quality and speed-up of the development cycle.