Prediction-Guided Sequential Data Learning Method

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High capacity machine learning models require large labeled datasets for training, which is inefficient due to the bottleneck of acquiring labels, especially in big data environments, and existing methods like crowdsourcing are expensive and scalable poorly.

Innovation Solution

A prediction-guided sequential data learning method that uses self-supervised learning to generate initial classifiers from unlabelled data sequences, allowing for efficient initial learning and subsequent update learning with a small number of labeled data for accurate semantic and data classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high capacity models are used to handle large variations in sequential data, then prediction accuracy is improved, but the requirement for large labeled datasets increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidlabeled dataset size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by performing self-supervised pre-training on large unlabeled sequential data before fine-tuning with small labeled datasets. The model learns temporal patterns and representations from unlabeled data in advance, which prepares it for subsequent supervised learning with minimal labeled data, thereby reducing the overall labeled data requirement while maintaining high prediction accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service through self-supervised learning where the model generates its own training signals from unlabeled sequential data by predicting future states or reconstructing past states. This self-generated supervision eliminates the need for external human labeling while still providing sufficient training signal for high-capacity models to learn effective representations

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If crowdsourcing is used to acquire large training sets, then data labeling quantity is improved, but cost and time consumption increase

Engineering Contradiction:
Improvelabeled data quantityVSAvoidlabeling time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs self-service by automatically generating training labels from unlabeled sequential data through self-supervised learning mechanisms. The model predicts future states or reconstructs past states to create supervisory signals without human intervention, completely eliminating the time-consuming crowdsourcing process while still producing sufficient labeled data for training

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If crowdsourcing is used to acquire large training sets, then data labeling quantity is improved, but labeling quality deteriorates

Engineering Contradiction:
Improvelabeled data quantityVSAvoidlabeling quality
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The model serves itself by generating training signals through self-supervised learning on unlabeled data. This automated label generation process eliminates human annotators entirely, avoiding the quality issues inherent in crowdsourcing while still providing sufficient supervisory signal for effective training of high-capacity models

Inventive Principle:
Principle #25Self-service

4Measurement precision

If high capacity models with >10^8 parameters are trained, then prediction accuracy is improved, but the bottleneck of acquiring labels is exacerbated

Engineering Contradiction:
Improveprediction accuracyVSAvoidlabeling efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by conducting self-supervised pre-training on large unlabeled datasets before supervised fine-tuning. This preliminary learning phase allows high-capacity models to learn effective representations from unlabeled data, reducing the amount of labeled data needed in subsequent stages and thereby improving overall labeling efficiency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The model performs self-service by generating its own training signals from unlabeled sequential data through self-supervised learning. This self-generated supervision enables high-capacity models to learn effectively without relying on expensive and inefficient human labeling processes

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11468286B2Prediction guided sequential data learning method
Publication Date: 2022.10.11 LEICA MICROSYSTEMS CMS GMBH
  • US11468286B2 patent drawing
  • US11468286B2 patent drawing

AI summary

A computerized prediction guided learning method for classification of sequential data performs a prediction learning and a prediction guided learning by a computer program of a computerized machine learning tool. The prediction learning uses an input data sequence to generate an initial classifier. The prediction guided learning may be a semantic learning, an update learning, or an update and semantic learning. The prediction guided semantic learning uses the input data sequence, the initial classifier and semantic label data to generate an output classifier and a semantic classification. The prediction guided update learning uses the input data sequence, the initial classifier and label data to generate an output classifier and a data classification. The prediction guided update and semantic learning uses the input data sequence, the initial classifier and semantic and label data to generate an output classifier, a semantic classification and a data classification.