Machine Learning Data Generation for Unlicensed Data Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Collecting a sufficient volume of machine learning data can be difficult when the data involves a third party right, leading to insufficient accuracy in estimation models.

Innovation Solution

A method and apparatus for generating machine learning data by reading target data with added license-requested portion information and replacement data, replacing unlicensed portions with replacement data based on license information, and storing this information in a blockchain ledger.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If specific data containing third party rights is collected for machine learning training, then the estimation model accuracy is improved, but the difficulty of data acquisition increases due to licensing requirements

Engineering Contradiction:
Improveestimation model accuracyVSAvoiddata acquisition difficulty
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent segments the target data into licensed portions and unlicensed portions, allowing the system to process and utilize the licensed portions for machine learning training while excluding or replacing unlicensed portions. This segmentation enables the system to maximize the use of available data without violating third-party rights, thereby improving model accuracy while maintaining ease of data acquisition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a license management system as an intermediary between data sources and machine learning training processes. This intermediary automatically verifies licensing status, manages permissions, and facilitates data access, thereby reducing the complexity and difficulty of acquiring licensed specific data for training.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a sufficient volume of machine learning data is collected, then the estimation model accuracy is improved, but the availability of licensed specific data is limited

Engineering Contradiction:
Improveestimation model accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates machine learning training data by copying and processing licensed portions of target data. The system extracts, replicates, and prepares licensed data segments for training purposes, generating sufficient training data volume from available licensed sources without needing to acquire additional licensed specific data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameters of data usage by transforming licensed target data into machine learning training data through processing, filtering, and formatting operations. This parameter transformation enables the system to maximize the utility of limited licensed data and generate sufficient training volume from smaller licensed datasets.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If target data is processed to identify and replace unlicensed portions, then the compliance with third party rights is improved, but the data processing complexity increases

Engineering Contradiction:
Improvecompliance with third party rightsVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary identification and marking of licensed and unlicensed portions in target data before the machine learning training process. By pre-processing the data to segment and tag licensed portions, the system simplifies subsequent processing steps and ensures compliance with third-party rights is built into the data pipeline from the beginning, reducing overall processing complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements an automated license management system that self-verifies licensing status, automatically identifies licensed portions, and autonomously processes data segmentation and replacement. This self-service approach reduces manual intervention and simplifies the overall data processing workflow while maintaining high compliance standards.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260080316A1Machine learning data generation method and machine learning data generation apparatus
Publication Date: 2026.03.19 ID HLDG CORP
  • US20260080316A1 patent drawing
  • US20260080316A1 patent drawing
  • US20260080316A1 patent drawing

AI summary

A sufficient volume of machine learning data can be prepared when a third party right is involved. License-requested portion information (23) and replacement data are added to target data (21) on a server. The license-requested portion information indicates a portion (license-requested portion (24)) to be licensed. The replacement data is to replace the license-requested portion of the target data that is not licensed. License information (30) indicating whether the license-requested portion is licensed is also produced. To generate machine learning data, the target data and the license information are read. The license-requested portion of the target data that is not licensed is replaced with the replacement data to generate the machine learning data. The machine learning data can thus be generated from target data that is not licensed, allowing a sufficient volume of machine learning data to be prepared easily.