Machine Learning Data Generation for Unlicensed Data Replacement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Collecting a sufficient volume of machine learning data can be difficult when the data involves a third party right, leading to insufficient accuracy in estimation models.
Innovation Solution
A method and apparatus for generating machine learning data by reading target data with added license-requested portion information and replacement data, replacing unlicensed portions with replacement data based on license information, and storing this information in a blockchain ledger.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If specific data containing third party rights is collected for machine learning training, then the estimation model accuracy is improved, but the difficulty of data acquisition increases due to licensing requirements
Solution Approach 1:
The patent segments the target data into licensed portions and unlicensed portions, allowing the system to process and utilize the licensed portions for machine learning training while excluding or replacing unlicensed portions. This segmentation enables the system to maximize the use of available data without violating third-party rights, thereby improving model accuracy while maintaining ease of data acquisition.
Solution Approach 2:
The patent introduces a license management system as an intermediary between data sources and machine learning training processes. This intermediary automatically verifies licensing status, manages permissions, and facilitates data access, thereby reducing the complexity and difficulty of acquiring licensed specific data for training.
2Measurement precision
If a sufficient volume of machine learning data is collected, then the estimation model accuracy is improved, but the availability of licensed specific data is limited
Solution Approach 1:
The patent creates machine learning training data by copying and processing licensed portions of target data. The system extracts, replicates, and prepares licensed data segments for training purposes, generating sufficient training data volume from available licensed sources without needing to acquire additional licensed specific data.
Solution Approach 2:
The patent changes the parameters of data usage by transforming licensed target data into machine learning training data through processing, filtering, and formatting operations. This parameter transformation enables the system to maximize the utility of limited licensed data and generate sufficient training volume from smaller licensed datasets.
3Reliability
If target data is processed to identify and replace unlicensed portions, then the compliance with third party rights is improved, but the data processing complexity increases
Solution Approach 1:
The patent performs preliminary identification and marking of licensed and unlicensed portions in target data before the machine learning training process. By pre-processing the data to segment and tag licensed portions, the system simplifies subsequent processing steps and ensures compliance with third-party rights is built into the data pipeline from the beginning, reducing overall processing complexity.
Solution Approach 2:
The patent implements an automated license management system that self-verifies licensing status, automatically identifies licensed portions, and autonomously processes data segmentation and replacement. This self-service approach reduces manual intervention and simplifies the overall data processing workflow while maintaining high compliance standards.
Data Source
AI summary
A sufficient volume of machine learning data can be prepared when a third party right is involved. License-requested portion information (23) and replacement data are added to target data (21) on a server. The license-requested portion information indicates a portion (license-requested portion (24)) to be licensed. The replacement data is to replace the license-requested portion of the target data that is not licensed. License information (30) indicating whether the license-requested portion is licensed is also produced. To generate machine learning data, the target data and the license information are read. The license-requested portion of the target data that is not licensed is replaced with the replacement data to generate the machine learning data. The machine learning data can thus be generated from target data that is not licensed, allowing a sufficient volume of machine learning data to be prepared easily.


