AI Training Data Licensing Using Vector-Based Access Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large organizations face legal risks and high transaction costs when obtaining specialized training data sets for large language models (LLMs) due to copyright infringement concerns, while smaller organizations lack economically viable ways to license their data, leading to a difficult choice between infringement or forgoing use.

Innovation Solution

A system and method for managing and licensing inference augmentation data using high-dimensional vector representations of media objects, enabling automated, cost-effective distribution and access control for LLMs, eliminating the need for traditional training and ensuring compliance with copyright laws.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual negotiation and collection of licensing fees is used for large quantities of general data, then licensing compliance is improved, but transaction costs increase significantly

Engineering Contradiction:
Improvelicensing complianceVSAvoidtransaction costs
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables automated self-service licensing where data owners can directly license their datasets to consumers through the marketplace interface, eliminating the need for manual negotiation. The automated licensing system handles contract generation, fee collection, and access control automatically, reducing transaction costs while maintaining compliance.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical licensing processes with an automated electronic system. The licensing marketplace uses software-based contract management, automated fee collection mechanisms, and digital access control to substitute human-mediated transactions, thereby reducing transaction costs and improving efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If scraping data from the Internet without permission is used, then data availability is improved, but legal risks increase due to copyright infringement

Engineering Contradiction:
Improvedata availabilityVSAvoidlegal risks
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The licensing marketplace acts as an intermediary between data owners and consumers. Instead of consumers directly scraping data, they must go through the marketplace to obtain licensed access. This intermediary system facilitates legal data acquisition while maintaining data availability, eliminating copyright infringement risks.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments data access into licensed portions rather than allowing unrestricted scraping. The marketplace divides data into individual datasets that can be licensed separately, enabling consumers to access specific data portions legally through automated licensing mechanisms.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If licensing data from smaller organizations and authors is used, then specialized training data quality is improved, but transaction costs exceed the value of the training data

Engineering Contradiction:
Improvetraining data qualityVSAvoidtransaction costs
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The marketplace enables small data owners to directly license their specialized datasets to consumers without requiring manual negotiation. The automated licensing system handles all transactional overhead, making the process as efficient as accessing public data while maintaining the quality benefits of specialized datasets from smaller organizations.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250363190A1Generation and management of training and tuning data for artificial intelligence systems
Publication Date: 2025.11.27 SUHADOLNIK THOMAS PATRICK
  • US20250363190A1 patent drawing
  • US20250363190A1 patent drawing
  • US20250363190A1 patent drawing

AI summary

Systems and methods are provided for management of inference augmentation data for artificial intelligence systems. An owner of a plurality of media objects is allowed to upload the plurality of media objects to a local data repository, and a dataset comprising the plurality of data objects each containing a high-dimensional vector representation of a content of a corresponding media object of the plurality of media objects. A user operating an artificial intelligence system licenses a dataset with a set of usage rights, and the artificial intelligence systems is selectively allowed to access the dataset according to a set of usage rights.