AI Training Data Segmentation for Transparent Output Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of AI systems lack control over the training process and desire greater insights into the factors influencing the training outcomes, leading to mistrust of 'black box' solutions.

Innovation Solution

A method and system for training AI systems using variable input data, allowing users to categorize and modify training data through interactive interfaces, creating multiple AI systems trained on different subsets of data to generate varied outputs, and iteratively refining these outputs based on user feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If AI systems are trained on large data-sets using traditional black box methods, then training efficiency and automation are improved, but user control and transparency over training outcomes deteriorate

Engineering Contradiction:
Improveautomation of training processVSAvoiduser control over training
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent segments the training data into multiple distinct data-sets, each representing different factors or characteristics. Users can selectively combine these segmented data-sets to train AI systems, allowing granular control over which factors influence the training outcome without sacrificing automation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically allows users to modify the composition of training data-sets during the training process. Users can add, remove, or adjust the weight of different data-sets based on desired outcomes, creating a flexible and adaptive training process that maintains both automation and user control.

Inventive Principle:
Principle #15Dynamics

2Productivity

If AI systems are trained on large data-sets using traditional black box methods, then productivity and output generation are improved, but user understanding and trust in training factors deteriorate

Engineering Contradiction:
ImproveAI output generationVSAvoidinsight into training factors
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system provides feedback mechanisms that allow users to observe how different data-set combinations affect AI training outcomes. Users can iteratively adjust their data-set selections and immediately see the impact on generated outputs, maintaining productivity while gaining understanding and trust in the training factors.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

By segmenting training data into distinct, labeled data-sets representing different factors, the system enables users to trace which specific data-sets contribute to which outputs. This segmentation preserves information about training factors while maintaining high productivity through efficient automated processing.

Inventive Principle:
Principle #1Segmentation

3Reliability

If users are given control over training data selection and modification, then transparency and trust are improved, but system complexity and operational difficulty increase

Engineering Contradiction:
Improveuser trust in AI systemVSAvoidtraining system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system employs a universal interface and standardized data-set structures that work across different AI training scenarios. This multi-functionality allows users to control training without dealing with scenario-specific complexities, maintaining trust through transparency while keeping the system accessible and relatively simple to operate.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12493818B2Method and system for generating variable training data for artificial intelligence systems
Publication Date: 2025.12.09 UNITY TECH APS
  • US12493818B2 patent drawing
  • US12493818B2 patent drawing
  • US12493818B2 patent drawing

AI summary

An artificial intelligence (AI) training method is disclosed. Training data associated with a training task is received. The training data is categorized into a plurality of categories. A set of category groups are generated, wherein each category group of the set of category groups includes one or more of the plurality of categories. A first AI system is trained for the training task using a first subset of the training data. The first subset of the training data corresponds to the one or more of the plurality of categories included in a first group of the set of category groups. A second AI system is trained for the training task using a second subset of the training data. The first AI system is used to generate a first output for the task and the second AI system to generate a second output for the task.