Bifurcated AI Data Labeling Workspaces for Secure Preprocessing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing artificial intelligence data labeling processes face challenges in maintaining data security, consistency, efficient storage, and workflow management, which can lead to increased storage needs, security risks, and reduced flexibility in preprocessing and data access.

Innovation Solution

Implementing a system that stores labeled data versions in separate workspaces with bifurcated management, using flat-file repositories and assigning security rights to metadata, allowing controlled access and preprocessing within segmented workspaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is copied to a new database for preprocessing, then data security is maintained, but storage capacity increases and security risks are exposed

Engineering Contradiction:
Improvedata securityVSAvoidstorage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments data management into separate workspaces (e.g., labeled data workspace, training data workspace) with distinct access controls. Each workspace contains specific datasets and associated metadata, allowing preprocessing operations to be performed in isolated environments without copying entire databases. This segmentation maintains security while optimizing storage utilization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces metadata as an intermediary layer between raw data and processed data. Metadata stores information about data transformations, preprocessing steps, and access permissions without requiring physical copies of the actual data. This intermediary approach enables security management and tracking without increasing storage capacity proportionally to data duplication.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If labeled data is preprocessed for specific applications, then data usability is improved, but security risks increase due to removal of encryption or supplementation with PII

Engineering Contradiction:
Improvedata usabilityVSAvoidsecurity risks
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent performs preliminary preprocessing operations within controlled workspaces before data leaves the secure environment. Feature engineering, data transformation, and supplementation with personally identifiable information (PII) are executed in advance within the labeled data workspace, which maintains encryption and access controls. The preprocessed results are then exported as training data with appropriate metadata tracking the transformations applied.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different security measures and preprocessing capabilities to different workspaces based on their specific requirements. The labeled data workspace maintains strict encryption and access controls for sensitive operations, while the training data workspace allows broader access for model development. This local quality approach enables tailored security and preprocessing for each data stage without compromising overall system security.

Inventive Principle:
Principle #3Local quality

3Stability of the object's composition

If previously labeled data is stored for future reference, then consistency can be maintained, but access is limited by security classifications

Engineering Contradiction:
Improvelabeling consistencyVSAvoiddata accessibility
Core Design Contradiction:
Stability of the object's compositionVSEase of operation

Solution Approach 1:

The patent creates workspaces that serve multiple functions: storing labeled data, maintaining version history, enabling preprocessing operations, and facilitating controlled access. The labeled data workspace acts as a universal repository that supports both security requirements and accessibility needs by implementing role-based access controls and metadata-driven retrieval, allowing different user roles to access appropriate data without compromising consistency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements metadata tracking that provides feedback about data usage, access patterns, and consistency requirements. Metadata records store information about labeling decisions, version history, and access permissions, enabling the system to maintain consistency across different access scenarios. This feedback mechanism allows the system to enforce security policies while enabling appropriate data retrieval for future labeling tasks.

Inventive Principle:
Principle #23Feedback

4Productivity

If data is collected and stored for AI training, then model development is enabled, but the data requires additional preprocessing that risks altering or exposing the data

Engineering Contradiction:
ImproveAI model developmentVSAvoiddata integrity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the data lifecycle into distinct workspaces: a labeled data workspace for secure preprocessing operations and a training data workspace for model development. This segmentation allows the original labeled data to remain intact and encrypted in the first workspace while enabling comprehensive preprocessing operations in the second workspace. The segmentation ensures data integrity is maintained throughout the AI model development process.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12386804B2Systems and methods for maintaining rights management while labeling data for artificial intelligence model development
Publication Date: 2025.08.12 CAPITAL ONE SERVICES LLC
  • US12386804B2 patent drawing
  • US12386804B2 patent drawing
  • US12386804B2 patent drawing

AI summary

Systems and methods are described for maintaining bifurcated data management while labeling data for artificial intelligence model development. For example, the system may receive a first label for a first sample from a first dataset, wherein the first dataset is accessible to a first subset of a plurality of users, and wherein the first subset comprises a first attribute. The system may receive first version metadata of the first label, wherein the first version metadata comprises a proposed label for the first sample assigned by a first user. The system may determine, based on a first user input from the first user, a first grouping of source code files for storing the first version metadata, wherein the first grouping of source code files is accessible to a second subset of the plurality of users, and wherein the second subset comprises a second attribute.