Bifurcated AI Data Labeling Workspaces for Secure Preprocessing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing artificial intelligence data labeling processes face challenges in maintaining data security, consistency, efficient storage, and workflow management, which can lead to increased storage needs, security risks, and reduced flexibility in preprocessing and data access.
Innovation Solution
Implementing a system that stores labeled data versions in separate workspaces with bifurcated management, using flat-file repositories and assigning security rights to metadata, allowing controlled access and preprocessing within segmented workspaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is copied to a new database for preprocessing, then data security is maintained, but storage capacity increases and security risks are exposed
Solution Approach 1:
The patent segments data management into separate workspaces (e.g., labeled data workspace, training data workspace) with distinct access controls. Each workspace contains specific datasets and associated metadata, allowing preprocessing operations to be performed in isolated environments without copying entire databases. This segmentation maintains security while optimizing storage utilization.
Solution Approach 2:
The patent introduces metadata as an intermediary layer between raw data and processed data. Metadata stores information about data transformations, preprocessing steps, and access permissions without requiring physical copies of the actual data. This intermediary approach enables security management and tracking without increasing storage capacity proportionally to data duplication.
2Adaptability or versatility
If labeled data is preprocessed for specific applications, then data usability is improved, but security risks increase due to removal of encryption or supplementation with PII
Solution Approach 1:
The patent performs preliminary preprocessing operations within controlled workspaces before data leaves the secure environment. Feature engineering, data transformation, and supplementation with personally identifiable information (PII) are executed in advance within the labeled data workspace, which maintains encryption and access controls. The preprocessed results are then exported as training data with appropriate metadata tracking the transformations applied.
Solution Approach 2:
The patent applies different security measures and preprocessing capabilities to different workspaces based on their specific requirements. The labeled data workspace maintains strict encryption and access controls for sensitive operations, while the training data workspace allows broader access for model development. This local quality approach enables tailored security and preprocessing for each data stage without compromising overall system security.
3Stability of the object's composition
If previously labeled data is stored for future reference, then consistency can be maintained, but access is limited by security classifications
Solution Approach 1:
The patent creates workspaces that serve multiple functions: storing labeled data, maintaining version history, enabling preprocessing operations, and facilitating controlled access. The labeled data workspace acts as a universal repository that supports both security requirements and accessibility needs by implementing role-based access controls and metadata-driven retrieval, allowing different user roles to access appropriate data without compromising consistency.
Solution Approach 2:
The patent implements metadata tracking that provides feedback about data usage, access patterns, and consistency requirements. Metadata records store information about labeling decisions, version history, and access permissions, enabling the system to maintain consistency across different access scenarios. This feedback mechanism allows the system to enforce security policies while enabling appropriate data retrieval for future labeling tasks.
4Productivity
If data is collected and stored for AI training, then model development is enabled, but the data requires additional preprocessing that risks altering or exposing the data
Solution Approach 1:
The patent segments the data lifecycle into distinct workspaces: a labeled data workspace for secure preprocessing operations and a training data workspace for model development. This segmentation allows the original labeled data to remain intact and encrypted in the first workspace while enabling comprehensive preprocessing operations in the second workspace. The segmentation ensures data integrity is maintained throughout the AI model development process.
Data Source
AI summary
Systems and methods are described for maintaining bifurcated data management while labeling data for artificial intelligence model development. For example, the system may receive a first label for a first sample from a first dataset, wherein the first dataset is accessible to a first subset of a plurality of users, and wherein the first subset comprises a first attribute. The system may receive first version metadata of the first label, wherein the first version metadata comprises a proposed label for the first sample assigned by a first user. The system may determine, based on a first user input from the first user, a first grouping of source code files for storing the first version metadata, wherein the first grouping of source code files is accessible to a second subset of the plurality of users, and wherein the second subset comprises a second attribute.


