Common Data Infrastructure for Multi-Source Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge in industries dealing with confidential information is the inefficiency and lack of security in data labeling processes for machine learning workflows, which hinders the effective use of machine learning classifiers for predictions and insights.

Innovation Solution

The implementation of a system that provides multiple data labeling interfaces with a common data infrastructure, allowing data from various sources in different formats to be converted into a common schema, and using a machine learning classifier to generate label suggestions, which are then processed by different data labeling tools.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data from multiple sources in different formats is processed by multiple specialized labeling tools, then labeling capability and versatility are improved, but system complexity and integration difficulty increase

Engineering Contradiction:
Improvelabeling capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a data infrastructure layer as an intermediary between diverse data sources and multiple labeling tools. This layer provides standardized data access, format conversion, and management capabilities, allowing labeling tools to work with unified data interfaces while the infrastructure handles the complexity of multiple formats and sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The data infrastructure is designed as a universal platform that serves multiple labeling tools simultaneously. It provides common data access, storage, and management functions that can be shared across different labeling tools, eliminating the need for each tool to independently handle data acquisition and preprocessing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of manufacture

If data is accessed and processed by multiple specialized labeling tools independently, then each tool can be optimized for its specific task, but data access efficiency and consistency deteriorate

Engineering Contradiction:
Improvetool optimizationVSAvoiddata access efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent merges data access, storage, and management functions into a single centralized data infrastructure. This consolidation allows all labeling tools to access the same data through unified interfaces, improving access efficiency and ensuring consistency across different labeling operations while still allowing individual tools to be optimized for their specific labeling tasks.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If confidential information is processed through multiple independent labeling tools, then specialized processing capability is improved, but security and privacy protection deteriorate

Engineering Contradiction:
Improveprocessing capabilityVSAvoidsecurity risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The data infrastructure acts as a secure intermediary between data sources and labeling tools. It implements centralized security controls, access management, and data protection mechanisms that ensure confidential information is handled securely while still allowing specialized labeling tools to process the data according to their specific capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If each data labeling tool has its own data access and storage methods, then tool independence and flexibility are improved, but overall system scalability and maintainability worsen

Engineering Contradiction:
Improvetool independenceVSAvoidintegration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the system into two independent layers: a data infrastructure layer that handles data access, storage, and management, and a labeling tools layer that provides specialized labeling capabilities. This segmentation allows labeling tools to remain independent and flexible in their processing logic while the infrastructure layer handles the complexity of data management, improving overall system scalability and maintainability.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12321858B2Multiple data labeling interfaces with a common data infrastructure
Publication Date: 2025.06.03 CAPITAL ONE SERVICES LLC
  • US12321858B2 patent drawing
  • US12321858B2 patent drawing
  • US12321858B2 patent drawing

AI summary

Systems as described herein may provide multiple data labeling interfaces with a common data infrastructure. An annotation system may retrieve data from a plurality of data sources and convert the data to a common schema. The annotation system may train a machine learning classifier to output a plurality of label suggestions, which may be sent to a plurality of data labeling tools. A plurality of labels may be received from the data labeling tools. The annotation system may accordingly export the plurality of labels and the converted data in the common schema to a label database.