Collaborative Dataset Management System for ML Data Versioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The machine learning (ML) development community lacks robust and collaborative tools for managing and curating datasets, leading to inefficient data management and limited collaboration among users, which hinders the growth of high-quality data repositories essential for ML applications.

Innovation Solution

A collaborative dataset management system (CDMS) is implemented to facilitate multi-user interaction and management of ML datasets, enabling dataset creation, review, and evolution through versioning, collaboration, data monetization, and reputation building, with features like graphical user interfaces for easy navigation and automated sanity checks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual data versioning and primitive review tools are used, then data management can be performed with simple processes, but collaboration efficiency and data quality improvement are limited

Engineering Contradiction:
Improvecollaboration efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments data management into distinct modules: version control subsystem, review subsystem, sanity check subsystem, and collaboration subsystem. Each module handles specific aspects of data management independently, enabling complex functionality through modular components that can be developed and maintained separately.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The platform is designed as a universal system that handles multiple functions through a single integrated architecture: versioning, reviewing, automated sanity checks, collaboration features, and data curation all operate within the same system framework, eliminating the need for separate tools and improving collaboration efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If automated sanity checks and versioning are implemented, then data quality and reliability improve, but the complexity of the management system increases

Engineering Contradiction:
Improvedata qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements self-service through automated sanity checks that automatically validate data quality, perform versioning without manual intervention, and maintain data integrity through automated processes. This reduces the need for manual quality control while maintaining high reliability standards.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms where automated sanity checks provide immediate feedback on data quality issues, version control tracks changes over time, and the review subsystem provides feedback loops for data validation. These feedback mechanisms ensure data quality while managing system complexity through automated monitoring and adjustment.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If collaborative tools are introduced to the ML community, then user community growth and data repository quality improve, but the complexity of data management processes increases

Engineering Contradiction:
Improvecollaboration capabilityVSAvoidprocess complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system merges multiple previously separate functions into a single unified platform: version control, review processes, automated sanity checks, collaboration features, and data curation are combined into one integrated system. This consolidation improves collaboration capability while managing complexity through unified architecture rather than multiple separate systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The platform serves as an intermediary layer between individual data managers and the broader community, providing standardized interfaces and protocols that facilitate collaboration. The system mediates between different user needs and data management requirements, enabling versatile collaboration while simplifying processes through standardized workflows.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If manual data curation and labeling are performed, then data quality can be controlled, but development time and labor requirements increase

Engineering Contradiction:
Improvedevelopment speedVSAvoiddata curation time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables continuous data curation through automated processes that operate continuously: versioning updates automatically, sanity checks run continuously to validate data quality, and the review process maintains continuous monitoring. This continuous action eliminates idle time and accelerates development by keeping data curation processes constantly active rather than batch-based.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system replaces manual mechanical data curation processes with automated electronic systems: automated sanity checks substitute for manual quality verification, version control automatically tracks changes instead of manual version management, and the review subsystem automates data validation processes. This substitution dramatically reduces labor requirements and accelerates development speed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11176154B1Collaborative dataset management system for machine learning data
Publication Date: 2021.11.16 AMAZON TECH INC
  • US11176154B1 patent drawing
  • US11176154B1 patent drawing
  • US11176154B1 patent drawing

AI summary

Computer systems and associated methods are disclosed to implement a collaborative dataset management system (CDMS) for machine learning (ML) data. In embodiments, CDMS allows many users to create, review, and collaboratively evolve ML datasets. In embodiments, dataset owners may make their datasets available to other users on CDMS for a fee and under specified licensing conditions. CDMS users can search for other users' datasets on the system to use in their own ML tasks. CDMS users may also create child datasets from existing datasets on the system. Parent and child datasets may be linked so that changes to one dataset are provided to the other via merge requests. A dataset owner may use CDMS to review an incoming merge request using one or more audit jobs before approving the request. In this manner, CDMS provides a shared repository and collaboration system for managing high-quality datasets to power machine learning processes.