Cloud Metadata Distribution for Scalable Data Sharing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data sharing methods are cumbersome, slow, and expensive, particularly for smaller entities, as they require manual data cleaning, de-identification, and aggregation, and do not allow scalable sharing, leading to latency and limited access to valuable data assets.

Innovation Solution

A data exchange system that enables secure, scalable sharing of data through a cloud computing service, allowing data providers to create listings with metadata descriptions, enabling data consumers to access and utilize data without copying or transferring the actual data, using a data dictionary generation system for automatic metadata generation and updates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data cleaning, de-identification, and aggregation are performed for data sharing, then data security and quality are improved, but the process becomes cumbersome, slow, and expensive

Engineering Contradiction:
Improvedata securityVSAvoiddata sharing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs data cleaning, de-identification, and aggregation in advance before data sharing requests are made. Data is preprocessed and stored in a ready-to-share format, eliminating the need for manual processing during actual sharing operations. This preliminary action resolves the contradiction by maintaining security and quality standards while enabling rapid data access when needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of manually processing and transferring actual data, the system creates and distributes metadata copies that describe the data. These metadata objects contain information about data structure, quality, and access requirements, allowing consumers to understand and utilize data without copying or transferring the actual data assets, thus maintaining security while improving sharing speed.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If manual data processing methods are used, then data quality control is maintained, but scalability is limited

Engineering Contradiction:
Improvedata qualityVSAvoidscalability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system implements automated metadata generation and distribution mechanisms that operate without manual intervention. Metadata is automatically created from data sources, validated against quality standards, and distributed to consumers through the platform. This self-service approach maintains data quality control through systematic processes while enabling the system to scale to handle numerous data sharing requests simultaneously.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the fundamental parameter of data sharing from transferring actual data to distributing metadata objects. This parameter change enables scalability because metadata is lightweight and can be replicated and distributed efficiently across the platform, while data quality is maintained through structured metadata schemas and validation rules that ensure consistent data description and access control.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If data is encrypted and stored securely, then unauthorized access is prevented, but data accessibility and sharing ease are reduced

Engineering Contradiction:
Improveaccess securityVSAvoiddata sharing ease
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system introduces metadata as an intermediary between encrypted data and consumers. Metadata objects describe the encrypted data without containing the actual data, allowing consumers to discover, understand, and request access to data while the actual data remains securely encrypted. This intermediary approach maintains strong security controls while improving ease of operation by enabling data exploration and sharing workflows without requiring data decryption or transfer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4345643A1Distributing large amounts of global metadata using object files
Publication Date: 2024.04.03 SNOWFLAKE INC
  • EP4345643A1 patent drawingFigure 1A
  • EP4345643A1 patent drawingFigure 1B
  • EP4345643A1 patent drawingFigure 2

AI summary

A method, a system and a computer-storage medium are provided comprising accessing, at a source deployment, a data dictionary generated for a listing object, the data dictionary describing shared data offered by the listing object and identifying, by at least one hardware processor, a destination deployment to which the listing object is to be replicated. In a further step the writing of the data dictionary and the shared data offered by the listing object to a global cloud object storage bucket associated with the destination deployment is carried out and based on the writing of the data dictionary and the shared data offered by the listing object to the global cloud object storage bucket, a notification is transmitted to the destination deployment, wherein the notification causes the destination deployment to replicate the shared data and the data dictionary at the destination deployment.