Distributing Metadata Descriptions via Data Exchange System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data sharing methods are cumbersome, slow, and expensive, particularly for smaller entities, as they require manual data cleaning, de-identification, and aggregation, and do not allow scalable sharing or real-time access to updated data, limiting the ability of smaller businesses to access valuable data assets.

Innovation Solution

A data exchange system facilitated by cloud computing services that enables data providers to share metadata descriptions of their data assets without copying them, allowing controlled access and updates, using a data dictionary generation system to automatically generate and update metadata for shared data listings, enabling seamless data collaboration and governance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is manually cleaned, de-identified, and aggregated for sharing, then data quality and security are improved, but time consumption and operational complexity increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs data cleaning, de-identification, and aggregation in advance during data ingestion, so that when data is shared, it is already prepared and ready for immediate use without requiring manual processing at the time of sharing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically performs data preparation tasks including cleaning, de-identification, and aggregation without human intervention, allowing the data infrastructure to serve itself and eliminate manual operational overhead

Inventive Principle:
Principle #25Self-service

2Speed

If data is copied and distributed to multiple locations for sharing, then access speed and availability are improved, but storage costs and data synchronization complexity increase

Engineering Contradiction:
Improveaccess speedVSAvoidstorage costs
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

Instead of copying actual data, the system copies only metadata descriptions (data dictionaries) that describe the data assets. This allows multiple consumers to access and query data from a single source location while having fast access to metadata information about the data structures, schemas, and available fields

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system separates metadata descriptions from actual data storage. Metadata is distributed and cached in multiple locations for fast access, while the actual data remains in a single centralized location, eliminating the need to duplicate large volumes of data across multiple storage locations

Inventive Principle:
Principle #1Segmentation

3Reliability

If data sharing is controlled through manual processes, then data security and governance are improved, but scalability and operational efficiency deteriorate

Engineering Contradiction:
Improvedata securityVSAvoidscalability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements automated access control where consumers self-register and self-provision access to data assets through automated workflows. The system automatically manages permissions, authentication, and data lineage tracking without requiring manual approval for each data sharing request, enabling scalable governance

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements automated monitoring and tracking of data access, usage, and lineage with feedback loops that automatically update governance policies, audit trails, and security configurations based on actual data consumption patterns and access requests

Inventive Principle:
Principle #23Feedback

4Loss of information

If complete metadata descriptions are distributed to all data consumers, then data discoverability and usability are improved, but network bandwidth consumption and distribution time increase

Engineering Contradiction:
Improvedata discoverabilityVSAvoidnetwork bandwidth
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The system distributes different portions of metadata to different consumers based on their specific needs and data access patterns. Each consumer receives only the metadata relevant to their queries and data consumption requirements, rather than a complete universal metadata set, optimizing network bandwidth utilization

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11775559B1Distributing large amounts of global metadata using object files
Publication Date: 2023.10.03 SNOWFLAKE INC
  • US11775559B1 patent drawing
  • US11775559B1 patent drawing
  • US11775559B1 patent drawing

AI summary

A data dictionary generation system automatically populates and updates a data dictionary for listings offering shared data. The data listing distribution component distributes the data dictionaries to various remote deployments in a data exchange by using a global messaging framework and replication method. For example, the data listing distribution component replicates a data dictionary generated for the listing and its shared data from a source deployment to one or more destination deployments associated with various geographic regions. The data listing distribution component distributes the listing to the various remote deployments to allow for the listing, including its shared data and data dictionary, to be accessed by users within the geographic region associated with the remote deployment.