Automated Data Dictionary Generation for Secure Cloud Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data sharing methods are cumbersome and inefficient, particularly for smaller entities, as they require time-consuming and costly data transfer processes, lack scalability, and introduce latency, making it difficult for smaller businesses to access valuable data sets due to unpolished and sensitive data formats.
Innovation Solution
A data exchange system facilitated by cloud computing services allows data providers to share data without copying it, using a 'share' object that encapsulates access privileges and metadata, enabling secure and controlled access, with an automated data dictionary generation system to provide comprehensive metadata descriptions for shared data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is shared using traditional data transfer methods, then data consumers can access data sets, but the process is time-consuming and costly, particularly for smaller entities
Solution Approach 1:
The patent extracts only the necessary metadata and access privileges from the complete data set, creating a lightweight 'share' object that enables data access without transferring the actual data. This resolves the contradiction by eliminating time-consuming data transfer while maintaining data accessibility through reference-based sharing.
Solution Approach 2:
The patent introduces a cloud-based data exchange platform as an intermediary that hosts the share object and coordinates data access between providers and consumers. This mediator enables seamless data sharing by managing access control and metadata distribution without requiring direct data transfer between parties.
2Reliability
If data is encrypted for secure storage and access, then unauthorized access is prevented, but data consumers cannot determine what shared data is valuable to them
Solution Approach 1:
The patent segments data information into two distinct components: encrypted data stored in data lakes and unencrypted metadata stored in the share object. This segmentation allows the metadata (data descriptions, schemas, data types) to remain accessible for evaluation while the actual data remains securely encrypted, resolving the contradiction between security and information availability.
Solution Approach 2:
The patent applies different quality characteristics to different parts of the data system: the metadata portion is kept unencrypted and highly accessible to enable data discovery and evaluation, while the actual data portion is encrypted for security. This local differentiation of security levels resolves the contradiction by allowing information access where needed while maintaining security where required.
3Ease of operation
If data is manually described for potential consumers, then data consumers can understand shared data, but the process is time-consuming and cumbersome
Solution Approach 1:
The patent implements self-service by enabling the automated generation of metadata (data dictionaries) that describe shared data. Instead of requiring manual description, the system automatically extracts and structures data documentation from the data sets themselves, significantly reducing the time and effort required to create comprehensive data descriptions while improving ease of data understanding for consumers.
4Productivity
If traditional data sharing methods are used, then data can be transferred, but the methods lack scalability and introduce latency
Solution Approach 1:
The patent creates a lightweight copy or representation of data access rights through the share object, which contains metadata and access privileges rather than the actual data. This copying approach enables scalable data sharing because the share object can be rapidly distributed and replicated without the overhead of transferring large data sets, reducing complexity and improving scalability.
Data Source
AI summary
A data dictionary generation system utilizes a background service that is programmed to automatically populate and update a data dictionary for listings offering shared data. A data dictionary includes metadata describing the shared data overall as well as the individual objects included in the listing, such as the individual tables, schemas, views, and functions. To generate the data dictionary, the data dictionary generation system analyzes the shared data to identify objects, identifies a set of data fields associated with each identified object and populates the set of data fields associated with each identified object based on the shared data offered by the listing. To ensure that a data dictionary for each listing remains up to date, the data dictionary generation system periodically scans the listings to identify any changes to share access granted to the listings.


