Crowdsourced Data Lake Onboarding via Metadata Governance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data lakes face challenges such as lack of semantic consistency, governance, data quality issues, and security risks due to unregulated data ingestion, leading to difficulties in data retrieval and analysis, as well as privacy and regulatory compliance concerns.

Innovation Solution

The Services Delivery Platform (SDP) architecture enables interactive, intuitive data onboarding with governance controls, transforming data into a unified schema and format, using connectors for ingestion from various sources, and applying quality rules and masking techniques to ensure data quality and security, while allowing for crowdsourced data management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data lakes store disparate sources of data in native format without governance, then data consolidation and storage flexibility are improved, but data quality and security deteriorate

Engineering Contradiction:
Improvedata storage flexibilityVSAvoiddata quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces a metadata layer as an intermediary between the raw data lake and users. This metadata layer provides governed access to data without restricting the underlying storage flexibility. The metadata includes data lineage, quality metrics, and access policies that enable reliable data consumption while maintaining the native format storage capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments data lake access into multiple layers: the raw data storage layer (maintaining native formats and storage flexibility) and the governed access layer (providing metadata-driven quality control and security). This segmentation allows both storage flexibility and data quality to coexist by separating concerns between storage and access.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If data lakes lack governed metadata and semantic consistency, then data ingestion ease is improved, but data retrieval and analysis difficulty increases

Engineering Contradiction:
Improvedata ingestion easeVSAvoiddata retrieval difficulty
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies preliminary action by automatically generating and maintaining metadata during the data ingestion process. Instead of requiring governance after data is stored, the system pre-establishes data lineage, quality metrics, and semantic information at ingestion time, making future retrieval and analysis straightforward.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where metadata quality metrics and data lineage information are continuously updated and fed back to improve data discovery and retrieval. This feedback loop ensures that as data is ingested and processed, the metadata becomes increasingly accurate and useful for finding and analyzing data.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If data lakes accept any data without restrictions, then data quantity and variety are improved, but security and privacy risks worsen

Engineering Contradiction:
Improvedata quantityVSAvoidsecurity risks
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent introduces metadata and governance policies as intermediaries between the data lake and external systems. These intermediaries enable the system to accept large quantities of diverse data while filtering and controlling access based on security requirements, data classification, and privacy policies embedded in the metadata layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If data lakes lack unified data schema, then data source adaptability is improved, but data correlation complexity increases

Engineering Contradiction:
Improvedata source adaptabilityVSAvoiddata correlation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses metadata as an intermediary layer that provides unified data description without imposing a rigid schema on the underlying data. The metadata layer correlates data from different sources by providing semantic information, data lineage, and relationship context, enabling data correlation while maintaining source adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11321337B2Crowdsourcing data into a data lake
Publication Date: 2022.05.03 CISCO TECHNOLOGY INC
  • US11321337B2 patent drawing
  • US11321337B2 patent drawing
  • US11321337B2 patent drawing

AI summary

A Services Delivery Platform (SDP) architecture is provided that is configured to onboard new data sets into an SDP data lake. The SDP enables the crowdsourcing of data on-boarding by configuring this process into an interactive, intuitive, step-by-step guided workflow while governing/controlling key functions like verification, acceptance and execution.