Metadata Elements for Data Lake Storage Organization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data lakes face inefficiencies in retrieving and processing unorganized raw data due to its native format storage, making it difficult to perform operations effectively.

Innovation Solution

Creating metadata elements with partitioning approaches that validate and categorize raw data into specific storage structures within a data lake, allowing for efficient streaming and analysis, and enabling actions such as billing, ordering, or controlling manufacturing processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If raw data is stored in native format in a data lake, then data storage capacity is improved, but data retrieval efficiency deteriorates

Engineering Contradiction:
Improvedata storage capacityVSAvoiddata retrieval efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent applies preliminary action by creating metadata elements and validation rules before data is stored in the data lake. This pre-organization of data through metadata classification enables efficient retrieval without compromising storage capacity, as the data remains in native format but is pre-categorized with descriptive metadata that facilitates quick access.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces metadata elements as an intermediary between raw data and retrieval operations. These metadata elements act as mediators that describe and categorize data without altering the native format of the stored data, enabling efficient queries while maintaining storage flexibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If raw data is stored without organization in native format, then data storage flexibility is improved, but data processing speed deteriorates

Engineering Contradiction:
Improvedata storage flexibilityVSAvoiddata processing speed
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The system performs preliminary organization by creating metadata elements that describe data characteristics before storage. This pre-classification enables faster processing during retrieval while maintaining the flexibility to store data in its native format without rigid structural constraints.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameter of data organization from structural organization (altering data format) to metadata-based organization (adding descriptive parameters). This allows data to remain flexible in native format while gaining processing speed through indexed metadata parameters that enable quick filtering and retrieval.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If metadata elements are created and validated against rules, then data organization is improved, but system complexity increases

Engineering Contradiction:
Improvedata organizationVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the data management system into distinct components: metadata elements, validation rules, and data storage. This segmentation allows each component to be independently managed and validated, making the overall system more manageable despite the added complexity of metadata processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes from direct data manipulation to parameter-based control through metadata. By validating metadata parameters against rules rather than directly validating data structure, the system achieves better organization while managing complexity through abstracted parameter validation rather than complex structural enforcement.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10846307B1System and method for managing interactions with a data lake
Publication Date: 2020.11.24 CSG SYSTEMS INC
  • US10846307B1 patent drawing
  • US10846307B1 patent drawing
  • US10846307B1 patent drawing

AI summary

Metadata elements are created and validated. Once the metadata element is validated it is applied to raw incoming data. If a match is obtained, then the raw data is sent to a designated storage structure. When there is no match, then the raw data is sent to a data structure designated for unorganized raw data.