Unified Metadata Service View for Multi-Engine Data Lakes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In big data scenarios, integrating different big data engines for unified data access in a data lake leads to significant ETL work, increasing costs and time due to the need for data processing and storage format conversion.

Innovation Solution

A data management method and apparatus that utilize a metadata storage module to store metadata of a data lake using different storage modes, allowing for a unified metadata service view across various engines. This enables engines to acquire target metadata and access corresponding data in the data lake efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple big data engines integrate on the same data lake, then data sharing capability is improved, but ETL work increases leading to higher cost and time consumption

Engineering Contradiction:
Improvedata sharing capabilityVSAvoidETL processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent introduces a metadata storage module as an intermediary layer between data lakes and multiple big data engines. This module stores metadata in different storage modes that correspond to different engine requirements, enabling engines to directly query needed information without full ETL processing. The metadata storage module acts as a mediator that translates diverse engine access requests into appropriate data retrieval operations, significantly reducing processing time while maintaining multi-engine compatibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The metadata storage module implements universality by supporting multiple storage modes simultaneously (first storage mode for certain engines, second storage mode for other engines). This allows a single metadata system to serve diverse big data engine requirements without requiring separate metadata systems for each engine type, thereby enabling universal data access across different engines while avoiding redundant ETL operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple big data engines integrate on the same data lake, then data sharing capability is improved, but processing cost increases

Engineering Contradiction:
Improvedata sharing capabilityVSAvoidprocessing cost
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The metadata storage module serves as an intermediary that optimizes resource utilization across multiple engines. By pre-organizing metadata in different storage modes, the system enables engines to retrieve only the necessary information with minimal processing, reducing overall computational resource consumption and associated costs while maintaining broad data sharing capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary action by pre-storing metadata in multiple storage modes before actual query operations. This advance preparation eliminates the need for real-time ETL transformations when engines access data, significantly reducing processing costs during actual data sharing operations while maintaining high adaptability across different engine types.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If unified metadata service view is implemented, then engine compatibility is improved, but metadata storage complexity increases

Engineering Contradiction:
Improveengine compatibilityVSAvoidmetadata storage complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The metadata storage module applies segmentation by dividing metadata into different storage modes (first storage mode, second storage mode, etc.), where each mode is optimized for specific engine requirements. This segmentation allows the system to maintain a unified service view while organizing metadata in a manageable, engine-specific manner, reducing the perceived complexity for each engine while preserving overall system compatibility.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12326906B2Data management method and apparatus, storage medium, and electronic device
Publication Date: 2025.06.10 BEIJING VOLCANO ENGINE TECH CO LTD
  • US12326906B2 patent drawing
  • US12326906B2 patent drawing
  • US12326906B2 patent drawing

AI summary

The present disclosure relates to a data management method and apparatus, a storage medium, and an electronic device. The method comprises: obtaining a data access request sent by an engine side, the data access request being used for requesting to perform an access operation on first target data in a data lake; determining, according to the data access request, target metadata corresponding to the first target data from a metadata storage module, the metadata storage module storing metadata of the data lake in different storage modes, respectively, and the metadata stored in the different storage modes having at least one type of the same information; and sending the first target data corresponding to the target metadata in the data lake to the engine side. By constructing a data lake metadata unified service view that meets various engine requirements, metadata intercommunication between different engines is realized.