Unified Metadata Storage for Data Lake and Warehouse Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data query solutions in data Lake environments suffer from poor user experience due to slow response times and high data volume requirements, leading to inefficient metadata management and increased operational costs.

Innovation Solution

A data processing method and apparatus that integrates data Lake and data Warehouse, pre-storing multiple protocol files for different data processing engines. These protocol files parse metadata acquisition requests and retrieve metadata from a unified metadata storage space, ensuring efficient metadata management and reduced storage capacity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple protocol files are pre-stored for different data processing engines, then the adaptability to different engines is improved, but the device complexity increases

Engineering Contradiction:
Improvecompatibility with different data processing enginesVSAvoidnumber of protocol files to manage
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal metadata storage space that serves multiple data processing engines (Spark, Flink, Presto, etc.) through a single interface. The unified metadata management system provides multi-functional support by accepting requests from different engines and routing them to the appropriate protocol handler, eliminating the need for separate metadata storage systems for each engine.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces a protocol file as an intermediary layer between the metadata storage space and different data processing engines. Each protocol file acts as a translator that converts engine-specific metadata requests into a unified format that the metadata storage space can process, thereby simplifying the overall system architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If one set of metadata is stored in a unified metadata storage space, then the data storage capacity is reduced, but the difficulty of managing metadata accessing rights increases

Engineering Contradiction:
Improvemetadata storage capacityVSAvoidmetadata access rights management
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the unified metadata storage space into logically isolated regions or views for different data processing engines. Each engine type can access only its designated metadata through protocol-specific interfaces, providing fine-grained access control while maintaining physical consolidation of metadata storage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary authentication and authorization mechanisms in the protocol files before metadata access requests are processed. Access rights are validated in advance based on engine type and user credentials, preventing unauthorized access while maintaining a unified storage backend.

Inventive Principle:
Principle #10Preliminary action

3Loss of energy

If data Lake and data Warehouse are integrated, then the operational costs are reduced, but the response time for metadata acquisition may be affected

Engineering Contradiction:
Improveoperational costVSAvoidmetadata acquisition response time
Core Design Contradiction:
Loss of energyVSSpeed

Solution Approach 1:

The patent pre-loads and caches frequently accessed metadata in the unified metadata storage space, and maintains protocol-specific metadata buffers in memory. This preliminary preparation ensures that when data processing engines request metadata, the information is already available or can be quickly retrieved, maintaining fast response times despite the integrated architecture.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an optimized protocol layer that acts as a mediator between data processing engines and the unified metadata storage space. This protocol layer includes caching mechanisms and query optimization capabilities that accelerate metadata retrieval while maintaining the benefits of unified storage.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250131005A1Data processing method and apparatus based on data lake and data warehouse integration, and electronic device
Publication Date: 2025.04.24 BEIJING VOLCANO ENGINE TECH CO LTD
  • US20250131005A1 patent drawing
  • US20250131005A1 patent drawing
  • US20250131005A1 patent drawing

AI summary

Embodiments of the present disclosure provide a data processing method, apparatus based on data Lake and data Warehouse integration, and an electronic device. The method includes: pre-storing at least two protocol files; receiving a first metadata acquisition request; determining a target protocol file for processing the first metadata acquisition request from the at least two protocol files according to an engine type of the data processing engine which sends the first metadata acquisition request; and parsing, based on the target protocol file, the first metadata acquisition request and acquiring metadata corresponding to the first metadata acquisition request from the metadata storage space. Therefore, for at least two sets of external protocols, one set of metadata is stored, accordingly, the data storage capacity can be reduced; with the setting of one set of metadata, the metadata accessing rights management can be unified, and the security of data can be improved.