Unified Metadata Storage for Data Lake and Warehouse Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data query solutions in data Lake environments suffer from poor user experience due to slow response times and high data volume requirements, leading to inefficient metadata management and increased operational costs.
Innovation Solution
A data processing method and apparatus that integrates data Lake and data Warehouse, pre-storing multiple protocol files for different data processing engines. These protocol files parse metadata acquisition requests and retrieve metadata from a unified metadata storage space, ensuring efficient metadata management and reduced storage capacity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple protocol files are pre-stored for different data processing engines, then the adaptability to different engines is improved, but the device complexity increases
Solution Approach 1:
The patent implements a universal metadata storage space that serves multiple data processing engines (Spark, Flink, Presto, etc.) through a single interface. The unified metadata management system provides multi-functional support by accepting requests from different engines and routing them to the appropriate protocol handler, eliminating the need for separate metadata storage systems for each engine.
Solution Approach 2:
The patent introduces a protocol file as an intermediary layer between the metadata storage space and different data processing engines. Each protocol file acts as a translator that converts engine-specific metadata requests into a unified format that the metadata storage space can process, thereby simplifying the overall system architecture.
2Quantity of substance
If one set of metadata is stored in a unified metadata storage space, then the data storage capacity is reduced, but the difficulty of managing metadata accessing rights increases
Solution Approach 1:
The patent segments the unified metadata storage space into logically isolated regions or views for different data processing engines. Each engine type can access only its designated metadata through protocol-specific interfaces, providing fine-grained access control while maintaining physical consolidation of metadata storage.
Solution Approach 2:
The patent implements preliminary authentication and authorization mechanisms in the protocol files before metadata access requests are processed. Access rights are validated in advance based on engine type and user credentials, preventing unauthorized access while maintaining a unified storage backend.
3Loss of energy
If data Lake and data Warehouse are integrated, then the operational costs are reduced, but the response time for metadata acquisition may be affected
Solution Approach 1:
The patent pre-loads and caches frequently accessed metadata in the unified metadata storage space, and maintains protocol-specific metadata buffers in memory. This preliminary preparation ensures that when data processing engines request metadata, the information is already available or can be quickly retrieved, maintaining fast response times despite the integrated architecture.
Solution Approach 2:
The patent introduces an optimized protocol layer that acts as a mediator between data processing engines and the unified metadata storage space. This protocol layer includes caching mechanisms and query optimization capabilities that accelerate metadata retrieval while maintaining the benefits of unified storage.
Data Source
AI summary
Embodiments of the present disclosure provide a data processing method, apparatus based on data Lake and data Warehouse integration, and an electronic device. The method includes: pre-storing at least two protocol files; receiving a first metadata acquisition request; determining a target protocol file for processing the first metadata acquisition request from the at least two protocol files according to an engine type of the data processing engine which sends the first metadata acquisition request; and parsing, based on the target protocol file, the first metadata acquisition request and acquiring metadata corresponding to the first metadata acquisition request from the metadata storage space. Therefore, for at least two sets of external protocols, one set of metadata is stored, accordingly, the data storage capacity can be reduced; with the setting of one set of metadata, the metadata accessing rights management can be unified, and the security of data can be improved.


