Unified Metadata Service View for Multi-Engine Data Lakes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In big data scenarios, integrating different big data engines for unified data access in a data lake leads to significant ETL work, increasing costs and time due to the need for data processing and storage format conversion.
Innovation Solution
A data management method and apparatus that utilize a metadata storage module to store metadata of a data lake using different storage modes, allowing for a unified metadata service view across various engines. This enables engines to acquire target metadata and access corresponding data in the data lake efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple big data engines integrate on the same data lake, then data sharing capability is improved, but ETL work increases leading to higher cost and time consumption
Solution Approach 1:
The patent introduces a metadata storage module as an intermediary layer between data lakes and multiple big data engines. This module stores metadata in different storage modes that correspond to different engine requirements, enabling engines to directly query needed information without full ETL processing. The metadata storage module acts as a mediator that translates diverse engine access requests into appropriate data retrieval operations, significantly reducing processing time while maintaining multi-engine compatibility.
Solution Approach 2:
The metadata storage module implements universality by supporting multiple storage modes simultaneously (first storage mode for certain engines, second storage mode for other engines). This allows a single metadata system to serve diverse big data engine requirements without requiring separate metadata systems for each engine type, thereby enabling universal data access across different engines while avoiding redundant ETL operations.
2Adaptability or versatility
If multiple big data engines integrate on the same data lake, then data sharing capability is improved, but processing cost increases
Solution Approach 1:
The metadata storage module serves as an intermediary that optimizes resource utilization across multiple engines. By pre-organizing metadata in different storage modes, the system enables engines to retrieve only the necessary information with minimal processing, reducing overall computational resource consumption and associated costs while maintaining broad data sharing capability.
Solution Approach 2:
The system performs preliminary action by pre-storing metadata in multiple storage modes before actual query operations. This advance preparation eliminates the need for real-time ETL transformations when engines access data, significantly reducing processing costs during actual data sharing operations while maintaining high adaptability across different engine types.
3Adaptability or versatility
If unified metadata service view is implemented, then engine compatibility is improved, but metadata storage complexity increases
Solution Approach 1:
The metadata storage module applies segmentation by dividing metadata into different storage modes (first storage mode, second storage mode, etc.), where each mode is optimized for specific engine requirements. This segmentation allows the system to maintain a unified service view while organizing metadata in a manageable, engine-specific manner, reducing the perceived complexity for each engine while preserving overall system compatibility.
Data Source
AI summary
The present disclosure relates to a data management method and apparatus, a storage medium, and an electronic device. The method comprises: obtaining a data access request sent by an engine side, the data access request being used for requesting to perform an access operation on first target data in a data lake; determining, according to the data access request, target metadata corresponding to the first target data from a metadata storage module, the metadata storage module storing metadata of the data lake in different storage modes, respectively, and the metadata stored in the different storage modes having at least one type of the same information; and sending the first target data corresponding to the target metadata in the data lake to the engine side. By constructing a data lake metadata unified service view that meets various engine requirements, metadata intercommunication between different engines is realized.


