Shared Metadata Access for Real-Time Cross-Cluster Data Warehousing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data warehouse systems face high storage costs and poor user experience due to separate metadata storage in a metadata cluster, which hinders real-time data access and large-scale data sharing.
Innovation Solution
A metadata processing method where a data production cluster generates and shares metadata with a data consumption cluster, allowing real-time metadata access and reducing storage overheads, enabling stable data analysis and better user experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If metadata is separately stored in a metadata cluster, then metadata storage is isolated, but storage costs increase and real-time access is hindered
Solution Approach 1:
The patent merges metadata storage with data storage by allowing data consumption clusters to directly access metadata from the shared storage system where data is stored, eliminating the need for a separate metadata cluster. This integration enables real-time metadata access while reducing storage costs.
Solution Approach 2:
The shared storage system is designed to serve multiple functions: it stores both data and metadata, and can be accessed by both data production clusters and data consumption clusters. This multi-functionality eliminates the need for dedicated metadata storage infrastructure.
2Reliability
If metadata is separately stored in a metadata cluster, then storage isolation is achieved, but storage costs increase
Solution Approach 1:
The patent combines metadata storage with data storage in the same shared storage system, eliminating duplicate storage infrastructure. By allowing the shared storage to hold both data and its associated metadata, the system reduces overall storage costs while maintaining data integrity.
Solution Approach 2:
The system creates logical copies of metadata access paths rather than physical separate storage. Data consumption clusters can access metadata directly from shared storage without requiring separate physical metadata clusters, reducing storage overhead while maintaining access capabilities.
3Ease of manufacture
If metadata is stored in a separate metadata cluster, then metadata management is simplified, but data sharing between clusters is hindered
Solution Approach 1:
The shared storage system serves as a universal access point for both data production clusters and data consumption clusters. It can store and provide access to both data and metadata for multiple clusters simultaneously, enabling large-scale multi-cluster data sharing without requiring separate metadata management infrastructure.
Solution Approach 2:
The shared storage system acts as an intermediary between data production clusters and data consumption clusters. It mediates access to both data and metadata, allowing clusters to share information efficiently without direct peer-to-peer metadata access requirements.
4Reliability
If metadata is accessed from a separate metadata cluster, then metadata retrieval is isolated, but user experience deteriorates
Solution Approach 1:
The patent combines metadata access with data access operations. Data consumption clusters can retrieve metadata and data from the same shared storage system in a single operation, eliminating the need for separate metadata retrieval steps and improving overall system responsiveness and user experience.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
A metadata processing method and system, and a computing device are provided. The method includes: A data production cluster generates metadata of shared data. The data production cluster stores the shared data into a shared storage. The data production cluster receives a metadata operation instruction, and determines target metadata based on the metadata operation instruction, where the target metadata is metadata of target shared data. The data production cluster sends the target metadata to a data consumption cluster. The data consumption cluster reads the target shared data from the shared storage based on the target metadata. In the method, the data consumption cluster may obtain the metadata of the shared data in real time, thereby providing better user experience, and storage costs of the metadata can also be reduced.