Data Catalog Service for Unknown Schema Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As data storage and processing systems evolve, accessing data objects with unknown data schemas becomes challenging, as existing technologies lack efficient methods to recognize and adapt to diverse data schemas, leading to blocked access and increased complexity and costs.
Innovation Solution
Implementing a data catalog service that recognizes unknown data objects by detecting their schemas, storing this information in a metadata store, and using machine learning techniques to classify and generate representations, allowing diverse data consumers to access and manipulate these objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data storage systems support diverse file types and data schemas, then data versatility and processing capabilities improve, but system complexity and maintenance costs increase
Solution Approach 1:
The patent introduces a data schema recognition service as an intermediary between data storage systems and data consumers. This service automatically detects, recognizes, and translates unknown data schemas into accessible formats, enabling diverse file types to be accessed without increasing the complexity of the core storage system. The intermediary handles the adaptability requirements while the storage system maintains its simplicity.
2Productivity
If data storage systems maintain specialized configurations for different data types, then data processing capabilities improve, but maintenance costs and operational complexity increase
Solution Approach 1:
The patent implements a self-service mechanism where the data schema recognition service automatically identifies and adapts to unknown data schemas without requiring manual configuration or specialized maintenance for each data type. The system autonomously processes diverse file formats, extracting schemas and making them accessible to consumers, thereby maintaining high processing capability while minimizing maintenance burden.
3Speed
If data consumers directly access unknown data objects without schema recognition, then access speed improves, but access reliability decreases due to blocked access
Solution Approach 1:
The patent implements preliminary action by pre-recognizing and storing data schemas in a metadata store before data consumers attempt to access the actual data objects. This upfront schema identification and translation work ensures that when consumers access data, they can do so reliably and efficiently without encountering blocked access, while the preliminary processing minimizes the impact on access speed.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Recognizing unknown data objects may be implemented for data objects stored in a data store. Data objects that are identified as unknown may be accessed to retrieve a portion of the data object. Different representations of the data object may be generated for recognizing different data schemas. An analysis of the representations may be performed to identify a data schema for the unknown data object. The data schema may be stored in a metadata store for the unknown data object.