Automated Schema Recommendation for Multi-Vendor Data Ingestion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large organizations face difficulties in identifying data schemas for ingested data from multiple vendors with different data formats and protocols, requiring domain expertise and manual effort.
Innovation Solution
A system and method that uses a schema recommendation program to access ingested objects, extract metadata, identify potential schemas, receive user selections, and publish schemas to a catalog store, applying natural language processing and entitlement-based restrictions, allowing for automated schema building and publishing based on data patterns and formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual schema identification with domain expertise is used, then schema accuracy is improved, but time consumption and operational complexity increase
Solution Approach 1:
The system enables self-service schema identification by automatically analyzing ingested objects and generating schema recommendations without requiring manual intervention. The data crawler extracts metadata, identifies potential schemas, and presents recommendations for automated schema selection, allowing the system to serve itself rather than requiring domain expertise for each schema identification task.
Solution Approach 2:
The patent replaces the mechanical process of manual schema identification with automated computational processes. The data crawler uses programmatic methods to extract metadata from objects, the schema recommender algorithmically identifies potential schemas, and the system automatically publishes schemas to the catalog store, substituting human expertise with automated systems.
2Productivity
If automated schema extraction is implemented, then productivity is improved, but system complexity increases
Solution Approach 1:
The system segments the schema identification process into distinct functional components: a data crawler for metadata extraction, a schema recommender for schema identification, and a publisher for schema deployment. This segmentation allows each component to be developed and maintained independently, managing overall system complexity while improving productivity through automation.
Solution Approach 2:
The data crawler serves multiple functions by extracting metadata from various types of ingested objects across different vendors and formats. The schema recommender handles multiple schema identification tasks universally, and the publisher consistently deploys schemas to the catalog store, providing multi-functionality that improves productivity across diverse data scenarios.
3Adaptability or versatility
If multiple vendor products with different data formats are supported, then adaptability is improved, but difficulty in identifying schemas increases
Solution Approach 1:
The system manages parameter changes by dynamically adapting to different data formats and structures from multiple vendors. The data crawler extracts metadata that captures the essential parameters of each data format, and the schema recommender identifies appropriate schemas based on these parameters, allowing the system to handle format variations without increasing identification difficulty.
Solution Approach 2:
The patent introduces an intermediary layer in the form of standardized schema representations that mediate between diverse vendor data formats and the unified data catalog. This intermediary schema language translates various data formats into a common representation, making schema identification easier despite the diversity of input formats from multiple vendors.
Data Source
AI summary
Systems and methods for building and publishing schemas based on data patterns and data formats are disclosed. According to one embodiment, a method for building and publishing schemas may include: (1) accessing, by a schema recommendation program executed by a computer processor, a plurality of ingested objects in an object store; (2) extracting, by a data crawler, metadata from each of the plurality of ingested objects, wherein the metadata is related to a schema for the object and comprises an object name, a field name, and a field type; (3) identifying, by the data crawler, a plurality of potential schemas for the plurality of ingested objects based on the metadata; (4) receiving, by the schema recommendation program, a selection of one of the plurality of potential schemas; and (5) publishing, by the schema recommendation program, the selected potential schema to a catalog store.

