Declarative Data Enrichment Service for HDFS and URL Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current big data systems face challenges in data onboarding due to noisy data, requiring substantial manual processes for cleaning and curation, which cannot scale with increasing data volumes, and struggle with varying protocols from different data sources, making manual connection and credential management burdensome.

Innovation Solution

A data enrichment service that automates data preparation and processing through a declarative approach, enabling users to specify data sources, manage connection metadata, and perform entity resolution and correlation, using a visual recommendation engine for data transformations and repairs, and supports various data sources including URL-based and HDFS resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual processes are used to clean and curate data, then data quality can be improved, but the process cannot scale with increasing data volumes

Engineering Contradiction:
Improvedata qualityVSAvoidscaling capability
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system enables self-service data enrichment by automatically discovering and applying transformations and repairs to datasets. The enrichment service autonomously profiles data, identifies issues, and applies corrections without requiring manual intervention for each dataset, thereby maintaining data quality while scaling to large volumes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the state of data processing from manual to automated by transforming raw datasets through applied transformations and repairs. Parameters such as data format, completeness, and accuracy are modified automatically through the enrichment pipeline, enabling scalable data quality improvement.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If users manually manage connection information and credentials for each data source, then access control can be maintained, but the burden on users increases significantly

Engineering Contradiction:
Improveaccess controlVSAvoiduser burden
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system introduces an intermediary enrichment service that manages connections to external data sources. This service acts as a mediator between users and data sources, handling authentication and connection management internally while presenting a simplified interface to users, thus maintaining security while reducing user burden.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The enrichment service provides universal access to multiple types of data sources (URL-based, HDFS, cloud storage) through a single unified interface. Users interact with one service that handles various connection protocols and credential management, eliminating the need to manually configure each data source connection.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If diverse data sources with different protocols are accessed, then data variety can be increased, but protocol differences introduce challenges for import

Engineering Contradiction:
Improvedata varietyVSAvoidimport complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The enrichment service implements universal functionality to access multiple data source types (URL-based resources, HDFS, cloud storage services) through a single interface. It automatically detects and adapts to different protocols, presenting a unified import mechanism that handles diverse data sources without increasing user-facing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The service acts as an intermediary layer between users and diverse data sources with different protocols. It translates various data source protocols into a unified internal format, managing protocol differences internally while providing simplified access to users, thus increasing data variety without proportionally increasing import complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11693549B2Declarative external data source importation, exportation, and metadata reflection utilizing HTTP and HDFS protocols
Publication Date: 2023.07.04 ORACLE INT CORP
  • US11693549B2 patent drawing
  • US11693549B2 patent drawing
  • US11693549B2 patent drawing

AI summary

Techniques are disclosure for a data enrichment system that enables declarative external data source importation and exportation. A user can specify via a user interface input for identifying different data sources from which to obtain input data. The data enrichment system is configured to import and export various types of sources storing resources such as URL-based resources and HDFS-based resources for high-speed bi-directional metadata and data interchange. Connection metadata (e.g., credentials, access paths, etc.) can be managed by the data enrichment system in a declarative format for managing and visualizing the connection metadata.