Data Catalog External Identifier Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data catalog systems face challenges in accurately identifying and retrieving data entities from diverse and heterogeneous data sources due to reliance on internal identifiers, which can lead to duplication and lack of metadata stitching, especially when data is ingested from multiple sources.

Innovation Solution

A data catalog system generates unique external identifiers for data assets and objects based on immutable configuration parameters and attributes, allowing for accurate identification and retrieval without relying on internal identifiers, and uses these identifiers to manage data state changes during metadata harvesting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data catalog systems use internal identifiers to manage data entities from diverse sources, then data can be ingested from multiple sources, but data duplication occurs and metadata stitching is insufficient

Engineering Contradiction:
Improveability to ingest data from diverse sourcesVSAvoiddata identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary identifier system that acts as a mediator between diverse data sources and the data catalog. This intermediary layer translates various internal identifiers from different sources into a unified external identifier format, enabling accurate matching and preventing duplication while maintaining the ability to ingest from multiple sources

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a universal identifier system that serves multiple functions: it uniquely identifies data entities across different sources, enables metadata stitching, prevents duplication, and maintains compatibility with various data source formats. This multi-functional approach resolves the contradiction between versatility and identification accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of manufacture

If data catalog systems rely on internal identifiers from data sources, then data ingestion is simplified, but metadata stitching and data entity reconciliation become inadequate

Engineering Contradiction:
Improvedata ingestion simplicityVSAvoidmetadata stitching capability
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent segments the identification process into two independent parts: internal identifiers that simplify data ingestion from various sources, and external identifiers that enable accurate metadata stitching and reconciliation. This segmentation allows each identifier type to optimize for its specific function without compromising the other

Inventive Principle:
Principle #1Segmentation

3Reliability

If data catalog systems use centralized external identifiers, then data entity identification accuracy improves and duplication is prevented, but system complexity increases

Engineering Contradiction:
Improvedata entity identification accuracyVSAvoididentifier management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a self-service identifier generation system that automatically creates external identifiers based on data source characteristics and metadata. This automation reduces manual intervention and operational complexity while maintaining high identification accuracy and preventing duplication

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240232175A1Generating external identifiers for data entities using a data catalog system
Publication Date: 2024.07.11 ORACLE INT CORP
  • US20240232175A1 patent drawing
  • US20240232175A1 patent drawing
  • US20240232175A1 patent drawing

AI summary

A data catalog system is disclosed that provides capabilities for uniquely identifying and retrieving data entities stored in diverse data sources managed by an organization. The data catalog system includes capabilities for generating a unique external identifier for a data entity (e.g., a data asset or a data object) by identifying a set of immutable configuration parameters associated with the data asset and identifying a set of data object attributes that uniquely identify data objects within the data asset. The generated unique external identifiers are stored as part of the metadata harvested by the data catalog system. The external identifiers are used to enforce a single representation of the data assets and the data objects in the data catalog system. The external object identifiers are used to perform data lookups and reconcile states of data entities during the metadata harvesting process.