Multi-layer Graph Labeling for Automated Data Catalog Maintenance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As data lakes grow in scale and complexity, manual processes for populating and maintaining data catalogs become inefficient, leading to decreased reliability and speed in data retrieval operations.

Innovation Solution

The use of natural language processing (NLP) methods, specifically neural network layers and ontology graphs, to predict labels for datasets, generate summarizations, and update data catalogs automatically, thereby improving data retrieval efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual processes are used to populate and maintain data catalogs, then accuracy and reliability can be maintained, but productivity and speed of data retrieval operations decrease

Engineering Contradiction:
Improvereliability of data retrieval operationsVSAvoidspeed of data retrieval operations
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables self-service by automatically populating and maintaining data catalogs using neural network models that process dataset metadata, generate summaries, and assign labels without human intervention. The neural network model independently performs catalog maintenance operations, allowing the system to serve itself rather than requiring manual administrative processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes with an automated neural network-based system. Instead of human operators manually entering and updating catalog records, the system uses neural networks to automatically process data, generate summaries, and maintain catalog entries, substituting human cognitive and manual operations with computational mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual processes are used for data catalog maintenance, then quality control can be maintained, but the time required for data retrieval operations increases

Engineering Contradiction:
Improvequality of data catalog entriesVSAvoidtime required for data retrieval operations
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-processing and automatically generating high-quality catalog entries before data retrieval operations are needed. The neural network model continuously processes dataset metadata, generates summaries, and maintains catalog records in advance, so that when data retrieval is needed, the catalog is already optimized and ready for rapid queries.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuity of useful action through continuous automated processing where the neural network model operates continuously to maintain the data catalog. Rather than periodic manual updates, the system continuously processes new data, updates summaries, and maintains catalog entries in real-time, ensuring the catalog is always current and optimized for retrieval operations.

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If automated NLP methods are used to populate data catalogs, then productivity and speed improve, but device complexity increases

Engineering Contradiction:
Improvespeed of data retrieval operationsVSAvoidcomplexity of the prediction model system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies universality by designing a multi-functional neural network model that performs multiple catalog maintenance tasks within a single system. The same neural network architecture handles data processing, summary generation, label assignment, and quality assessment, eliminating the need for separate specialized systems for each function and thereby managing complexity through consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system segments the complex automated catalog maintenance process into distinct functional modules within the neural network architecture. Different layers or components of the neural network handle specific tasks such as data processing, summary generation, and quality assessment, allowing the complex system to be managed through modular segmentation of functions.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12223264B2Multi-layer graph-based categorization
Publication Date: 2025.02.11 CAPITAL ONE SERVICES LLC
  • US12223264B2 patent drawing
  • US12223264B2 patent drawing
  • US12223264B2 patent drawing

AI summary

A method may include a obtaining a first data model instance comprising an identifier string and. a set of attributes associated with a set of attribute name strings. The method may include obtaining an ontology graph that includes a first label, a second label, and an association between them. The method may include using a prediction model to select the first label based on the first data model instance and determining the second label based on the relationship. The method may include determining a selected set of labels that includes the first label and the second label to associate with the first data model instance. The method may include associating the selected set of labels with the first data model instance in a dataset that includes a plurality of records, where each record is associated with a different data model instance.