Automated Data Domain Identification via Dependency Graph Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modernizing legacy IT systems is hindered by the tedious and manual process of identifying data domains, which is crucial for determining related datasets, associated code, and other incremental artifacts.

Innovation Solution

The technique involves extracting entities from data, code, and user interface artifacts, generating dependency graphs, and performing lexical and semantic analyses to automatically identify the data domain of one or more datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual process is used to identify data domains, then identification accuracy can be maintained through human judgment, but the process is tedious and time-consuming

Engineering Contradiction:
Improvedata domain identification accuracyVSAvoidtime required for data domain identification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of data domain identification with an automated computer-based system that performs lexical analysis and semantic analysis on dependency graphs, thereby eliminating the time-consuming manual effort while maintaining identification accuracy through systematic algorithmic processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables automated self-identification of data domains by extracting entities from data artifacts, code artifacts, and user interface artifacts, generating dependency graphs, and performing analyses without requiring manual human intervention in the identification process

Inventive Principle:
Principle #25Self-service

2Extent of automation

If automated methods are used to identify data domains, then manual effort is reduced, but the system complexity increases

Engineering Contradiction:
Improveautomation level of data domain identificationVSAvoidsystem complexity for automated identification
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent segments the automated identification process into distinct modular components: entity extraction from data artifacts, code artifacts, and user interface artifacts; dependency graph generation; lexical analysis; and semantic analysis. This segmentation manages system complexity by organizing the automated process into manageable, independent modules that can be developed and maintained separately

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If comprehensive entity extraction from multiple artifact types is performed, then data domain identification accuracy is improved, but the processing complexity and time increase

Engineering Contradiction:
Improvedata domain identification accuracyVSAvoidprocessing complexity of entity extraction
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs a universal entity extraction approach that handles multiple types of artifacts (data artifacts, code artifacts, user interface artifacts) through a unified processing framework. This multi-functional extraction mechanism improves identification accuracy by comprehensively analyzing all relevant artifact types while managing processing complexity through a standardized universal approach

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12333253B2Automatic data domain identification
Publication Date: 2025.06.17 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12333253B2 patent drawing
  • US12333253B2 patent drawing
  • US12333253B2 patent drawing

AI summary

An apparatus is disclosed which includes at least one processing device comprising a processor coupled to a memory. The at least one processing device, when executing program code, is configured to: extract one or more entities identified in a plurality of data artifacts based at least in part on one or more datasets, extract one or more entities identified in a plurality of code artifacts based at least in part on the one or more datasets, extract one or more entities identified in a plurality of user interface artifacts based at least in part on the one or more datasets, generate a set of dependency graphs each based at least in part on one or more relationships among the respective extracted one or more entities, and perform one or more of a lexical analysis and a semantic analysis on the set of dependency graphs to identify a data domain of the one or more datasets.