Data Set Fingerprinting for Automatic Domain Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data integration tools require manual effort and specialized algorithms to classify data domains, making it time-consuming to understand and document data sources, especially for non-standard domains like postal codes or enterprise-specific data, which are not well-supported by existing tools.

Innovation Solution

The method computes 'fingerprints' for data sets using metric algorithms that capture various data characteristics, allowing for automatic classification and comparison of data sets without requiring domain-specific algorithms, enabling semi-automatic data classification and detection of invalid values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If specialized algorithms are used to classify data domains, then classification accuracy is improved, but device complexity and development time increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidalgorithm complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by creating a single metric algorithm that can classify multiple different data domains (US addresses, Belgian postal codes, product references, enterprise codes) without requiring separate specialized algorithms for each domain. The algorithm achieves this by computing characteristics like data length, character composition, and format patterns that are applicable across diverse data types, thereby reducing algorithmic complexity while maintaining classification accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter changes by computing multiple characteristics (length, character composition, format patterns) of data sets and comparing these parameters to classify data domains. Instead of using complex domain-specific logic, the algorithm changes the approach to using measurable parameters that can be computed and compared systematically across different data types, simplifying the classification process.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual classification is performed for each column, then classification accuracy is maintained, but productivity decreases

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata classification speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies self-service by enabling the system to automatically classify data domains without requiring manual expert intervention for each column. The metric algorithm autonomously computes characteristics of data sets, compares them against known domain patterns, and performs classification automatically. This eliminates the need for users to manually evaluate each column while maintaining classification quality, thereby significantly improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses copying by creating characteristic profiles (fingerprints) of data sets that capture essential domain-specific patterns. These characteristic copies can be stored and reused for classification, allowing the system to quickly compare new data against established patterns without re-analyzing the entire classification logic, thus speeding up the classification process while maintaining accuracy.

Inventive Principle:
Principle #26Copying

3Measurement precision

If domain-specific algorithms are developed for each data type, then measurement precision is improved, but loss of time increases

Engineering Contradiction:
Improvedomain detection accuracyVSAvoidalgorithm development time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent eliminates the need to develop separate algorithms for each data type by creating a universal metric algorithm that handles multiple domains (US addresses, Belgian postal codes, product references, enterprise codes) through a single unified approach. This universal algorithm computes characteristics and applies pattern matching that works across diverse data types, reducing development time while maintaining domain detection accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies preliminary action by pre-computing and storing characteristic profiles of known data domains during system initialization or setup. These pre-computed characteristics serve as reference patterns that the metric algorithm can quickly compare against new data, eliminating the need to develop and test new algorithms for each domain while maintaining high detection accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8666998B2Handling data sets
Publication Date: 2014.03.04 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8666998B2 patent drawing
  • US8666998B2 patent drawing
  • US8666998B2 patent drawing

AI summary

A method, system and computer program product provides a first characteristic associated with a first data set and a single data value, and a second characteristic associated with a second data set; and calculates at least one of: 1) the similarity of the first data set with the second data set based on the first and second characteristics, 2) the similarity of the first data set with the single data value based on the first characteristic and the single data value, 3) confidence indicating how well the first characteristic reflects properties of the first data set based on the first characteristic, and 4) confidence indicating how well the similarity of the first data set with the single data value reflects properties of the single data value based on the first characteristic and the single data value.