Data Set Fingerprinting for Automatic Domain Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data integration tools require manual effort and specialized algorithms to classify data domains, making it time-consuming to understand and document data sources, especially for non-standard domains like postal codes or enterprise-specific data, which are not well-supported by existing tools.
Innovation Solution
The method computes 'fingerprints' for data sets using metric algorithms that capture various data characteristics, allowing for automatic classification and comparison of data sets without requiring domain-specific algorithms, enabling semi-automatic data classification and detection of invalid values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If specialized algorithms are used to classify data domains, then classification accuracy is improved, but device complexity and development time increase
Solution Approach 1:
The patent applies universality by creating a single metric algorithm that can classify multiple different data domains (US addresses, Belgian postal codes, product references, enterprise codes) without requiring separate specialized algorithms for each domain. The algorithm achieves this by computing characteristics like data length, character composition, and format patterns that are applicable across diverse data types, thereby reducing algorithmic complexity while maintaining classification accuracy.
Solution Approach 2:
The patent uses parameter changes by computing multiple characteristics (length, character composition, format patterns) of data sets and comparing these parameters to classify data domains. Instead of using complex domain-specific logic, the algorithm changes the approach to using measurable parameters that can be computed and compared systematically across different data types, simplifying the classification process.
2Measurement precision
If manual classification is performed for each column, then classification accuracy is maintained, but productivity decreases
Solution Approach 1:
The patent applies self-service by enabling the system to automatically classify data domains without requiring manual expert intervention for each column. The metric algorithm autonomously computes characteristics of data sets, compares them against known domain patterns, and performs classification automatically. This eliminates the need for users to manually evaluate each column while maintaining classification quality, thereby significantly improving productivity.
Solution Approach 2:
The patent uses copying by creating characteristic profiles (fingerprints) of data sets that capture essential domain-specific patterns. These characteristic copies can be stored and reused for classification, allowing the system to quickly compare new data against established patterns without re-analyzing the entire classification logic, thus speeding up the classification process while maintaining accuracy.
3Measurement precision
If domain-specific algorithms are developed for each data type, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The patent eliminates the need to develop separate algorithms for each data type by creating a universal metric algorithm that handles multiple domains (US addresses, Belgian postal codes, product references, enterprise codes) through a single unified approach. This universal algorithm computes characteristics and applies pattern matching that works across diverse data types, reducing development time while maintaining domain detection accuracy.
Solution Approach 2:
The patent applies preliminary action by pre-computing and storing characteristic profiles of known data domains during system initialization or setup. These pre-computed characteristics serve as reference patterns that the metric algorithm can quickly compare against new data, eliminating the need to develop and test new algorithms for each domain while maintaining high detection accuracy.
Data Source
AI summary
A method, system and computer program product provides a first characteristic associated with a first data set and a single data value, and a second characteristic associated with a second data set; and calculates at least one of: 1) the similarity of the first data set with the second data set based on the first and second characteristics, 2) the similarity of the first data set with the single data value based on the first characteristic and the single data value, 3) confidence indicating how well the first characteristic reflects properties of the first data set based on the first characteristic, and 4) confidence indicating how well the similarity of the first data set with the single data value reflects properties of the single data value based on the first characteristic and the single data value.


