Data Field Semantic Labeling from Automated Profile Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data sets are often large, undocumented, and labeled inconsistently, making it difficult to ascertain business information and requiring manual annotation, especially in legacy systems with versioning issues.
Innovation Solution
A system that automatically classifies and labels data fields by analyzing metadata and data content, using statistical checks and machine learning to generate meaningful labels, and builds an infrastructure for applying data standards across different data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation is used to label data fields, then labeling accuracy can be ensured, but time consumption and labor costs increase significantly
Solution Approach 1:
The system enables data fields to be automatically labeled through self-service mechanisms. The labeling system performs statistical checks on metadata and data content, automatically generates label proposals, and assigns labels without requiring manual intervention for each field, thus reducing time consumption while maintaining accuracy through automated validation
Solution Approach 2:
The patent replaces the mechanical manual annotation process with an automated computing system. The system uses statistical analysis, machine learning models, and algorithmic processing to substitute human labor in the labeling task, significantly reducing time consumption while maintaining or improving consistency and accuracy
2Productivity
If existing labels are used from legacy systems, then data can be quickly processed, but label consistency and business meaning are compromised
Solution Approach 1:
The system performs preliminary statistical analysis and profiling of data fields before final labeling. By pre-processing the data to extract metadata, analyze data content distributions, and generate candidate labels in advance, the system ensures that labels are assigned based on comprehensive analysis, preserving business meaning while enabling efficient processing
Solution Approach 2:
The system incorporates feedback mechanisms where label proposals are evaluated based on multiple criteria including statistical significance, consistency with existing data, and alignment with business context. The system iteratively refines label assignments by evaluating feedback from statistical checks and validation rules, ensuring both consistency and preservation of business information
3Measurement precision
If comprehensive statistical checks are performed on all data fields, then labeling accuracy improves, but system complexity and processing overhead increase
Solution Approach 1:
The system segments the labeling process into distinct modular components: data profiling, statistical checking, label proposal generation, similarity determination, and final label selection. Each module performs a specific function with well-defined inputs and outputs, reducing overall system complexity while enabling comprehensive statistical analysis through organized, manageable stages
Solution Approach 2:
The system employs universal statistical checking mechanisms and standardized labeling procedures that can be applied across different data fields and data types. By creating a multi-functional framework that handles various data formats and structures through common statistical methods and unified label proposals, the system maintains high accuracy without proportionally increasing complexity
Data Source
AI summary
A data processing system for discovering a semantic meaning of a field included in one or more data sets is configured to identify a field included in one or more data sets, with the field having an identifier. For that field, the system profiles data values of the field to generate a data profile, accesses a plurality of label proposal tests, and generates a set of label proposals by applying the plurality of label proposal tests to the data profile. The system determines a similarity among the label proposals and selects a classification. The system identifies one of the label proposals as identifying the semantic meaning. The system stores the identifier of the field with the identified one of the label proposals that identifies the semantic meaning.


