Data Cleaning Template Automating Format Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The labor-intensive process of data cleaning across multiple subsidiaries with different payroll service applications, requiring manual data set construction and repeated normalization, validation, and enrichment, becomes increasingly cumbersome as the number of subsidiaries and data sets grows.
Innovation Solution
A data cleaning template is constructed by specifying data types and formats, allowing for automatic normalization, validation, and enrichment of data across various formats, enabling reusability across different data sets and access parameter management for user visibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data cleaning processes are used for each subsidiary, then data can be normalized and validated, but the labor intensity and time required increase significantly
Solution Approach 1:
The patent creates a universal data cleaning template that can be applied across multiple subsidiaries and data sources. The template defines reusable data type specifications (e.g., currency, date, percentage) that automatically normalize data from different payroll service applications, eliminating the need to manually recreate cleaning processes for each subsidiary while maintaining consistent data quality standards.
Solution Approach 2:
The patent implements preliminary action by pre-defining data cleaning templates with specified data types, formats, and validation rules before data processing begins. These templates are constructed in advance and stored for reuse, allowing the system to automatically apply predefined normalization and validation logic when data is imported, rather than performing manual cleaning operations each time new data arrives.
2Adaptability or versatility
If data cleaning templates are constructed for each subsidiary, then specific data formats can be achieved, but the complexity of template construction and maintenance increases
Solution Approach 1:
The patent implements universality by creating a single data cleaning template that handles multiple data types (currency, date, percentage, text) and can be applied to data from various subsidiaries. The template uses standardized data type definitions that automatically adapt to different data sources, reducing the need to construct separate templates for each subsidiary while maintaining format adaptability through parameterized data type specifications.
3Reliability
If repeated data cleaning operations are performed when underlying data changes, then data remains current and accurate, but the workload increases with more divisions and data sets
Solution Approach 1:
The patent applies preliminary action by pre-configuring data cleaning templates with all necessary normalization rules, data type definitions, and validation logic before data processing begins. When underlying data changes or new data arrives from additional divisions, the preconstructed templates can be automatically reapplied without requiring manual intervention, maintaining data accuracy while improving efficiency through automated reuse of established cleaning patterns.
Data Source
AI summary
Described herein are various technologies pertaining to construction and application of a data cleaning template. A data cleaning tool, when applying the data cleaning template to a data set, is configured to identify a column in the data set that has data entries of a data type specified in the data cleaning template. In response to identifying the column in the data set, the data cleaning tool, when applying the data cleaning template to the data set, alters a format of the data entries in the column from a first format to a second format, the second format specified in the data cleaning template.


