Data Cleansing Step Analysis for Automated Quality Rule Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large enterprises face challenges in managing and maintaining accurate data quality due to the volume of data columns, manual data quality rule writing, and the complexity of automated data cleansing processes, particularly in toolchains with conflicting codes.

Innovation Solution

A method and system that uses AI and ML to automatically detect and correct data cleansing steps by analyzing transformation assets, identifying redundant rules, and applying data quality policies across similar data domains, shifting from manual to automated data quality management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data quality rules are written to manage data quality, then data quality accuracy can be maintained, but the time and effort required increases significantly

Engineering Contradiction:
Improvedata quality accuracyVSAvoidtime for writing and testing rules
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables automated generation of data quality rules by analyzing data cleansing code transformations. The code itself serves as the source for rule creation, eliminating the need for manual rule writing and testing while maintaining accuracy through programmatic extraction of transformation logic.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical processes of writing and testing data quality rules are replaced with an automated computational system that analyzes data cleansing code and generates corresponding rules programmatically, significantly reducing time investment while preserving rule accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated machine learning methods are used for data cleansing, then productivity increases, but code conflicts and maintenance complexity increase

Engineering Contradiction:
Improvedata cleansing efficiencyVSAvoidtoolchain code complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system analyzes data cleansing code transformations and feeds this information back to automatically generate data quality rules. This feedback loop enables the system to understand and manage code conflicts by extracting transformation logic, thereby reducing maintenance complexity while preserving productivity gains from automated cleansing.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary analysis layer between the data cleansing code and the data quality rules. This intermediary automatically translates code transformations into quality rules, serving as a mediator that resolves conflicts and simplifies maintenance without reducing the productivity of automated cleansing operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If thousands of data quality rules are created to cover all data columns, then comprehensive data quality coverage is achieved, but rule management and maintenance become overwhelming

Engineering Contradiction:
Improvedata quality coverageVSAvoidrule maintenance difficulty
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically generates data quality rules by analyzing data cleansing code, eliminating the need for manual creation and maintenance of thousands of rules. The code itself serves as the source of truth, and the system self-generates appropriate quality rules, maintaining comprehensive coverage while simplifying management through automation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a universal system that can generate data quality rules across all data columns and transformation types through a single automated process. This multi-functional approach replaces the need for individual rule creation and maintenance for each data column, achieving comprehensive coverage while making the system easy to operate and maintain.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12554689B2Detecting high-impact data quality rules and policies from data cleansing steps
Publication Date: 2026.02.17 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12554689B2 patent drawing
  • US12554689B2 patent drawing

AI summary

A method, computer system, and a computer program product are provided for cleansing steps. These are in accordance with existing data quality and rules in existence for using a plurality of different transformation assets. Information is obtained about a plurality of different transformation assets and their associated data quality and rules are extracted. A plurality of possible cleansing steps to be performed are identified for the plurality of different transformation assets. An analysis is performed for the identified cleansing steps, on impact on the different transformation assets. It is then determined when more than one identified step has a similar semantics across the plurality of different transformation assets and when more than any two of them need to perform a similar step across the same dataset. The relevance of each cleansing step to be performed is then determined and a cleansing step order of performance is provided.