Data Quality Analysis via Natural Language Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data quality engines require technical expertise to access and utilize data quality rules, limiting their utility for non-technical users and hindering efficient data quality analysis in organizations with large datasets.

Innovation Solution

A data processing system that allows specification of natural language data quality requirements and rules, enabling non-technical users to interact with the system through a metadata repository, which associates business-oriented data elements and metrics with technical data quality rules, facilitating automated data quality analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data quality engines use predefined technical data quality rules, then data quality measurement precision is improved, but ease of operation deteriorates because technical expertise is required to access and utilize these rules

Engineering Contradiction:
Improvedata quality measurement precisionVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces natural language requirements as an intermediary layer between non-technical users and technical data quality rules. Users can specify data quality requirements in natural language without needing to understand the underlying technical rules, while the system automatically maps these requirements to appropriate technical rules for execution. This mediator resolves the contradiction by shielding users from technical complexity while maintaining measurement precision through rule-based evaluation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated data quality analysis is implemented, then productivity is improved, but device complexity increases due to the need for technical infrastructure and rule management

Engineering Contradiction:
ImproveproductivityVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent enables self-service by allowing the system to automatically map natural language requirements to technical data quality rules without requiring manual configuration or technical expertise. The automated analysis engine independently performs rule selection, data evaluation, and result generation, reducing the operational complexity burden on users while maintaining high productivity through automation.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If technical data quality rules are used for analysis, then measurement precision is improved, but adaptability deteriorates because non-technical users cannot easily specify or modify requirements

Engineering Contradiction:
Improvemeasurement precisionVSAvoidadaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the parameter of user interaction from technical rule specification to natural language requirement description. By changing the interface parameter from technical jargon to everyday language, the system maintains measurement precision through underlying rule execution while dramatically improving adaptability, allowing non-technical users to easily specify and modify data quality requirements without understanding the technical rule structure.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11341116B2Techniques for automated data analysis
Publication Date: 2022.05.24 AB INITIO TECHNOLOGY LLC
  • US11341116B2 patent drawing
  • US11341116B2 patent drawing
  • US11341116B2 patent drawing

AI summary

According to some aspects, a data processing system is provided, the data processing system comprising at least one computer readable medium comprising processor-executable instructions that, when executed, cause the at least one processor to receive, through at least one user interface, input indicating a data element and one or more data quality metrics, identify, based on relationship information associated with the data element and/or the one or more data quality metrics, one or more datasets, one or more fields of the one or more datasets, and one or more data quality rules, each of the data quality rules being associated with at least one of the one or more fields, and perform an analysis of data quality of the one or more fields based at least in part on the one or more data quality rules associated with the one or more fields.