Inline Data Quality Schema Management for Enterprise Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large enterprise organizations face challenges in managing data quality rules across disparate database technologies, making it difficult to ensure data consistency and identify invalid data.
Innovation Solution
The system uses a data fabric to extend data schema metadata, automatically inferring data quality rules using AI/ML algorithms. These rules are added as metadata to the data schema, creating data pipelines that continuously identify and enforce data quality rules across relational and noSQL datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple databases based on different technologies are used to support complex enterprise operations, then the organization can provide diverse products and services, but data quality management becomes difficult and inconsistent
Solution Approach 1:
The patent creates a universal data quality validation framework that works across multiple database technologies (relational, NoSQL, etc.) by defining a common validation language and standardized data quality rules that can be applied uniformly regardless of the underlying database system, enabling consistent data quality management across diverse enterprise operations
2Ease of manufacture
If standardized validation techniques are used across different database technologies, then implementation becomes simpler, but customization to specific business unit needs is lost
Solution Approach 1:
The patent enables local quality by allowing business units to define customized data quality rules and validation criteria specific to their domain requirements while using the same standardized validation framework. Each business unit can tailor validation parameters, data quality thresholds, and business logic rules to their specific needs without affecting other units, achieving both standardization and customization
3Measurement precision
If manual data quality validation is performed across disparate database systems, then data accuracy can be monitored, but the process becomes time-consuming and inefficient
Solution Approach 1:
The patent implements self-service data quality validation by enabling the system to automatically infer data quality rules from existing data patterns, automatically generate validation logic, and continuously monitor data quality across databases without requiring manual intervention. The system learns from data characteristics and autonomously maintains data quality standards, significantly improving validation efficiency while preserving accuracy
4Adaptability or versatility
If customized validation rules are defined for each business unit, then data quality meets specific needs, but the complexity of managing rules across the enterprise increases
Solution Approach 1:
The patent introduces an intermediary validation management layer that sits between business units and database systems. This layer provides a unified interface for defining, storing, and managing data quality rules, automatically translates business-specific validation requirements into standardized validation logic, and coordinates rule execution across multiple databases, thereby reducing the complexity of managing customized validation rules enterprise-wide
Data Source
AI summary
Various aspects of the disclosure relate to automatically inferring data quality rules for relational and/or non-SQL datasets. The data quality rules may be added as additional metadata for a data schema to improve and add validations for data to ensure data consistency and/or to identify invalid data. Rules may be automatically inferred based on data of data elements in one or more datasets and an identified significance of data in data points to extract common characteristics of data to create, teach and train one or more data quality (DQ) models for data elements across all data stores of the enterprise network.


