Database Query Generation for Automated Data Outlier Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data quality analysis systems rely heavily on manual rule generation and validation, which is time-consuming and inefficient, lacking automation and accuracy in detecting data outliers.
Innovation Solution
A system that uses machine learning to automatically generate database query statements, apply outlier detection models, and visualize anomalies through graph networks, optimizing the rule generation and detection process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If manual rule generation and validation is used, then rule accuracy can be maintained, but the process is time-consuming and inefficient
Solution Approach 1:
The system enables self-service by allowing the machine learning model to automatically generate, validate, and optimize data quality rules without requiring manual intervention from domain experts. The model autonomously learns from historical data and business context to create accurate rules, eliminating the time-consuming manual rule generation process while maintaining high rule accuracy through continuous learning and adaptation.
Solution Approach 2:
The patent replaces the mechanical manual process of rule generation with an automated machine learning system. The ML model substitutes human analysts in the rule creation workflow, using algorithms to automatically generate rules from historical data patterns, validate them against business logic, and optimize them for performance, thereby eliminating manual labor while maintaining or improving rule accuracy.
2Productivity
If manual validation of rules is performed, then rule reliability can be ensured, but productivity is reduced
Solution Approach 1:
The system implements feedback mechanisms where the machine learning model continuously monitors rule performance against actual data quality issues and business outcomes. Historical data and validation results feed back into the model to refine and optimize rule generation, ensuring that rules remain reliable and effective while enabling rapid iteration and deployment without manual validation bottlenecks.
Solution Approach 2:
The patent replaces manual validation with automated machine learning validation processes. The system uses ML algorithms to automatically validate rules against historical data patterns, business logic, and known data quality issues, ensuring rule reliability while dramatically improving productivity through automated rather than manual validation workflows.
3Measurement precision
If comprehensive rule coverage is achieved through manual analysis, then detection accuracy improves, but the process becomes more complex and time-consuming
Solution Approach 1:
The system segments the complex rule management task into manageable components handled by the machine learning model. The model automatically divides and organizes rules by data type, business context, and detection priority, making the system more scalable and less complex. This segmentation allows comprehensive coverage of data quality issues while reducing manual management complexity through automated organization and prioritization.
Solution Approach 2:
The patent applies parameter changes by using the machine learning model to dynamically adjust rule parameters such as sensitivity thresholds, detection priorities, and coverage criteria based on historical data patterns and business requirements. This allows the system to achieve high detection accuracy across diverse data scenarios while automatically optimizing parameters to reduce overall system complexity and improve performance.
4Productivity
If existing manual systems are used, then implementation is simpler, but computing resources are wasted and efficiency is low
Solution Approach 1:
The patent replaces inefficient manual processing with optimized machine learning systems that achieve higher productivity and better computing resource utilization. The ML model processes data quality analysis tasks more efficiently than manual methods, improving throughput while using computing resources more effectively through automated optimization and intelligent processing algorithms.
Data Source
AI summary
Systems, computer program products, and methods are described herein for combinatorial data outlier detection via database query statement generation. A table comprising rows and columns is received, wherein each column comprises records and a corresponding data element. A datatype to be assigned to each data element is determined, based on each column and using a first machine learning model. Conditions are generated for columns corresponding to each data element. The conditions are combined into a predetermined number of condition combinations, wherein the condition combinations comprise combinatorial sequencing of the conditions. A query statement is generated for each of the condition combinations. The table in queried with each query statement to determine query results. A data element quantity and a record quantity in the query results are determined for each query statement.


