Database Query Generation for Automated Data Outlier Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data quality analysis systems rely heavily on manual rule generation and validation, which is time-consuming and inefficient, lacking automation and accuracy in detecting data outliers.

Innovation Solution

A system that uses machine learning to automatically generate database query statements, apply outlier detection models, and visualize anomalies through graph networks, optimizing the rule generation and detection process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If manual rule generation and validation is used, then rule accuracy can be maintained, but the process is time-consuming and inefficient

Engineering Contradiction:
Improverule generation timeVSAvoidautomation level
Core Design Contradiction:
Loss of timeVSExtent of automation

Solution Approach 1:

The system enables self-service by allowing the machine learning model to automatically generate, validate, and optimize data quality rules without requiring manual intervention from domain experts. The model autonomously learns from historical data and business context to create accurate rules, eliminating the time-consuming manual rule generation process while maintaining high rule accuracy through continuous learning and adaptation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of rule generation with an automated machine learning system. The ML model substitutes human analysts in the rule creation workflow, using algorithms to automatically generate rules from historical data patterns, validate them against business logic, and optimize them for performance, thereby eliminating manual labor while maintaining or improving rule accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If manual validation of rules is performed, then rule reliability can be ensured, but productivity is reduced

Engineering Contradiction:
Improvedata quality analysis efficiencyVSAvoidrule validity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system implements feedback mechanisms where the machine learning model continuously monitors rule performance against actual data quality issues and business outcomes. Historical data and validation results feed back into the model to refine and optimize rule generation, ensuring that rules remain reliable and effective while enabling rapid iteration and deployment without manual validation bottlenecks.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces manual validation with automated machine learning validation processes. The system uses ML algorithms to automatically validate rules against historical data patterns, business logic, and known data quality issues, ensuring rule reliability while dramatically improving productivity through automated rather than manual validation workflows.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If comprehensive rule coverage is achieved through manual analysis, then detection accuracy improves, but the process becomes more complex and time-consuming

Engineering Contradiction:
Improveoutlier detection accuracyVSAvoidrule management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex rule management task into manageable components handled by the machine learning model. The model automatically divides and organizes rules by data type, business context, and detection priority, making the system more scalable and less complex. This segmentation allows comprehensive coverage of data quality issues while reducing manual management complexity through automated organization and prioritization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies parameter changes by using the machine learning model to dynamically adjust rule parameters such as sensitivity thresholds, detection priorities, and coverage criteria based on historical data patterns and business requirements. This allows the system to achieve high detection accuracy across diverse data scenarios while automatically optimizing parameters to reduce overall system complexity and improve performance.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If existing manual systems are used, then implementation is simpler, but computing resources are wasted and efficiency is low

Engineering Contradiction:
Improvedata quality analysis throughputVSAvoidcomputing resource utilization
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent replaces inefficient manual processing with optimized machine learning systems that achieve higher productivity and better computing resource utilization. The ML model processes data quality analysis tasks more efficiently than manual methods, improving throughput while using computing resources more effectively through automated optimization and intelligent processing algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12517914B2System and method for combinatorial data outlier detection via database query statement generation
Publication Date: 2026.01.06 BANK OF AMERICA CORP
  • US12517914B2 patent drawing
  • US12517914B2 patent drawing
  • US12517914B2 patent drawing

AI summary

Systems, computer program products, and methods are described herein for combinatorial data outlier detection via database query statement generation. A table comprising rows and columns is received, wherein each column comprises records and a corresponding data element. A datatype to be assigned to each data element is determined, based on each column and using a first machine learning model. Conditions are generated for columns corresponding to each data element. The conditions are combined into a predetermined number of condition combinations, wherein the condition combinations comprise combinatorial sequencing of the conditions. A query statement is generated for each of the condition combinations. The table in queried with each query statement to determine query results. A data element quantity and a record quantity in the query results are determined for each query statement.