Data Mining Algorithm for Automated Rule Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Developing data rules for large datasets with millions of records and hundreds of columns is time-consuming and requires significant user effort, often leading to reactive rule creation after business problems arise, rather than proactive issue prevention.
Innovation Solution
A method and system using data mining algorithms, such as association rules and tree classifications, to automatically generate data rules, which are then edited and stored in a repository for validation, allowing for the identification and correction of deviant records.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual coding of data rules is used, then user control and rule accuracy are improved, but user time and effort increase significantly
Solution Approach 1:
The system performs self-service by automatically generating data rules through data mining algorithms without requiring manual user input. The algorithm autonomously analyzes data patterns, generates candidate rules, and presents them for user approval, eliminating the time-consuming manual rule creation process while maintaining rule quality through automated pattern recognition.
Solution Approach 2:
A data mining algorithm acts as an intermediary between the raw data and the user. Instead of the user directly creating rules from data analysis, the algorithm processes the data, extracts patterns, and generates draft rules that the user then reviews and refines. This intermediary step significantly reduces user effort while preserving user control over the final rules.
2Loss of information
If manual analysis of large datasets is performed, then understanding of data patterns is improved, but processing time and computational resources increase
Solution Approach 1:
The patent replaces the mechanical process of manual data analysis with an automated data mining algorithm. The algorithm computationally processes large datasets to identify patterns, correlations, and anomalies that would be time-consuming for users to discover manually. This substitution maintains comprehensive pattern understanding while dramatically reducing processing time through automated computational methods.
3Reliability
If data rules are created reactively after problems occur, then rules address immediate issues, but proactive prevention of future problems is reduced
Solution Approach 1:
The system enables preliminary action by automatically generating data rules before problems occur. By continuously analyzing data patterns and proactively creating rules that prevent anomalies, the system addresses potential issues before they manifest as business problems. This shifts the approach from reactive problem-solving to proactive prevention while maintaining rule effectiveness through ongoing data mining.
4Productivity
If automated data mining is used, then user effort is reduced, but rule accuracy and user control may decrease
Solution Approach 1:
The system implements a dynamic rule generation process where the level of automation adapts based on user needs. Users can control the degree of automation by adjusting parameters such as rule confidence thresholds, pattern complexity filters, and review requirements. This dynamic approach allows users to balance between automated efficiency and manual control, ensuring rule accuracy while maintaining high productivity through selective automation.
Data Source
AI summary
Provided are a method, system, and article of manufacture for using a data mining algorithm to discover data rules. A data set including multiple records is processed to generate data rules for the data set. Each record has a record format including a plurality of fields and each rule provides a predicted condition for one field based on at least one predictor condition in at least one other field. The generated data rules are provided to a user interface to enable a user to edit the generated data rules. The data rules are stored in a rule repository to be available to use to validate data sets having the record format.


