Association Rule Search Using Ordinal Ordering and Convex Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for searching association rules in databases, such as the PRIM process, face challenges with complex and time-consuming calculations, incomplete database exploration, and the need for post-management of redundant variables, leading to reduced precision and increased processing time.
Innovation Solution
A method that orders numerical values into ordinal values, explores right convex regions in subspaces, and selects regions of interest based on density and size thresholds, reducing combinatorial complexity from exponential to polynomial and enabling parallelized calculations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the PRIM process is used to search for association rules, then the method can identify regions of interest in the database, but the calculation complexity becomes exponential and processing time increases significantly
Solution Approach 1:
The patent segments the database exploration into distinct subspaces, where each subspace corresponds to a specific combination of input variables. By dividing the overall search space into manageable subspaces and processing them independently, the method reduces the exponential complexity of exploring the entire space at once, while maintaining comprehensive coverage for precise association rule discovery
Solution Approach 2:
The patent performs preliminary ordering of numerical values into ordinal values before the actual association rule search. This preprocessing step transforms continuous numerical data into discrete ordinal categories, which simplifies subsequent calculations and enables more efficient exploration of the database without losing important relationships, thereby reducing processing time while maintaining precision
2Productivity
If the PRIM process is used to search for association rules, then regions can be identified, but the database exploration remains incomplete and precision is reduced
Solution Approach 1:
By segmenting the database into subspaces based on combinations of input variables, the method ensures that each subspace is explored completely and independently. This segmentation allows the algorithm to systematically cover the entire database without missing important patterns, thereby improving both exploration efficiency and result precision simultaneously
Solution Approach 2:
The patent introduces a new dimension by ordering numerical values into ordinal values, creating an additional layer of data organization. This dimensional transformation allows the method to explore the database more systematically and completely, capturing relationships that might be missed in the original numerical space, thus improving precision without sacrificing exploration efficiency
3Device complexity
If redundant variables are managed in the PRIM process, then variable selection can be performed, but post-management is required and precision is reduced
Solution Approach 1:
The patent performs variable ordering and redundancy management as preliminary actions before the association rule search begins. By organizing variables into ordinal categories and identifying redundant relationships in advance, the method eliminates the need for post-processing variable management, and ensures that the association rules are discovered with maximum precision from the outset
Data Source
Figure 1~5
Figure 2
Figure 3
AI summary
The invention pertains to a method, implemented by a processing unit of a computer and/or by a programmable or reprogrammable logic circuit, for searching for rules of association in a database. The database comprises a list of instances each exhibiting a set of real numerical values taken by a predetermined number of variables. The method comprises the following steps: - the selection (20) of a set of NI input variables from among the variables of the list of instances, said set of NI variables defining a space of dimension NI; and the selection (22) of an output variable from among the remaining variables; - the ordering (24), for each input variable of the selected list of instances, of the numerical values of the instances for this variable, each instance then being defined by a set of ranks, and being represented by a point in said space of dimension NI; - the definition (26) of at least one modality for the selected output variable; - the exploration (32), in sub-spaces of the space of dimension NI, of right convex regions; - the selection (34), from among said explored right convex regions, of regions of interest.