Interactive Data Mining System Rule Context Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data mining techniques, particularly in the context of class-labeled data, face challenges such as the completeness problem, context loss, and the limitations of long rules, which hinder the identification of interesting and actionable knowledge. Existing methods often fragment the knowledge space, losing contextual information and failing to systematically explore rule spaces effectively.

Innovation Solution

The development of data mining software that retains all rule counts regardless of minimum support or confidence levels, allowing for the presentation of rules in context and enabling the extraction of 'General Impressions' through processing related sets of rules. This approach facilitates the identification of trends and discriminative power, providing users with an intuitive understanding of the data by visualizing rule cubes and histograms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional data mining techniques are used to generate classification rules, then a large number of rules are produced, but it becomes very difficult to identify interesting rules by manual inspection

Engineering Contradiction:
Improverule generation volumeVSAvoidrule identification difficulty
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent extracts and highlights only the most interesting and relevant rules from the complete rule set by computing interestingness scores based on user profiles and domain knowledge. This selective extraction approach presents a manageable subset of rules to users while maintaining access to the complete underlying rule set, resolving the contradiction between generating comprehensive rules and enabling easy identification of interesting ones.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of rule presentation by transforming raw rules into scored, ranked results based on multiple criteria including user needs, domain knowledge, and rule quality metrics. This parameter transformation converts an overwhelming volume of rules into a prioritized list that is easy to navigate and interpret.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If existing techniques propose methods to deal with the interestingness problem, then some techniques are developed, but interestingness remains a difficult problem with limited success in real life applications

Engineering Contradiction:
Improvetechnique varietyVSAvoidapplication success rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal framework that integrates multiple existing interestingness techniques into a single cohesive system. The framework can accommodate different user profiles, domain knowledge representations, and rule scoring methods, making it adaptable to various applications while providing consistent reliable results across different datasets and use cases.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent incorporates feedback mechanisms where the system learns from user interactions with the rule set. User preferences, selection patterns, and feedback on rule usefulness are fed back into the interestingness computation to refine and personalize the rule ranking, improving both adaptability to individual users and overall reliability of the system.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If data mining software follows the current rule mining paradigm, then the knowledge space is fragmented and massive rules are generated, but a large number of holes are created in the space of useful knowledge

Engineering Contradiction:
Improverule quantityVSAvoidknowledge context loss
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent merges fragmented rules into cohesive knowledge structures by grouping related rules and presenting them in contextual clusters. Instead of displaying isolated rules, the system combines rules that share common patterns, attributes, or relationships, thereby reducing the perception of fragmentation while maintaining the completeness of the underlying rule set.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary organization and contextualization of rules before presentation to users. By pre-processing the rule set to identify relationships, patterns, and contextual groupings, the system prepares the knowledge structure in advance to minimize information loss and make the complete rule set more accessible without requiring users to navigate fragmented individual rules.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If minimum support and confidence thresholds are applied to filter rules, then computational efficiency is improved, but rules with potential contextual importance are excluded

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidcontextual information loss
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies minimum support and confidence thresholds as partial filtering rather than complete exclusion criteria. The system computes rules with relaxed thresholds to maintain contextual completeness, then uses interestingness scoring to prioritize and present the most relevant rules. This partial application of filtering maintains efficiency while preserving contextual information that would be lost with strict threshold enforcement.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS7979362B2Interactive data mining system
Publication Date: 2011.07.12 MOTOROLA SOLUTIONS INC
  • US7979362B2 patent drawing
  • US7979362B2 patent drawing
  • US7979362B2 patent drawing

AI summary

An interactive data mining system (100, 3000) that is suitable for data mining large high dimensional (e.g., 200 dimension) data sets is provided. The system graphically presents rules in a context allowing users to readily gain an intuitive appreciation of the significance of important attributes (data fields) in the data. The system (100, 3000) uses metrics to quantify the importance of the various data attributes, data values, attribute/value pairs, ranks them according to the metrics and displays histograms and lists of attributes and values in order according to the metric, thereby allowing the user to rapidly find the most interesting aspects of the data. The system explores the impact of user defined constraints and presents histograms and rule cubes including superposed and interleaved rule cubes showing the effect of the constraints.