Visual Data Mining for Interactive Decision Tree Modification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data classifiers face challenges in achieving high accuracy for high-dimensional data sets, commonly referred to as big data, due to their poor performance and cost-prohibitive iteration processes, making it difficult to efficiently improve classification accuracy.

Innovation Solution

The implementation of visual data mining techniques that generate and visualize composite decision tree structures, such as random forests, allowing users to interact with the visualization to modify and improve classification accuracy by adding or removing decision trees, thereby enhancing the classification process for high-dimensional data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data classification methods are used for high-dimensional data sets, then the process can be completed with existing tools, but the classification accuracy is poor and the iteration process becomes cost-prohibitive

Engineering Contradiction:
Improveclassification accuracyVSAvoiditeration cost and time
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the complex classification problem into multiple decision tree data structures that can be independently generated, visualized, and modified. Each decision tree represents a segment of the overall classification model, allowing users to interact with and modify individual segments without redesigning the entire model, thereby reducing iteration costs and improving accuracy for high-dimensional data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a visualization interface as an intermediary between the complex composite data structure and the user. This visualization mediator transforms the abstract composite structure into an intuitive visual representation, enabling users to easily identify, select, and modify specific decision trees without needing to understand the underlying complexity, thus reducing iteration time and improving classification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If multiple iterations of data preprocessing and model training are performed to improve accuracy, then classification accuracy improves, but the process becomes cost-prohibitive and time-consuming

Engineering Contradiction:
Improveclassification accuracyVSAvoiditeration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary visualization of the composite data structure before final classification. By pre-visualizing the decision trees and their relationships, users can identify which trees need modification or removal in advance, avoiding the need for multiple time-consuming iterations of training and retraining. This preliminary action allows for targeted improvements that reduce overall iteration time while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a dynamic visualization interface that allows real-time modification of the composite data structure. Users can add, remove, or modify individual decision trees dynamically without triggering a complete retraining process. This dynamic capability enables iterative improvement of classification accuracy while significantly reducing the time required for each iteration, as only the affected portions need to be reprocessed rather than the entire model.

Inventive Principle:
Principle #15Dynamics

3Reliability

If existing data classification tools are used for big data, then the tools can process the data, but the high dimensionality results in poor performance

Engineering Contradiction:
Improveclassification performanceVSAvoiddata structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the complex high-dimensional data structure into a visual representation that operates in a lower dimension. By visualizing the composite data structure as a graphical representation of decision trees and their relationships, the system reduces the dimensional complexity from the raw high-dimensional data space to a two-dimensional or three-dimensional visual space. This dimensional transformation makes the complex structure more manageable and easier to interpret, improving classification performance while reducing perceived complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS8996583B2Interactive visual data mining for increasing classification accuracy
Publication Date: 2015.03.31 EMC IP HLDG CO LLC
  • US8996583B2 patent drawing
  • US8996583B2 patent drawing
  • US8996583B2 patent drawing

AI summary

Visual data mining techniques for increasing data classification accuracy are disclosed. For example, a method comprises the following steps. At least two decision tree data structures from a high-dimensional data set are generated. A composite data structure comprising the at least two decision tree data structures is generated. The composite data structure is generated based on a correlation computed between the at least two decision tree data structures. The composite data structure is visualized on a display. Modification of the composite data structure is enabled via interaction with the visualization of the composite data structure on the display.