Visual Data Mining for Interactive Decision Tree Modification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data classifiers face challenges in achieving high accuracy for high-dimensional data sets, commonly referred to as big data, due to their poor performance and cost-prohibitive iteration processes, making it difficult to efficiently improve classification accuracy.
Innovation Solution
The implementation of visual data mining techniques that generate and visualize composite decision tree structures, such as random forests, allowing users to interact with the visualization to modify and improve classification accuracy by adding or removing decision trees, thereby enhancing the classification process for high-dimensional data sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data classification methods are used for high-dimensional data sets, then the process can be completed with existing tools, but the classification accuracy is poor and the iteration process becomes cost-prohibitive
Solution Approach 1:
The patent segments the complex classification problem into multiple decision tree data structures that can be independently generated, visualized, and modified. Each decision tree represents a segment of the overall classification model, allowing users to interact with and modify individual segments without redesigning the entire model, thereby reducing iteration costs and improving accuracy for high-dimensional data.
Solution Approach 2:
The patent introduces a visualization interface as an intermediary between the complex composite data structure and the user. This visualization mediator transforms the abstract composite structure into an intuitive visual representation, enabling users to easily identify, select, and modify specific decision trees without needing to understand the underlying complexity, thus reducing iteration time and improving classification accuracy.
2Reliability
If multiple iterations of data preprocessing and model training are performed to improve accuracy, then classification accuracy improves, but the process becomes cost-prohibitive and time-consuming
Solution Approach 1:
The patent performs preliminary visualization of the composite data structure before final classification. By pre-visualizing the decision trees and their relationships, users can identify which trees need modification or removal in advance, avoiding the need for multiple time-consuming iterations of training and retraining. This preliminary action allows for targeted improvements that reduce overall iteration time while maintaining high accuracy.
Solution Approach 2:
The patent creates a dynamic visualization interface that allows real-time modification of the composite data structure. Users can add, remove, or modify individual decision trees dynamically without triggering a complete retraining process. This dynamic capability enables iterative improvement of classification accuracy while significantly reducing the time required for each iteration, as only the affected portions need to be reprocessed rather than the entire model.
3Reliability
If existing data classification tools are used for big data, then the tools can process the data, but the high dimensionality results in poor performance
Solution Approach 1:
The patent transforms the complex high-dimensional data structure into a visual representation that operates in a lower dimension. By visualizing the composite data structure as a graphical representation of decision trees and their relationships, the system reduces the dimensional complexity from the raw high-dimensional data space to a two-dimensional or three-dimensional visual space. This dimensional transformation makes the complex structure more manageable and easier to interpret, improving classification performance while reducing perceived complexity.
Data Source
AI summary
Visual data mining techniques for increasing data classification accuracy are disclosed. For example, a method comprises the following steps. At least two decision tree data structures from a high-dimensional data set are generated. A composite data structure comprising the at least two decision tree data structures is generated. The composite data structure is generated based on a correlation computed between the at least two decision tree data structures. The composite data structure is visualized on a display. Modification of the composite data structure is enabled via interaction with the visualization of the composite data structure on the display.


