Automated Graphing System for Massive Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Off-the-shelf graphing tools are inadequate for handling real-world datasets due to issues like corrupted data, lack of variable type documentation, and presence of special values, leading to distorted graphs and inefficient manual adjustments.
Innovation Solution
A method and apparatus for automated graphing that selects a graphing range with minimal scale distortion, chooses an appropriate graphing style, and detects and filters special values, allowing for efficient plotting of mixed variables with continuous and categorical values on the same scale.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If off-the-shelf graphing tools are used to graph every variable in a dataset, then complete visualization coverage is achieved, but the time and effort required for manual adjustment becomes intractable
Solution Approach 1:
The system performs self-service by automatically detecting variable types, selecting appropriate graphing styles, and generating graphs without requiring manual analyst intervention for each variable. The automated graphing system evaluates dataset characteristics and configures graphs autonomously, eliminating the intractable manual adjustment process while maintaining complete visualization coverage.
Solution Approach 2:
The system changes parameters automatically by detecting variable types (continuous, categorical, mixed) and dynamically selecting graphing styles and configurations based on these detected parameters. This automated parameter adjustment allows the system to generate appropriate graphs for all variables without manual intervention, resolving the contradiction between complete coverage and time consumption.
2Reliability
If real-world datasets are graphed as-is, then data completeness is maintained, but graph distortion occurs due to bad or corrupted data
Solution Approach 1:
The system performs preliminary actions by detecting and filtering bad or corrupted data before graphing occurs. It identifies outliers, handles missing values, and cleans the dataset in advance, ensuring that only reliable data points are used in graph generation. This preliminary data cleaning prevents graph distortion while maintaining the integrity of valid data.
Solution Approach 2:
The system extracts and removes harmful elements from the dataset by identifying and filtering out bad or corrupted data points, outliers, and improperly encoded special values. This extraction process separates reliable data from corrupted data, allowing accurate graphing of only the valid portions while maintaining overall data completeness where possible.
3Ease of operation
If traditional graphing tools are used without variable type documentation, then tool simplicity is maintained, but graphing accuracy deteriorates due to unknown variable types
Solution Approach 1:
The system performs self-service by automatically detecting variable types without requiring external documentation. It analyzes data patterns, ranges, and distributions to infer whether variables are continuous, categorical, or mixed types. This automated type detection maintains tool simplicity while achieving accurate variable type identification, resolving the contradiction between ease of operation and measurement precision.
4Loss of information
If arbitrarily encoded special values are included in graphs, then data completeness is preserved, but graph interpretability is compromised
Solution Approach 1:
The system extracts and separates arbitrarily encoded special values from the main data stream. It identifies these special values through pattern recognition and filters them out during graphing, preventing them from distorting the visual representation. The system handles these values appropriately (exclusion, separate categorization, or transformation) to maintain both data information and graph interpretability.
Data Source
AI summary
A method and apparatus for leveraging the inherent massiveness of real-world data sets to solve the problems typically associated with graphing the data is provided. Three particular areas of concern are as follows: a high likelihood of containing instances of bad or corrupted data that could distort the graph; little or no documentation about the type of each variable; and the presence of arbitrarily encoded missing or special values. One embodiment of the invention provides a methodology for automatically selecting a graphing range with minimal scale distortion. Another embodiment of the invention provides a methodology for automatically choosing an appropriate graphing style. Another embodiment of the invention provides a methodology for automatically detecting and filtering special values in data.


