Automated Graphing System for Massive Datasets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Off-the-shelf graphing tools are inadequate for handling real-world datasets due to issues like corrupted data, lack of variable type documentation, and presence of special values, leading to distorted graphs and inefficient manual adjustments.

Innovation Solution

A method and apparatus for automated graphing that selects a graphing range with minimal scale distortion, chooses an appropriate graphing style, and detects and filters special values, allowing for efficient plotting of mixed variables with continuous and categorical values on the same scale.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If off-the-shelf graphing tools are used to graph every variable in a dataset, then complete visualization coverage is achieved, but the time and effort required for manual adjustment becomes intractable

Engineering Contradiction:
Improvegraphing efficiencyVSAvoidmanual adjustment time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically detecting variable types, selecting appropriate graphing styles, and generating graphs without requiring manual analyst intervention for each variable. The automated graphing system evaluates dataset characteristics and configures graphs autonomously, eliminating the intractable manual adjustment process while maintaining complete visualization coverage.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes parameters automatically by detecting variable types (continuous, categorical, mixed) and dynamically selecting graphing styles and configurations based on these detected parameters. This automated parameter adjustment allows the system to generate appropriate graphs for all variables without manual intervention, resolving the contradiction between complete coverage and time consumption.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If real-world datasets are graphed as-is, then data completeness is maintained, but graph distortion occurs due to bad or corrupted data

Engineering Contradiction:
Improvegraph accuracyVSAvoiddata corruption impact
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary actions by detecting and filtering bad or corrupted data before graphing occurs. It identifies outliers, handles missing values, and cleans the dataset in advance, ensuring that only reliable data points are used in graph generation. This preliminary data cleaning prevents graph distortion while maintaining the integrity of valid data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts and removes harmful elements from the dataset by identifying and filtering out bad or corrupted data points, outliers, and improperly encoded special values. This extraction process separates reliable data from corrupted data, allowing accurate graphing of only the valid portions while maintaining overall data completeness where possible.

Inventive Principle:
Principle #2Taking out (Extraction)

3Ease of operation

If traditional graphing tools are used without variable type documentation, then tool simplicity is maintained, but graphing accuracy deteriorates due to unknown variable types

Engineering Contradiction:
Improvegraphing tool simplicityVSAvoidvariable type identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system performs self-service by automatically detecting variable types without requiring external documentation. It analyzes data patterns, ranges, and distributions to infer whether variables are continuous, categorical, or mixed types. This automated type detection maintains tool simplicity while achieving accurate variable type identification, resolving the contradiction between ease of operation and measurement precision.

Inventive Principle:
Principle #25Self-service

4Loss of information

If arbitrarily encoded special values are included in graphs, then data completeness is preserved, but graph interpretability is compromised

Engineering Contradiction:
Improvedata information retentionVSAvoidgraph interpretability
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system extracts and separates arbitrarily encoded special values from the main data stream. It identifies these special values through pattern recognition and filters them out during graphing, preventing them from distorting the visual representation. The system handles these values appropriately (exclusion, separate categorization, or transformation) to maintain both data information and graph interpretability.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS7830382B2Method and apparatus for automated graphing of trends in massive, real-world databases
Publication Date: 2010.11.09 FAIR ISAAC & CO INC
  • US7830382B2 patent drawing
  • US7830382B2 patent drawing
  • US7830382B2 patent drawing

AI summary

A method and apparatus for leveraging the inherent massiveness of real-world data sets to solve the problems typically associated with graphing the data is provided. Three particular areas of concern are as follows: a high likelihood of containing instances of bad or corrupted data that could distort the graph; little or no documentation about the type of each variable; and the presence of arbitrarily encoded missing or special values. One embodiment of the invention provides a methodology for automatically selecting a graphing range with minimal scale distortion. Another embodiment of the invention provides a methodology for automatically choosing an appropriate graphing style. Another embodiment of the invention provides a methodology for automatically detecting and filtering special values in data.