Composite Dataset Visualization for Data Preparation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data visualization applications face challenges in efficiently preparing and comparing large or complex datasets, as they often require manipulation and formatting to be usable, and lack effective tools for analyzing and understanding distributions, trends, and outliers across multiple data fields.
Innovation Solution
A data preparation method and user interface that allows users to compare datasets at different nodes in a process flow by forming a composite dataset, grouping data values into bins, and displaying distributions using an unstacked overlapping bar chart, with each bin representing data from multiple sources and visually indicating the origin of data values through color and shading.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If data sets are manually manipulated and formatted to be usable by data visualization applications, then data preparation accuracy is improved, but time consumption and operational complexity increase
Solution Approach 1:
The system performs preliminary data profiling and validation automatically during the data import stage, creating a structured understanding of data characteristics, types, and relationships before visualization is needed. This preliminary analysis includes generating data dictionaries, detecting outliers, and establishing data quality metrics, thereby reducing the need for manual manipulation later in the process.
Solution Approach 2:
The patent introduces an intelligent data preparation layer that acts as an intermediary between raw data sources and visualization applications. This layer automatically performs data cleaning, transformation, and formatting operations using algorithms that adapt to the specific data characteristics, eliminating the need for manual intervention while ensuring data quality and compatibility with visualization tools.
2Loss of information
If multiple data sets from different nodes are compared individually, then data analysis thoroughness is improved, but system complexity and operational difficulty increase
Solution Approach 1:
The system merges multiple data sets from different process nodes into a unified comparative view, automatically aligning data structures and normalizing formats. The unified interface presents side-by-side comparisons of data distributions, transformations, and quality metrics across nodes, enabling comprehensive analysis without requiring users to manually configure complex comparison workflows for each data set pair.
Solution Approach 2:
The patent creates a universal data comparison framework that can handle any combination of data sets from any number of process nodes through a single standardized interface. The system automatically detects data relationships, applies appropriate comparison methods based on data types, and provides consistent visualization formats, thereby simplifying the user experience while maintaining analytical rigor across diverse data sources.
3Ease of operation
If data distributions are visualized using traditional methods, then visualization simplicity is improved, but ability to identify outliers and understand trends across multiple data fields deteriorates
Solution Approach 1:
The system extends traditional distribution visualizations by adding comparative dimensions that display data from multiple process nodes simultaneously. Instead of showing single-node histograms or box plots, the patent implements multi-layered visualizations where each node's data distribution is overlaid or juxtaposed with others, enabling users to spot outliers and trends by comparing positions, spreads, and shapes across dimensions while maintaining visual accessibility through consistent graphical encodings.
Solution Approach 2:
The patent employs color-coded visual indicators to represent different process nodes and their data characteristics. Each node is assigned a distinctive color that is consistently applied across all visualization elements (histograms, box plots, scatter plots), allowing users to quickly trace data origins and compare distributions. Outliers and significant deviations are highlighted with contrasting colors or intensity variations, making them immediately visible without complicating the overall visual presentation.
Data Source
AI summary
A method compares data sets in a data preparation application. The method displays a user interface including a flow diagram having a plurality of nodes. Each of the nodes corresponds to a data set having a plurality of data fields. A user selects two nodes from the flow diagram. In response to the user selection, the method forms a composite data set comprising a union of two data sets corresponding to the two nodes and groups data values for each of a plurality of data fields in the composite data set to form a respective set of bins. The method then displays distributions of data values for the plurality of data fields in the composite data set. Each distribution comprises the respective set of bins for a respective data field. Each displayed bin depicts counts of data values in the respective bin originating from each of the two data sets.


