Data Profiling System for Big Data Preprocessing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Big data analysis systems face inefficiencies due to poor quality raw data, which often contains missing fields, anomalies, and formatting issues, limiting the amount of data that can be processed and affecting the accuracy of analysis results.

Innovation Solution

A data preprocessing system that configures user interfaces for visual analysis of datasets, processes attributes, and transforms raw data into a suitable format for big data analysis systems by generating transformation scripts based on user interactions, improving data quality and quantity processed by big data analysis systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If raw data is directly processed by big data analysis systems, then processing speed is maintained, but data quality and processing completeness deteriorate due to formatting issues and anomalies

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements a data profiling system that performs preliminary analysis of raw data characteristics, statistics, and quality metrics before the big data analysis system processes the data. This preliminary action identifies formatting issues, anomalies, and data quality problems in advance, allowing the system to prepare appropriate preprocessing operations that improve data quality without significantly impacting processing throughput.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual data preprocessing is performed to improve data quality, then data quality improves, but time consumption and cost increase

Engineering Contradiction:
Improvedata qualityVSAvoidpreprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements an automated data profiling system that performs self-service data analysis by automatically examining raw data, generating statistical profiles, identifying quality issues, and recommending preprocessing operations without requiring manual expert intervention. The system autonomously completes tasks that traditionally required data experts and developers, significantly reducing preprocessing time while maintaining improved data quality.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If extensive data preprocessing is performed to increase processable data volume, then the amount of processable data increases, but system complexity and resource requirements increase

Engineering Contradiction:
Improveprocessable data volumeVSAvoidpreprocessing system complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements a data profiling system that analyzes and determines optimal parameters for preprocessing operations based on the actual characteristics of the raw data. By dynamically adjusting preprocessing parameters and operations according to data-specific profiles, the system increases the volume of processable data while avoiding unnecessary complex preprocessing steps, thus controlling system complexity and resource requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10346421B1Data profiling of large datasets
Publication Date: 2019.07.09 ALTERYX INC
  • US10346421B1 patent drawing
  • US10346421B1 patent drawing
  • US10346421B1 patent drawing

AI summary

A system provides data profile information describing attributes of a dataset. The system determines relative frequency of occurrences of attribute values with respect to a set of bins from a histogram of another attribute. The system presents a user interface that presents statistical information describing attributes of a dataset based on the relative frequency of occurrences of attribute values. The system generates a transformation script based on the user interactions for transforming records of the dataset. The transformation script is configured to preprocess data of the dataset for further analysis.