Data Profiling System for Big Data Preprocessing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Big data analysis systems face inefficiencies due to poor quality raw data, which often contains missing fields, anomalies, and formatting issues, limiting the amount of data that can be processed and affecting the accuracy of analysis results.
Innovation Solution
A data preprocessing system that configures user interfaces for visual analysis of datasets, processes attributes, and transforms raw data into a suitable format for big data analysis systems by generating transformation scripts based on user interactions, improving data quality and quantity processed by big data analysis systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If raw data is directly processed by big data analysis systems, then processing speed is maintained, but data quality and processing completeness deteriorate due to formatting issues and anomalies
Solution Approach 1:
The patent implements a data profiling system that performs preliminary analysis of raw data characteristics, statistics, and quality metrics before the big data analysis system processes the data. This preliminary action identifies formatting issues, anomalies, and data quality problems in advance, allowing the system to prepare appropriate preprocessing operations that improve data quality without significantly impacting processing throughput.
2Reliability
If manual data preprocessing is performed to improve data quality, then data quality improves, but time consumption and cost increase
Solution Approach 1:
The patent implements an automated data profiling system that performs self-service data analysis by automatically examining raw data, generating statistical profiles, identifying quality issues, and recommending preprocessing operations without requiring manual expert intervention. The system autonomously completes tasks that traditionally required data experts and developers, significantly reducing preprocessing time while maintaining improved data quality.
3Quantity of substance
If extensive data preprocessing is performed to increase processable data volume, then the amount of processable data increases, but system complexity and resource requirements increase
Solution Approach 1:
The patent implements a data profiling system that analyzes and determines optimal parameters for preprocessing operations based on the actual characteristics of the raw data. By dynamically adjusting preprocessing parameters and operations according to data-specific profiles, the system increases the volume of processable data while avoiding unnecessary complex preprocessing steps, thus controlling system complexity and resource requirements.
Data Source
AI summary
A system provides data profile information describing attributes of a dataset. The system determines relative frequency of occurrences of attribute values with respect to a set of bins from a histogram of another attribute. The system presents a user interface that presents statistical information describing attributes of a dataset based on the relative frequency of occurrences of attribute values. The system generates a transformation script based on the user interactions for transforming records of the dataset. The transformation script is configured to preprocess data of the dataset for further analysis.


