Automated Dataset Preprocessing for Heterogeneous Data Streams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data processing methods for preparing datasets are time-consuming, cumbersome, and lack automation, especially in industrial contexts, leading to inefficiencies and potential security risks due to the heterogeneity and complexity of data sources, which can result in divergent analysis results and obsolete predictive models.
Innovation Solution
A data processing method and system that standardizes, encodes, and preprocesses data streams using interconnected modules, including data augmentation and outlier removal, to generate a reliable and relevant dataset, capable of handling diverse data sources and adapting to changes over time, with automated processes for data quality control and feedback mechanisms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data preprocessing is performed manually by Data Scientists using multiple tools, then data quality can be improved, but time consumption increases significantly
Solution Approach 1:
The system performs self-service by automatically detecting data quality issues and applying appropriate preprocessing transformations without manual intervention. The automated preprocessing system analyzes data characteristics and executes cleaning, normalization, and validation operations autonomously, eliminating the need for Data Scientists to manually perform these time-consuming tasks while maintaining high data quality standards.
Solution Approach 2:
The patent replaces the mechanical manual process of data preprocessing with an automated computational system. Instead of Data Scientists manually using multiple tools like Excel and statistical software, the system employs automated algorithms and software agents that perform data cleaning, validation, and transformation operations, dramatically reducing preprocessing time while maintaining or improving data quality.
2Adaptability or versatility
If multiple preprocessing tools are used to handle heterogeneous data, then data processing capability is improved, but system complexity increases
Solution Approach 1:
The patent implements a universal preprocessing system that can handle multiple data types and formats through a single integrated platform. The system employs versatile data quality agents and transformation modules that adapt to different data characteristics (structured, unstructured, semi-structured) and perform various preprocessing operations (cleaning, normalization, validation) within one cohesive architecture, eliminating the need to manage multiple separate tools while maintaining high adaptability.
Solution Approach 2:
The system introduces intermediary components such as data quality agents and standardized interface layers that mediate between diverse data sources and the core processing engine. These intermediaries translate and normalize different data formats and structures into a unified representation, enabling the system to handle heterogeneous data without requiring complex point-to-point integration between multiple specialized tools.
3Measurement precision
If advanced statistical analysis techniques are applied to increase analysis depth, then insight quality improves, but processing time increases
Solution Approach 1:
The system performs preliminary data preprocessing and quality improvement actions before conducting advanced statistical analysis. By automatically cleaning, normalizing, and validating data in advance, the system prepares high-quality input data that enables faster and more accurate analysis. This preliminary preparation reduces the computational burden during the actual analysis phase, allowing advanced techniques to execute more efficiently.
Solution Approach 2:
The patent replaces time-consuming manual statistical analysis with automated computational algorithms and machine learning models. The system employs automated statistical agents that can perform complex analyses (correlation analysis, anomaly detection, predictive modeling) much faster than manual methods, maintaining or improving analysis depth while dramatically reducing processing time through computational automation.
Data Source
AI summary
The invention relates to a data processing method for preparing a dataset that includes a processor that receives a first plurality of data input streams to prepare an output dataset. The plurality of data input streams and the output dataset are different. The method includes standardizing the plurality of data input streams, encoding the normalized data, preprocessing missing data, and transmitting a preprocessed dataset. The invention further relates to a data processing system, and a recording medium on which the data processing program is recorded.


