Modular Data Preprocessing System for Automated Cleaning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual data preprocessing consumes a significant portion of time in data projects, and existing methods are not efficient for automating data cleaning across various types of data, leading to inaccurate and incomplete analyses.

Innovation Solution

The development of a modular data preprocessing system that utilizes Natural Language Processing (NLP) and Machine Learning (ML) methods to automate data cleaning tasks, including data agnostic preprocessing, allowing for minimal human intervention and providing automatic analytical results such as spell checking, clustering, and outlier detection through a graphical user interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data preprocessing is performed, then data cleaning accuracy can be maintained through human review, but time consumption increases significantly (up to 80% of total project time)

Engineering Contradiction:
Improvedata cleaning accuracyVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical data cleaning processes with an automated computational system that uses machine learning models and natural language processing to perform data preprocessing tasks, thereby eliminating the trade-off between accuracy and time consumption

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service data cleaning through automated anomaly detection, duplicate identification, and data validation mechanisms that operate without human intervention, maintaining high accuracy while dramatically reducing preparation time

Inventive Principle:
Principle #25Self-service

2Productivity

If existing automated data cleaning methods are used, then time consumption is reduced, but they lack efficiency and adaptability across various types of data

Engineering Contradiction:
Improvedata processing efficiencyVSAvoiddata type flexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal data preprocessing system that can handle multiple data types (structured, unstructured, semi-structured) and formats through a single integrated platform using NLP and machine learning techniques, achieving both high productivity and broad adaptability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adjusts processing parameters and algorithms based on the characteristics of the input data, allowing it to efficiently adapt to different data types and formats while maintaining high processing speeds

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If comprehensive data analysis is performed, then analytical depth and insights are improved, but data preparation time increases

Engineering Contradiction:
Improveanalytical depthVSAvoidfront-end preparation time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs comprehensive data cleaning, validation, and transformation in advance through automated preprocessing pipelines, so that when analysis begins, the data is already optimized for deep analytical processing, thereby achieving both analytical depth and time efficiency

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220342583A1Methods and apparatus for data preprocessing
Publication Date: 2022.10.27 THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY OF THE NAVY
  • US20220342583A1 patent drawing
  • US20220342583A1 patent drawing
  • US20220342583A1 patent drawing

AI summary

Methods, apparatus, and/or computer program products for data preprocessing are provided. The invention utilizes selective modular processes that automate data cleaning tasks and provide an initial analysis of datasets using, among other things, Natural Language Processing and Machine Learning methods. The apparatus and methods preprocess text data, clean database data with minimal human intervention, and provide automatic analytical results that include basic data cleaning, spell check, clustering, outlier detection, and natural language processing. Each modular process can be selectively switched on or off based on user input, preprogrammed instructions, or AI inputs. Additionally, any data format may be input as the apparatus and methods are data agnostic and, therefore, are not designed specifically for one type of data or theme of data. The inventive apparatus and methods dramatically reduce preparation time on the front end of data projects.