Machine Learning Data Processing Model for Automated ETL

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing process of data analytics is labor-intensive and time-consuming due to the need for custom map-reduce programs to cleanse and structure unstructured data, which is exacerbated by the continuous changes in data schematics and usage patterns, making it difficult to create a standard data processing pipeline.

Innovation Solution

A method and apparatus that utilize machine learning to determine data patterns and create standard data processing models, allowing for the training of these models to reflect the patterns, thereby streamlining the ETL process by standardizing data processing across disparate systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If custom map-reduce programs are used to cleanse and structure unstructured data through disparate systems, then data processing can be performed, but the process becomes labor-intensive and time-consuming

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidtime for data preparation
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system employs machine learning models that automatically learn data patterns and schemas from the data itself, eliminating the need for manual custom programming. The models self-adjust to changing data schematics by continuously learning from incoming data streams, making the ETL process autonomous and adaptive without requiring continuous human intervention to create custom map-reduce programs.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent utilizes machine learning models that can dynamically adjust their parameters and structures based on the learned data patterns. This allows the system to adapt to changing data schematics by modifying model parameters rather than requiring complete reprogramming, thereby maintaining high productivity while reducing time consumption for data preparation.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If custom work is provided to create pipelines through disparate systems for each data source, then data can be processed, but the ETL process becomes the biggest obstacle and time-consuming area

Engineering Contradiction:
Improveability to handle different data sourcesVSAvoidcomplexity of ETL pipeline
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal machine learning-based data processing model that can handle multiple data sources and schemas through a single standardized interface. The model learns the specific patterns of each data source automatically, providing multi-functionality without requiring separate custom pipelines for each source, thus reducing ETL complexity while maintaining adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The machine learning model acts as an intermediary layer between disparate data sources and the analytics engine. This mediator automatically learns and adapts to different data schemas, translating various data formats into a standardized structure without requiring custom pipeline development for each source, thereby simplifying the overall system architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If standard data models are used to streamline the ETL process, then data processing can be standardized, but standard data models are very hard to figure out because schematics of the data changes continuously

Engineering Contradiction:
Improveease of creating data pipelineVSAvoidability to handle changing data schematics
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent employs dynamic machine learning models that continuously learn and adapt to changing data schematics rather than relying on static standard data models. The models evolve with the data by learning new patterns as they emerge, maintaining ease of pipeline creation while adapting to continuous changes in data structures without requiring manual model updates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms where the machine learning models continuously learn from incoming data and adjust their parameters accordingly. This feedback loop enables the models to automatically adapt to changing data schematics, maintaining standardized processing ease while responding to evolving data patterns without requiring manual intervention.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9324033B2Method and apparatus for providing standard data processing model through machine learning
Publication Date: 2016.04.26 NOKIA TECHNOLOGIES OY
  • US9324033B2 patent drawing
  • US9324033B2 patent drawing
  • US9324033B2 patent drawing

AI summary

An approach for providing a standard data processing model through machine learning is described. A machine learning data processing platform may process and/or facilitate a processing of the at least one data set associated with one or more computation closures to determine at least one data pattern. The machine learning data processing platform may also determine one or more data processing models associated with the one or more computation closures, the at least one data set, or a combination thereof. The machine learning data processing platform may further cause, at least in part, a training of the one or more data processing models to reflect the at least one data pattern.