Distributed Execution Engine for Non-Uniform Data Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently processing large amounts of non-uniform datasets into structured datasets, requiring significant time, manpower, and funding, and often necessitating highly skilled teams to manage data preparation and processing across distributed computing systems.
Innovation Solution
A distributed execution engine is initialized to combine, interpret, and distribute non-uniform datasets into structured datasets, utilizing a data centric language for script files that allows for parallel execution across computational clusters, building a tree of execution with checkpoints, and generating structured datasets with standardized formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data processing is performed using distributed computing systems, then processing capacity and efficiency are improved, but the complexity of data preparation and system management increases significantly
Solution Approach 1:
The patent introduces a data preparation service as an intermediary component that sits between the data sources and the distributed computing system. This service automatically performs data validation, format conversion, and quality checks, eliminating the need for manual data preparation by skilled programmers while enabling the distributed system to process data efficiently.
Solution Approach 2:
The system implements self-service capabilities through automated data preparation services that can independently validate, transform, and prepare data without human intervention. The service automatically detects data issues, applies appropriate transformations based on predefined schemas, and prepares data for processing, allowing the system to serve itself rather than requiring external expert intervention.
2Measurement precision
If highly skilled teams are employed to manage data preparation, then data processing accuracy is improved, but time consumption and labor costs increase
Solution Approach 1:
The patent implements preliminary action by pre-defining data schemas, validation rules, and transformation logic before data processing begins. Data preparation services are configured in advance with knowledge graphs and relationship models that automatically guide the preparation process, eliminating the need for time-consuming manual analysis by skilled professionals while maintaining high accuracy standards.
Solution Approach 2:
The system replaces the mechanical process of manual data preparation by skilled programmers with an automated computational system. The data preparation service uses algorithms, validation rules, and predefined schemas to automatically perform tasks that previously required human expertise, dramatically reducing both time consumption and labor costs while maintaining or improving accuracy.
3Quantity of substance
If data is pulled from multiple independent sources without uniform methods, then data comprehensiveness is improved, but data integration difficulty increases
Solution Approach 1:
The patent implements a universal data preparation service that can handle multiple data sources with different formats, structures, and quality characteristics through a single unified interface. The service uses configurable schemas and transformation rules that can adapt to various data sources, eliminating the need for separate integration processes for each source while maintaining comprehensive data collection.
Solution Approach 2:
The system employs parameter changes by dynamically adjusting data transformation parameters based on the characteristics of each data source. The data preparation service can modify data formats, validation thresholds, and transformation logic according to the specific properties of incoming data, enabling seamless integration of diverse data sources without requiring uniform methods at the source level.
Data Source
AI summary
Systems, computer program products, and methods are described herein for a system for combining, interpreting, and distributing non-uniform datasets into structured datasets, the system comprising: a memory device with computer-readable program code stored thereon, a communication device, and a processing device operatively coupled to the memory device and the communication device. The processing device identifies a configuration file with location information for at least one dataset or script file, retrieves the dataset, builds a tree of execution comprising at least one node wherein each node comprises data, executes the script file on two or more datasets in parallel on a computational cluster comprising at least one computing component, wherein the at least one computing component comprises an at least one node and generates an output comprising as a structured dataset, wherein the structured dataset comprises a standardized format associated with the at least one instruction of the script file.


