Structured Data Processing Workflow Engine
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for structured big data processing struggle to balance processing speed, resource utilization, flexibility, scalability, and configurability, often requiring multiple programs, being platform-specific, and not adaptable to varying computing resources or new data formats, leading to inefficiencies and reduced dataset integrity.
Innovation Solution
A method and system that processes structured data by accessing a dataset, pre-processing it to generate reference and metadata files, dynamically allocating processing tasks to computer nodes, and performing data processing with polymorphic plugin connections and automatic workflow distribution, enabling efficient handling of large datasets across diverse computing platforms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If existing solutions process structured big data, then processing speed can be improved, but resource utilization deteriorates due to inability to adapt to varying computing resources
Solution Approach 1:
The system dynamically adapts processing parameters and resource allocation based on available computing resources. The workflow engine adjusts task distribution, parallelization degree, and processing strategies in real-time to match resource availability, preventing both over-provisioning and under-utilization of computing power.
Solution Approach 2:
The system changes processing parameters such as block size, memory allocation, and computational complexity based on detected resource conditions. When resources are abundant, more aggressive processing parameters are used; when resources are constrained, parameters are adjusted to maintain efficiency while reducing resource demands.
2Productivity
If existing solutions are designed for specific platforms, then processing efficiency can be optimized, but adaptability to new platforms deteriorates
Solution Approach 1:
The system employs a universal workflow engine that can execute processing tasks across diverse computing platforms including supercomputers, cloud environments, and edge devices. The platform-agnostic architecture allows the same processing workflows to be deployed efficiently on different hardware configurations without requiring platform-specific redesign.
Solution Approach 2:
The processing system is segmented into independent, modular workflow components that can be individually configured and executed. This segmentation allows the system to adapt to different platform capabilities by selecting and combining appropriate processing modules for each specific computing environment.
3Adaptability or versatility
If multiple programs are used to process different data aspects, then functional completeness can be improved, but system complexity deteriorates
Solution Approach 1:
The system merges multiple processing functions into a single integrated workflow engine that can handle diverse data processing tasks through configurable workflow definitions. Instead of requiring separate programs for different processing aspects, the unified engine executes coordinated sequences of processing steps defined in workflow files, reducing system complexity while maintaining functional completeness.
Solution Approach 2:
The workflow engine acts as an intermediary layer between data sources and processing algorithms. It manages the coordination, data flow, and execution of multiple processing tasks without requiring direct integration between individual processing programs, thereby simplifying the overall system architecture.
4Speed
If datasets are thinned or dissociated to fit processing capacity, then processing speed can be improved, but data integrity deteriorates
Solution Approach 1:
The system performs preliminary assessment of available computing resources and processing capacity before initiating data processing. Based on this pre-evaluation, it configures appropriate processing strategies that can handle complete datasets without requiring thinning or dissociation, thereby maintaining data integrity while optimizing processing speed for the given resource constraints.
Data Source
AI summary
Described herein is a method and system for flexible, high performance structured data processing. The method and system contains techniques for balancing and jointly optimising processing speed, resource utilisation, flexibility, scalability, and configurability in one workflow. A prime example of its application is the analysis of spatial data, e.g. LiDAR and imagery. However, the invention is applicable to a wide range of structured data problems in a variety of dimensions and settings.


