Distributed Data Analysis Workflow Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analysis systems face inefficiencies in processing large and complex data sets from disparate sources, as they require moving data to a central location for analysis, leading to increased data analysis time.
Innovation Solution
A data analysis system that allows data sources to execute analysis operations locally by transmitting instructions for data operations directly to each data source, reducing data movement and enhancing computational speed through a visual objected-oriented data analysis tool and graphical user interface for creating data analytic workflows.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is retrieved from disparate data sources and combined in local memory for analysis, then data analysis can be performed using conventional systems, but data analysis time increases due to the large volume of data movement and computation required
Solution Approach 1:
Instead of the conventional approach where data is moved to a central analysis system, this patent inverts the process by sending analysis instructions to the data sources themselves. The data sources execute the analysis operations locally and return only the results, fundamentally reversing the data flow direction and eliminating the bottleneck of moving large volumes of data across the network.
Solution Approach 2:
The patent extracts the analysis computation from the central system and places it at the data sources. By separating the analysis function from the data storage location and executing it where the data resides, the system eliminates the need to move and store large datasets in local memory, thereby reducing data analysis time while maintaining analysis capability.
2Adaptability or versatility
If data from multiple disparate data sources is moved to a central location for analysis, then comprehensive data analysis can be performed, but data movement and storage resources are inefficiently utilized
Solution Approach 1:
The patent inverts the traditional data warehouse architecture by instead of centralizing data, it centralizes the analysis instructions while distributing the execution across multiple data sources. This approach maintains the ability to analyze data from diverse sources while eliminating the energy-intensive data movement and central storage requirements.
Solution Approach 2:
The system employs a universal instruction set that can be executed by different types of data sources (relational databases, NoSQL databases, data lakes, etc.). This multi-functional approach allows the same analysis workflow to operate across disparate data sources without requiring data movement or source-specific adapters, thereby improving data movement efficiency while maintaining broad data source integration capability.
3Reliability
If large data sets are retrieved and stored in local memory before analysis, then complete data analysis can be performed, but computational speed decreases due to the time required for data retrieval and processing
Solution Approach 1:
The patent applies preliminary action by having the data sources prepare and execute analysis operations in advance, before any data movement occurs. The analysis instructions are formulated and sent to the data sources, which then execute the computations locally and return results ready for immediate use, eliminating the sequential bottleneck of retrieve-then-analyze and enabling parallel processing across multiple data sources.
Solution Approach 2:
The patent inverts the conventional compute-where-data-resides model by implementing data-where-compute-should-happen. Instead of bringing data to the compute resource, the compute instructions are brought to the data, allowing analysis to be performed in-place at the data sources and significantly improving computational speed while maintaining complete analysis capability.
Data Source
AI summary
Embodiments are described for a system and method to analyze data at a plurality of data sources. A data analytic workflow may be received. The data analytic workflow may include at least one operation to be performed on a plurality of data sets stored at a plurality of data sources. Instructions may be created based on the operation to be performed and a type of platform that operates the data sources. Furthermore, the instructions may be transmitted to the data sources such that the data sources may execute the operations on the data sets stored at the data sources.


