Distributed Data Analysis Engine Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data analysis methods require manual intervention and are inefficient and error-prone, especially as business scale and complexity increase, making it difficult to automatically transfer data analysis results between various products and services.
Innovation Solution
A method and device for processing distributed data by integrating and configuring data analysis services into a distributed computing engine program, using a distributed scheduler to monitor message content and generate a data execution plan, allowing for automatic execution of data analysis services without manual intervention, thereby reducing business complexity and improving efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual intervention by data analysts is used for data analysis services, then flexibility in handling complex business requirements is maintained, but data analysis efficiency decreases and error rates increase
Solution Approach 1:
The system enables self-service automation where the distributed computing engine automatically executes data analysis services based on configured parameters and triggers, eliminating the need for manual analyst intervention in routine operations while maintaining the ability to handle complex requirements through configurable service parameters
Solution Approach 2:
Data analysis services are pre-configured with parameters, triggers, and execution logic before deployment. The distributed scheduler monitors predefined triggers and automatically initiates execution when conditions are met, allowing the system to respond to business requirements without manual intervention
2Adaptability or versatility
If multiple data analysis services are executed separately, then each service can be customized for specific requirements, but business complexity increases and transfer between services becomes difficult
Solution Approach 1:
The distributed computing engine provides a universal platform that can execute multiple different data analysis services through a common interface and configuration mechanism. Different services are distinguished by class files but managed through the same scheduling and execution framework, reducing operational complexity while maintaining service diversity
Solution Approach 2:
Different data analysis services are segmented into distinct class files within the analysis service data package, allowing each service to be independently configured and customized while being managed as part of a unified distributed execution system
3Productivity
If data analysis services are integrated into distributed computing engine, then automatic execution is enabled and efficiency improves, but system complexity increases
Solution Approach 1:
The distributed scheduler acts as an intermediary between the message middleware and the computing engine. It monitors triggers, manages execution plans, and coordinates service execution, thereby automating the process while managing system complexity through a dedicated coordination layer
Data Source
AI summary
Disclosed are a method and a device for processing distributed data. The method includes: integrating and configuring data analysis services of multiple users with different data analysis requirements into a distributed computing engine program to obtain an analysis service data package; configuring a distributed scheduler in the cluster server according to the analysis service data package, and calling the distributed scheduler to monitor a message content transmitted by a message middleware including multiple data analysis services to be executed; and generating a distributed data execution plan according to the message content, and performing distributed scheduling calculation on the distributed data execution plan to obtain a distributed calculation result.


