一种大数据批处理任务运行时间的预测方法和装置
By breaking down big data batch processing tasks into operations and combining similarity and causal relationships, and using custom benchmark programs and benchmark data to generate runtime logs, machine learning models are trained, solving the problem of traditional methods relying on historical logs and achieving high-precision runtime prediction and resource optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ACT TECH DEV CO LTD
- Filing Date
- 2023-08-21
- Publication Date
- 2026-07-17
AI Technical Summary
Existing methods for predicting the runtime of big data batch processing tasks rely on historical log data, which cannot accurately predict in the initialization environment. Furthermore, traditional methods do not consider all factors, resulting in insufficient prediction accuracy.
Batch processing tasks are broken down into multiple operations. By combining similarity and causal relationships, and running them on the target environment using a custom benchmark program and benchmark data, runtime log data is generated. Machine learning algorithms are then used to train a model to predict task runtime.
It improves the prediction accuracy of big data batch processing task execution time, ensures the rationality of resource allocation and the minimization of execution time, and achieves a prediction accuracy of over 95%.
Smart Images

Figure CN117271281B_ABST