物理计划生成方法、装置、设备及存储介质

By constructing dataset dependencies and generating full-order physical plans using full-order code optimization algorithms, replacing traditional caching methods and applying lazy execution mode, the problem of slow execution and long resource consumption time of batch processing jobs in cloud Spark clusters is solved, achieving efficient execution of SQL jobs.

CN117591533BActive Publication Date: 2026-07-17CHINA MERCHANTS BANK

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MERCHANTS BANK
Filing Date
2023-11-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Cloud-based Spark clusters experience slow batch processing jobs with long resource consumption times. This is especially true for systems with long chains, large and complex relationships, where resource consumption is high, running time is long, and trial-and-error costs are high, affecting the timely completion of important downstream systems.

Method used

By obtaining a set of SQL statements, the dataset dependencies are constructed based on preset subquery instructions, and a full-order physical plan is generated using a pre-built full-order code optimization algorithm to replace the Cache Table command. The caching method of select + temporary table/view name is adopted, the show function is canceled, the Catalog API is called to register cache information, the cache level and method are configured, and the lazy execution mode is applied to build the full-process execution plan.

Benefits of technology

It enables batch execution of SQL jobs, reducing time consumption and resource usage, solving the problem of slow batch processing jobs in cloud Spark clusters, and improving resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117591533B_ABST
    Figure CN117591533B_ABST
Patent Text Reader

Abstract

本发明公开了一种物理计划生成方法、装置、设备及存储介质,属于计算机技术领域。该方法包括:获取SQL语句集合;基于预设的子查询指令,根据所述SQL语句集合,构建数据集依赖关系;基于预先构建的全阶代码优化算法,根据所述数据集依赖关系,获得全阶物理计划。通过子查询指令和全阶代码优化算法,得出了全阶物理计划。由此,实现了SQL作业的批量执行,解决了现有技术中云上Spark集群批处理作业运行缓慢、资源占用时间长的技术问题。相较于现有技术,具有耗时短、资源消耗少的优势。
Need to check novelty before this filing date? Find Prior Art