A Batch Application Eviction Method Based on Spark Cost Awareness under Hybrid Deployment

By monitoring the recalculation and remaining time cost of Spark computing tasks, as well as the resource usage requirements of LC applications, and optimizing the eviction strategy, the recalculation problem of Spark tasks under hybrid deployment is solved, and the resource utilization and throughput of the data center is improved.

CN115562858BActive Publication Date: 2025-07-04UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211180785.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-07-04
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

In hybrid deployment data centers, existing eviction strategies may cause Spark computing tasks to be recalculated, resulting in a decline in overall data center throughput, and cannot effectively solve the resource competition problem when LC application burst loads.

Method used

By monitoring the recalculation cost, remaining time cost and resource usage requirements of the Spark computing task, the Spark computing task with the smallest eviction cost is used to evict, and the resource utilization and throughput are optimized.

Benefits of technology

It improves the resource utilization and overall throughput of the data center, is compatible with existing resource management platforms, does not require modification of LC and Spark applications, and is versatile.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115562858B_ABST
    Figure CN115562858B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of cloud computing technology, and discloses a batch processing application eviction method based on Spark cost awareness under hybrid deployment, including defining the recalculation cost of Spark computing tasks, estimating the remaining time cost, and predicting the resource usage requirements of LC applications; First, since the present invention only involves the modification of the container orchestration platform and is compatible with existing Spark applications and all LC applications, it has good versatility; And since the algorithm adopts a trigger type, the normal components only consume extremely small amounts of resources and have almost no impact on LC and BE applications; Finally, since the recalculation cost can be defined as two parts, the computing cost and the transmission cost, and the changes in the resource usage requirements of LC can be dynamically sensed, and a large number of customizable parameters are provided, this method is applicable to various scenarios and has good versatility; The present invention is applicable to the situation where memory contention occurs due to the increase in LC application load under hybrid deployment, improving the resource utilization rate and the total throughput of the cluster at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and particularly to a method for evicting batch applications based on Spark cost awareness under hybrid deployment. Background Art

[0002] Data centers are the foundation of cloud computing services. Applications in data centers can be mainly divided into two categories. One category is latency-critical (LC) applications represented by online services (such as Nignx, MongoDB, etc.). LC applications consume low resources most of the time, but will experience bursty loads and consume a large amount of resources during a small period of time. When resources are insufficient, request latency will increase, resulting in a decline in user experience. Another category of applications is best-effort (BE) applications represented by batch applications (such as Spark, MapReduce, etc.). BE applications usually consume a large amount of resources during operation, but can tolerate higher latency and task restart due to errors. Mixing LC applications and BE applications in the same cluster has become the mainstream method to improve the resource utilization rate of data centers. However, hybrid deployment will lead to resource competition and cause a decline in application performance. When LC faces bursty loads, in order to ensure the quality of service (QoS) target of LC applications, low-priority BE applications are usually evicted. Existing eviction strategies usually mainly rely on random eviction or evicting the BE application that occupies the most resources. However, big data processing tasks represented by Spark computing tasks have the characteristic of multi-stage computing. The above eviction methods may cause tasks to be recalculated, resulting in a longer completion time of computing tasks and a decline in the overall throughput of the data center. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a method for evicting batch applications based on Spark cost awareness under hybrid deployment, which can improve the maximum performance and throughput of the data center. The method in the present invention has good versatility and can cooperate with mainstream resource management platforms on the market to monitor data, calculate the optimal eviction object, and achieve the best overall eviction effect.

[0004] To solve the above technical problems, the present invention adopts the following technical solutions:

[0005] First step, quantify the recalculation cost of each Spark computing task:

[0006] By monitoring the information recorded during the operation of each Spark computing task, the total number of stages X, the total number of partitions N of each stage i and the number of completed partitions C i, where i is the stage number, then the completion rate R of stage i i = C i / N i ). The present invention defines S i,j as the data size of the j-th partition in stage i, T i,j as the completion time of the j-th partition in stage i, and F i as the data size transmitted during the shuffle stage (between the i-th stage and the (i + 1)-th stage).

[0007] The recalculation cost of the Spark computing task is strongly correlated with the following factors: the amount of recalculated data, the operator complexity, and the transmission cost. The amount of recalculated data is the sum of the partition data amounts of all calculations within the completed stages in the Spark computing task. The operator complexity is the average processing duration per unit data of all partitions of the calculations within the completed stages in the Spark computing task. The transmission cost is the amount of data transmission completed within the stage.

[0008] Among them, the amount of recalculated data in the i-th stage is The operator complexity of the i-th stage is The transmission cost R of the i-th stage i+1 F i .

[0009] Based on the above analysis, the present invention defines the following total calculated cost, that is, the calculation cost

[0010]

[0011] and the transmission cost

[0012]

[0013] Finally, by assigning weights α and β to the calculation cost and the transmission cost respectively, the recalculation cost

[0014] RecalculationCost = α·ComputationCost + β·TransmissionCost.

[0015] Second step, estimate the remaining time cost of each Spark computing task:

[0016] The entire calculation process of the Spark computing task includes completed stages, the current processing stage, and unprocessed stages; the present invention first calculates the completed workload of the Spark computing task where x - 1 represents the number of completed stages, r x represents the completion rate of the current stage, and x + y represents the total number of stages; according to the record of the calculation time (i.e., CalculatedTime) of the Spark computing task, estimate the remaining time cost

[0017]

[0018] Among them, the positive coefficient γ is set according to the workload complexity in different stages.

[0019] Step 3: Predict the resource usage requirements of the LC application:

[0020] The present invention predicts the resource usage requirements of the LC application in the next period by monitoring the resource consumption of the LC application within a past time window (window size is w); since most evictions are caused by memory contention, the present invention monitors the memory usage of the LC application regularly (once every p minutes), and uses this information for eviction prediction. w monitoring samples are taken in the time window with window size w. Assuming the sampled memory usage amounts are M1, M2, …, M w , the average memory usage amount can be obtained and the variance where the variance Var is used to predict the memory usage trend of the LC application.

[0021] Step 4: Select the Spark computing task with the minimum eviction cost for eviction:

[0022] The present invention determines and evicts the Spark computing task with the minimum eviction cost according to the recalculation cost, remaining time cost of the Spark computing task obtained in the first three steps, and the resource usage requirements of the LC application. When an eviction occurs (host resources are insufficient), calculate the eviction cost of each Spark computing task. The eviction cost

[0023] EvictionCost = RecalculationCost - δ·RTCost;

[0024] where, δ represents the standard deviation, that is, δ 2 = Var. If δ is less than the set value, that is, the growth rate of the resource usage requirements of the LC application is small, the weight of the recalculation cost is larger; if δ is greater than or equal to the set value, that is, the resource usage requirements of the LC application increase sharply, the weight of the remaining time cost is larger. Once the system detects resource contention, eviction will be triggered.

[0025] Therefore, the specific approach is to first calculate the recalculation cost and remaining time cost of Spark computing tasks, as well as predict the memory requirements of LC applications. Then, by calculating the eviction cost of each Spark computing task and selecting the Spark computing task with the minimum eviction cost for eviction. This process will continue until the resource usage requirements of the LC application are met. In particular, the present invention preferentially evicts the worker containers of Spark computing tasks, and then the master containers of Spark computing tasks, so that after a part of the worker containers are evicted, the Spark computing tasks can still continue to run.

[0026] Compared with the prior art, the beneficial technical effects of the present invention are:

[0027] The method for evicting batch applications based on Spark cost awareness of the present invention includes steps such as Spark task recalculation cost, Spark task remaining time cost, prediction of LC application resource usage requirements, and minimum eviction cost. First, since the present invention does not involve modifications to the resource management platform and the underlying code of containers, it can be compatible with a variety of LC applications, so it has good versatility. And since there is no need to know in advance the entire process resource consumption of BE applications and LC applications, the main overhead is to monitor the shuffle size and time processing information of BE jobs, as well as the memory usage information of LC applications, and the impact of the monitoring overhead on the performance of LC and BE jobs can be ignored. Compared with traditional eviction strategies, this method can improve the overall resource utilization rate and system throughput. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a schematic diagram of the operation process of the method for evicting batch applications based on Spark cost awareness of the present invention;

[0029] Figure 2 is a schematic diagram of the Spark architecture and process division;

[0030] Figure 3 is a schematic diagram of the definition of Spark recalculation cost;

[0031] Figure 4 is a schematic diagram of the estimation of Spark remaining time cost;

[0032] Figure 5 is a schematic diagram of estimating the resource usage requirements of LC applications according to the resource usage requirement diagram;

[0033] Figure 6 is a schematic diagram of implementing the architecture on the container orchestration platform. DETAILED DESCRIPTION OF THE INVENTION

[0034] The following describes in detail a preferred embodiment of the present invention with reference to the accompanying drawings.

[0035] The following further elaborates in detail the eviction method based on the recalculation cost of Spark computing tasks of the present invention through specific embodiments in conjunction with the accompanying drawings.

[0036] Embodiment 1

[0037] In this embodiment, the eviction method for batch processing applications based on Spark cost awareness will illustrate the specific implementation manner according to the specific running situation of a Spark computing task and the actual resource usage requirements of the foreground LC application.

[0038] Figure 2 The figure shows the division of the cluster architecture and computing tasks in a Spark application. A Spark application consists of a master node and several worker nodes. An executor process will be started in each worker node, and the executor process will start several task threads according to the allocated resources to be responsible for this computing task. After submitting the computing task to the Spark application, the master node will divide this computing task into multiple stages. The basis for the division of stages is the dependency between partitions. A partition is the smallest data structure in Spark. When a partition depends on all its parent partitions, it is called a wide dependency. A large number of shuffle operations will be triggered during the calculation process of a wide dependency. Each time a shuffle operation occurs, a new stage will be divided; when a partition only depends on one partition, it is called a narrow dependency. For a partition within a stage, all its calculations are narrow dependency calculations. The eviction method for batch processing applications based on Spark cost awareness in this embodiment specifically includes the following steps:

[0039] Step 1: Quantify the recalculation cost of each Spark computing task:

[0040] Figure 3 The figure shows the specific execution process of a Spark computing task that is being calculated. The map-then-flatten (flatMap) and filter operations for each data partition are completed by a task thread using a logical core, and then the intermediate stage is written to the local disk. Figure 3The left part shows that two data partitions have completed all narrow dependency calculations within this stage. The recomputation data volume is the sum of the data volumes of the two data partitions. The operator complexity can be expressed as the average processing duration per KB of data for these two partitions, and the product of the two is the computational cost. After all the calculations in the previous stage are completed, the shuffle operation between stages will cause a large amount of data transmission, and its data transmission volume is the transmission cost. The recomputation cost of a Spark computing task consists of two parts: computational cost and transmission cost. The weights of the computational cost and the transmission cost are determined by the specific state of the current server.

[0041] Step 2: Estimate the remaining time cost of each Spark computing task:

[0042] As Figure 4 shown, define the stages of a Spark computing task as three types: one is the stage that has completed the calculation, the second is the stage that is currently being calculated, and the third is the stage that has not been calculated. Figure 4 There are a total of three stages. One stage has been completed, the stage that is currently being calculated has completed 50% of the calculation, and there is still one stage that has not started to be calculated. It can be defined that Figure 4 the Spark computing task in has completed 1.5 stages, accounting for 50% of the overall progress. Therefore, the remaining time is γ times the calculated time. The parameter γ here is positively correlated with the complexity of the subsequent calculation.

[0043] Step 3: Predict the resource usage requirements of the LC application:

[0044] Figure 5 is the resource usage requirement diagram of the LC application. The resource usage requirements of the LC application are recorded every minute. Select two time windows in the diagram for analysis. The standard deviation within time window 1 (window 1) is small, and the required resources only increase slightly within a period of time. It can be predicted that this LC application will be relatively stable in the subsequent period of time and only requires a small amount of resource improvement in the long term. While the standard deviation within time window 2 (window 2) is large, it can be predicted that this application will require a large amount of resources to meet the service quality in the subsequent period of time.

[0045] Step 4: Select the Spark computing task with the minimum eviction cost for eviction:

[0046] Based on the results obtained from the previous three steps, the eviction cost of each Spark computing task can be calculated. As Figure 6As shown in the figure, the resource monitoring component monitors whether resource contention occurs. When resource contention occurs, the eviction controller obtains the eviction cost of each Spark computing task from the monitoring component, the computing monitoring component, and the transmission monitoring component, and performs eviction at the container granularity until the resource contention disappears. Figure 1 The implementation operation process of the entire process is given.

[0047] In the eviction method for batch processing applications based on Spark cost awareness in this embodiment, on the one hand, by defining the recalculation cost of each Spark computing task, estimating the remaining time cost, and predicting the resource usage requirements of LC applications, the eviction cost of each computing task is obtained, so that the overall recalculation cost can be minimized under different resource usage requirements of LC applications, improving the resource utilization rate and throughput of the cluster; on the other hand, only the eviction method of the container scheduling platform is modified without modifying the LC application and the Spark application itself, and the versatility is relatively good.

[0048] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0049] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A batch application eviction method based on Spark cost awareness under hybrid deployment, comprising the following steps: Step 1: Quantify the recomputation cost of each Spark computing task: The recomputation cost of a Spark computing task is equal to the weighted sum of the computing cost and the transmission cost; where the computing cost is equal to the product of the recomputed data volume and the operator complexity, the recomputed data volume is the sum of the partition data volumes of all computations in the completed stages of the Spark computing task, and the operator complexity is the average processing duration per unit data of all partitions of the computations in the completed stages of the Spark computing task; the transmission cost is the data transmission volume completed within the stage; Step 2: Estimate the remaining time cost of each Spark computing task: The computing process of a Spark computing task includes completed stages, the current processing stage, and unprocessed stages. Obtain the completed workload of the Spark computing task based on the number of completed stages, the current processing stage, and unprocessed stages, and estimate the total computing time and the remaining time cost RTCost of the Spark computing task in combination with the completed workload: Among them, JobProgress represents the completed workload of the Spark computing task, CalculatedTime represents the computing time of the Spark computing task, and the positive coefficient γ of the remaining time cost is set according to the workload complexity of different stages; Step 3: Predict the resource usage requirements of latency-sensitive applications: Predict the resource usage requirements of latency-sensitive applications in the next period by monitoring the resource consumption of latency-sensitive applications within a past time window; Step 4: Select the Spark computing task with the minimum eviction cost for eviction Determine and evict the Spark computing task with the minimum eviction cost based on the recomputation cost and the remaining time cost of the Spark computing task; when determining the Spark computing task with the minimum eviction cost, set the weights of the recomputation cost and the remaining time according to the resource usage requirements of latency-sensitive applications; Hybrid deployment means that latency-sensitive applications and best-effort applications are deployed in the same cluster.

Citation Information

Patent Citations

  • Spark distributed computing data processing method and system

    CN107526546A

  • Dynamic load balancing of operations for real-time deep learning analytics

    US20220035684A1