Task resource allocation method and related device
By predicting the target data volume and resource volume of tasks in the computing cluster and using a regression model to optimize resource allocation, the problem of unreasonable resource allocation in the computing cluster is solved, and the task execution efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202410405255.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2025-10-14
AI Technical Summary
The existing dynamic resource allocation scheme for computing clusters leads to irrational resource allocation, resulting in long task waiting times, resource shortages and high task failure rates.
By obtaining the data screening conditions of the task and the total data volume of the computing cluster, the target data volume and resource volume are predicted. The regression model is used to predict the target resource volume required for task execution, and the appropriate resource volume is specified when the task is delivered to avoid resource waste.
It improves the efficiency of computing cluster task execution, ensures that each task is executed efficiently under reasonable resource allocation, and reduces task waiting time and resource waste.
Smart Images

Figure CN120780447A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cluster computing technology, and in particular to a method for allocating task resources and related devices. Background Art
[0002] With the advancement of technology, vast amounts of data are generated constantly across all sectors of society, such as transaction data in e-commerce, multimedia data in streaming media, and patient data in healthcare. In most fields, users often desire to analyze and process the vast amounts of data maintained in databases. Computing clusters have emerged to facilitate large-scale data analysis and processing. Spark, a typical computing cluster, is a powerful distributed computing framework capable of processing large datasets and accomplishing various tasks.
[0003] In related technologies, when executing tasks based on Spark computing clusters, dynamic resource allocation schemes are often used to automatically allocate resources for tasks. In this dynamic resource allocation scheme, the scheduling mechanism allocates as many resources as possible to the tasks in the queue in order to complete them faster, which can easily lead to resource waste.
[0004] In general, existing dynamic resource allocation solutions in the industry suffer from irrational resource allocation, which can lead to long task wait times and a shortage of cluster computing resources, ultimately resulting in high task failure rates. Therefore, there is an urgent need for a better resource allocation solution to address this issue of irrational computing resource allocation in computing clusters. Summary of the Invention
[0005] This application provides a task resource allocation method that can improve the rationality of task resource allocation, thereby improving the efficiency of computing clusters when executing tasks.
[0006] In a first aspect, the present application provides a method for allocating resources for tasks to be executed in a computing cluster. The method comprises: first, obtaining a first task, the first task being used to instruct the computing cluster to process data that meets a data screening condition, so that the computing cluster can clearly determine which portion of the data in the data storage system is processed when executing the first task.
[0007] Then, based on the data screening conditions and the total amount of data accessed by the computing cluster, the target data volume that the computing cluster needs to process when executing the first task is predicted. That is, the target data volume indicates the total amount of data that the computing cluster needs to process when executing the first task.
[0008] Next, based on the target data volume, a target resource volume required for executing the first task is predicted, i.e., the processing resources that the computing cluster needs to allocate to the first task. The target resource volume indicates the amount of processor resources and / or memory, i.e., at least one of the two. Furthermore, the target resource volume predicted based on the target data volume is a resource volume that maximizes the execution efficiency of the first task. That is, when processing resources are allocated to the first task according to the target resource volume, the execution efficiency of the first task can be maximized.
[0009] Finally, a task processing request is sent to the computing cluster. The task processing request is used to request the computing cluster to process the first task and specify the amount of resources allocated by the computing cluster to the first task as the target resource amount, ensuring that the computing cluster can allocate processing resources to the first task according to the specified resource amount, and avoiding the computing cluster allocating too many resources to the first task.
[0010] In this solution, the target data volume involved in task execution is predicted based on the data filtering conditions specified in the task and the total amount of data accessed by the computing cluster. This in turn predicts the target resource volume required for task execution based on the target data volume. This allows you to specify the target resource volume required when delivering a task to the computing cluster, ensuring that the computing cluster allocates the appropriate amount of resources for the task, ensuring that each task executes with the highest possible efficiency, and avoiding the impact of improper task resource allocation on normal execution. This effectively improves the task execution efficiency of the computing cluster.
[0011] In one possible implementation, the method further includes determining, based on the data screening conditions, a data redistribution operation involved in executing the first task, where the data redistribution operation is used to redistribute data processed by multiple computing nodes in the computing cluster. Because the first task actually processes data that meets the data screening conditions, after the data that actually needs to be processed is determined based on the data screening conditions, it can be determined whether the data redistribution operation is necessary.
[0012] When predicting the target amount of resources required for executing the first task, the target amount of resources required for executing the first task can be specifically predicted based on the target data volume and the data redistribution operation. Generally, if the task processing involves data redistribution, the amount of resources required for the task will tend to increase compared to a task with the same data volume but not involving data redistribution. That is, if the task involves data redistribution, the task processing process will tend to require more resources to complete.
[0013] In this solution, whether the task execution process involves data redistribution operations is determined based on the data screening conditions indicated by the task, and the amount of resources required for the task execution process is further predicted based on the amount of data involved in the task and the data redistribution operations, which can effectively improve the accuracy of the predicted resource amount.
[0014] In one possible implementation, a regression model can be used to predict the target resource requirements for executing the first task based on the target data volume and the data redistribution operation. The regression model is fitted based on the computing cluster's historical task processing data. By substituting the target data volume and the data redistribution operation into the fitted regression model, the target resource requirements for executing the first task, as output by the regression model, can be obtained.
[0015] Specifically, the computing cluster's historical task processing data can include the actual data volume involved in tasks executed during the historical time period, the data redistribution operations involved in the tasks, and the resource allocation during task execution. Therefore, based on the computing cluster's historical task processing data, we can obtain the relationship between the data volume involved in the tasks, the data redistribution operations involved in the tasks, and the resource allocation during task execution, thereby fitting the corresponding regression model.
[0016] In this solution, by fitting the regression model using the historical task processing data of the computing cluster, the regression model can be effectively used to characterize the relationship between the amount of data involved in the task, the data redistribution operation, and the amount of resources required for the task, thereby accurately and effectively predicting the amount of resources required for task execution.
[0017] In one possible implementation, the task processing efficiency corresponding to the historical task processing data satisfies a preset condition, for example, the task processing efficiency is greater than or equal to a preset threshold. The task processing efficiency is related to the amount of task data, the amount of resources allocated to the task, and the execution time of the task.
[0018] In this solution, by fitting the regression model by screening task processing data corresponding to higher execution efficiency, it can be ensured that when resources are allocated to the first task based on the resource amount predicted by the regression model, the first task can also have a higher execution efficiency, thereby ensuring the rationality of resource allocation.
[0019] In one possible implementation, the computing cluster includes multiple task queues for processing tasks, and the task processing request further specifies a target queue for processing the first task, where the target queue is one of the multiple task queues. In other words, the task processing request simultaneously specifies the target amount of resources to be allocated to the first task and the target queue for processing the first task. Thus, when processing the first task, the computing cluster places the first task in the target queue and allocates the target amount of resources to execute the first task.
[0020] In this solution, when a computing cluster includes multiple task queues, by specifying a target queue for processing the first task, the loads of the multiple task queues can be kept balanced as much as possible, avoiding the situation where some task queues are overloaded while others are underloaded, thereby ensuring that the resources in the computing cluster can be reasonably utilized.
[0021] In one possible implementation, to determine the target queue corresponding to the first task, the method further includes the following steps: based on the resource occupancy of each of the multiple task queues at the current time, retrieving, for each task queue, multiple target samples with the closest resource occupancy from historical samples of the multiple task queues, where the historical samples are used to indicate the resource occupancy of the task queue in the historical time period. The multiple target samples corresponding to each task queue all belong to the historical samples of each task queue.
[0022] Then, based on the multiple target samples corresponding to each task queue, the expected resource usage of each task queue in the future is predicted. By averaging the resource usage of multiple target samples for a certain period of time after each, a resource usage figure is obtained, which can be used as the expected resource usage of the task queue in the future.
[0023] Secondly, based on the expected resource usage and the target data volume, a target queue is determined from among the multiple task queues for processing the first task. Since other tasks may need to be delivered simultaneously when the first task is delivered, the task queue corresponding to each task can be determined based on the amount of resources required to be allocated to each task and the expected resource usage of each task queue.
[0024] In this solution, the task queue to which the first task is assigned is determined by predicting the resource occupancy of multiple task queues in the computing cluster in the future. This can effectively ensure that the loads between the task queues are in a relatively balanced state after task assignment, thereby effectively utilizing the resources of the computing cluster.
[0025] In a possible implementation, the task processing request is further used to indicate a processing priority of the first task, where the processing priority is determined based on the target resource amount.
[0026] Specifically, the smaller the amount of resources a task requires, the smaller the resources consumed when the task is running, and the faster the task will be completed. Therefore, the priority of tasks that require less resources can be increased so that the tasks can be completed as soon as possible to free up more resources to run other tasks.
[0027] In this solution, by determining the priority of tasks based on the amount of resources required for the tasks, the execution priority of each task in the task queue can be clarified, ensuring that the task queue executes tasks in order of execution priority when processing them, and meeting the diverse needs of users as much as possible.
[0028] In one possible implementation, the processing priority is also related to the expected runtime of the first task, which is determined based on the target data volume and the target resource volume. That is, the processing priority is related to both the target resource volume and the expected runtime of the first task.
[0029] Specifically, when the target data volume and target resource volume corresponding to the first task have been determined, the expected running time of the first task can be predicted based on the regression model fitted in advance to determine the processing priority of the first task.
[0030] In one possible implementation, the first task is a task that is periodically executed in the computing cluster. For example, in the financial field, the first task may be a sales data analysis task that requires the computing cluster to periodically extract and analyze past daily, weekly, or monthly sales data.
[0031] In a possible implementation, the first task is a data query task based on Structured Query Language (SQL), and the data screening conditions include one or more of customized data selection conditions, data statistical fields, data statistical dimensions, and data screening time periods.
[0032] A second aspect of the present application provides a task resource allocation device, comprising: an acquisition module for acquiring a first task, the first task being used to instruct a computing cluster to process data that meets a data screening condition; a processing module for predicting a target amount of data to be processed by the computing cluster when executing the first task based on the data screening condition and the total amount of data accessed by the computing cluster; the processing module is also used to predict a target amount of resources required for executing the first task based on the target amount of data, the target amount of resources being used to indicate the amount of processor resources and / or memory; the processing module is also used to send a task processing request to the computing cluster, the task processing request being used to request the computing cluster to process the first task and specifying the amount of resources allocated by the computing cluster to the first task as the target amount of resources.
[0033] In one possible implementation, the processing module is further used to: determine, based on the data screening conditions, a data redistribution operation involved in the execution of the first task, where the data redistribution operation is used to redistribute the data processed by multiple computing nodes on the computing cluster; and predict, based on the target data volume and the data redistribution operation, a target amount of resources required for the execution of the first task.
[0034] In one possible implementation, the processing module is further used to: predict the target amount of resources required for executing the first task through a regression model based on the target data volume and the data redistribution operation; the regression model is obtained by fitting based on historical task processing data of the computing cluster.
[0035] In a possible implementation, the task processing efficiency corresponding to the historical task processing data meets a preset condition, and the task processing efficiency is related to the data volume of the task, the amount of resources allocated to the task, and the execution time of the task.
[0036] In a possible implementation, the computing cluster includes multiple task queues for processing tasks. The task processing request is further used to specify a target queue for processing the first task, where the target queue is one of the multiple task queues.
[0037] In one possible implementation, the processing module is further used to: based on the resource occupancy of each task queue in a plurality of task queues at the current time, retrieve, for each task queue, a plurality of target samples with the closest resource occupancy from historical samples of the plurality of task queues, where the historical samples are used to indicate the resource occupancy of the task queue in a historical time period; based on the plurality of target samples corresponding to each task queue, predict the expected resource occupancy of each task queue in the future; and based on the expected resource occupancy and the target data volume, determine a target queue among the plurality of task queues for processing the first task.
[0038] In a possible implementation, the task processing request is further used to indicate a processing priority of the first task, where the processing priority is determined based on the target resource amount.
[0039] In a possible implementation, the processing priority is further related to an expected running time of the first task, where the expected running time is determined based on a target data volume and a target resource volume.
[0040] In a possible implementation, the first task is a task that is periodically executed in the computing cluster.
[0041] In a possible implementation, the first task is a SQL-based data query task, and the data screening conditions include one or more of a customized data selection condition, a data statistical field, a data statistical dimension, and a data screening time period.
[0042] A third aspect of the present application provides a computing device comprising a processor and a memory. The processor of the computing device is configured to execute instructions stored in the memory of the computing device, causing the computing device to perform the method described in the first aspect or any implementation of the first aspect. For details regarding the steps in each possible implementation of the first aspect performed by the computer device, please refer to the first aspect and will not be repeated here.
[0043] In a fourth aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device including a processor and memory. The processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method described in the first aspect or any of the implementations of the first aspect. For details regarding the steps in each possible implementation of the first aspect performed by the computing device cluster, please refer to the first aspect and will not be repeated here.
[0044] In a fifth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a computer, the computer can execute any of the methods described above.
[0045] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the methods described above.
[0046] In the seventh aspect of the present application, a chip system is provided, which includes a processor and a communication interface, wherein the communication interface is used to communicate with modules outside the chip system, and the processor is used to run computer programs or instructions so that the device installed with the chip system can execute any of the methods in the above aspects.
[0047] Among them, the technical effects brought about by any design method in the second to seventh aspects can refer to the technical effects brought about by different implementation methods in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of a system architecture 100 provided in an embodiment of the present application;
[0049] Figure 2 A flowchart of a method for allocating resources for a task provided in an embodiment of the present application;
[0050] Figure 3 A schematic diagram of determining a target amount of resources required for executing a first task provided in an embodiment of the present application;
[0051] Figure 4 A schematic diagram of determining the expected resource occupancy of a task queue provided in an embodiment of the present application;
[0052] Figure 5 A schematic diagram of the operational architecture of a method for allocating resources for a task provided in an embodiment of the present application;
[0053] Figure 6 A schematic diagram of a process for generating a task processing request based on a scheduled task provided in an embodiment of the present application;
[0054] Figure 7 A schematic diagram of selecting a task queue provided in an embodiment of the present application;
[0055] Figure 8 A schematic diagram of determining task priority provided in an embodiment of the present application;
[0056] Figure 9 A schematic diagram of the structure of a task resource allocation device provided in an embodiment of the present application;
[0057] Figure 10 A schematic diagram of the structure of a computing device 1000 provided in an embodiment of the present application;
[0058] Figure 11 A schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0059] Figure 12 A schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application;
[0060] Figure 13 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0062] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.
[0063] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.
[0064] (1) Computing cluster
[0065] A computing cluster is a computer system primarily connected by a loosely integrated set of computer software or hardware to collaborate closely to complete computing tasks. In a sense, a computing cluster can be thought of as a single computer. Individual computers in a computing cluster are typically called nodes, and nodes are typically connected via a network.
[0066] (2) Spark
[0067] Spark is a computing engine designed specifically for large-scale data processing. Essentially, it's a computing cluster. Generally, Spark can be applied to a variety of data processing scenarios to perform tasks. The following briefly introduces some of the scenarios in which Spark can be used.
[0068] 1. Big data processing and analysis scenarios: Spark is a powerful distributed computing framework that can process large-scale data sets and supports various data sources and formats, such as text, images, audio, video, etc.
[0069] 2. Machine learning scenarios: Spark supports various machine learning algorithms and tools, such as classification, regression, clustering, recommendation systems, etc., and can be used in various application scenarios such as finance, medical care, retail, etc.
[0070] 3. Data mining and visualization scenarios: Spark provides a variety of data mining and visualization tools that can be used to explore data, discover data patterns, and visualize data results.
[0071] 4. Stream processing and real-time analysis scenarios: Spark supports stream processing and real-time analysis. It can be used to process real-time data streams and build real-time analysis applications, such as streaming media processing and Internet of Things applications.
[0072] 5. Big data warehouse and data lake scenarios: Spark can be used to build big data warehouses and data lakes, integrating data from multiple data sources and formats into a unified dataset to facilitate data sharing and analysis.
[0073] (3) Structured Query Language (SQL)
[0074] SQL is a database language with diverse capabilities, including data manipulation and definition. Databases are typically searched using SQL syntax, so most data retrieval tasks require SQL. SQL can not only be used independently on terminals but can also serve as a sub-language to effectively enhance other programming. This means that SQL can be used alongside other programming languages to optimize program functionality and provide users with more comprehensive information.
[0075] (4) Shuffle operator
[0076] The Shuffle operator is a data partitioning operator in Spark that shuffles and repartitions data. In Spark, data processing is based on Resilient Distributed Datasets (RDDs). RDDs are distributed, immutable data sets divided into multiple partitions, each of which is processed on a different node. The Shuffle operator repartitions data so that data with the same key value is processed on the same node, improving data processing efficiency.
[0077] (5) Regression model
[0078] A regression model is a mathematical model that quantitatively describes statistical relationships. Specifically, a regression model is a predictive modeling technique that studies the relationship between dependent and independent variables.
[0079] For example, the mathematical model of multiple linear regression can be expressed as y = β0 + β1*x1 + β2*x2… + βp*xp; where β0, β1,…, βp are p+1 parameters to be estimated, y is the dependent variable; x1-xp is the independent variable, and βi is called the regression coefficient, which represents the degree of influence of the independent variable on the dependent variable.
[0080] When executing tasks based on Spark computing clusters, a dynamic resource allocation scheme is usually used to automatically allocate resources for the tasks. In this dynamic resource allocation scheme, in order to complete the tasks in the queue more quickly, the scheduling mechanism will allocate as many resources as possible to the tasks in the queue, which easily leads to resource waste. For example, assuming that a task actually only requires 8 cores (Central Processing Unit, CPU) and 16 gigabytes (G) of memory to be completed efficiently, the current dynamic resource allocation scheme may allocate more than 100 cores of CPU and 100GB of memory to the task, resulting in an extremely large amount of resources occupied by the task. Moreover, the extra resources allocated to the task will not effectively increase the execution speed of the task, but may only slightly increase the execution speed of the task, ultimately resulting in extremely low actual execution efficiency of the task.
[0081] In general, existing dynamic resource allocation solutions in the industry suffer from irrational resource allocation, which can lead to long task wait times and a shortage of cluster computing resources, ultimately resulting in high task failure rates. Therefore, there is an urgent need for a better resource allocation solution to address this issue of irrational computing resource allocation in computing clusters.
[0082] Based on this, an embodiment of the present application provides a method for allocating resources for tasks. This method predicts the target amount of data involved in task execution based on the data screening conditions indicated in the task and the total amount of data accessed by the computing cluster, and then predicts the target amount of resources required for task execution based on the target amount of data. In this way, when delivering a task to the computing cluster, the target amount of resources required for the task can be specified, thereby ensuring that the computing cluster allocates the appropriate amount of resources to the task, ensuring that each task can have the highest possible execution efficiency when executed, avoiding the impact of unreasonable task resource allocation on the normal execution of the task, and effectively improving the task execution efficiency of the computing cluster.
[0083] See also Figure 1 , Figure 1 A schematic diagram of a system architecture 100 provided in an embodiment of the present application. Figure 1 As shown, the system architecture 100 includes an execution device 110 , a data storage system 120 and a computing cluster 130 .
[0084] Execution device 110 can be implemented as a computing instance of at least one of a physical host (computing device), a virtual machine, or a container. When implemented as a virtual machine or container, execution device 110 effectively exists as a cloud computing product, capable of providing cloud services. Optionally, execution device 110 can be coordinated with other computing devices, such as data storage devices and load balancers. Execution device 110 can be deployed at a single physical site or distributed across multiple physical sites.
[0085] Computing cluster 130 includes multiple computing devices that can perform analysis and processing on the data stored in data storage system 120 to complete the tasks assigned by execution device 110. Optionally, execution device 110 can be a device independent of computing cluster 130 (e.g., a server independent of computing cluster 130); execution device 110 can also be a device within computing cluster 130, that is, execution device 110 itself can be used to execute tasks.
[0086] Optionally, the data storage system 120 may be located external to the execution device 110 and exchange data with the execution device 110 via a network. Alternatively, if the execution device 110 is a physical host, the data storage system 120 may be located internal to the execution device 110, such as if the data storage system 120 exchanges data with the processor via a bus. In this case, the data storage system 120 is represented by a hard disk. With the data storage system 120, the execution device 110 may use the data in the data storage system 120 or call program code in the data storage system 120 to implement the resource allocation method for tasks provided in the embodiments of the present application.
[0087] In a specific implementation, the execution device 110 is used to implement the resource allocation method for tasks provided in the embodiments of the present application to determine the amount of resources required for task execution, and then specify the amount of resources to be allocated to the task when delivering the task to the computing cluster 130.
[0088] See also Figure 2 , Figure 2 This is a flow chart of a method for allocating resources for a task provided in an embodiment of the present application. Figure 2 As shown, the resource allocation method for tasks provided in the embodiment of the present application includes the following steps 201-204.
[0089] Step 201: The execution device obtains a first task, where the first task is used to instruct the computing cluster to process data that meets a data screening condition.
[0090] In this embodiment, the first task is a task that requires processing by the computing cluster, and the first task includes data filtering conditions to instruct the computing cluster to process data that meets the specific data filtering conditions. Simply put, since the computing cluster can be used to process a large amount of data stored in the data storage system, the first task often instructs the computing cluster to process a specific portion of the data stored in the data storage system. Because the first task can include data filtering conditions, the computing cluster can clearly determine which portion of the data in the data storage system is being processed when executing the first task.
[0091] For example, the first task can be a SQL-based data query task, where the data screening conditions include one or more of a custom data selection condition (such as the WHERE condition in SQL), a data statistical field (such as the SELECT field in SQL), a data statistical dimension (such as the GROUP BY field in SQL), and a data screening time period (such as an event window specified in SQL). In addition, the first task can also be other tasks supported by the computing cluster, such as big data analysis tasks, machine learning tasks, data flow analysis tasks, and the like. This embodiment does not limit the specific type of the first task.
[0092] There are multiple ways for the execution device to obtain the first task.
[0093] In a possible implementation, the first task is a task pre-stored on the execution device, and the execution device can actively obtain the first task.
[0094] For example, the first task may be a task that needs to be executed periodically in the computing cluster. Therefore, the execution device needs to periodically obtain the first task and deliver it to the computing cluster so that the first task can be executed periodically in the computing cluster. For example, in the financial field, the first task may be a sales data analysis task, requiring the computing cluster to periodically extract and analyze past daily, weekly, or monthly sales data.
[0095] In another possible implementation, the first task is sent by the user to the execution device via the client. For example, if a user temporarily creates a first task that requires execution by a computing cluster, the user can send the first task to the execution device via the client, and the execution device then delivers the first task to the computing cluster.
[0096] Step 202 : Based on the data screening condition and the total amount of data accessed by the computing cluster, the execution device predicts the target amount of data that the computing cluster needs to process when executing the first task.
[0097] Because the total amount of data accessed by the computing cluster is fixed at a given point in time, the amount of data that can be filtered when filtering all the data accessed by the computing cluster based on the data filtering criteria is often also fixed. Thus, based on the total amount of data accessed by the computing cluster and the data filtering criteria specified in the first task, the execution device can predict the target data volume corresponding to the data that the computing cluster needs to process when executing the first task. In other words, the target data volume indicates the total amount of data that the computing cluster needs to process when executing the first task.
[0098] For example, assuming that the data stored in the data storage system connected to the computing cluster is the company's sales details for each day in the past ten years, and the total amount of data stored in the data storage system (that is, the total amount of data connected to the computing cluster) is 100G, then when the data filtering condition is to filter the profit field and sales amount field of a certain year and month in all sales details data, the specific amount of data that meets the data filtering condition can be effectively predicted based on the total amount of sales details data.
[0099] Optionally, to facilitate the execution device's ability to accurately predict the target data volume required for the computing cluster to process the first task, the execution device may predict the target data volume based on a pre-built regression model (e.g., a linear regression model). In the case of a pre-built regression model, the execution device may substitute the data screening criteria and the total amount of data accessed by the computing cluster into the regression model to obtain the target data volume output by the regression model.
[0100] When building a regression model for predicting the amount of data required for task execution, various data screening conditions can be pre-constructed. Then, based on these pre-constructed data screening conditions, data can be searched in the data storage system connected to the computing cluster to determine the amount of data required for each of these data screening conditions. In this way, based on the relationship between the pre-constructed data screening conditions and the corresponding data amounts, a regression model can be constructed. This regression model indicates the relationship between the total amount of data connected to the computing cluster, the data screening conditions, and the amount of data required for task execution.
[0101] Step 203: Based on the target data volume, the execution device predicts the target resource volume required for executing the first task. The target resource volume is used to indicate the processor resource volume and / or memory volume.
[0102] When the amount of data involved in executing the first task is determined, the amount of resources required to execute the first task can often also be determined. Therefore, based on the target amount of data required to execute the first task, the execution device can predict the target amount of resources required to execute the first task, that is, the processing resources that the computing cluster needs to allocate to the first task. Specifically, the target amount of resources can be used to indicate the amount of processor resources and / or memory, that is, to indicate at least one of the amount of processor resources and the amount of memory.
[0103] For example, if the computing cluster is a Spark computing cluster, the target resource quantity can specifically indicate the number of driver CPU cores, driver memory, number of executors, number of executor CPU cores, and executor memory. Driver and executor are two different roles. The driver is the master node of a Spark application (i.e., an application running on a Spark computing cluster) and is responsible for controlling and coordinating the entire Spark application. The driver node breaks down the Spark application into multiple tasks and assigns these tasks to executor nodes for execution. The driver node is also responsible for maintaining application state information and processing user requests. The executor is the worker node of the Spark application, responsible for executing tasks assigned by the driver node. Each executor node has its own virtual machine process and can run on different physical machines. The executor node receives tasks by communicating with the driver node and returns the task results to the driver node. In general, the driver node and executor node play different roles in a Spark application. The driver node is the control center of the application, while the executor node is the worker node of the application.
[0104] It should be noted that in this embodiment, the target resource amount predicted by the execution device based on the target data amount is a resource amount that can maximize the execution efficiency of the first task. That is, when processing resources are allocated to the first task according to the target resource amount, the execution efficiency of the first task can be maximized.
[0105] Specifically, for any task, if too few processing resources are allocated to the task, the execution time of the task will be too long, resulting in low execution efficiency of the task; if too many processing resources are allocated to the task, most of the processing resources allocated to the task will not be effectively utilized, and the execution time of the task will not be significantly shortened due to the large number of processing resources allocated, resulting in relatively low execution efficiency of the task. Generally speaking, when the amount of resources allocated to a task is within a certain range, the execution efficiency of the task will often reach a relatively high value, and the range corresponding to higher execution efficiency is often related to the amount of data of the task itself. Therefore, in this embodiment, the target amount of resources that need to be allocated to the first task is predicted based on the target data amount of the first task, which can effectively determine the amount of resources that can make the execution efficiency of the first task as high as possible.
[0106] Step 204 : The execution device sends a task processing request to the computing cluster. The task processing request is used to request the computing cluster to process the first task and specifies the amount of resources allocated by the computing cluster to the first task as the target amount of resources.
[0107] After predicting the target amount of resources required for executing the first task, the execution device can send a task processing request to the computing cluster, requesting the computing cluster to process the first task. Furthermore, in the task processing request, the execution device specifies the target amount of resources to be allocated by the computing cluster when processing the first task. This ensures that the computing cluster allocates processing resources to the first task according to the specified amount, preventing the computing cluster from allocating excessive resources to the first task.
[0108] In some embodiments, when executing a task, the computing cluster may also involve data redistribution operations, that is, repartitioning the data, thereby adjusting the data processed by multiple computing nodes on the computing cluster to maximize data processing efficiency. In this case, the amount of resources required for task execution may also change to a certain extent. That is, for two tasks with the same amount of data, the amount of resources required for the task involving data redistribution operations and the task not involving data redistribution operations are usually different, that is, the data redistribution operation will affect the amount of resources required for task execution. Based on this, this embodiment proposes that when predicting the amount of resources required for task execution, whether the task involves data redistribution operations is also considered.
[0109] Optionally, in the above embodiment, after obtaining the first task, a data redistribution operation involved in executing the first task may be determined based on the data screening conditions, where the data redistribution operation is used to redistribute data processed by multiple computing nodes in the computing cluster. Since the first task actually processes data that meets the data screening conditions, after determining the data that actually needs to be processed based on the data screening conditions, it can be determined whether the data redistribution operation needs to be performed. Therefore, in this embodiment, the data redistribution operation involved in executing the first task can be determined based on the data screening conditions.
[0110] Thus, in step 203, the execution device may specifically predict the target amount of resources required for executing the first task based on the target data volume and the data redistribution operation. Specifically, the data redistribution operation is equivalent to a means in the first task processing process, which can adjust the distribution of data to maximize the speed of first task processing. Generally speaking, if the task processing process involves a data redistribution operation, the amount of resources required for the task will often increase compared to a task with the same data volume but not involving a data redistribution operation. That is, when a task involves a data redistribution operation, the task processing process often requires more resources to complete.
[0111] In this solution, whether the task execution process involves data redistribution operations is determined based on the data screening conditions indicated by the task, and the amount of resources required for the task execution process is further predicted based on the amount of data involved in the task and the data redistribution operations, which can effectively improve the accuracy of the predicted resource amount.
[0112] For example, if the first task is a SQL-based data query task, the data redistribution operation in the Spark computing cluster is performed by the shuffle operator. Generally speaking, the situations in which the shuffle operator is triggered are roughly as follows. The triggering of the shuffle operator can be determined by determining whether the SQL statement corresponding to the first task (i.e., the data filtering condition indicated by the first task) contains relevant operation statements.
[0113] 1. Aggregation operations: When using aggregation operations such as group by, distinct, count, and sum, the shuffle operator is triggered, and the hash shuffle operator is triggered. For example: SELECT department, AVG(salary) FROM employee GROUP BY department.
[0114] 2. Join operations: When using join operations such as join and union, the shuffle operator is triggered, and the Sort Shuffle operator is triggered. For example: SELECT * FROM employee JOIN department ON employee.dept_id = department.dept_id.
[0115] 3. Sorting operations: When using sorting operations such as order by or sort, the shuffle operator is triggered, and the Sort Shuffle operator is triggered. For example: SELECT * FROM employee ORDER BY salary DESC.
[0116] 4. Repartitioning: When the repartition operation is used, the shuffle operator is triggered, and the HashShuffle operator is triggered. For example: SELECT * FROM employee.repartition(10).
[0117] 5. Window function operation: When using a window function, the shuffle operator is triggered, and the SortShuffle operator is triggered. For example, SELECT department,salary,RANK() OVER (PARTITION BY department ORDER BY salary DESC) as rank FROM employee.
[0118] In general, when the first task is essentially a task represented by an SQL statement, by detecting whether the SQL statement includes a specific operation statement, it can be determined whether the execution of the first task involves a shuffle operator (ie, a data redistribution operation).
[0119] Optionally, when predicting the amount of resources required for executing the first task, a regression model may be used to predict the target amount of resources required for executing the first task based on the target amount of data and data redistribution operations involved in executing the first task. The regression model is fitted based on historical task processing data from the computing cluster. By substituting the target amount of data and data redistribution operations into the fitted regression model, the target amount of resources required for executing the first task, as output by the regression model, can be obtained.
[0120] Specifically, since the computing cluster is continuously running and processes a large number of similar tasks every day (such as data query tasks with similar SQL statements), the historical task processing data of the computing cluster can be effectively fitted to obtain a regression model. Among them, the historical task processing data of the computing cluster can include the actual amount of data involved in the tasks executed in the historical time period, the data redistribution operations involved in the tasks, and the amount of resources allocated when the tasks are executed. Therefore, based on the historical task processing data of the computing cluster, it is possible to obtain the relationship between the amount of data involved in the tasks, the data redistribution operations involved in the tasks, and the amount of resources allocated when the tasks are executed, so as to fit the corresponding regression model.
[0121] Among them, the regression model used to predict the target resource quantity required for the execution of the first task can be, for example, a linear regression model, an auto-regressive moving average model (ARIMA), a Prophet model or a random forest model. This embodiment does not specifically limit the specific type of the regression model.
[0122] For example, see Figure 3 , Figure 3 This is a schematic diagram of determining the target resource amount required for executing the first task provided by the embodiment of the present application. Figure 3 As shown in the figure, if the target data volume of the first task is 1GB and the Shuffle operator is triggered when the first task is executed, the target data volume of 1GB and the triggering of the Shuffle operator are input into the regression model to obtain the target resource volume output by the regression model. The target resource volume is specifically: 1 Driver CPU core, 1GB Driver memory, 2 Executors, 4 Executor CPU cores, and 5GB Executor memory.
[0123] Optionally, when fitting the regression model, the task processing efficiency corresponding to the historical task processing data used to fit the regression model satisfies a preset condition, where the task processing efficiency corresponding to the historical task processing data is related to the amount of task data, the amount of resources allocated to the task, and the execution time of the task. The task processing efficiency satisfying the preset condition may, for example, be that the task processing efficiency is greater than or equal to a preset threshold.
[0124] Generally speaking, when the amount of data and the amount of resources allocated to a task are relatively fixed, the longer the task's execution time, the lower the task's processing efficiency. When the amount of data and the amount of resources allocated to a task are relatively fixed, the more resources allocated to a task, the lower the task's processing efficiency. When the amount of resources allocated to a task and the amount of time allocated to a task are fixed, the larger the amount of data, the higher the task's processing efficiency. Therefore, when the amount of data for a task is fixed, minimizing the amount of resources allocated to the task and minimizing the task's execution time can both improve the task's execution efficiency. However, there is a conflict between the amount of resources allocated to a task and the task's execution time: the more resources allocated to a task, the shorter the task's execution time. Furthermore, after the amount of resources allocated to a task reaches a certain level, the reduction in the task's execution time significantly decreases. Therefore, by rationally determining the amount of resources allocated to a task based on the amount of data, a balance can be achieved between the amount of resources allocated to the task and the task's execution time, thereby ensuring that the task's execution efficiency meets the preset requirements.
[0125] Specifically, since the historical task processing data itself indicates the task execution status in the historical time period, it will include specific content such as the task data volume, the amount of resources allocated to the task, and the task execution time. Therefore, based on the task data volume, the amount of resources allocated to the task, and the task execution time, the execution efficiency of each task in all the historical task processing data of the computing cluster can be calculated, and then the historical task processing data whose task processing efficiency meets the preset conditions can be screened out to fit the regression model, thereby ensuring that the fitted regression model predicts the amount of resources corresponding to higher execution efficiency based on the task data volume. That is, by screening the task processing data corresponding to higher execution efficiency to fit the regression model, it can be ensured that when resources are allocated to the first task based on the amount of resources predicted by the regression model, the first task can also have a higher execution efficiency, thereby ensuring the rationality of resource allocation.
[0126] The preceding describes the process of estimating the resource allocation required for tasks delivered to a compute cluster. However, in most scenarios, multiple task queues are configured within a compute cluster to process tasks. The task queue in which a task is placed often affects its execution. Therefore, the following describes how to specify the task queue to which a task is delivered.
[0127] Optionally, in the above embodiment, the computing cluster may include multiple task queues for processing tasks. In this case, in addition to specifying the target amount of resources to be allocated to the first task, the task processing request also specifies a target queue for processing the first task, where the target queue is one of the multiple task queues. In other words, the task processing request simultaneously specifies the target amount of resources to be allocated to the first task and the target queue for processing the first task. Thus, when processing the first task, the computing cluster will place the first task in the target queue and allocate the target amount of resources to execute the first task.
[0128] Each of the multiple task queues included in the computing cluster is pre-allocated a fixed amount of resources, and the resource amounts allocated to different task queues can be the same or different. For example, suppose there are 10 task queues in the computing cluster, the first task queue is allocated 60% of the resources of the entire computing cluster, the second task queue is allocated 20% of the resources of the entire computing cluster, and the remaining 8 task queues are evenly distributed among the remaining 20% of the resources of the computing cluster.
[0129] In this solution, when a computing cluster includes multiple task queues, by specifying a target queue for processing the first task, the loads of the multiple task queues can be kept balanced as much as possible, avoiding the situation where some task queues are overloaded while others are underloaded, thereby ensuring that the resources in the computing cluster can be reasonably utilized.
[0130] Optionally, the process of determining a target queue for processing the first task among multiple task queues may include the following steps.
[0131] First, based on the resource usage of each of the multiple task queues at the current time, multiple target samples with the closest resource usage are retrieved for each task queue from the historical samples of the multiple task queues. The historical samples are used to indicate the resource usage of the task queues in the historical time period. The multiple target samples corresponding to each task queue are all historical samples of each task queue.
[0132] Specifically, since the task queues in the computing cluster exist for a long time after being divided, that is, each task queue will run for a long time, each task queue can find corresponding multiple historical samples. For example, if the computing cluster has been running continuously for 100 hours, these 100 hours can be divided into 100 time periods of 1 hour, and the resource usage of each queue in the computing cluster during these 100 time periods constitutes the 100 historical samples corresponding to each queue. Based on the current resource usage of each task queue, multiple target samples with the closest resource usage can be retrieved for each task queue from the 100 historical samples corresponding to each task queue.
[0133] Then, based on multiple target samples corresponding to each task queue, the expected resource occupancy of each task queue in the future is predicted.
[0134] Since the first task is not delivered in real time, it can only be delivered to the computing cluster after the queue to which the first task needs to be delivered and the amount of resources that need to be allocated are determined, and it often takes a certain amount of time to determine the queue to which the first task needs to be delivered and the amount of resources that need to be allocated. Therefore, in this embodiment, the expected resource occupancy of the task queue in the future is predicted to ensure that the resource occupancy of the task queue when the first task is actually delivered is close to the resource occupancy based on which the task queue selection is performed, thereby ensuring the accuracy of selecting the task queue for the first task.
[0135] Since the multiple target samples corresponding to each task queue actually represent resource usage within a specific time period, determining multiple target samples is equivalent to determining multiple time periods in history. This way, by moving forward a certain amount of time based on the multiple time periods represented by the target samples, we can determine the resource usage for a certain amount of time after the target samples. By averaging the resource usage for a certain amount of time after the target samples, we can obtain a single resource usage profile, which can be used as the expected resource usage for the task queue in the future.
[0136] For example, see Figure 4 , Figure 4 This is a schematic diagram of determining the expected resource occupancy of a task queue provided in an embodiment of the present application. Figure 4As shown, the resource usage of task queue A in the current time period is: CPU usage 75% and memory usage 80%. The 75% CPU usage means that the CPU resources currently occupied by task queue A are 75% of the total CPU resources available to the task queue, and the 80% memory usage means that the memory resources currently occupied by task queue A are 80% of the total memory resources available to task queue A. Using the K-Nearest Neighbor (KNN) algorithm, we can search for the resource usage of task queue A in each time period in the past to find the K time periods whose resource usage is closest to that of the current time period. In this example, K is 3. The three time periods include time periods 1 through 3 in the past. The resource usage in time period 1 is: CPU usage 73% and memory usage 79%; the resource usage in time period 2 is: CPU usage 76% and memory usage 82%; and the resource usage in time period 3 is: CPU usage 75% and memory usage 81%.
[0137] Based on the retrieved three time periods, the resource usage for each of the three time periods can be determined from the historical resource usage of task queue A, where X can be determined based on actual circumstances, for example, 1 or 2. Specifically, the resource usage for X hours after time period 1 is: CPU usage 80%, memory usage 85%; the resource usage for X hours after time period 2 is: CPU usage 78%, memory usage 83%; and the resource usage for X hours after time period 3 is: CPU usage 76%, memory usage 80%. By averaging the resource usage for each of these three time periods after X hours, we can obtain the mean resource usage: CPU usage 78%, memory usage 84%. At this point, 78% CPU usage and 84% memory usage can be used as the expected resource usage for task queue A X hours after the current time.
[0138] Finally, based on the expected resource occupancy and the target resource amount, a target queue is determined among the multiple task queues for processing the first task.
[0139] After predicting the expected resource occupancy of each task queue in the future, it is equivalent to predicting the resource occupancy of each task queue when the first task is delivered, so that the corresponding task queue can be selected for the first task based on the expected resource occupancy of each task queue. Since there may be other tasks that need to be delivered at the same time when the first task is delivered, the task queue corresponding to each task can be determined based on the amount of resources required to be allocated to each task and the expected resource occupancy of each task queue in this embodiment.
[0140] Specifically, when determining the task queue corresponding to the first task, it is necessary to determine, based on the expected resource occupancy of the task queue, that the remaining resources of the task queue are no less than the target resource amount required for the first task, thereby ensuring that the first task can be normally executed. In addition, when determining the corresponding task queue for each task, the selection of the task queue can also be based on the principle of load balancing, that is, ensuring that the load of each task queue is as balanced as possible, avoiding the phenomenon that some task queues are overloaded while some task queues are underloaded.
[0141] Optionally, the task processing request is further used to indicate a processing priority of the first task, where the processing priority is determined based on the target resource amount.
[0142] Specifically, the smaller the amount of resources a task requires, the smaller the resources consumed when the task is running, and the faster the task will be completed. Therefore, the priority of tasks that require less resources can be increased so that the tasks can be completed as soon as possible to free up more resources to run other tasks.
[0143] Optionally, the processing priority is also related to the expected runtime of the first task, which is determined based on the target data volume and target resource volume. That is, the processing priority is related to both the target resource volume and the expected runtime of the first task. Specifically, if the target data volume and target resource volume corresponding to the first task have been determined, the expected runtime of the first task can be predicted based on a pre-fitted regression model to determine the processing priority of the first task.
[0144] Specifically, when determining the processing priority of a task, the processing priority of the task can be determined based on the amount of resources required by the task and the expected running time of the task. The smaller the amount of resources required by the task and the shorter the expected running time of the task, the higher the processing priority of the task; the larger the amount of resources required by the task and the longer the expected running time of the task, the lower the processing priority of the task.
[0145] In general, the embodiments of the present application tend to prioritize the execution of tasks in the task queue that occupy as little resources as possible and have the shortest running time, to ensure that the computing cluster can complete the execution of tasks as quickly and efficiently as possible.
[0146] To facilitate understanding, the resource allocation method for tasks provided by this embodiment will be described in detail below with reference to specific examples.
[0147] See also Figure 5 , Figure 5 This is a schematic diagram of the operation architecture of a method for allocating resources for a task provided in an embodiment of the present application. Figure 5As shown in the figure, the operational architecture of the task resource allocation method includes two parts: a big data platform and an algorithm platform. The big data platform is deployed with a computing cluster and a database. The computing cluster retrieves data from the database and processes the retrieved data to complete the offline tasks delivered by the algorithm platform. In addition, the database also stores various cluster resource conditions generated when the computing cluster processes real-time tasks, such as the resource usage of each of the multiple queues in the computing cluster during historical time periods, the amount of resources allocated to each task by the computing cluster when processing each task, and the running time of each task. Among them, offline tasks refer to tasks that have been delivered to the computing cluster and are waiting for the computing cluster to process; real-time tasks refer to tasks that are already being processed by the computing cluster.
[0148] The algorithm platform is a resource allocation method for executing the tasks provided in the embodiments of this application. It determines the amount of resources required for task execution and the queue selected for task execution, thereby instructing the computing cluster to process the task based on the specified amount of resources and queue. Specifically, the algorithm platform can be deployed in the execution device introduced in the above embodiments.
[0149] One or more scheduled tasks can be stored in the algorithm platform. The algorithm platform will periodically obtain the scheduled tasks (for example, when the scheduled task is executed once a day, the algorithm platform will obtain the scheduled task once a day) and calculate the resources required for the scheduled task, the selected task queue and the task priority, so as to deliver the scheduled task to the computing cluster and specify the resources required for the scheduled task, the selected task queue and the task priority.
[0150] Specifically, after acquiring a scheduled task, the algorithm platform obtains task information to determine the data volume of the scheduled task and whether the scheduled task involves data redistribution. Based on the data volume of the scheduled task and the data redistribution operations involved, the task resource prediction model can be used to predict the resources required for the scheduled task's execution.
[0151] Furthermore, the algorithm platform retrieves the current and historical resource usage of the computing cluster from the big data platform's database. Based on this information, the queue resource prediction model can predict the future resource usage of the task queues within the computing cluster. Based on the resources required for task execution and the future resource usage of the task queues within the computing cluster, the appropriate task queue can be selected for the task and its corresponding priority can be determined.
[0152] Finally, based on the resources required for task execution, the task's corresponding task queue, and the task's priority, a task processing request is generated. This request is then sent to the computing cluster in the big data platform to complete the task delivery. When processing the task, the computing cluster performs the task processing based on the resources required for task execution, the task's corresponding task queue, and the task's priority as indicated in the task processing request.
[0153] For example, see Figure 6 , Figure 6 The present invention provides a flow chart of a method for generating a task processing request based on a scheduled task. Figure 6 As shown in Figure 1, the process of generating a task processing request based on a scheduled task can include four stages: data input, task resource usage prediction, queue resource usage prediction, and queue selection. The following uses the Spark computing cluster as an example to explain these four stages in detail.
[0154] Phase 1: data input.
[0155] First, by setting a scheduled task on the execution device, the execution device periodically retrieves and parses the task. During task parsing, the task's data filtering conditions must be analyzed, including the task's selection criteria (WHERE conditions), statistical fields (SELECT fields), statistical dimensions (GROUP BY fields), and time windows (data filtering time periods). Then, based on the task's data filtering conditions and the total amount of data accessed by the Spark computing cluster, a linear regression prediction is performed to accurately determine the data magnitude (i.e., the amount of data involved during task execution). Furthermore, the data statistical dimensions and selection criteria are used to determine the Shuffle operator involved in the task's execution.
[0156] When the parsing task is triggered, the current and historical cluster resource status are obtained as the current and historical cluster inputs. The cluster resource status includes the number of tasks, CPU usage, memory usage, number of idle CPU cores, and idle memory resources for each queue in different time periods. This is used to predict the cluster resource usage in a subsequent time period.
[0157] In addition, during the execution of the task, the running status of each task is recorded in real time and the CPU resources, memory resources, task execution queue and task running time used by each task are recorded in the database to predict the time occupancy of subsequent tasks.
[0158] Phase 2: Task resource occupancy prediction
[0159] Based on the task data level obtained in Phase 1 and the Shuffle operator involved in the task, a linear regression prediction model is used to fit the CPU resource parameters and memory resource parameters required for the current task (such as the number of Driver CPU cores, Driver memory resources, number of Executors, number of Executor CPU cores, and Executor memory). A separate linear regression prediction model is constructed for each resource parameter required for the task, and the linear regression prediction model is constructed and adjusted based on historical task execution data. Furthermore, after predicting the amount of resources required for the task to run, the expected running time of the task can be predicted based on the amount of resources required for the task to run and the data level of the task.
[0160] Phase 3: Queue Resource Occupancy Prediction
[0161] Because data preparation and model scheduling take time, we usually start predicting queue resource occupancy X hours before task delivery (X can be determined or adjusted based on the specific scenario, for example, X is 2).
[0162] Specifically, based on the resource usage of each task queue in the current time period, the KNN algorithm is used to retrieve the resource usage of task queues in historical time periods, thereby finding the K samples with the most similar resource usage for each task queue. Then, for each task queue, the average resource usage of these K samples after X hours can be used as a reference to predict the resource usage of the current task queue X hours in the future. In this way, by performing the merchant's queue resource usage prediction operation on all task queues, the overall resource usage of the cluster during future task delivery runtimes can be derived.
[0163] Phase 4: Queue Selection
[0164] Based on the future resource occupancy of the task queues predicted in stage three, the rationality of the operation of the current task in different queues can be gradually fitted. If the remaining resources of the current task queue are less than the resources required by the current task, the fitting of the current task queue is skipped, that is, the current task cannot select this task queue. If a task queue will maintain high occupancy for a long time during the task delivery period, directly fit the task with a small resource consumption and a data volume suitable for this task queue. In general, the goal of selecting the corresponding task queue for each task is to balance the load between the various task queues as much as possible to avoid the phenomenon of excessively unbalanced loads between task queues.
[0165] For example, when selecting a task queue for a batch task, you can combine different tasks and select the corresponding task queue. Compare the load balancing between the task queues after combining different tasks, and then select a combination that maximizes the load balancing between the task queues.
[0166] For example, see Figure 7 , Figure 7 This is a diagram of a task queue selection provided by an embodiment of the present application. Figure 7 As shown, the tasks to be delivered include four tasks: Task 1, Task 2, Task 3, and Task 4. Furthermore, the computing cluster currently has two task queues: Queue A and Queue B, and Queue A and Queue B have the same amount of resources. Task 1 requires 60% of the total resources of a queue, Task 2 requires 25% of the total resources of a queue, Task 3 requires 23% of the total resources of a queue, and Task 4 requires 10% of the total resources of a queue. Queue A has 80% of its idle resources during the future task delivery period, while Queue B has 60% of its idle resources during the future task delivery period. By comparing the status of each task after delivery to its corresponding task queue, it can be determined that Task 1 and Task 4 will be delivered to Queue A, and Task 2 and Task 3 will be delivered to Queue B. After tasks 1 and 4 are delivered to queue A, the remaining resources in queue A are 10%; after tasks 2 and 3 are delivered to queue B, the remaining resources in queue B are 12%. It is obvious that the loads of queues A and B are basically balanced.
[0167] Furthermore, after selecting a corresponding task queue for each task, the task priority corresponding to each task may be determined so that the task queue processes the assigned tasks in a certain priority order.
[0168] For each task queue, a regression prediction model may be used to perform fitting based on the amount of resources and expected running time of each task to be processed in the task queue, thereby obtaining the task priority corresponding to each task.
[0169] For example, see Figure 8 , Figure 8 A schematic diagram of determining task priority provided in an embodiment of the present application. Figure 8 As shown, in Figure 7Based on the example shown, after determining the task queues corresponding to Tasks 1 to 4, the task priority corresponding to each task can be further determined. In particular, in Queue A, Task 1 (required resources are 60%), which requires a large amount of resources and has a longer expected running time, can be set to a low task priority, while Task 4 (required resources are 10%), which requires a smaller amount of resources and has a shorter expected running time, can be set to a low task priority. In addition, in Queue B, since Tasks 2 and 3 have similar amounts of resources required, Task 3, which has a shorter expected running time, can be set to a high task priority, while Task 2, which has a longer expected running time, can be set to a low task priority.
[0170] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.
[0171] See also Figure 9 , Figure 9 This is a schematic diagram of the structure of a task resource allocation device provided in an embodiment of the present application. Figure 9 As shown, the resource allocation device for tasks provided in an embodiment of the present application includes: an acquisition module 901, used to acquire a first task, the first task being used to instruct the computing cluster to process data that meets a data screening condition; a processing module 902, used to predict, based on the data screening condition and the total amount of data accessed by the computing cluster, a target amount of data to be processed when the computing cluster executes the first task; the processing module 902 is also used to predict, based on the target amount of data, a target amount of resources required for the execution of the first task, the target amount of resources being used to indicate the amount of processor resources and / or memory; the processing module 902 is also used to send a task processing request to the computing cluster, the task processing request being used to request the computing cluster to process the first task and specifying the amount of resources allocated by the computing cluster to the first task as the target amount of resources.
[0172] In one possible implementation, the processing module 902 is further used to: determine the data redistribution operation involved in the execution of the first task based on the data screening conditions, where the data redistribution operation is used to redistribute the data processed by multiple computing nodes on the computing cluster; and predict the target resource amount required for the execution of the first task based on the target data volume and the data redistribution operation.
[0173] In one possible implementation, the processing module 902 is further configured to: predict the target amount of resources required for executing the first task through a regression model based on the target data volume and the data redistribution operation; the regression model is obtained by fitting the historical task processing data of the computing cluster.
[0174] In a possible implementation, the task processing efficiency corresponding to the historical task processing data meets a preset condition, and the task processing efficiency is related to the data volume of the task, the amount of resources allocated to the task, and the execution time of the task.
[0175] In a possible implementation, the computing cluster includes multiple task queues for processing tasks. The task processing request is further used to specify a target queue for processing the first task, where the target queue is one of the multiple task queues.
[0176] In one possible implementation, the processing module 902 is further used to: based on the resource occupancy of each task queue in the multiple task queues at the current time, retrieve multiple target samples with the closest resource occupancy for each task queue from the historical samples of the multiple task queues, and the historical samples are used to indicate the resource occupancy of the task queue in the historical time period; based on the multiple target samples corresponding to each task queue, predict the expected resource occupancy of each task queue in the future; based on the expected resource occupancy and the target data volume, determine the target queue in the multiple task queues for processing the first task.
[0177] In a possible implementation, the task processing request is further used to indicate a processing priority of the first task, where the processing priority is determined based on the target resource amount.
[0178] In a possible implementation, the processing priority is further related to an expected running time of the first task, where the expected running time is determined based on a target data volume and a target resource volume.
[0179] In a possible implementation, the first task is a task that is periodically executed in the computing cluster.
[0180] In a possible implementation, the first task is a SQL-based data query task, and the data screening conditions include one or more of a customized data selection condition, a data statistical field, a data statistical dimension, and a data screening time period.
[0181] The acquisition module 901 and the processing module 902 can be implemented by software or hardware. For example, the implementation of the processing module 902 will be described below using the processing module 902 as an example. Similarly, the implementation of the acquisition module 901 can refer to the implementation of the processing module 902.
[0182] The processing module 902 is an example of a software functional unit. The processing module 902 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing module 902 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0183] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0184] As an example of a hardware functional unit, processing module 902 may include at least one computing device, such as a server. Alternatively, processing module 902 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0185] The multiple computing devices included in processing module 902 can be distributed in the same region or in different regions. The multiple computing devices included in processing module 902 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in processing module 902 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0186] It should be noted that the information interaction, implementation process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and no further details will be given here.
[0187] The present application also provides a computing device 1000. Figure 10 , Figure 10 This is a schematic diagram of the structure of a computing device 1000 provided in an embodiment of the present application. Figure 10 As shown, computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. Computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1000.
[0188] The bus 1002 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 The bus 1002 may include a path for transmitting information between various components of the computing device 1000 (eg, the memory 1006, the processor 1004, and the communication interface 1008).
[0189] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0190] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0191] The memory 1006 stores executable program code, and the processor 1004 executes the executable program code to implement the functions of the aforementioned acquisition module and processing module, thereby implementing the aforementioned task resource allocation method. In other words, the memory 1006 stores instructions for executing the task resource allocation method.
[0192] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0193] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0194] See also Figure 11 , Figure 11 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. Figure 11 As shown, the computing device cluster includes at least one computing device 1000. The memory 1006 of one or more computing devices 1000 in the computing device cluster may store the same instructions for the resource allocation method for executing tasks.
[0195] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the resource allocation method for the task. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the resource allocation method for the task.
[0196] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster can store different instructions, each for executing a portion of the functions of the data processing apparatus. In other words, the instructions stored in the memory 1006 in different computing devices 1000 can implement the functions of one or more of the aforementioned acquisition module and processing module.
[0197] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 12 A possible implementation is shown. Figure 12 This is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. Figure 12 As shown, in computing device cluster 1200, two computing devices 1000A and 1000B are connected via a network. Specifically, the connection to the network is achieved through a communication interface within each computing device. In this possible implementation, the memory 1006 within computing device 1000A stores instructions for executing the functions of the acquisition module. Simultaneously, the memory 1006 within computing device 1000B stores instructions for executing the functions of the processing module.
[0198] It should be understood that Figure 12 The functionality of the computing device 1000A shown in FIG. 1 may also be implemented by multiple computing devices 1000. Similarly, the functionality of the computing device 1000B may also be implemented by multiple computing devices 1000.
[0199] See Figure 13 , Figure 13 This is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. This application also provides a computer-readable storage medium. In some embodiments, the execution process described in the above embodiments can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.
[0200] Figure 13 Schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.
[0201] In one embodiment, the computer-readable storage medium 1300 is provided using a signal-bearing medium 1301. The signal-bearing medium 1301 may include one or more program instructions 1302, which when executed by one or more processors may provide the functions or part of the functions described above for the database system.
[0202] In some examples, signal bearing medium 1301 may include computer readable medium 1303 such as, but not limited to, a hard drive, compact disk (CD), digital video disk (DVD), digital tape, memory, ROM or RAM, and the like.
[0203] In some embodiments, the signal-bearing medium 1301 may include a computer-recordable medium 1304, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1301 may include a communication medium 1305, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1301 may be communicated via a wireless form of the communication medium 1305 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocol).
[0204] The one or more program instructions 1302 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1302 communicated to the computing device via one or more of computer-readable media 1303, computer-recordable media 1304, and / or communication media 1305.
[0205] The present application also provides a computer program product including instructions. This computer program product can be software or a program product including instructions that can be executed on a computing device or stored on any available medium. When the computer program product is executed on at least one computing device, it causes the at least one computing device to perform the resource allocation method described in the above embodiments.
[0206] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.
[0207] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
Claims
1. A method for allocating resources for a task, characterized in that: include: Obtaining a first task, where the first task is used to instruct the computing cluster to process data that meets a data screening condition; Based on the data screening condition and the total amount of data accessed by the computing cluster, predicting a target amount of data that needs to be processed by the computing cluster when executing the first task; Based on the target data volume, predicting a target resource volume required for executing the first task, where the target resource volume indicates a processor resource volume and / or a memory volume; A task processing request is sent to the computing cluster, where the task processing request is used to request the computing cluster to process the first task and specify the amount of resources allocated by the computing cluster to the first task as the target amount of resources.
2. The method according to claim 1, characterized in that The method further comprises: Determining, based on the data screening condition, a data redistribution operation involved in the execution of the first task, wherein the data redistribution operation is used to redistribute data processed by multiple computing nodes on the computing cluster; The predicting, based on the target data volume, a target resource volume required for executing the first task includes: Based on the target data volume and the data redistribution operation, a target resource volume required for execution of the first task is predicted.
3. The method according to claim 2, characterized in that The predicting, based on the target data volume, a target resource volume required for executing the first task includes: Based on the target data volume and the data redistribution operation, predicting a target resource volume required for executing the first task through a regression model; The regression model is obtained by fitting based on historical task processing data of the computing cluster.
4. The method according to claim 3, characterized in that The task processing efficiency corresponding to the historical task processing data meets a preset condition, and the task processing efficiency is related to the data volume of the task, the amount of resources allocated to the task, and the execution time of the task.
5. The method according to any one of claims 1 to 4, characterized in that The computing cluster includes multiple task queues for processing tasks. The task processing request is further used to specify a target queue for processing the first task, and the target queue is one of the multiple task queues.
6. The method according to claim 5, characterized in that The method further comprises: Based on the resource occupancy of each of the multiple task queues at the current time, retrieving a plurality of target samples with the closest resource occupancy for each task queue from historical samples of the multiple task queues, the historical samples being used to indicate the resource occupancy of the task queue in the historical time period; Based on the multiple target samples corresponding to each task queue, predict the expected resource occupancy of each task queue in the future; Based on the expected resource occupancy and the target data volume, the target queue is determined among the multiple task queues for processing the first task.
7. The method according to claim 5 or 6, characterized in that The task processing request is further used to indicate a processing priority of the first task, where the processing priority is determined based on the target resource amount.
8. The method according to claim 7, characterized in that The processing priority is further related to an expected execution time of the first task, where the expected execution time is determined based on the target data volume and the target resource volume.
9. The method according to any one of claims 1 to 8, characterized in that The first task is a task that is periodically executed in the computing cluster.
10. The method according to any one of claims 1 to 9, characterized in that: The first task is a data query task based on the structured query language SQL, and the data screening conditions include one or more of a customized data selection condition, a data statistical field, a data statistical dimension, and a data screening time period.
11. A task resource allocation device, characterized in that: include: an acquisition module, configured to acquire a first task, wherein the first task is configured to instruct the computing cluster to process data that meets a data screening condition; a processing module, configured to predict a target amount of data to be processed by the computing cluster when executing the first task based on the data screening condition and the total amount of data accessed by the computing cluster; The processing module is further configured to predict a target amount of resources required for executing the first task based on the target data amount, where the target amount of resources indicates an amount of processor resources and / or an amount of memory; The processing module is further configured to send a task processing request to the computing cluster, wherein the task processing request is configured to request the computing cluster to process the first task and specify that the amount of resources allocated by the computing cluster to the first task is the target amount of resources.
12. The device according to claim 11, characterized in that The processing module is further configured to: Determining, based on the data screening condition, a data redistribution operation involved in the execution of the first task, wherein the data redistribution operation is used to redistribute data processed by multiple computing nodes on the computing cluster; Based on the target data volume and the data redistribution operation, a target resource volume required for execution of the first task is predicted.
13. The device according to claim 12, characterized in that The processing module is further configured to: Based on the target data volume and the data redistribution operation, predicting a target resource volume required for executing the first task through a regression model; The regression model is obtained by fitting based on historical task processing data of the computing cluster.
14. The device according to claim 13, characterized in that The task processing efficiency corresponding to the historical task processing data meets a preset condition, and the task processing efficiency is related to the data volume of the task, the amount of resources allocated to the task, and the execution time of the task.
15. The device according to any one of claims 11 to 14, characterized in that The computing cluster includes multiple task queues for processing tasks. The task processing request is further used to specify a target queue for processing the first task, and the target queue is one of the multiple task queues.
16. The device according to claim 15, characterized in that The processing module is further configured to: Based on the resource occupancy of each of the multiple task queues at the current time, retrieving a plurality of target samples with the closest resource occupancy for each task queue from historical samples of the multiple task queues, the historical samples being used to indicate the resource occupancy of the task queue in the historical time period; Based on the multiple target samples corresponding to each task queue, predict the expected resource occupancy of each task queue in the future; Based on the expected resource occupancy and the target data volume, the target queue is determined among the multiple task queues for processing the first task.
17. The device according to claim 15 or 16, characterized in that The task processing request is further used to indicate a processing priority of the first task, where the processing priority is determined based on the target resource amount.
18. The device according to claim 17, characterized in that The processing priority is further related to an expected execution time of the first task, where the expected execution time is determined based on the target data volume and the target resource volume.
19. The device according to any one of claims 11 to 18, characterized in that The first task is a task that is periodically executed in the computing cluster.
20. The device according to any one of claims 11 to 19, characterized in that The first task is a data query task based on the structured query language SQL, and the data screening conditions include one or more of a customized data selection condition, a data statistical field, a data statistical dimension, and a data screening time period.
21. A computing device, characterized in that The computer comprises a processor and a memory; the processor is configured to execute instructions stored in the memory, so that the computing device performs the operating steps of the method according to any one of claims 1 to 10.
22. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster performs the operating steps of the method according to any one of claims 1 to 10.
23. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 10.
24. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 10.
Citation Information
Cited By
Task execution time adjustment method and device, medium and product
CN121350084A